Get Started
Menu

Graphical processing unit (GPU) acceleration has become the standard for high-performance computing, including leading-edge scientific computing. D2S recognized the opportunity presented by the rise of GPU computing in 2009, shortly after NVIDIA introduced general-purpose GPUs (GPGPUs) with CUDA, and focused exclusively since then on solutions for semiconductor design and manufacturing that harnessed the power of GPU acceleration.

D2S understood that a mere “port” to GPU would not be enough to reap the full benefits, so D2S solutions have been built from the ground up on algorithms specifically designed for GPU acceleration. With GPU acceleration, D2S solutions have made full-reticle mask design, correction, and analysis of any shapes including Entirely Manufacturable™ curvilinear masks a practical reality.

GPUs Are Ideal for Scientific Computing

GPUs are a single-instruction multiple data (SIMD) computing architecture, meaning that at any given time, all tens of thousands of cores are executing the same instructions in parallel. On CPUs, there are tens and maybe hundreds of processors on a chip, but they can operate independently of each other.

For example, if there is an IF-THEN-ELSE clause in the software, two cores on the CPU can execute different code paths at the same time. One core can be executing the THEN code for an execution path, while another can be executing the ELSE code. In fact, all cores would be executing totally different parts of the code. On a SIMD machine like GPUs, all cores execute the THEN code or wait and do nothing and then all cores execute the ELSE code or wait and do nothing. So, logical processing like reading input formats is better suited for CPUs. But scientific computing tends to be better on GPUs. This is because nature is inherently SIMD. There isn’t an IF-THEN-ELSE in nature. That’s a human “reasoning” construct.

The different atoms of any given material act on the same math of the physics and chemistry that dictate their particular stimulus-response. The same math acts on all data. What produces different results is the data, not the program. A resist coated on a mask or wafer reacting to sources of energy like 193i, EUV, or eBeam is SIMD. The math of the physics and chemistry is exactly the same everywhere. Again, what produces different results is the differences in data. This is why scientific computing, including in semiconductor manufacturing, is very well suited for GPU acceleration.

D2S Rethinks the Problem and Uses Pixel-Based Computing

A simple port to GPU-based computing from central-processing unit (CPU)-based computing typically realizes 2x-4x speedup by adding GPUs to the servers. To realize the >10X acceleration potential of GPU-based computing, the computing approach needs to be reconsidered from the ground up.

With the NVIDIA Ampere generation of GPGPUs in 2016, physics- and chemistry-based computing at the full-reticle scale with GPU acceleration became practical and cost-effective for the first time. This enabled the computing paradigm to shift from rule-based, abbreviated “reasoning” to mathematical and uniform processing of the whole surface across the reticle. The specific approach taken by D2S was to foundationally change the computing paradigm to be in the pixel domain rather than manipulating edges of polygons.

Once data is rasterized into pixels, and the pixel data is transferred onto a GPU, computing – even very sophisticated scientific computing – has very little overhead. But getting the data in and out of the GPUs, doing “IF-THEN-ELSE” work like reading and outputting files, or rasterizing and contouring then become the bottleneck in overall computing. This is why the actual simulation for a lithography simulation or a multi-Gaussian mask simulation may be 100-1000x faster on GPUs once you get the data to the GPUs, but the overall runtime of GPU-accelerated computing may only be 10-20x faster. However, that is significantly faster than the 2-4x that can be achieved by porting code that was conceived for CPUs and just accelerating that algorithm. As Jensen Huang, CEO of NVIDIA, has said, you must fundamentally rethink the problem in order to realize the maximum potential of GPU acceleration. For D2S, pixel-based computing is that rethink.

D2S GPU+CPU Solutions Produce Optimal Overall Performance

Optimal GPU acceleration is not the result of a simple replacement of GPUs for CPUs. To realize the >10X acceleration potential of GPU-based computing, you must understand when to deploy GPUs and when to use CPUs. The D2S GPU-acceleration approach employs sophisticated software engineering to combine the strength of each to the benefit of the whole system, deciding what to put on CPUs, what to put on GPUs, then scheduling and load-balancing the two to achieve optimal overall performance.

GPU Computing Makes Curvilinear Processing Practical

For more than a decade, the semiconductor industry has recognized that curvilinear shapes on photomasks computed by inverse lithography technology (ILT) produce the best wafer quality, but adoption was hindered by long mask write times using conventional variable-shaped beam (VSB) writing, as well as long ILT runtimes on CPU-based computing platforms. The latest-generation D2S Computational Design Platform (CDP) with GPU acceleration makes implementing and verifying curvilinear ILT a practical reality, enabling TrueMask® ILT to output Entirely Manufacturable™ mask designs.

D2S software applications are based on NVIDIA CUDA, a parallel computing platform and programming model for GPUs. The D2S CDP with GPU acceleration enables simulation-based, accurate manipulation and analysis, particularly for curvilinear shapes, which are not practical with CPU-only applications.

From creating and processing complex mask shapes to helping to write the masks and analyzing mask SEM data to providing deep learning engines, D2S GPU-accelerated solutions help customers to achieve manufacturing success on their leading-edge mask and chip designs.