Method, system, computer program, and computing device for determining geometries for semiconductor or flat panel display manufacturing

By using neural networks to adjust parameters and employing advanced lithography techniques, the method addresses the challenges of accurately transferring small critical dimensions in optical lithography, improving manufacturing accuracy and efficiency.

JP7672615B2Active Publication Date: 2025-05-08D2S INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023524691
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-22
Filing Date
2021-10-15
Publication Date
2025-05-08
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

Current optical lithography techniques face challenges in accurately transferring small critical dimensions of circuit patterns onto substrates due to resolution limits and manufacturing variations, leading to increased complexity and cost in reticle pattern design and mask writing.

Method used

The method involves inputting a physical design pattern and generating multiple potential mask designs and substrate patterns by calculating variations in manufacturing steps, using neural networks to adjust parameters and reduce manufacturing variations, and employing techniques like OPC and ILT to enhance pattern accuracy.

Benefits of technology

This approach improves pattern manufacturing accuracy and reduces calculation time by modeling multiple parameters simultaneously, allowing for real-time visualization and modification of design variations, thereby enhancing the reliability and efficiency of semiconductor manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672615000004
    Figure 0007672615000004
  • Figure 0007672615000005
    Figure 0007672615000005
  • Figure 0007672615000006
    Figure 0007672615000006
Patent Text Reader

Abstract

A method for calculating a pattern to be fabricated on a substrate includes inputting a physical design pattern, determining a plurality of possible neighborhoods for the physical design pattern, generating a plurality of possible mask designs for the physical design pattern, calculating the plurality of possible patterns on the substrate, calculating a variation band from the plurality of possible patterns, and modifying the physical design pattern to reduce the variation band. An embodiment includes inputting a set of parameters for a neural network to calculate a pattern to be fabricated on the substrate, calculating a plurality of patterns to be fabricated on the substrate for the physical design in each of a plurality of possible neighborhoods, training the neural network with the calculated plurality of patterns, and adjusting the set of parameters to reduce manufacturing variation of the calculated plurality of patterns to be fabricated on the substrate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to lithography, and more particularly to the design and manufacture of a surface, be it a reticle, wafer, or any other surface, using charged particle beam lithography. [Background technology]

[0002] Three common types of charged particle beam lithography are unshaped (Gaussian) beam lithography, shaped charged particle beam lithography, and multi-beam lithography. In all types of charged particle beam lithography, a charged particle beam delivers energy to a resist-coated surface, exposing the resist.

[0003] In the production or manufacturing of semiconductor devices such as integrated circuits, optical lithography is used to fabricate the semiconductor devices. Optical lithography is a printing process that uses a lithographic mask or photomask made from a reticle to form a pattern on a substrate such as a semiconductor or silicon wafer to fabricate an integrated circuit. Other substrates include flat panel displays and other reticles. Extreme ultraviolet (EUV) and X-ray lithography are also considered types of optical lithography. One or more reticles contain a circuit pattern that corresponds to each layer of the integrated circuit. The pattern is imaged onto an area on a substrate that is coated with a layer of radiation-sensitive material known as photoresist or resist. Once a pattern layer is formed, the layer is treated with various other processes such as etching, ion implantation (doping), metallization, oxidation, and polishing. These processes are employed to finish each layer in the substrate. If several layers are required, the entire process or a variation of it is repeated for each new layer. Eventually, a combination of multiple devices or integrated circuits will be present on the substrate. These integrated circuits are then separated from one another by dicing or sawing before being mounted in their respective packages. In the more general case, the patterns on the substrate can be used to define artifacts such as display pixels or magnetic recording heads.

[0004] In the production or manufacturing of semiconductor devices such as integrated circuits, maskless direct writing can also be used to manufacture semiconductor devices. Maskless direct writing is a printing process in which charged particle beam lithography is used to form patterns on substrates such as semiconductor or silicon wafers to manufacture integrated circuits. Other substrates can include flat panel displays, imprint masks for nanoimprinting, or reticles. The desired pattern of a layer is written directly onto a surface, which in this case is also the substrate. Once a pattern layer is created, the layer is treated with various other processes such as etching, ion implantation (doping), metallization, oxidation, and polishing. These processes are used to finish each layer in the substrate. If several layers are required, the entire process or a variation of it is repeated for each new layer. Some of the layers may be written using optical lithography, and others may be written using maskless direct writing to manufacture the same substrate. Eventually, a combination of multiple devices or integrated circuits will be present on the substrate. These integrated circuits are then separated from each other by dicing or sawing, and then mounted in their respective packages. More often, patterns on a surface can be used to define artifacts such as display pixels or magnetic recording heads.

[0005] In optical lithography, a lithographic mask or reticle comprises a geometric pattern corresponding to the circuit components to be integrated on a substrate. The patterns used to manufacture the reticle can be generated using computer-aided design (CAD) software or programs. In designing the pattern, the CAD program can follow certain design rules to create the reticle. These rules are set by process, design, and end-use limitations. An example of an end-use limitation is to dictate the geometry of a transistor so that it cannot operate satisfactorily at a required supply voltage. In particular, design rules can define the tolerance between circuit devices or interconnect lines. Design rules are used, for example, to ensure that circuit devices or lines do not interact with each other in an undesirable way. For example, design rules are used to ensure that lines are not too close together that would cause a short circuit. Design rule limitations reflect, among other things, the minimum dimensions that can be manufactured reliably. When referring to small dimensions, the idea of ​​a critical dimension is usually introduced. A critical dimension is defined, for example, as the critical width or area of ​​a feature, or the critical space or space area between two features. These dimensions require exquisite control. Due to the nature of integrated circuit design, many patterns in a design are repeated in different locations. A pattern may be repeated hundreds or thousands of times, and each copy of the pattern is called an instance. When a design rule violation is found in such a pattern, hundreds or thousands of violations are reported (one for each instance of the pattern).

[0006] One goal in integrated circuit manufacturing by optical lithography is to reproduce an original circuit design on a substrate through the use of a reticle. A reticle, also called a mask or photomask, is the surface that is exposed to light during manufacturing using charged particle beam lithography. Integrated circuit manufacturers are always trying to use the real estate of semiconductor wafers as efficiently as possible. Engineers continue to reduce the size of circuits so that integrated circuits contain more circuit elements and use less power. As the size of the critical dimensions of integrated circuits is reduced and the circuit density increases, the critical dimensions of the circuit pattern or physical design approach the resolution limit of the optical exposure tools used in conventional optical lithography. As the critical dimensions of the circuit pattern become smaller and approach the resolution value of the exposure tools, it becomes more difficult to accurately transfer the physical design into the actual circuit pattern developed on the resist layer. To further the use of optical lithography to form patterns with features smaller than the wavelength of light used in the optical lithography process, a process known as optical proximity correction (OPC) has been developed. OPC alters the physical design to compensate for distortions caused by, for example, optical diffraction and optical interaction of the feature with the nearest feature. Resolution enhancement techniques (RET) performed with a reticle include, for example, OPC and inverse lithography techniques (ILT).

[0007] OPC can add sub-resolution lithography features to a mask pattern to reduce the difference between the pattern of the original physical design, i.e., between the design and the circuit pattern that will ultimately be formed on the substrate. The sub-resolution lithography features interact with and interact with the original pattern in the physical design to compensate for proximity effects and thereby improve the final circuit pattern. One feature that is added to improve pattern formation is called a "serif." A serif is a small feature that increases the accuracy or resilience to manufacturing variations in the printing of a particular feature. An example of a serif is a small feature that is placed at the corner of a pattern to sharpen the corner in the final image. The pattern that is intended to be printed on the substrate is called the primary feature. There is a long discussion of OPC features, including OPC-decorated patterns written on a reticle relative to the primary features, i.e., features that reflect the design before OPC decoration, serifs, jogs, sub-resolution assist features (SRAFs), and negative features. SRAFs are isolated shapes that are not attached to the main feature and are small enough not to be printed on the substrate, while serifs, jogs, and negative features modify the main feature. OPC features follow various design rules, including rules based on the size of the smallest feature that can be formed on a wafer using optical lithography. Other design rules come from the mask manufacturing process, or from the stencil manufacturing process if a character projection (CP) type charged particle beam writing system is used to form the pattern on the reticle. Summary of the Invention

[0008] In an embodiment, a method for calculating a pattern to be manufactured on a substrate includes inputting a physical design pattern and determining a plurality of possible neighborhoods for the physical design pattern. A plurality of possible mask designs are generated for the physical design pattern, the plurality of possible mask designs corresponding to the plurality of possible neighborhoods. A plurality of possible patterns on the substrate are calculated, the plurality of possible patterns on the substrate corresponding to the plurality of possible mask designs. Variation bands from the plurality of possible patterns on the substrate are calculated, and the physical design pattern is modified to reduce the variation bands.

[0009] In an embodiment, a method for computing a pattern to be fabricated on a substrate includes inputting a physical design, inputting a set of parameters for a neural network to compute a pattern to be fabricated on the substrate, and generating a plurality of possible neighborhoods for the physical design. A plurality of patterns to be fabricated on the substrate are computed for the physical design in each of the plurality of possible neighborhoods. The neural network is trained on the computed plurality of patterns, the training being performed using a computing hardware processor. The set of parameters is adjusted to reduce manufacturing variations of the computed plurality of patterns to be fabricated on the substrate. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 illustrates an example of a variable shaped beam system known in the art. [Diagram 2] FIG. 1 shows an example electro-optical schematic of a multi-beam exposure system known in the art. [Figure 3A] FIG. 1 illustrates an example of a rectangular shot as known in the art. [Figure 3B] FIG. 1 illustrates an example of a circular character projection shot as known in the art. [Figure 3C] FIG. 1 illustrates an example of a trapezoid shot as known in the art. [Figure 3D] FIG. 1 shows an example of a drag shot known in the art. [Figure 3E] FIG. 1 shows an example of a shot of an array of circular patterns as known in the art. [Figure 3F] FIG. 1 illustrates an example of a shot of a rectangular pattern sparse array as known in the art. [Figure 4] 1 illustrates an example of a multi-beam charged particle beam system known in the art. [Figure 5A] FIG. 2 shows an example of a cross-sectional dose graph illustrating resist pattern width for each of two resist thresholds known in the art. [Figure 5B] FIG. 5B shows an example of a cross-sectional dose graph similar to FIG. 5A, but with a higher dose edge gradient than FIG. 5A, known in the art. [Figure 6] FIG. 1 shows an example of an orientation variation of a standard cell design known in the art. [Figure 7] FIG. 1 illustrates an example of a physical design flow in accordance with some embodiments. [Figure 8] FIG. 1 illustrates an example of a single input / output neural network in accordance with some embodiments. [Figure 9] FIG. 1 illustrates details of a single input / output neural network in accordance with some embodiments. [Figure 10] FIG. 1 illustrates an example of a multiple input / output neural network in accordance with some embodiments. [Figure 11] FIG. 1 illustrates an example of an input physical design, a calculated mask image, and a generated post-deep learning image in accordance with some embodiments. [Figure 12] FIG. 1 illustrates an example of a neural network with post-processing according to some embodiments. [Figure 13] FIG. 13 illustrates an example of a computed mask image and a generated post-deep learning image according to some embodiments. [Figure 14] FIG. 13 illustrates an example of a computed mask image and a generated post-deep learning image according to some embodiments. [Figure 15]FIG. 1 illustrates a single multi-corner neural network and post-processing steps according to some embodiments. [Figure 16] FIG. 1 illustrates an example of a neural network with multiple output channels in accordance with some embodiments. [Figure 17] FIG. 1 illustrates a schematic of a GPU system in accordance with some embodiments. [Figure 18] FIG. 1 illustrates a schematic of a GPU system in accordance with some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] This disclosure describes a method and system that improves the manufacturing accuracy and computation time of a pattern. The embodiments allow for simultaneous modeling of multiple parameters at different stages of the manufacturing process, e.g., physical design, mask and substrate stages. The results of multiple scenarios are output, such as visualized schematics, for the user to view and modify in near real-time. The embodiments estimate the variations in the mask design and wafer manufacturing steps and utilize statistical methods to improve the physical design of the pattern.

[0012] A typical RET method has OPC verification to identify and correct hot spots. Hot spots are areas that require ideal conditions to print properly. Therefore, they are not resilient to manufacturing variations or may not print properly even in ideal conditions. Hot spots reduce yield. In lithography, features that are needed on a substrate, called main features, are found to print with higher fidelity and improved process windows when SRAFs are added that are too small to print themselves but have a positive effect on how nearby main features are printed.

[0013] However, adding OPC features such as SRAF is a very tedious task, requires costly computation time, and results in a more expensive reticle. Not only are OPC patterns complex, but the exact OPC pattern at a given location depends heavily on what other feature arrangements are nearby, as optical proximity effects occur at distances greater than the minimum line and space dimensions. Thus, for example, a line end has different size serifs depending on what is close to it on the reticle. This is so even when the goal is to form identical shapes on the wafer. These slight but critical variations are significant and have prevented others from forming reticle patterns that precisely form the desired shapes on the wafer. To quantify what is meant by slight variations, typical slight variations in OPC decoration from neighborhood to neighborhood are 5% to 80% of the main feature size. If these OPC variations are to form substantially identical patterns on the wafer, the goal is for the feature arrangements on the wafer to be the same within a certain tolerance specified depending on the details of the function that the feature arrangement is designed to perform, e.g., transistor or wire details. Nevertheless, typical specifications are 2% to 50% of the main feature range.

[0014] Inverse Lithography Technology (ILT) is one of the OPC techniques. ILT is a process of directly calculating the pattern to be formed on a reticle from the pattern formed on a substrate, such as a silicon wafer. It involves simulating the optical lithography process in reverse using the desired pattern on the substrate as an input. ILT calculated reticle patterns may be purely curvilinear, i.e. completely non-linear, such as circular, near-circular, annular, near-annular, elliptical and / or near-elliptical patterns. These patterns are impractical for conventional variable shaped beam (VSB) mask writers with fracturing, as so many variable shaped beam (VSB) shots would be required to expose the curved patterns. A linear approximation or linearization of the curved patterns may be used. However, the linear approximation reduces accuracy compared to the ideal ILT curve patterns. Furthermore, if a linear approximation is generated from an ideal ILT curve pattern, the overall calculation time increases compared to the ideal ILT curve pattern. Mask writing time is a critical business factor, and VSB writing time is proportional to the number of VSB shots that need to be printed. Model-based mask data preparation using overlapping shots can significantly reduce the write-time impact of curved ILT mask designs. However, curved features generally take longer to write than straight features.

[0015] Multi-beam writing eliminates the need to perform linearization to convert curved features for VSB writing. However, mask printability and resilience to manufacturing variations remain important considerations for mask features output by ILT. For example, features that are too small, too close together, or have contour bends that are too sharp will make the mask difficult to manufacture reliably, especially in terms of manufacturing variations. The remaining problem with ILT is the enormous computational demands of dense simulation of the full design, especially the full mask layer of the full reticle size design, which for semiconductor manufacturing typically has wafer dimensions of about 3.0 cm x 2.5 cm.

[0016] Referring now to the drawings, like numbers refer to like features. FIG. 1 illustrates one embodiment of a lithography system, such as a charged particle beam writing system, here an electron beam writing system 10, that uses a variable shaped beam (VSB) to produce a surface 12. The electron beam writing system 10 includes an electron beam source 14 that projects an electron beam 16 toward an aperture plate 18. The plate 18 has an aperture 20 formed therein for the electron beam 16 to pass through. Once the electron beam 16 passes through the aperture 20, it is directed or deflected by a lens system (not shown) as an electron beam 22 toward another rectangular aperture plate or stencil mask 24. The stencil 24 has a number of apertures or openings 26 formed therein that define a variety of simple shapes, such as rectangles and triangles. Each opening 26 formed in the stencil 24 is used to form a pattern on a surface 12 of a substrate 34, such as a silicon wafer, reticle, or other substrate. After the electron beam 30 exits one of the apertures 26, it passes through an electromagnetic or electrostatic demagnification lens 38 which reduces the size of the pattern emerging from the aperture 26. In commonly available charged particle beam writing systems, the reduction factor is between 10 and 60. The demagnified electron beam 40 exits the demagnification lens 38 and is directed onto the surface 12 as a pattern 28 by a series of deflectors 42. The surface 12 is coated with a resist (not shown) which reacts with the electron beam 40. The electron beam 22 may be directed to overlap a variable portion of the aperture 26 which affects the size and shape of the pattern 28. A blanking plate (not shown) may be used to deflect the beam 16 or the shaped beam 22 to prevent the electron beam from reaching the surface 12 during the period after each shot when the lenses and deflectors 42 directing the beam 22 are realigned for the subsequent shot. Conventionally, the blanking period may be a fixed length of time or may vary depending, for example, on how much the deflectors 42 need to be realigned for the position of the next shot.

[0017] In the electron beam writing system 10, the substrate 34 is mounted on a movable platform or stage 32. The stage 32 allows the substrate 34 to be repositioned so that a pattern larger than the maximum deflection capability or field size of the charged particle beam 40 can be written on the surface 12 in a series of subfields that are within the capability of the deflector 42 to deflect the beam 40. In one embodiment, the substrate 34 may be a reticle. In this embodiment, the reticle is exposed with a pattern and then undergoes various manufacturing steps to become a lithography mask or photomask. This mask is then used in an optical lithography apparatus to project an image of the reticle pattern 28, typically at a reduced size, onto a silicon wafer to manufacture integrated circuits. More typically, the mask is used in another device or apparatus to form the pattern 28 on a substrate (not shown).

[0018] Charged particle beam systems can expose a surface with multiple individually controllable beams or beamlets. FIG. 2 shows an electro-optical schematic in which there are three charged particle beamlets 210. Each beamlet 210 has an associated beam controller 220. Each beam controller 220, for example, allows the associated beamlet 210 to hit the surface 230 or prevents the beamlet 210 from hitting the surface 230. In some embodiments, the beam controller 220 may control the beam blur, magnification, size and / or shape of the beamlet 210. In this disclosure, a charged particle beam system with multiple individually controllable beamlets is referred to as a multi-beam system. In some embodiments, charged particles from a single source may be subdivided to form multiple beamlets 210. In other embodiments, multiple sources may be used to generate multiple beamlets 210. In some embodiments, the beamlets 210 may be shaped by one or more apertures, and in other embodiments, there may be no apertures to shape the beamlets. Each beam controller 220 allows for individual control of the exposure duration of the associated beamlet. Typically, the beamlets are reduced in size by one or more lenses (not shown) before striking a surface 230, which is typically coated with resist. In some embodiments, each beamlet may have a separate electro-optic lens, while in other embodiments, multiple beamlets, possibly including all beamlets, share an electro-optic lens.

[0019] For purposes of this disclosure, a shot is an exposure of a surface area over a period of time. The surface area may be comprised of multiple areas that are discontinuous and smaller than the surface area. A shot may be comprised of multiple other shots. The multiple shots may or may not overlap and may or may not be exposed simultaneously. A shot may include a specified dose. The dose may not be specified. A shot may use a shaped beam, an unshaped beam, or a combination of shaped and unshaped beams. Figures 3A-3F show various different shots. Figure 3A shows an example of a rectangular shot 310. A VSB charged particle beam system can, for example, produce rectangular shots with various X and Y dimensions. Figure 3B shows an example of a character projection (CP) shot 320, which in this example is circular. Figure 3C shows an example of a trapezoidal shot 330. In one embodiment, the shot 330 can be generated using a raster scanned charged particle beam, where the beam is scanned in the X direction, for example, as shown by scan line 332. 3D shows an example of a dragged shot 340 as disclosed in U.S. Patent Application Publication No. 2011 / 0089345. Shot 340 is formed by exposing a surface with a curved shaped beam 342 at an initial nominal position 344 and then moving shaped beam 342 across the surface from initial position 344 to position 346. The path of the dragged shot may be, for example, linear, piecewise linear, or curvilinear.

[0020] FIG. 3E shows an example of a shot 350 that is an array of circular patterns 352. Shot 350 can be formed in a variety of ways, including multiple shots of a single circular CP character, one or more shots of a CP character that is an array of circular apertures, and one or more multi-beam shots using a circular aperture. FIG. 3F shows an example of a shot 360 that is a sparse array of rectangular patterns 362 and 364. Shot 360 can be formed in a variety of ways, including multiple VSB shots, a CP shot, and one or more multi-beam shots using a rectangular aperture. In some multi-beam embodiments, shot 360 can include multiple alternating groups of other multi-beam shots. For example, pattern 362 can be shot simultaneously followed by pattern 364 at a different time than pattern 362.

[0021] Many techniques are used to form patterns on a reticle, some using optical lithography and others using charged particle beam lithography. The most commonly used system is the variable shaped beam (VSB), where a dose of electrons exposes a resist-coated reticle surface with simple shapes such as Manhattan rectangles and 45 degree right triangles, as mentioned above. In conventional mask writing, the dose or shot of electrons is traditionally designed to avoid overlap as much as possible, greatly simplifying the calculation of how the resist on the reticle will record the pattern. Similarly, a set of shots is designed to completely cover the pattern area to be formed on the reticle. U.S. Patent No. 7,754,401, owned by the assignee of this patent application, discloses a mask writing method in which deliberate shot overlap is used to write the pattern. When overlapping shots are used, charged particle beam simulation can be used to define the pattern that the resist on the reticle will record. By using overlapping shots, the pattern can be written with fewer shots, or with greater accuracy, or both. U.S. Patent No. 7,754,401 also discloses the use of dose variation, where the assigned dose of a shot is different from the dose of other shots. The term model-based fracturing is used to describe the process of defining shots using the techniques of U.S. Patent No. 7,754,401.

[0022] FIG. 4 illustrates an embodiment of a charged particle beam exposure system 400. The charged particle beam system 400 is a multi-beam system in which multiple individually controllable shaped beams can simultaneously expose a surface. The multi-beam system 400 includes an electron beam source 402 that generates an electron beam 404. The electron beam 404 is directed to an aperture plate 408 by a capacitor 406 that includes electrostatic and / or magnetic elements. The aperture plate 408 has multiple openings 410 through which the electron beam 404 is projected. The electron beam 404 passes through these openings 410 to form multiple shaped beamlets 436. In some embodiments, the aperture plate 408 has hundreds or thousands of openings 410. FIG. 4 illustrates an embodiment with a single electron beam source 402. In other embodiments, the openings 410 are projected with electrons from multiple electron beam sources. The openings 410 may be rectangular or may be a different shape, for example, circular. The set of beamlets 436 then illuminates a blanking controller plate 432. The blanking controller plate 432 has a number of blanking controllers 434, each of which is aligned with a beamlet 436. Each blanking controller 434 can control its associated beamlet 436 to either cause the beamlet 436 to strike the surface 424 or to prevent the beamlet 436 from striking the surface 424. The amount of time the beam strikes the surface controls the total energy or "dose" applied by the beamlet. Thus, the dose of each beamlet is independently controlled. The area where the beam strikes the surface may encompass a portion of an entire pixel.

[0023] The ability of a multi-beam system to modify the dose at each pixel to bias the edges of a feature is described in U.S. Patent No. 10,444,629, owned by the assignee of the present patent application, which also discloses improving the dose margin so that edges are less susceptible to manufacturing variations. This method of modifying the dose on a pixel-by-pixel basis is referred to as pixel-level dose correction (PLDC).

[0024] In FIG. 4, four beamlets that can strike the surface 424 are illustrated as beamlets 412. In one embodiment, a blanking controller 434 prevents the beamlet 436 from striking the surface 424 by deflecting the beamlet 436 so that it is stopped by an aperture plate 416 that includes an aperture 418. In some embodiments, the blanking plate 432 may be directly adjacent to the aperture plate 408. In other embodiments, the relative positions of the aperture plate 408 and the blanking controller 432 may be reversed from that shown in FIG. 4, such that the beam 404 strikes multiple blanking controllers 434. The lens system, including elements 414, 420, and 422, allows multiple beamlets 412 to be projected onto the surface 424 of the substrate 426, typically at a size smaller than the multiple apertures 410. The reduced beamlets form a beamlet group 440 that strikes the surface 424 and produces a pattern that matches a subset of the apertures 410. A subset is an aperture 410 that allows a beamlet 436 to impinge on surface 424 via a corresponding blanking controller 434. A beamlet group 440 having four beamlets is shown in FIG.

[0025] The substrate 426 is positioned on a movable platform or stage 428 that can be repositioned by an actuator 430. By moving the stage 428, the beam 440 can be exposed in multiple exposures or shots over an area larger than the dimensions of the largest pattern formed by the beamlets 440. In some embodiments, the stage 428 remains stationary during an exposure and is then repositioned for the next exposure. In other embodiments, the stage 428 moves continuously and at a variable speed. In yet other embodiments, the stage 428 moves continuously but at a constant speed. This can increase the accuracy of the stage positioning. In embodiments where the stage 428 moves continuously, a set of deflectors (not shown) can be used to move the beam in a direction and speed consistent with the stage 428, allowing the beamlets 440 to remain stationary relative to the surface 424 during exposure. In yet other embodiments of a multi-beam system, each beamlet in a beamlet group can be deflected across the surface 424 independently of other beamlets in the beamlet group. In some embodiments, stage 428 can move in a single direction across the exposure area to expose portions of the area, referred to as stripes. Thus, the entire exposure area is exposed as multiple stripes. In some embodiments, stage 428 moves in opposite directions on adjacent or alternating stripes.

[0026] Other types of multi-beam systems can generate multiple unshaped beamlets 436, such as by using multiple charged particle beam sources to form an array of Gaussian beamlets.

[0027] 1, the minimum size of a pattern that can be projected with reasonable accuracy onto a surface 12 is limited by a variety of short-range physical effects associated with the electron beam writing system 10 and the surface 12, which typically includes a resist coating on a substrate 34. These effects include forward scattering, the Coulomb effect, and resist diffusion. fBeam blur, or β, is a term used to include all of these short-range effects. Modern electron beam writing systems have effective beam blur radii, or β, in the range of 20 nm to 30 nm. f can be achieved. Forward scattering may constitute ¼ to ½ of the total beam blur. Modern electron beam writing systems are equipped with multiple mechanisms to minimize each of the multiple components of beam blur. Since some components of beam blur are a function of the calibration level of the particle beam writer, the β f may vary. The diffusion characteristics of the resist may also change. f The variations in can be simulated and systematically explained. However, there are other effects that cannot or will not be explained, and which appear as random variations.

[0028] The shot dose of a charged particle beam writer, such as an electron beam writing system, is a function of the intensity of the beam source 14 and the exposure time of each shot. Typically, the beam intensity is fixed and the exposure time is varied to obtain a variable shot dose. Different regions within a shot may have different exposure times, such as a multi-beam shot. The exposure time is varied to compensate for various long-range effects, such as backscatter, fogging, and loading effects, in a process called proximity effect correction (PEC). An electron beam writing system can set an overall dose, usually called a base dose, that affects all shots in an exposure pass. Some electron beam writing systems perform dose compensation calculations within the electron beam writing system itself, so that the dose of each shot is not individually assigned as part of the input shot list. Thus, the input shots have unassigned shot doses. In such electron beam writing systems, all shots have a base dose before PEC. In other electron beam writing systems, the dose can be assigned on a shot-by-shot basis. In an electron beam writing system capable of shot-by-shot dose allocation, the number of available dose levels may range from 64 to 4096 or more, or there may be a relatively small number of available dose levels, for example 3 to 8 levels.

[0029] The mechanisms in e-beam writing systems have a relatively coarse resolution for the calculations, so the mid-range corrections required for EUV masks in the 2 μm range cannot be accurately calculated by current e-beam writing systems.

[0030] For example, when exposing a repeating pattern on a surface using charged particle beam lithography, the size of each pattern instance measured on the final fabricated surface will be slightly different due to manufacturing variations. The amount of size variation is an essential manufacturing optimization criterion. In current mask masking, a root mean square (RMS) variation of 1 nm (1 sigma) or less in pattern size is desired. More size variation leads to more variation in circuit performance, requiring higher design margins and making it increasingly difficult to design faster and lower power integrated circuits. This variation is referred to as critical dimension (CD) variation. Low CD variation is desirable, indicating that manufacturing variations lead to relatively small size variations on the final fabricated surface. At smaller scales, the effect of high CD variation is observed as line edge roughness (LER). LER is caused by each part of the line edge being fabricated slightly differently, resulting in waviness in lines that are intended to have straight edges. CD variation is, among other things, inversely proportional to the slope of the dose curve at the resist threshold, which is referred to as edge slope. Therefore, edge slope or dose margin becomes an important optimization factor for particle beam writing of a surface. In this disclosure, the terms edge slope and dose margin are used interchangeably.

[0031] 5A-5B show how critical dimension (CD) variations are reduced by exposing a pattern on the resist to produce a relatively high edge gradient in the exposure or dose curve, as described in U.S. Pat. No. 8,473,875, “Method and System for Forming High Accuracy Patterns Using Charged Particle Beam Lithography,” owned by the assignee of this patent application. FIG. 5A shows a cross-sectional dose curve 502. The x-axis shows the cross-sectional distance through the exposed pattern (e.g., the distance perpendicular to two of the edges of the pattern) and the y-axis shows the dose received by the resist. A pattern is recorded by the resist when the dose received is higher than a threshold. FIG. 5A shows two thresholds. FIG. 5A shows the effect of variations in resist sensitivity. A higher threshold 504 causes a pattern of width 514 to be recorded by the resist. A lower threshold 506 causes a pattern of width 516 to be recorded by the resist, where width 516 is greater than width 514. FIG. 5B shows another cross-sectional dose curve 522. Two thresholds are shown, where threshold 524 is the same as threshold 504 in FIG. 5A. Threshold 526 is the same as threshold 506 in FIG. 5A. The slope of dose curve 522 is higher than that of dose curve 502 near the two thresholds. For dose curve 522, the higher threshold 524 causes a pattern of width 534 to be recorded by the resist. The lower threshold 526 causes a pattern of width 536 to be recorded by the resist. As can be seen, the difference between width 536 and width 534 is smaller than the difference between width 516 and width 514 due to the higher edge slope of dose curve 522 compared to dose curve 502. If the resist-coated surface is a reticle, the lower sensitivity of curve 522 to variations in resist thresholds can allow the pattern width on a photomask produced from the reticle to be closer to the target pattern width of the photomask. Thereby, when the photomask is used to form a pattern on a substrate, such as a silicon wafer, it increases the yield of usable integrated circuits.Improved tolerance to dose variations with each shot is observed for dose curves with higher edge slopes. Therefore, it is desirable to achieve a relatively high edge slope, such as dose curve 522.

[0032] A design cell in semiconductor manufacturing (e.g., a memory cell or a standard cell from a library) represents an abstracted representation of an electronic component in a physical design. A cell-based approach allows designers to reuse components from relatively simple to complex designs. A cell is composed of several layers that contain shapes of varying size and orientation. A cell, i.e., a set of shapes from a given layer in a cell, is placed relatively isolated in the design with no neighboring shapes in its vicinity, resulting in a different pattern on the substrate than if the cell were placed with other cells and / or shapes in its immediate vicinity, i.e., with different neighboring shapes in close proximity on the same layer. FIG. 6 shows an example of a standard cell that contains two cells, i.e., cells A and B, in various legal orientations. Due to the proximity of placement within cells that are adjacent to each other (i.e., in the same vicinity), each orientation results in a variation in the mask design calculated for each cell. As mentioned earlier, optical proximity correction (OPC) varies to account for light diffraction and light interaction with nearby features. In a proximity effect correction (PEC) refinement step, the shot dose is adjusted as necessary due to different long-range effects for each neighborhood.

[0033] Manufacturing process variations and neighborhood induced variations have a significant impact on design performance and manufacturing reliability. It is desirable for circuit and / or mask designers to visualize the impact of various sources of variation in the context of the actual design. For example, process variations can cause the pattern width on a photomask to vary from the intended or target width. The variation in pattern width on the photomask causes variation in the pattern width on a wafer exposed using the photomask in an optical lithography process. The sensitivity of the wafer pattern width to the variation in the photomask pattern width is referred to as the Mask Edge Error Factor, or MEEF. In an optical lithography system using a 4x photomask, if a 4x reduced version of the photomask pattern is projected onto the wafer by the optical lithography process, for example, a MEEF of 1 means that for a 1 nm error in the pattern width on the photomask, the pattern width on the wafer changes by 0.25 nm. A MEEF of 2 means that for a 1 nm error in the pattern width on the photomask, the pattern width on the wafer changes by 0.5 mm. For the smallest integrated circuit processes, the MEEF is greater than 2. By better visualizing and understanding these sources and effects of variation, a designer can modify the design itself (i.e., the geometry that comprises the design) to be more robust against such variations.

[0034] FIG. 7 is a flow 700 for calculating a pattern to be manufactured on a substrate such as a silicon wafer, according to some embodiments. In a first step, a physical design pattern 702, such as a physical design of an integrated circuit, is input. In one embodiment, the pattern to be manufactured on the substrate is calculated from the physical design pattern. These calculations include determining manufacturable shapes for logic gates, transistors, metal layers, and other components that need to be found in the physical design, such as an integrated circuit. The physical design may be straight, piecewise straight, partially curved, or fully curved. Curved patterns, in particular, are very computationally intensive. Therefore, being able to optimize the pattern by calculating the cumulative effect of variations from multiple manufacturing stages, as in the present embodiment, is very beneficial.

[0035] Step 704 includes generating a plurality of possible neighborhoods for the physical design. In some embodiments, the physical design pattern is part of an overall design. The plurality of possible neighborhoods generated in step 704 are a plurality of actual neighborhoods used for the physical design pattern. The neighborhood variations can be compounded. For example, one method is to randomly place the cell in all possible neighborhoods where it may end up, i.e., all possible neighborhoods that are surrounded by the various neighboring cells that are most likely to be surrounded by the actual circuit design. In some embodiments, the portion of the physical design pattern is an instance of the physical design pattern, and the plurality of possible neighborhoods includes all the neighborhoods of each instantiation. Thus, an instance of the cell of interest is placed in its various legal orientations, with various neighboring instances placed facing up, down, left or right, and with various offsets in its placement, parallel to various orientations of the various neighboring cells. In some embodiments, the portion of the overall design is a standard cell design including a plurality of standard cells, and the plurality of possible neighborhoods includes all legal orientations of the standard cells.

[0036] In step 706, a composite of the substrate layer, some of which are separated into mask layers, is formed from the physical design. This step may also be referred to as a coloring step or colorization, in which each feature on the reticle layer is colored to reflect the assignment of the feature to a particular mask layer. The colorization step 706 may be performed on the physical design pattern prior to optical proximity correction (OPC). In step 708, OPC is performed on the physical design pattern to generate a plurality of possible mask designs 710, each mask design in the plurality of mask designs corresponding to a plurality of possible neighborhoods generated in step 704. The plurality of possible mask designs 710 may be combined to form a nominal mask design having variations. Conventionally, the nominal mask design is determined by calculating a nominal contour of the mask design using a nominal dose, such as 1.0, and a threshold value, such as 0.5. In one embodiment, the nominal contour of the mask design is calculated from the plurality of possible mask designs 710. The variations are calculated for all possible neighborhoods generated in step 704.

[0037] In one embodiment of the present disclosure, the OPC step 708 includes an ILT that creates an ideal curved ILT pattern, while in another embodiment, an ILT with linearization of the curved pattern is used.

[0038] OPC features or ILT patterns for the same physical design pattern vary from neighborhood to neighborhood. Multiple possible mask images can be calculated from multiple possible mask designs in each of many possible neighborhoods. In one embodiment, a nominal mask design is calculated from the calculated OPC features or ILT patterns in many possible neighborhoods. In some embodiments, the multiple possible mask designs are stored in a file system 726 on disk, memory, or any other storage device.

[0039] In some embodiments, mask process simulation step 716 includes mask data preparation (MDP) to prepare the mask design for mask writing. This step includes "fracturing" the data into trapezoids, rectangles, or triangles. Mask process correction (MPC) can also be included in step 716. MPC geometrically corrects shapes and / or assigns dose to shapes to make the resulting shape on the mask closer to the desired shape. MDP can use potential mask designs 710 or the results of MPC as input. MPC is performed as part of the fracturing or other MDP operation. Other corrections are also performed as part of the fracturing or other MDP operation. Possible corrections include forward scattering, resist diffusion, Coulomb effect, etching, back scattering, fogging, loading, resist charging, and EUV mid-range scattering. Pixel level dose correction (PLDC) can also be applied in step 716. In other embodiments, a multi-beam VSB shot list or exposure information can be generated to generate multiple potential mask images 718 from potential mask design 710. In some embodiments, a set of VSB shots is generated for a calculated mask pattern in the multiple calculated mask patterns. In some embodiments, MPC and / or MDP is performed on potential mask design 710.

[0040] In step 716, calculating a plurality of possible mask images 718 includes a charged particle beam simulation. In some embodiments, the plurality of possible mask images are stored on a file system 726. Effects that are simulated include forward scattering, back scattering, resist diffusion, Coulomb effect, fogging, loading, and resist charging. Step 716 also includes a mask process simulation in which the effects of various post-exposure processes are calculated. These post-exposure processes include resist baking, resist developing, and etching. If a charged particle beam simulation is performed for a mask on a given layer, the simulation is performed over a range of process variations to establish a manufacturable contour for the mask itself. The contour may be an extension of a nominal contour. The nominal contour in this case may be based on a pattern generated at a particular resist threshold, e.g., a threshold of 0.5. In some embodiments, a mask image with variations is created for display in a viewport 728 that includes upper and lower process variation bands surrounding the nominal contour by calculating a percentage difference in exposure dose, e.g., + / - 10% dose variation. In some embodiments, the positive and negative variations may be different from each other, for example +10% and -8%. The charged particle beam simulation and the mask process simulation may be performed separately from each other in step 716.

[0041] In a substrate simulation step 720, calculating potential substrate patterns 722 includes lithography simulation using the calculated mask image 718. A plurality of potential patterns on the substrate are calculated from the plurality of mask images. Each pattern in the plurality of potential patterns on the substrate corresponds to a set of manufacturing variation parameters. Calculating substrate patterns from the calculated mask images is described in U.S. Pat. No. 8,719,739, entitled "Method and System for Forming High Accuracy Patterns Using Charged Particle Beam Lithography," owned by the assignee of this patent application. The plurality of potential patterns on the substrate 722 may be combined to form a nominal substrate pattern with variations. In some embodiments, sources of substrate pattern variations include some variations in exposure (dose) combined with some variations in depth of focus, e.g., + / - 10% in exposure, and + / - 30 nm in depth of focus. In some embodiments, the positive and negative variations may be different from each other, e.g., +5% / -7% and 30 nm / -28 nm. Traditional statistical methods are used to generate a 3 sigma variation from the nominal contour. The variation includes a lower 3 sigma limit less than the nominal contour for a minimum value and an upper 3 sigma limit greater than the nominal contour for a maximum value. In some embodiments, instead of calculating the 3 sigma variation extending from the nominal contour, a mask image with variation is created by combining multiple mask images 718 that include process variation bands with lower and upper limits. In some embodiments, a substrate pattern is formed on a wafer using an optical lithography process with a mask image with variation. In some embodiments, multiple possible patterns on a substrate are stored in a file system 726. In some embodiments, a wafer process simulation is performed on the substrate pattern.Wafer process simulations include simulations of resist baking, resist developing, and etching. Lithography simulation 720 and wafer process simulations may be separate steps, and optionally each step has process variations. In other embodiments, lithography simulation 720 includes flat panel display (FPD) simulation, microelectromechanical system (MEMS) simulation, other process simulations, or other things fabricated on the substrate.

[0042] In each step of FIG. 7, the variations are statistically accumulated, taking into account the variations from the previous step, so that the substrate pattern in the final step incorporates not only the variations in determining the possible patterns on the substrate 722, but also the variations in the mask process 716 and mask design 710. In step 724, process variation bands are calculated from the possible substrate patterns. To more efficiently calculate many possible combinations of variations, the variations are accumulated using insight into how certain variations and pattern parameters may affect each other. For example, rather than simply feeding the minimum and maximum 3 sigma values ​​from one step to the next, the worst case variations fed to the next step can take into account the distance of one pattern from another. This is because features that are close to each other affect each other more than features that are far apart. Because of the impact of these variations on design performance and manufacturing reliability, it is desirable to allow the designer to visualize the impact of different variations in the context of an actual circuit design. Visualizing the impact of the statistically accumulated variations predicted on the substrate can be shown after calculating the variation bands in step 724, or by visualizing the impact of different variations at each step. If the variations are not acceptable in step 725, the designer can modify the physical design 702 to create an improved physical design to ensure that the improved physical design is more robust to manufacturing variations. Modifications to the physical design can include modifying the possible neighborhoods of the physical design or modifying the coloring, e.g., modifying the shape assignment to any particular layer. In a design environment where curvilinear designs are allowed, providing the calculated nominal contour as a new manufacturable physical design has the advantage that manufacturing variations are reduced. This is because a manufacturable design has less variation than a non-manufacturable design (e.g., a shape with a 90 degree corner that is essentially unmanufacturable). It should be noted that the manufacturing variations predicted in the above steps need to be repeated with the improved physical design to estimate the manufacturing variations of the modified physical design.In some embodiments, the variations at each step may be shown simultaneously in a single viewport 728 with the nominal contour with the variations overlaid with the corresponding design, image, or pattern, or the variations may be shown in multiple viewports 728.

[0043] Calculating the pattern to be manufactured on the substrate includes calculating a plurality of substrate patterns from a plurality of mask images calculated from a plurality of mask designs. These calculations may take a significant amount of time and may take a long time to retrieve even if pre-computed and stored. In one embodiment, the calculation of the pattern to be manufactured on the substrate may be trained in a neural network. A neural network is a framework of machine learning algorithms that work together to predict a pattern based on a previous training process. An embodiment includes training a neural network to calculate a pattern to be manufactured on the substrate using an input physical design 702 and any combination of one or more outputs shown in FIG. 7 such as potential mask designs 710, potential mask images 718, potential substrate patterns 722, etc. Also, step 725 includes tuning a set of parameters of the neural network to reduce manufacturing variations of the calculated plurality of patterns as part of the process of training the neural network. Training of the neural network is performed using a computing hardware processor. Such training achieves a similar goal as the previous embodiment, but once trained, translation with a trained neural network can be much faster, for example 10 times faster, than with simulation alone. In one embodiment, a trained neural network or group of trained neural networks can convert the physical design pattern into a pattern to be fabricated on a substrate, i.e., in some embodiments, computing the pattern on the substrate includes a neural network having the physical design as an input.

[0044] In one embodiment, each of the outputs 710, 718, and 722 are generated by a trained neural network. Digital twins replicate physical entities. Traditionally, digital twins model the properties, conditions, and attributes of their real-world counterparts. This is accomplished through rigorous simulation. In the present application, the simulation results can be used to train a neural network, resulting in a neural network digital twin that performs much faster than the simulation alone. At any stage or combination of stages, the neural network digital twin trained with the simulated data is used to perform image-to-image translation. In one embodiment, a deep convolutional neural network (CNN) architecture, such as a fully convolutional network (FCN), is trained with paired image data, each of which represents the input and output of any of the computational steps in FIG. 7. In FIG. 8, an image 800 representing the physical design or CAD data is provided as an input to a CNN 810, such as a FCN, and an image 820 representing the manufactured output shape is generated by the CNN 810. Other neural network architectures, such as U-Net, a type of FCN, or Generative Adversarial Networks (GANs), are also used. In other embodiments, a neural network may be trained to generate OPC / ILT features or shapes for various neighborhoods, generate optimized images for mask process correction or data preparation, calculate patterns on a substrate, or any combination of steps. In an embodiment, any one or more of the steps of FIG. 7 may be replaced with a digital twin, a neural network, or a group of digital twins or neural networks in combination.

[0045] In an embodiment, a method for calculating a pattern to be manufactured on a substrate includes inputting a physical design pattern 702, determining a plurality of possible neighborhoods of the physical design pattern (step 704), and generating a plurality of possible mask designs 710 of the physical design pattern, where the plurality of possible mask designs correspond to the plurality of possible neighborhoods. The method also includes calculating a plurality of possible patterns on the substrate corresponding to the plurality of possible mask designs (722), calculating a variation band from the plurality of possible patterns on the substrate (step 724), and modifying the physical design pattern to reduce the variation band (step 725 looping back to the physical design 702).

[0046] In some embodiments, the method includes calculating a plurality of calculated mask images from a plurality of possible mask designs (step 718). In some embodiments, calculating the plurality of possible mask images includes a charged particle beam simulation (step 716). In some embodiments, modifying the physical design pattern includes modifying a plurality of possible neighborhoods of the physical design pattern (step 704). In some embodiments, the variation band of step 724 corresponds to a set of manufacturing variation parameters. In some embodiments, the variation band of step 724 includes a process variation having lower and upper limits that encompass a nominal substrate pattern. In some embodiments, the method includes performing a coloring step 706 that separates features of the physical design pattern into layers, and in further embodiments, modifying the physical design pattern includes modifying the coloring step.

[0047] In some embodiments, the physical design 702 includes optical proximity correction of the physical design pattern (step 708). In some embodiments, determining the multiple possible neighborhoods 704, generating the multiple possible mask designs 710, or calculating the multiple possible patterns on the substrate 722 includes using a neural network. In some embodiments, calculating the multiple possible patterns on the substrate includes a lithography simulation (step 720).

[0048] In some embodiments, the physical design pattern 702 comprises a portion of an overall design, and the method further includes determining a set of actual neighborhoods where the physical design pattern is used in the overall design, step 704. The portion of the overall design may be an instance of the physical design pattern, and the multiple possible neighborhoods include all neighborhoods of each instantiation.

[0049] U-Net applications such as FCN are used for predicting process variability bands associated with semiconductor manufacturing. The first U-Net architecture was deployed for biomedical image segmentation problems. In the first U-Net model architecture, each layer features a multi-channel feature map with multiple channels varying at each layer. In the final layer, a 1x1 convolution is used to map each 64-component feature vector to the desired number of classes. In total, a typical network has 23 convolution layers.

[0050] In one embodiment, the main neural network architecture for FCN is essentially the encoder-decoder network shown in Figure 9, where the left encoding side and bottleneck layer 910 guides the model to learn a low-dimensional encoding of the input image 900. The decoder network, including layers 912, 914, 916, and 918, then decodes the low-dimensional representation of the image back to the full output resolution, with both sides working together during training to learn the transformation from the input image 900 to the output image 920. The copy and crop operations, indicated by the horizontal arrows going from the encoder layers to their corresponding decoder layers, provide additional information from the encoder side of the network and act as skip connections that are concatenated with the decoder side information to help localize the information in x,y space.

[0051] If the input image is too large to be processed at once, it is divided into a set of image tiles. The image tiles may overlap each other. Each of the smaller tiles may then be processed by the network, and the output tiles may be assembled and reconstructed into a final output image. To reduce artifacts at tile boundaries, the FCN includes a halo of neighboring pixels. The halo may overlap with neighboring tiles and may be used to reconstruct the large input image.

[0052] In a semiconductor manufacturing application, the input 900 to the neural network represents an input image, or tiles from an input image that represent the design intent, i.e., as would be manufactured under "ideal" rather than realistic manufacturing process assumptions. In one embodiment, the output image 920 represents what would actually be manufactured by a realistic manufacturing process, where sharp corners are rounded, small squares are manufactured as circles or ellipses, etc. A set of model weights is determined, and after training the FCN on the semiconductor manufacturing image data, the model weights are significantly different than the model weights used in other applications.

[0053] In one embodiment, the FCN architecture shown in FIG. 9 is a multi-resolution U-Net with a reduced initial number of filters from 64 to 8 in the first layer 902, followed by filter doubling after each max pooling operation in each encoder layer 902, 904, 906, 908 and bottleneck layer 910. This has the effect of significantly reducing the overall number of trainable parameters for the network while maintaining a sufficient level of accuracy for semiconductor manufacturing applications. In another embodiment, there may be 16 filters in the first layer 902. The final encoder layer 908 and bottleneck layer 910 may each employ dropout regularization. In one embodiment, the input and output tile sizes may be, for example, 256×256 pixels (with a 128×128 inner core tile surrounded by a 64 pixel wide halo). In another embodiment, the network is further modified by removing some of the layers (shorter U depth) or adding additional layers (greater U depth) as needed for accuracy. In another embodiment, rather than doubling the number of filters after each downsampling (max pooling) or upsampling convolution, a different ratio is used. In one embodiment, a fixed ratio (e.g., 2.0) may be used at each layer, and in an alternative embodiment, a different layer-specific ratio may be used at each layer. For example, the ratio gradually increases as it gets lower and closer to the bottom bottleneck layer of the U-shape. Then, the ratio decreases again as it moves away from the bottleneck layer and up toward the output. These ratios and other network parameters are adjusted during the training phase. That is, an initial set of parameters is input for the neural network, and the set of parameters is adjusted as the neural network is trained. In one embodiment, the adjustments may be repeated for different manufacturing processes and / or for different layers in the manufacturing process.

[0054] In one embodiment, the network has a single input and a single output representing a manufactured output image corresponding to a single set of process conditions, such as a process corner. The input to the network consists of an image corresponding to computer-aided design (CAD) data (tiles from a physical design drawn by a circuit designer), and the output consists of an image corresponding to silicon manufactured for a unique set of process conditions.

[0055] In another embodiment, multiple sets of process conditions are represented via multiple copies of a single-output network as shown in Figure 10, with one network for each unique set of process conditions. Each of these single-output networks 1001, 1002-1010 are trained in parallel. After training, each of these networks can be used to infer, for a given CAD data input image 1000, the output for a unique set of process conditions 1021, 1022-1030, i.e., a particular process corner.

[0056] An example of the inferred output is shown in FIG. 11. Reassembled tiles are shown that represent images representative of the manufactured shape of a D-type flip-flop (DFF) design image 1101 under three different unique process conditions. Although similar at first glance, upon closer inspection it is clear that the three images are different, e.g., different amounts of corner rounding are evident in each. The shape in image 1102 is closest to the rectilinear CAD shape drawn from image 1101. The shape in image 1104 is perhaps the furthest away, with a greater degree of rounded corners and narrowing of the shape. Image 1103 is somewhere between these two extremes. For simplicity, in this example, only three examples are shown as representative of semiconductor manufacturing process conditions. A more comprehensive set would include dozens, representing different extremes of dose variation in mask manufacturing, and different extremes of both dose and depth of focus variation in semiconductor manufacturing.

[0057] In another embodiment, FIG. 12 shows a process in which one output network 1201, 1202-1210 is used to infer output manufacturing images for each process corner 1211, 1212-1220, and then post-processing 1230 is used to combine and aggregate the images for each corner to produce an average image 1233 representative of a typical set of manufacturing conditions, a maximum image 1231 representative of the most extreme outcome where the most material is deposited on the silicon, and a minimum image 1232 representative of the most extreme outcome where the least material is deposited on the silicon.

[0058] The output images produced by combining the per-corner image tiles are shown in detail in Figure 13. The maximum image 1301 is calculated by taking the maximum per pixel value across all per-corner output images. The minimum image 1302 is calculated by taking the minimum per pixel value across all per-corner output images. Comparing the minimum ellipse shapes 1311 and 1312 near the center of both images, it is clear that the ellipse 1312 for the minimum image 1302 is shown to be much smaller than the ellipse 1311 for the maximum image 1301.

[0059] The average image 1303 is calculated by taking the pixel-wise sum divided by the number of process corners, or pixel-wise average over all corner-wise output images. The process variation band or PV band image 1241 shown in FIG. 12 is calculated by a post-processing step by subtracting the minimum image from the maximum image. The PV band image 1304 in FIG. 13 is shown in detail. The white pixels indicate places where metal may or may not be deposited on the silicon during manufacturing. That is, each white pixel represents an area of ​​uncertainty due to process variations. The more white pixels there are, the more sensitive the design is to manufacturing process variations.

[0060] Image thresholding compares each pixel value to a predefined threshold (e.g., 0.5) so that pixel values ​​above the threshold are converted to white (1.0) while pixel values ​​below the threshold are converted to black (0.0). In another embodiment, image thresholding is performed before calculating the maximum, minimum, or average. This refers, for example, for a metal fabrication step, to determining a single binary value for each pixel (1 or 0, corresponding to white or black, respectively) regardless of whether metal is present at each pixel location. In a further embodiment, the maximum, minimum, and average for each pixel may be calculated first and then image thresholding may be performed.

[0061] As shown in Figure 12, it may be desirable to generate two additional images to calculate a metric that represents the sensitivity or tolerance of the design / process combination to process variations: false positives 1242, i.e., manufactured image pixel locations where material is deposited on the silicon but was not set in the original CAD data (i.e., unintended material), and false negatives 1243, i.e., output image pixel locations that were set as the intended material in the original CAD data image but were not deposited during manufacturing.

[0062] An example of a false negative occurs at a 90 degree corner of a drawn CAD polygon, where a sharp corner is drawn by the circuit designer, but during manufacturing a form of corner rounding and / or line pullback occurs, effectively cutting or shortening the corner of the deposited material. An example of a false positive is excess material generated, for example, at a 270 degree corner, or excess material generated via pinching. In one embodiment, false positive and false negative images are shown in FIG. 14 and are generated as a post-processing step. The false positive image 1401 is calculated by subtracting the original CAD data image from the maximum image. The false negative image 1402 is calculated by taking the product (logical AND) of the minimum image and the original CAD data image, thresholding the result, and then subtracting the thresholded result from the thresholded original CAD data image.

[0063] To reduce the post-processing burden, in one embodiment, a CNN architecture with multi-channel output is shown in FIG. 15. In one embodiment, the first N channels may be reserved for each of the N process corner conditions 1502. This is accomplished by forming an output layer consisting of 1×1 convolution operations with a filter depth of N, where N is the number of process corners. The idea is for a single trained multi-output network 1501 to generate images corresponding to the outputs produced for each of the individual process corners 1502.

[0064] Also, as shown in Figure 15, the calculation of maximum image 1504, minimum image 1505, average image 1506, PV band image 1507, false positive image 1508, and false negative image 1509 are accomplished as a post-processing step 1503 described in connection with Figure 12 after a single trained CNN 1501 is used to generate outputs 1502 for each of a number of process conditions. It will be appreciated that any of the aggregate output images or any combination of these images are obtained via post-processing rather than being directly inferred by the network.

[0065] As previously mentioned, these images may be calculated through post-processing of the minimum and maximum images and the input CAD data images. In one embodiment shown in Figure 16, these images may be generated directly by a deep neural network 1601 through the introduction of additional output image channels such as process corners 1602, and minimum, maximum, and average values ​​1610. PV bands, false negatives, and false positives 1620 may be generated directly.

[0066] If the maximum, minimum, and average images, etc. are directly generated by the trained network, the corner images 1602 for each process may not need to be learned / inferred by the network. In this case, the network is trained to directly output aggregate images 1610 (maximum, minimum, average) and 1620 (PV band, false positive, false negative) without outputting the corner images 1602 for each process. When the number of process corners to be considered is large, it is desirable (to reduce computational and / or GPU resources such as memory) not to output the images 1602 for each process corner, but instead to output only the remaining aggregated images. In this case, the filters for each corner are removed from the CNN output layer, and their corresponding images are removed during training. In one embodiment, the user can choose to have the network output all, some, or none of the images for each corner before training. And the network architecture and parameters of the neural network are adjusted accordingly.

[0067] Although the shapes that are fabricated on silicon depend heavily on the immediate location or neighborhood of the input shapes, there are also long-range effects such as local pattern density. Simply put, the shapes that are fabricated for a CDA data image tile of an image will contain some differences if the tile is from a dense part of a larger design compared to when it is from a relatively isolated part of the design. To allow the CNN model to learn these density effects, embodiments extend the input to include multiple channels. In such an embodiment, the local pattern density is coded into a single number between 0.0 (totally isolated) and 1.0 (fully surrounded by metal), and a grayscale image is generated with all pixels set to the same number. The grayscale image dimensions are set to be the same as the CAD data tile dimensions, and are represented as an additional channel in the input image, just as color images are represented as R, G, B channels for normal image processing. The CNN architecture is then extended to handle a two-channel input instead of a single-channel input. During the training process, the network parameters learn the relationship between the grayscale color levels and the corresponding effects in the output fabrication image.

[0068] In one embodiment, the input image may be composed of two channels, each of which may itself be represented as a grayscale image, one for the CAD data and one for a lower resolution image of the larger area from which the patch tiles representing the local density information were obtained. In some embodiments, the output image may include multiple channels with different grayscale images per channel (e.g., channels representing maximum, minimum, average, PV band, false positive, or false negative images). In additional embodiments, the output image may also include additional channels, e.g., one channel per process corner, where the image for each corner represents the expected manufactured shape for a particular process corner-specific combination of process variables.

[0069] In an embodiment, a method for computing a pattern to be fabricated on a substrate includes inputting a physical design 900, inputting a set of parameters for a neural network to compute a pattern to be fabricated on the substrate, generating a plurality of possible neighborhoods for the physical design (step 704 of FIG. 7), and computing a plurality of patterns to be fabricated on the substrate for the physical design in each of the plurality of possible neighborhoods (step 722). The method also includes training the neural network with the plurality of computed patterns (e.g., in a loop from step 725 to the physical design 702), where the training is performed using a computing hardware processor, and adjusting the set of parameters (e.g., in step 725) to reduce manufacturing variation of the computed plurality of patterns to be fabricated on the substrate.

[0070] In some embodiments, the neural network includes using post processing to aggregate the variations within the variation bands. The neural network includes multiple output channels to aggregate the variations within the variation bands. In some embodiments, the method includes calculating false negatives and false positives for the patterns on the substrate.

[0071] In some embodiments, the neural network includes a single fully convolutional network (FCN) architecture (e.g., FIG. 9). The FCN includes a first encoding layer, a second encoding layer, a final encoding layer, and a bottleneck layer, where the final encoding layer and the bottleneck layer each employ dropout regularization. In some embodiments, the FCN includes a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer, where the decoding layers employ concatenation with additional information from the fourth encoding layer, the third encoding layer, the second encoding layer, and the first encoding layer, respectively.

[0072] In some embodiments, the physical design and the plurality of patterns to be calculated are each divided into tiles. For example, each of the tiles comprises a 256×256 pixel tile with an inner core of 128×128 pixels and a halo that is 64 pixels wide. In some embodiments, calculating the pattern to be fabricated on the substrate includes a charged particle beam simulation. In some embodiments, calculating the pattern to be fabricated on the substrate includes a lithography simulation 720. In some embodiments, the method includes inputting a local pattern density of the physical design 702. Design Variability Metrics Various aggregate images across the variability can be used to generate a scalar design variability metric.

[0073] Let TP (True Positives) be the number of white pixels in the CAD design that represent where metal would ideally be deposited in silicon manufacturing, and TN (True Negatives) be the number of black pixels in the same image. Let VB (Variation Bands) be the number of white pixels in the variation band plot that serves as an upper bound on the uncertainty associated with metal deposition due to process variation.

[0074] Let FN (false negative) be the number of white pixels in a false negative design image, which represents the amount of metal where metal was ideally intended to be deposited during silicon fabrication, but was actually found not to be deposited due to corner rounding, line end pullback, etc. FN is a metric that serves as an upper bound on the amount of missing metal found after fabrication.

[0075] Let FP (false positive) be the number of white pixels in the false positive design image, which represents how much metal was inadvertently deposited in locations where no metal was ideally intended to be deposited during silicon manufacturing. FP is a metric that acts as an upper bound on the measurement of unwanted material deposited during manufacturing.

[0076] The Matthew Correlation Coefficient (MCC), defined as follows, is often used as the single metric by which a classification algorithm is measured when using the TP, FP, TN, and FN measurements from a confusion matrix:

[0077]

number

[0078] In this semiconductor manufacturing scenario, the MCC formula has a different meaning than traditional usage because the four variables TP, FP, TN, and FN have different meanings for the semiconductor manufacturing application of the present disclosure. In this case, MCC is a function of the amount of intended metal (TP), unintended metal (TN), the upper limit of metal inadvertently removed where it was originally intended (FN), and the upper limit of metal inadvertently deposited where it was not originally intended (FP). The MCC score calculated via this formula serves as a single scalar valued measure of how process variability tends to produce an image on silicon that differs from the intended image (per the originally drawn CAD data). A large value of MCC (close to 1.0) indicates good correlation between the CAD data image and the manufactured silicon image, i.e., high resistance to process variations. A very small value of MCC (close to 0) indicates very little correlation between the intended image and the manufactured image. The MCC value can be improved by modifying the manufacturing process to reduce the variations, which is difficult and expensive, or by modifying the manufacturing process by design or a combination of both approaches. Integrated device manufacturers (IDMs) are in a position to modify their processes for critical designs that require high levels of repeatability (yield).

[0079] Precision is defined as TP divided by the sum of TP and FP, and recall is defined as TP divided by the sum of TP and FN. Precision and recall are used to calculate another metric, F1.

[0080]

number

[0081] The formula for calculating the F1 score from these images is given below:

[0082]

number

[0083] Note that the MCC value is considered a more useful metric than the F1 formula, which does not include the TN measure, by taking into account all four quantities (TP, TN, FP, FN), and therefore can be asymmetric for imbalanced class problems (where TP and TN are significantly different), which is one of the reasons why MCC is traditionally the preferred quantity to use in classification algorithms.

[0084] Let TP2 be the number of white pixels in the average image after the grayscale average image has been thresholded. This represents the number of pixels a designer can realistically expect to have metal deposited by a realistic process. The designer is aware that the process is not ideal and effects such as rounded corners occur during manufacturing. However, the designer continues to draw images consisting of straight lines with right angle corners during the circuit design process simply for drawing convenience. By calling TP the number of (ideal) white pixels drawn initially, and TP2 the number of white pixels a designer can more realistically expect, two further quantities are defined.

[0085] Let VBI=VB / TP2, which is the ratio of the number of variation band white pixels to the average (realistically expected) image white pixels. This is a more realistic assessment of how sensitive the design is to a given process variation since the numerator of the fraction, VB, still contains the uncertainty term and includes the amount of pixels whose manufacturing output is uncertain. VB is then normalized by the denominator, TP2, which is the amount of pixels that can realistically be expected to have metal on average across the process variation.

[0086] The second quantity, VBI′=VB / TP, serves as the ratio of manufacturing uncertainty to the initially rendered number of white pixels (an unrealistic but expected outcome in an ideal manufacturing scenario).

[0087] Using these definitions as appropriate, different designs / cells or design candidates generated by the designer are processed by the trained neural network according to an embodiment to generate their various aggregated manufacturing output images and the white pixels of those images are counted as described above, and the designs are later scored using a metric related to their resistance to process variations (MCC) or their susceptibility to process variations (VBI). Deep Learning Challenge In deep convolutional neural networks, or deep learning, computer models learn to perform classification or regression tasks directly from images, text, or sounds. Deep learning models can achieve state-of-the-art accuracy and even exceed human-level performance in some cognitive applications. The models are trained by using a large set of labeled data and neural network architectures that contain many layers. Most deep learning methods use network architectures. This is why deep learning is referred to as deep neural networks. The term "deep" usually refers to the number of hidden layers in a neural network. While traditional neural networks contain only 2-3 hidden layers, deep networks can have as many as 150.

[0088] Deep learning models are trained by using a large set of labeled data and neural network architectures. One of the most common types of deep neural networks is the CNN architecture. CNN architectures convolve learned features with input data, typically using 2D convolutional layers, making this architecture well suited for processing 2D data such as images.

[0089] CNNs eliminate the need for manual feature extraction, i.e., pre-identifying features that will be used to classify or predict images. CNNs work by extracting features directly from images. Relevant features are not pre-trained. Relevant features are learned while the network is training on a sufficiently large set of images. Such automated feature extraction makes deep learning models highly accurate for general computer vision tasks such as object classification, and for semiconductor manufacturing image-to-image translation tasks such as the present invention.

[0090] There are a few key reasons why deep learning has only recently become useful. Deep learning requires large amounts of labeled data. Deep learning requires significant computing power.

[0091] Deep learning is an iterative process. Deep learning requires large amounts of labeled data. For example, the development of driverless cars requires millions of images and thousands of hours of video. In the context of this disclosure, obtaining labeled data refers to assembling a large collection of thousands to millions of images that represent the physical design to be manufactured, and refers to the image-based output of the various computational steps in FIG. 7, such as OPC / ILT, mask process simulation, substrate simulation, etc.

[0092] Some of the data can be collected by actual fabrication of dedicated test chips, but given the cost of mask set production and fabrication for today's high density processes, such fabrication-based data collection approaches are prohibitively expensive. An alternative is to use computational simulations instead of fabrication, but the computational costs for any of the steps in Figure 7 are enormous. Until recently, this option was prohibitively expensive in terms of the computational power required. In particular, the OPC / ILT calculations required to calculate mask shapes that result in manufacturable designs required huge amounts of computational power. ILT / OPC computational tools coupled with highly parallel and GPU-accelerated computational design platforms now finally make it possible to use computational software to determine mask shapes for full reticle-sized IC designs in a time frame that no longer prohibits the generation of sufficient amounts of labeled data required for deep learning.

[0093] Deep learning requires significant computational power. High-performance GPUs have an efficient parallel architecture for deep learning. Combined with cluster or cloud computing, this may allow development teams to reduce the training time for deep learning networks from weeks to hours or less, depending on the problem and the complexity of the deep learning neural network architecture. Dedicated architectures for advanced computations such as those described in Figures 17 and 18 can facilitate the application of deep learning to previously difficult problems.

[0094] The sequential process for deep learning layer training includes loading / pre-processing data, fitting a model to make predictions. While this sequential approach is certainly reasonable and useful to understand, in reality, deep learning is not so linear. Instead, the actual deep learning that produces the trained image as in Figure 9 has a well-defined cyclical nature that requires constant iteration, tuning, and improvement. The cycle begins with an iteration from the input physical design 900, computing a mask image, and then comparing that image to the output image 920 produced by deep learning. After each process, the impact on how the model performs is measured and adjustments are made to improve performance in the next cycle.

[0095] Deep learning practitioners must deal with the following iterative process: Model level: fitting of model parameters Micro-level: Hyperparameter tuning Macro Level: Problem Solving Meta-level: Improving training / test data Model level: Fitting parameters The first level where iteration plays a big role is the model level. Any model, whether it is a regression model, a decision tree, or a neural network, is defined by many (sometimes millions) of model parameters. For example, a regression model is defined by its feature coefficients, a decision tree is defined by its branching positions, and a neural network is defined by the weights connecting its layers. In deep learning, the model parameters are learned via iterative techniques such as gradient descent, which is an iterative method for finding the minimum of a function. In deep learning, that function is typically a loss (or cost) function. "Loss" is a metric that quantifies the cost of an incorrect prediction, such as mean squared error, mean absolute error, cross entropy, etc. Gradient descent calculates the loss achieved by a model with a given set of parameters, and then adjusts the parameters to reduce the loss. This process is repeated until the loss cannot be substantially reduced any further.

[0096] Micro-level: Hyperparameter tuning Hyperparameters are "high-level" parameters that cannot be learned directly from data using gradient descent or other optimization algorithms. For example, dropout is a regularization method that approximates training many neural networks in parallel with different architectures. During training, some layer outputs are randomly ignored or "dropped out". This has the effect of making the layer look and be treated as a layer with a different number of nodes and connectivity to previous layers. In fact, each update to a layer during training is performed with a different "view" of the constructed layer. Conceptually, dropout breaks down the situation where network layers adapt to each other to correct errors from previous layers, which in turn makes the model more robust. Hyperparameters describe structural information about the model that must be determined before fitting model parameters, such as whether dropout or other forms of regularization should be included in the model, whether batch normalization should be performed before computing the layer's outputs, the number of epochs (outer iterations to use in the model parameter fitting process), the specific optimization algorithm to use during model parameter fitting, and whether cross-validation is used to validate the model during fitting. Determining appropriate values ​​for each of these various parameters / decisions is an iterative process and may require many iterations of the model parameter fitting process described above.

[0097] Macro Level: Problem Solving There is no one model architecture / family that works best for all problems: different model families perform better than others depending on various factors such as the type of data, the problem domain, the sparsity of the data, and even the amount of data collected.

[0098] Thus, one way to improve a candidate solution for a given problem is to try several different model families or model architectures, such as the shape of the network itself, the number of filter layers, and the size of the convolution kernels used in the convolution layers, as well as whether or not to use skip layer techniques. Determining appropriate values ​​for each of these various parameters / decisions is an iterative process, requiring many iterations of the model parameter fitting and hyperparameter tuning processes described above.

[0099] Another way to improve a deep learned solution is by combining multiple deep learned models into an ensemble. This is a direct extension from the iterative process required to fit multiple models. A common form of forming an ensemble is to average predictions from multiple trained models. There are more advanced methods for combining multiple models, but the iterations required to fit multiple models are the same. Determining an appropriate combination / ensemble for each of the various deep learned models is an iterative process.

[0100] Meta-level: Improving training / test data When it comes to machine learning, better data generally imparts more value than a better algorithm. However, better data is not the same as more data. Better data means having less missing data and lower measurement error (e.g., more accurate data). The data also needs to be representative and avoid problems known to those skilled in the art such as data imbalance. The overall process of obtaining a sufficient set of such clean and accurately labeled data is itself often an iterative process. In the case of the present disclosure, the iterative process includes running more simulations of the various steps of FIG. 7 under different conditions to ensure that a sufficiently wide sampling of possible neighborhoods is performed. The learning process includes determining the various types of shape combinations that will be present in the cell design, such as the physical design input to FIG. 7, to make the deep learned model perform well on shape combinations that were not previously surfaced.

[0101] The various iterations of the above process to actually train a deep learned model to a sufficient level of accuracy require enormous amounts of computing power and vast amounts of data associated with the semiconductor manufacturing process. Prior to recent developments in the semiconductor manufacturing computing industry and the recent ability to run computational software simulators on dedicated GPU-accelerated hardware, such deeply nested iterative processes for deep learning of patterns to be manufactured on substrates such as silicon wafers were not tractable, and there was no motivation to even attempt such an approach.

[0102] Neural networks such as CNNs need to be trained with training data. Typically, the higher the capacity of a network (the amount of ability to generalize to unseen data) and the more parameters it has, the more data samples are needed to train the network without overfitting. The networks considered here typically contain hundreds of thousands of learnable parameters, requiring an extremely large number of training data samples.

[0103] The supervised training paradigm feeds a large number (hundreds of thousands to millions) of input / output pairs to the CNN. In our case, the input items consist of patches of CAD data, i.e., sets of CAD data representing the physical design drawn by the circuit designer. The sets of CAD data are rasterized and divided into patches or tiles. For each layer to be fabricated, the input image is a single channel image with a particular width and height. The input images may be binary images (each pixel is either black or white) or grayscale images, with each pixel taking on a continuous value from 0.0 (black) to 1.0 (white). The output items for each pair consist of the corresponding predicted image after fabrication at a particular process corner. In one embodiment, the output items are single channel output images, containing either binary or grayscale (continuous value) pixels. The idea is to train the network to be able to infer or predict an output image given only the input image. The idea is also to train the network to be able to infer or predict an output image for design input images that have not been seen before.

[0104] It should be noted that while the generation of sufficient input image data (CAD data) is relatively fast, the generation of corresponding predicted output image data representing the manufacturing result is a very long-term problem for a real semiconductor manufacturing process with state-of-the-art process nodes, including huge amounts of computing hardware resources. The CAD data needs to be simulated using various computationally intensive algorithms, including but not limited to OPC and ILT, as well as wafer manufacturing simulation using calibrated mask models. Such simulation tools and models are used together with dedicated GPU-based hardware in the form of high-performance computing clusters (HPC) or computational data platforms (CDP) to accelerate the simulation. Only after a series of such tools have been run can the output image be obtained. Moreover, the generation of corresponding corner images for each process entails significant additional costs when process variations are considered. Only recently have semiconductor process manufacturing simulation tools, especially ILT, become fast enough to generate the required amount of data within a realistic time frame.

[0105] FIG. 17 illustrates an example of a computing hardware device 1700 that may be used to perform the computations described in this disclosure. The computing hardware device 1700 includes a central processing unit (CPU) 1702 with a main memory 1704 attached. The CPU may include, for example, eight processing cores, thereby improving the performance of any portion of computer software that is multi-threaded. The size of the main memory 1704 may be, for example, 64 GB. The CPU 1702 is connected to a Peripheral Component Interconnect Express (PCIe) bus 1720. A graphics processing unit (GPU) 1714 is also connected to the PCIe bus. In the computing hardware device 1700, the GPU 1714 may or may not be connected to a graphics output device such as a video monitor. When not connected to a graphics output device, the GPU 1714 is simply used as a high-speed parallel computation engine. By using the GPU for part of the computation, the computing software can obtain significantly higher performance than if it were to use the CPU 1702 for all of the computation. CPU 1702 communicates with GPU 1714 via PCIe bus 1720. In other embodiments (not shown), GPU 1714 may be integrated into CPU 1702 rather than connected to PCIe bus 1720. Disk controller 1708 may also be attached to the PCIe bus, with, for example, two disks 1710 connected to disk controller 1708. Finally, a local area network (LAN) controller 1712 may be connected to the PCIe bus, providing Gigabit Ethernet (GbE) connectivity to other computers. In some embodiments, computer software and / or design data are stored on disk 1710. In other embodiments, either or both of the computer programs and design data may be accessed from other computers or file serving hardware via GbE Ethernet.

[0106] FIG. 18 is another embodiment of a system for performing the computations of the present embodiment. The system 1800, sometimes referred to as a CDP, includes a master node 1810, an optional viewing node 1820, an optional network file system 1830, and a GPU-enabled computing node 1840. The viewing node 1820 may be absent, may have only one node, or may have other numbers of nodes. The GPU-enabled computing node 1840 may include one or more GPU-enabled nodes forming a cluster. Each GPU-enabled computing node 1840 includes, for example, a GPU, a CPU, a pair of a GPU and a CPU, multiple GPUs for a CPU, or other combinations of a GPU and a CPU. The GPU and / or CPU may be on a single chip, such as a GPU chip with a CPU accelerated by the GPU on the chip, or a CPU chip with a GPU accelerating the CPU. The GPU may be replaced by another coprocessor.

[0107] The master node 1810 and the viewing nodes 1820 are connected to the network file system 1830 and the GPU-enabled computing nodes 1840 via switches and high-speed networks, such as networks 1850, 1852, and 1854. In one embodiment, the network 1850 may be a 56 Gbps network, 1852 may be a 1 Gps network, and 1854 may be a management network. In various embodiments, there may be fewer or more networks and a combination of different types of networks, such as high speed and low speed. The master node 1810 controls the CDP 1800. An external system may connect to the master node 1810 from an external network 1860. In some embodiments, a job may be initiated from an external system. Data for a job is loaded onto the network file system 1830 before launching the job, and a program is used to dispatch and monitor tasks on the GPU-enabled computing nodes 1840. The progress of the job can be viewed via a graphical interface, such as the viewing node 1820, or by a user on the master node 1810. Tasks are executed on the CPU using scripts that launch the appropriate executables on the CPU, which connect to the GPU, perform various computational tasks, and then disconnect from the GPU. The master node 1810 is also used to disable a failed GPU-enabled compute node 1840 and then operate as if the node never existed.

[0108] Although the present specification has been described in detail with respect to certain embodiments, it will be understood that those skilled in the art, upon understanding the foregoing, may easily conceive of alternatives, variations, and equivalents of these embodiments. These and other modifications and variations to the method may be implemented by those skilled in the art without departing from the scope of the present subject matter, which is more fully described in the appended claims. Moreover, those skilled in the art will appreciate that the foregoing description is merely exemplary and is not intended to be limiting. Steps may be added, steps may be removed, or steps may be modified from those described herein without departing from the scope of the present invention. In general, the presented flow charts are intended only to illustrate one possible sequence of basic operations to achieve a function, and many variations are possible. Thus, the present subject matter is intended to include such modifications and variations that come within the scope of the appended claims and their equivalents.

Claims

1. 1. A method for calculating a pattern to be manufactured on a substrate, comprising the steps of: A physical design pattern represented by CAD data is input; generating a plurality of neighborhoods for the physical design pattern, each of the plurality of neighborhoods being of different shapes adjacent to the physical design pattern, i.e., shapes having different orientations; generating a plurality of mask designs corresponding to the plurality of neighborhoods for the input physical design pattern; calculating a number of patterns expected to be produced on the substrate for the number of mask designs; calculating a variation band representative of the variation between the patterns expected to be produced on the substrate; modifying the physical design pattern to reduce the variation band; A method comprising:

2. The method of claim 1 further comprising computing a plurality of mask images from the plurality of mask designs.

3. 3. The method of claim 2, The method, wherein calculating the multiple mask images includes a charged particle beam simulation.

4. 10. The method of claim 1 , The method, wherein modifying the physical design pattern includes modifying a plurality of neighborhoods of the physical design pattern.

5. 10. The method of claim 1 , The method of claim 1, wherein the variation band comprises a set of manufacturing variation parameters.

6. The method of claim 1 further comprising: The method includes performing a coloring step that assigns at least two different shapes of the physical design pattern to different masks.

7. 7. The method of claim 6, Modifying the physical design pattern includes modifying the coloring step. A method comprising:

8. The method of claim 1 further comprising: The method includes optical proximity correction (OPC) of the physical design pattern.

9. 10. The method of claim 1 , The method, wherein generating the plurality of neighborhoods, generating the plurality of mask designs, or calculating the plurality of patterns expected to be generated on the substrate comprises using a neural network.

10. The method according to claim 9 further comprises: using post-processing to aggregate variability within the variability bands.

11. 10. The method of claim 9, The method of claim 1, wherein the neural network further comprises multiple output channels for aggregating variability within the variability bands.

12. The method of claim 1 further comprising: calculating false negatives and false positives for the patterns on the substrate.

13. 10. The method of claim 1 , The method, wherein calculating the plurality of patterns expected to be produced on the substrate comprises a lithography simulation.

14. 10. The method of claim 1 , the physical design pattern comprises a portion of an overall design; The method further includes generating a set of actual neighborhoods in which the physical design patterns are used throughout the design.

15. 15. The method of claim 14, a portion of the overall design is an instance of the physical design pattern; The method, wherein the plurality of neighborhoods includes all neighborhoods of each instantiation.

16. 10. The method of claim 1 , The method of claim 1, wherein each of the plurality of neighborhoods is a neighborhood surrounding the physical design pattern.

17. 1. A method for modifying a physical design pattern to be fabricated on a substrate, the physical design pattern being part of an integrated circuit (IC) design and represented by CAD data, comprising: calculating, for the physical design pattern, a number of patterns that are expected to be fabricated on the substrate based on various proximity to the physical design pattern of the integrated circuit design; forming a variation band representative of the variation between the plurality of patterns on the substrate; modifying the physical design pattern to reduce the variation band; A method comprising:

18. 18. The method of claim 17, The method, wherein the plurality of patterns is further based on a variation in at least one manufacturing process parameter.

19. 18. The method of claim 17, The method of claim 1, wherein the variation band represents a variation between the plurality of patterns expected to be produced on the substrate.

20. The method of claim 18 further comprising: generating a plurality of mask designs for the physical design pattern; The method, wherein the multiple mask designs correspond to the different neighborhoods.

21. 18. The method of claim 17, The method, wherein each of the different neighborhoods is a neighborhood surrounding the physical design pattern and is comprised of different shapes, i.e., shapes of different orientations.

22. A computer program comprising: A computer program executed by at least one processing unit to carry out the method according to any one of claims 1 to 21.

23. 1. A computing device comprising: A set of processing units; A machine-readable medium storing a program executed by at least one processing unit to perform the method according to any one of claims 1 to 21. A computing device comprising:

Citation Information

Patent Citations

  • Evalulating method of pattern in photomask, photomask, production of photomask, forming method of pattern in photomask and exposing method

    JP1996202020A

  • Correction method and correction device for mask pattern

    JP1996248614A

  • Method for correcting optical proximity effect and light intensity simulation method

    JP2001174974A

  • Mask for exposure, optical proximity effect correction apparatus, optical proximity effect correction method, method for manufacturing semiconductor device, and optical proximity effect correction program

    JP2004341160A

  • Method for creating design pattern data, method for creating mask pattern data, method for manufacturing mask, and method and program for manufacturing semiconductor device

    JP2006053248A