Method, system, computer program and computing device for determining shape for manufacturing semiconductor or flat panel display
Neural networks and advanced charged particle beam systems address the challenges of accuracy and computation in lithography by simulating and optimizing manufacturing processes, improving the precision and efficiency of pattern formation on substrates.
Patent Information
- Application Number
- JP2025061528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-10-22
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
Current charged particle beam lithography methods face challenges in accurately forming small critical dimensions on substrates due to manufacturing variations and computational complexity, particularly in inverse lithography technology (ILT) and multi-beam writing, which require extensive computation time and are prone to hotspots and variations in OPC features.
A method involving neural networks and charged particle beam systems, such as multi-beam and variable shaping beam systems, to model and simulate multiple manufacturing scenarios, including OPC verification and ILT, to reduce variations and improve accuracy by statistically simulating and visualizing manufacturing processes, using computational hardware to train neural networks for faster pattern calculation.
Enhances manufacturing accuracy and reduces computational time by modeling multiple parameters in real-time, improving the resilience of patterns to manufacturing variations and reducing hotspots, thus enhancing the yield and reliability of integrated circuits.
Smart Images

Figure 2025102930000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to lithography, and more particularly, to surface design and manufacturing of a reticle, a wafer, or any other surface using charged particle beam lithography.
Background Art
[0002] There are three general types of charged particle beam lithography: non-shaping (Gaussian) beam lithography, shaped charged particle beam lithography, and multi-beam lithography. In all types of charged particle beam lithography, the charged particle beam emits energy onto a surface coated with a resist to expose the resist.
[0003] In the production or manufacturing of semiconductor devices such as integrated circuits, photolithography is used to manufacture semiconductor devices. Photolithography is a printing process that manufactures integrated circuits by forming patterns on a substrate such as a semiconductor or silicon wafer using a lithography mask or photomask manufactured from a reticle. Other substrates include flat panel displays and other reticles. Also, extreme ultraviolet (EUV) and X-ray lithography are considered to be types of photolithography. One or more reticles contain circuit patterns corresponding to each layer of the integrated circuit. This pattern is imaged onto a certain area on the substrate coated with a layer of a radiation-sensitive material known as photoresist or resist. Once the pattern layer is formed, the layer is processed by various other processes such as etching, ion implantation (doping), metallization, oxidation, and polishing. These processes are employed to finish each layer within the substrate. If several layers are required, the entire process or a variation thereof is repeated for each new layer. Eventually, a combination of multiple devices or integrated circuits will exist on the substrate. These integrated circuits are then separated from each other by dicing or sawing and then mounted in each package. More generally, the patterns on the substrate can be used to define artifacts such as display pixels or magnetic recording heads.
[0004] In the production or manufacture of semiconductor devices such as integrated circuits, semiconductor devices can also be manufactured using maskless direct writing. Maskless direct drawing is a printing process in which charged particle beam lithography is used to form patterns on a substrate such as a semiconductor or silicon wafer to manufacture integrated circuits. Other substrates can include flat panel displays, imprint masks for nanoimprinting, or reticles. The desired pattern of the layer is written directly onto the surface, which in this case is also the substrate. Once the pattern layer is generated, the layer is processed in various other processes such as etching, ion implantation (doping), metallization, oxidation, and polishing. These processes are used to finish each layer within the substrate. If multiple layers are required, the entire process or a variation thereof is repeated for each new layer. Some of the layers may be written using photolithography, while other layers may be written using maskless direct writing for the same substrate manufacture. Eventually, a combination of multiple devices or integrated circuits will be present on the substrate. These integrated circuits are then separated from each other by dicing or sawing and then mounted in each package. More often, artifacts such as display pixels or magnetic recording heads can be defined using the patterns on the surface.
[0005] In optical lithography, a lithography mask or reticle has a geometric pattern corresponding to circuit components integrated on a substrate. The pattern used to manufacture the reticle can be generated using computer-aided design (CAD) software or a program. When designing the pattern, the CAD program can follow certain design rules for fabricating the reticle. These rules are set by processing, design, and end-use limitations. An example of an end-use limitation may be to define the geometry of a transistor so that it cannot operate properly at the required supply voltage. In particular, design rules can define tolerances between circuit devices or interconnect lines. Design rules are used, for example, to ensure that circuit devices or lines do not interact in an undesirable way. For example, design rules are used to prevent lines from getting too close to each other in a way that could cause a short circuit. The limitations of design rules particularly reflect the minimum dimensions that can be manufactured while ensuring reliability. When referring to small dimensions, the concept of critical dimensions is usually introduced. Critical dimensions are defined, for example, as the important width or area of a feature, or the important space or space area between two features. These dimensions require precise control. Due to the nature of integrated circuit design, many patterns in the design are repeated at different positions. A pattern may be repeated hundreds or thousands of times, and each copy of the pattern is referred to as an instance. If a design rule violation is found in such a pattern, hundreds or thousands of violations are reported (one for each instance of the pattern).
[0006] One goal in integrated circuit manufacturing by photolithography is to reproduce a prototype circuit design on a substrate using a reticle. A reticle, also known as a mask or photomask, is the surface that is exposed during manufacturing using charged particle beam lithography. Integrated circuit manufacturers are constantly trying to use the area of semiconductor wafers as efficiently as possible. Engineers are continuously reducing the size of circuits so that the integrated circuits can contain more circuit elements and be used with less power. As the size of the critical dimensions of the integrated circuit is reduced and the circuit density increases, the critical dimensions of the circuit pattern or physical design approach the resolution limit of the optical exposure tools used in conventional photolithography. As the critical dimensions of the circuit pattern become smaller and approach the value of the resolution of the exposure tool, it becomes difficult to accurately transfer the physical design to the actual circuit pattern developed on the resist layer. To further advance the use of photolithography to form patterns with features smaller than the light wavelength used in the photolithography process, a process known as optical proximity correction (OPC) has been developed. OPC modifies the physical design to compensate for distortions caused by, for example, the light diffraction and optical interactions of features along with the nearest features. Resolution enhancement techniques (RET) executed using a reticle include OPC and inverse lithography technology (ILT), etc.
[0007] OPC can add sub-resolution lithography features to the mask pattern to reduce the difference between the initial physical design pattern, i.e., the design, and the circuit pattern ultimately formed on the substrate. The features of sub-resolution lithography interact with and act on each other in the physical design, compensating for the proximity effect to improve the ultimately formed circuit pattern. One feature added to improve pattern formation is called a "serif". A serif is a small feature to enhance the accuracy or resilience against manufacturing variations in the printing of a specific feature. An example of a serif is a small feature placed at the corner of a pattern to sharpen the corner in the ultimately formed image. The pattern intended to be printed on the substrate is called the main feature. Discussions about OPC features, including the OPC-decorated pattern written on the reticle in relation to the main feature, i.e., features reflecting the design before OPC decoration, serifs, jogs, sub-resolution assist features (SRAFs), and negative features, etc., have been conducted conventionally. An SRAF is an isolated shape not attached to the main feature and is small enough not to be printed on the substrate, while serifs, jogs, and negative features change the main feature. OPC features follow various design rules, such as rules based on the size of the smallest feature that can be formed on a wafer using optical lithography. Other design rules are derived from the mask manufacturing process or, in the case of forming patterns on a reticle using a CP (character projection) type charged particle beam writing system, from the stencil manufacturing process. Summary of the Invention
[0008] In an embodiment, a method for calculating a pattern to be manufactured on a substrate includes inputting a physical design pattern and determining a plurality of possible neighborhoods for the physical design pattern. A plurality of possible mask designs are generated for the physical design pattern, and the plurality of possible mask designs correspond to the plurality of possible neighborhoods. A plurality of possible patterns on the substrate are calculated, and the plurality of possible patterns on the substrate correspond to the plurality of possible mask designs. A variation band is calculated from the plurality of possible patterns on the substrate, and the physical design pattern is modified to reduce the variation band.
[0009] In an embodiment, a method for calculating a pattern to be manufactured on a substrate includes inputting a physical design, inputting a set of neural network parameters for calculating the pattern to be manufactured on the substrate, and generating a plurality of possible neighborhoods for the physical design. A plurality of patterns to be manufactured on the substrate are calculated for the physical design in each of the plurality of possible neighborhoods. The neural network is trained with the plurality of calculated patterns, and the training is executed using a computing hardware processor. The set of parameters is adjusted to reduce manufacturing variations of the plurality of calculated patterns to be manufactured on the substrate.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 3F
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
[0011] The present disclosure describes methods and systems for improving the manufacturing accuracy and calculation time of patterns. Embodiments enable simultaneous modeling of multiple parameters at different stages of the manufacturing process, such as physical design, mask, and substrate stages. The results of multiple scenarios are output by means of visualized schematics and the like so that the user can view and change them almost in real time. Embodiments estimate variations in mask design and wafer manufacturing steps and utilize statistical methods to improve the physical design of patterns.
[0012] Typical RET methods have OPC verification for identifying and correcting hotspots. A hotspot is a region that requires ideal conditions for proper printing. Therefore, it may not be resilient to manufacturing variations or may not even be printed properly under ideal conditions. Hotspots reduce the yield. In lithography, a feature required on a substrate, called a main feature, is printed with higher fidelity and an improved process window when an SRAF is added that is too small to print itself but has a favorable effect on the way neighboring main features are printed.
[0013] However, adding OPC features such as SRAF is a very cumbersome task, requiring costly computation time and resulting in more expensive reticles. Not only is the OPC pattern complex, but because the optical proximity effect extends over a longer distance than the minimum line and space dimensions, the exact OPC pattern at a given location depends heavily on what other shape configurations are in the vicinity. Thus, for example, line ends have serifs of different sizes depending on what is near them on the reticle. This is true even when the goal is to form the same shape on the wafer. These slight but critical variations are important and have prevented others from forming reticle patterns that accurately form the desired shape on the wafer. To quantify what is meant by these slight variations, the typical slight variations in OPC decoration from neighborhood to neighborhood are between 5% and 80% of the main feature size. When these OPC variations form substantially the same pattern on the wafer, the shape configuration on the wafer is targeted to be the same within a specific error specified according to the details of the function that the shape configuration is designed to perform, e.g., the details of a transistor or wire. Nevertheless, the typical specification is between 2% and 50% of the main feature range.
[0014] Inverse lithography technology (ILT) is one of the OPC technologies. ILT is a process that directly calculates the pattern to be formed on a reticle from the pattern formed on a substrate such as a silicon wafer. This involves using the desired pattern on the substrate as input to simulate the optical lithography process in the reverse direction. The ILT-calculated reticle pattern may be purely curved, i.e., completely non-linear, and examples include circular, substantially circular, annular, substantially annular, elliptical and / or substantially elliptical patterns. These patterns require a very large number of variable shaped beam (VSB) shots to expose the curved pattern, and thus are not practical for a variable shaped beam (VSB) mask writer with conventional fracturing. A linear approximation or straightening of the curved pattern may be used. However, the linear approximation reduces the accuracy compared to the ideal ILT curved pattern. Further, when the linear approximation is generated from the ideal ILT curved pattern, the overall calculation time increases compared to the ideal ILT curved pattern. Mask writing time is an important business factor, and the VSB writing time is proportional to the number of VSB shots required for printing. Model-based mask data preparation using overlapping shots can significantly reduce the impact of the writing time for curved ILT mask design. However, generally, curved shapes take longer to write than linear shapes.
[0015] In multi-beam writing, the need to perform straightening to convert the curved shape for VSB writing is eliminated. However, the resilience to mask printability and manufacturing variations remains an important consideration for the mask shapes output by ILT. For example, if the shapes are too small, too close to each other, or the bends in the shape contours are too sharp, it becomes difficult to reliably manufacture the mask, especially from the perspective of manufacturing variations. The remaining problem associated with ILT is the enormous computational requirements for full design, especially for full mask layer high-density simulations of full reticle size designs where the wafer size is typically about 3.0 cm × 2.5 cm in semiconductor manufacturing.
[0016] Referring now to the drawings, like numerals refer to like features. FIG. 1 is a charged particle beam writing system that uses a variable shaping beam (VSB) to fabricate surface 12, and here shows one embodiment of a lithography system such as electron beam writing system 10. Electron beam writing system 10 includes an electron beam source 14 that projects an electron beam 16 toward an aperture plate 18. The plate 18 has an opening 20 formed therein for passing the electron beam 16. When the electron beam 16 passes through the opening 20, the electron beam 16 is directed or deflected by a lens system (not shown) as electron beam 22 toward another rectangular aperture plate or stencil mask 24. The stencil 24 has a number of openings or apertures 26 that define various simple shapes such as rectangles and triangles. Each aperture 26 formed in the stencil 24 is used to form a pattern on the surface 12 of a substrate 34 such as a silicon wafer, reticle, or other substrate. The electron beam 30, after exiting one of the apertures 26, passes through an electromagnetic or electrostatic reduction lens 38 that reduces the size of the pattern exiting the aperture 26. In a generally available charged particle beam writing system, the reduction factor is between 10 and 60. The reduced electron beam 40 exits the reduction lens 38 and is directed as pattern 28 onto the surface 12 by a series of deflectors 42. The surface 12 is coated with a resist (not shown) that reacts with the electron beam 40. The electron beam 22 may be directed to overlap the variable portion of the aperture 26 that affects the size and shape of the pattern 28. A blanking plate (not shown) may be used to deflect the beam 16 or shaped beam 22 to prevent the electron beam from reaching the surface 12 during the period after each shot while the lenses and deflectors 42 that direct the beam 22 are being readjusted for subsequent shots. Conventionally, the blanking period may be a fixed length of time or may be varied, for example, depending on how much the deflector 42 needs to be readjusted for the position of the next shot.
[0017] In an electron beam writing system 10, a substrate 34 is mounted on a movable platform or stage 32. The stage 32 enables repositioning of the substrate 34 such that a pattern larger than the maximum deflection ability or field size of the charged particle beam 40 can be written onto the surface 12 in a series of sub - fields within the capabilities of a deflector 42 that deflects the beam 40. In one embodiment, the substrate 34 may be a reticle. In this embodiment, after being exposed with a pattern, the reticle goes through various manufacturing steps to become a lithography mask or a photomask. This mask is then used in an optical lithography apparatus to generally project an image of the reticle pattern 28, which is reduced in size, onto a silicon wafer to fabricate an integrated circuit. More generally, the mask is used in another device or apparatus to form the pattern 28 on a substrate (not shown).
[0018] A charged particle beam system can expose a surface using a plurality of individually controllable beams or beamlets. FIG. 2 shows an electro-optical schematic diagram in which there are three charged particle beamlets 210. Each beamlet 210 is associated with a beam controller 220. Each beam controller 220 allows, for example, the associated beamlet 210 to strike the surface 230 or prevents the beamlet 210 from striking the surface 230. In some embodiments, the beam controller 220 may control the beam blur, magnification, size, and / or shape of the beamlet 210. In the present disclosure, a charged particle beam system having a plurality of individually controllable beamlets is referred to as a multi-beam system. In some embodiments, charged particles from a single source may be subdivided to form a plurality of beamlets 210. In other embodiments, a plurality of sources may be used to generate a plurality of beamlets 210. In some embodiments, the beamlets 210 may be shaped by one or more apertures, and in other embodiments, there may be no apertures for shaping the beamlets. Each beam controller 220 enables individual control of the exposure period of the associated beamlet. Generally, the beamlets are reduced in size by one or more lenses (not shown) before striking a surface 230 that is typically coated with a resist. In some embodiments, each beamlet may have a separate electro-optical lens, and in other embodiments, in some cases, a plurality of beamlets, including all the beamlets, may share an electro-optical lens.
[0019] For the purposes of the present disclosure, a shot is an exposure of a surface area over a period of time. The surface area may be discontinuous and may be composed of a plurality of regions that are smaller than the surface area. A shot may be composed of a plurality of other shots. The plurality of shots may or may not overlap. Also, they may or may not be exposed simultaneously. A shot may include a specified dose. The dose may not be specified. A shot can use a shaped beam, an unshaped beam, or a combination of a shaped beam and an unshaped beam. FIGS. 3A - 3F show various different shots. FIG. 3A shows an example of a rectangular shot 310. A VSB charged particle beam system can form, for example, rectangular shots having various X and Y dimensions. FIG. 3B shows an example of a character projection (CP) shot 320, which is circular in this example. FIG. 3C shows an example of a trapezoidal shot 330. In one embodiment, the shot 330 can be generated using a raster-scanned charged particle beam. Here, the beam is scanned in the X direction, for example, as indicated by the scan line 332. FIG. 3D shows an example of a drag shot 340 disclosed in U.S. Patent Application Publication No. 2011 / 0089345. The shot 340 is formed by exposing the surface with a curved shaped beam 342 at an initial reference position 344 and then moving the shaped beam 342 from the initial position 344 across the surface to a position 346. The path of the dragged shot may be, for example, linear, piecewise linear, or curved.
[0020] FIG. 3E shows an example of a shot 350 that is an array of circular patterns 352. The shot 350 can be formed in various ways, including a plurality of shots of one circular CP character, one or more shots of a CP character that is an array of circular apertures, and one or more multi-beam shots that use circular apertures. FIG. 3F shows an example of a shot 360 that is a sparse array of rectangular patterns 362 and 364. The shot 360 can be formed in various ways, including a plurality of VSB shots, CP shots, and one or more multi-beam shots that use rectangular apertures. In some embodiments of the multi-beam, the shot 360 can include a plurality of alternating groups of other multi-beam shots. For example, after simultaneously shooting pattern 362, pattern 364 can be simultaneously shot at a different time from pattern 362.
[0021] Many techniques are used to form patterns on a reticle, some using optical lithography or charged particle beam lithography. The most commonly used system is the variable shaped beam (VSB), which, as described above, exposes a resist-coated reticle surface with a dose of electrons defined by simple shapes such as Manhattan rectangles and 45-degree right triangles. In conventional mask writing, the dose or shot of electrons is designed to avoid overlap as much as possible in order to greatly simplify the calculation of how the resist on the reticle records a pattern. Similarly, the set of shots is designed to completely cover the pattern area to be formed on the reticle. U.S. Patent No. 7,754,401, owned by the assignee of this patent application, discloses a mask writing method in which intentional shot overlap is used to write a pattern. When overlapping shots are used, a charged particle beam simulation can be used to determine the pattern recorded by the resist on the reticle. By using overlapping shots, a pattern can be written with fewer shots, or with higher accuracy, or both. U.S. Patent No. 7,754,401 also discloses the use of dose modulation, where the assigned dose of a shot is different from the dose of other shots. The term model-based fracturing is used to describe the process of determining shots using the technology of U.S. Patent No. 7,754,401.
[0022] FIG. 4 shows an embodiment of a charged particle beam exposure system 400. The charged particle beam system 400 is a multi-beam system in which a plurality of individually controllable shaped beams can simultaneously expose a surface. The multi-beam system 400 has an electron beam source 402 that generates an electron beam 404. The electron beam 404 is directed towards an aperture plate 408 by a condenser 406 that includes electrostatic and / or magnetic elements. The aperture plate 408 has a plurality of apertures 410 through which the electron beam 404 is irradiated. The electron beam 404 forms a plurality of shaped beamlets 436 by passing through these apertures 410. In some embodiments, the aperture plate 408 has hundreds or thousands of apertures 410. FIG. 4 shows an embodiment having a single electron beam source 402. In other embodiments, the apertures 410 are irradiated with electrons from a plurality of electron beam sources. The apertures 410 may be rectangular or may have different shapes such as, for example, circular. Next, a set of beamlets 436 illuminates a blanking controller plate 432. The blanking controller plate 432 has a plurality of blanking controllers 434, and each of the plurality of blanking controllers 434 is aligned with a beamlet 436. Each blanking controller 434 can control the associated beamlet 436 to cause the beamlet 436 to strike the surface 424 or to prevent the beamlet 436 from striking the surface 424. The total energy or "dose" of the beamlet applied is controlled by the amount of time the beam strikes the surface. Thus, the dose of each beamlet is independently controlled. The area where the beam strikes the surface may surround a portion of the entire pixel.
[0023] The possibility of a multi-beam system for modifying the dose of each pixel to apply a bias to the edges of a shape is described in U.S. Patent No. 10,444,629, owned by the assignee of this patent application. Further, U.S. Patent No. 10,444,629 discloses an improvement in the dose margin to make the edges less susceptible to manufacturing variations. This method of modifying the dose for each pixel is referred to as pixel-level dose correction (PLDC).
[0024] In FIG. 4, four beamlets capable of hitting the surface 424 are illustrated as beamlet 412. In one embodiment, the blanking controller 434 prevents the beamlet 436 from hitting the surface 424 by deflecting the beamlet 436 such that the beamlet 436 is stopped by the aperture plate 416 including the aperture 418. In some embodiments, the blanking plate 432 may be directly adjacent to the aperture plate 408. In other embodiments, the relative positions of the aperture plate 408 and the blanking controller 432 may be reversed from the positions shown in FIG. 4 such that the beam 404 hits a plurality of blanking controllers 434. The lens system including elements 414, 420, 422 is typically sized smaller than the plurality of apertures 410 and enables projecting a plurality of beamlets 412 onto the surface 424 of the substrate 426. The reduced beamlets form a beamlet group 440 that hits the surface 424 and generates a pattern that coincides with a subset of the apertures 410. The subset is the aperture 410 that enables the beamlet 436 to hit the surface 424 by the corresponding blanking controller 434. In FIG. 4, a beamlet group 440 having four beamlets is illustrated as forming a pattern on the surface 424.
[0025] The substrate 426 is positioned on a movable platform or stage 428 that can be repositioned by the actuator 430. By moving the stage 428, the beam 440 can expose an area larger than the dimensions of the maximum pattern formed by the beamlet group 440 through a plurality of exposures or shots. In some embodiments, the stage 428 remains stationary during exposure and is then repositioned for the next exposure. In other embodiments, the stage 428 moves continuously and at a variable speed. In still other embodiments, the stage 428 is continuous but moves at a constant speed. Thereby, the accuracy of stage positioning can be improved. In embodiments where the stage 428 moves continuously, by using a set of deflectors (not shown), the beam can be moved to match the direction and speed of the stage 428, and the beamlet group 440 can be kept stationary with respect to the surface 424 during exposure. In yet other embodiments of the multi-beam system, each beamlet of the beamlet group may be deflected to cross the surface 424 independently of the other beamlets within the beamlet group. In some embodiments, the stage 428 can be moved in a single direction across the entire exposure area to expose a portion called a stripe of the entire area. Thus, the entire exposure area is exposed as a plurality of stripes. In some embodiments, the stage 428 moves in opposite directions on adjacent stripes or alternate stripes.
[0026] Other types of multi-beam systems can generate a plurality of unshaped beamlets 436, such as by using a plurality of charged particle beam sources for forming an array of Gaussian beamlets.
[0027] As shown in FIG. 1, the minimum size pattern that can be projected onto the surface 12 with reasonable accuracy is limited by various short-distance physical effects associated with the electron beam writing system 10 and the surface 12, which usually includes a resist coating on the substrate 34. These effects include forward scattering, Coulomb effects, and resist diffusion, etc. β fThe term beam blur, which is what we call it, is used to encompass all of these short-range effects. The latest electron beam writing systems can achieve an effective beam blur radius or β in the range of 20 nm to 30 nm. f Forward scattering may constitute 1 / 4 to 1 / 2 of the total beam blur. Recent electron beam writing systems have multiple mechanisms for minimizing each of the multiple components of beam blur. Since some components of beam blur are a function of the calibration level of particle beam writing, the β of two particle beam writings of the same design may be different. The diffusion characteristics of the resist may also vary. The variation of β based on shot size or shot dose is simulated and can be systematically explained. However, there are other effects that cannot be explained or are not accounted for, and they appear as random variations. f f
[0028] The shot dose of a charged particle beam writing apparatus such as an electron beam writing system is a function of the intensity of the beam source 14 and the exposure time of each shot. Usually, the beam intensity is fixed and the exposure time is changed to obtain a variable shot dose. Different regions within a shot may have different exposure times, such as in a multi-beam shot. The exposure time is changed to compensate for various long-range effects such as backscattering, fogging, and loading effects in a process called proximity effect correction (PEC). An electron beam writing system can usually set an overall dose called the base dose that affects all shots in the exposure path. Some electron beam writing systems perform dose compensation calculations within the electron beam writing system itself so that the dose of each shot is not individually assigned as part of an input shot list. Thus, the input shots have an unassigned shot dose. In such an electron beam writing system, all shots have a base dose before PEC. In other electron beam writing systems, the dose can be assigned on a per-shot basis. In an electron beam writing system capable of per-shot dose assignment, the number of available dose levels can be 64 to 4096 or more, or relatively few available dose levels, such as 3 to 8 levels.
[0029] The mechanisms within an electron beam writing system have a relatively coarse resolution for calculations. Therefore, the intermediate range corrections required for EUV masks in the 2μm range cannot be accurately calculated by current electron beam writing systems.
[0030] For example, when exposing a repetitive pattern on a surface using charged particle beam lithography, the size of each pattern instance measured on the ultimately fabricated surface will vary slightly due to manufacturing variations. The amount of size variation is an essential manufacturing optimization criterion. In current mask masking, a root mean square (RMS) variation of less than 1 nm (1 sigma) in pattern size is desirable. More size variation results in more variation in circuit performance, requires a higher design margin, and makes it increasingly difficult to design faster and lower power integrated circuits. This variation is referred to as critical dimension (CD) variation. Low CD variation is desirable, indicating that manufacturing variations result in relatively small size variations on the final fabricated surface. At smaller scales, the effect of high CD variation is observed as line edge roughness (LER). LER is caused by each part of the line edge being fabricated slightly differently, resulting in undulations in a line intended to have a straight edge. CD variation is inversely proportional, inter alia, to the slope of the dose curve at the resist threshold, which is referred to as the edge slope. Thus, the edge slope or dose margin becomes an important optimization factor for particle beam writing on the surface. In the present disclosure, the terms edge slope and dose margin are used interchangeably.
[0031] As described in U.S. Patent No. 8,473,875, "Method and System for Forming High Accuracy Patterns Using Charged Particle Beam Lithography," owned by the assignee of the present patent application, FIGS. 5A - 5B show how critical dimension (CD) variations are reduced by exposing a pattern on a resist to generate a relatively high edge gradient in the exposure or dose curve. FIG. 5A shows a cross - sectional dose curve 502. The x - axis represents the cross - sectional distance through the exposed pattern (e.g., the distance perpendicular to two of the edges of the pattern), and the y - axis represents the dose received by the resist. The pattern is recorded by the resist when the received dose is higher than a threshold. FIG. 5A shows two thresholds. FIG. 5A shows the effect of variations in resist sensitivity. The higher threshold 504 records a pattern of width 514 by the resist. The lower threshold 506 records a pattern of width 516 by the resist. Here, width 516 is larger than width 514. FIG. 5B shows another cross - sectional dose curve 522. Two thresholds are illustrated, and threshold 524 is the same as threshold 504 in FIG. 5A. Threshold 526 is the same as threshold 506 in FIG. 5A. The slope of dose curve 522 is higher than the slope of dose curve 502 in the vicinity of the two thresholds. For dose curve 522, a pattern of width 534 is recorded by the resist with the higher threshold 524. A pattern of width 536 is recorded by the resist with the lower threshold 526. As can be seen from the figure, since the edge gradient of dose curve 522 is larger compared to dose curve 502, the difference between width 536 and width 534 is smaller than the difference between width 516 and width 514. When the resist - coated surface is a reticle, the lower sensitivity of curve 522 to variations in the resist threshold can bring the pattern width on the photomask manufactured from the reticle closer to the target pattern width of the photomask. Thereby, when the photomask is used to form a pattern on a substrate such as a silicon wafer, the yield of the integrated circuit that can be used as a product is improved.Improved resistance to dose variations in each shot is observed for dose profiles with a higher edge gradient. Thus, it is desirable to achieve a relatively high edge gradient, such as that of dose profile 522.
[0032] Design cells in semiconductor manufacturing (e.g., memory cells or standard cells from a library) represent an abstraction of electronic components in physical design. The cell-based approach enables designers to reuse components from relatively simple designs to complex designs. A cell is composed of several layers including shapes whose size and orientation vary. A cell, i.e., a set of shapes from a given layer within the cell, is arranged relatively separately in the design with no adjacent shapes nearby, and results in a different pattern on the substrate when the cell is arranged with other cells and / or shapes in its immediate vicinity, i.e., with different adjacent shapes in proximity on the same layer. FIG. 6 shows an example of a standard cell including two cells, i.e., cells A and B, in various legal orientations. Due to the proximity of the arrangements within cells adjacent to each other (i.e., in the same neighborhood), each orientation results in variations in the mask design calculated for each cell. As described above, optical proximity correction (OPC) varies considering light diffraction and the light interaction with adjacent features. In the refinement step of proximity effect correction (PEC), the shot dose is adjusted as needed due to various long-range effects for each neighborhood.
[0033] Manufacturing process variations and proximity-induced variations have a significant impact on design performance and manufacturing reliability. It is desirable for circuit and / or mask designers to visualize the impact of various sources of variations in relation to the actual design. For example, due to process variations, the pattern width on a photomask may vary from the intended or target width. Variations in the pattern width on the photomask cause variations in the pattern width on the wafer exposed using the photomask in the photolithography process. The sensitivity of the wafer pattern width to the variation in the photomask pattern width is referred to as the mask edge error factor, i.e., MEEF. In a photolithography system using a 4x photomask, when a pattern reduced 4 times from the photomask pattern is projected onto the wafer by the photolithography process, for example, an MEEF of 1 means that for a 1 nm error in the pattern width on the photomask, the pattern width on the wafer changes by 0.25 nm. An MEEF of 2 means that for a 1 nm error in the pattern width of the photomask, the pattern width on the wafer changes by 0.5 mm. In the smallest integrated circuit processes, the MEEF is greater than 2. By visualizing and understanding these sources of variations and their effects well, designers can modify the design itself (i.e., the shape including the design) to be more robust against such variations.
[0034] FIG. 7 is a flow 700 for calculating a pattern manufactured on a substrate such as a silicon wafer according to some embodiments. In a first step, for example, a physical design pattern 702 such as a physical design of an integrated circuit is input. In one embodiment, the pattern manufactured on the substrate is calculated from the physical design pattern. These calculations include determining manufacturable shapes for other components found in the physical design such as logic gates, transistors, metal layers, and integrated circuits. The physical design may be linear, piecewise linear, partially curved, or fully curved. In particular, curved patterns are very computationally intensive. Therefore, it is very beneficial that the pattern can be optimized by calculating the cumulative effect of variations from multiple manufacturing steps as in this embodiment.
[0035] Step 704 includes generating a plurality of possible neighborhoods for the physical design. In some embodiments, the physical design pattern is part of the overall design. The plurality of possible neighborhoods generated in step 704 are the plurality of actual neighborhoods used for the physical design pattern. Neighborhood variations can be synthesized. For example, one method is to randomly place cells in all possible neighborhoods where they may ultimately end up, i.e., all possible neighborhoods surrounded by various neighboring cells that are most likely to be surrounded by the actual circuit design. In some embodiments, a part of the physical design pattern is an instance of the physical design pattern, and the plurality of possible neighborhoods include all neighborhoods of each instantiation. Thus, instances of the cell of interest are arranged with instances of various neighborhoods facing up, down, left, or right in their various legal orientations, and with various offsets in that arrangement, parallel to the various orientations of the various neighboring cells. In some embodiments, a part of the overall design is a standard cell design including a plurality of standard cells, and the plurality of possible neighborhoods include all legal orientations of the standard cells.
[0036] In step 706, a composite of substrate layers, a portion of which is separated into mask layers, is formed from the physical design. This step may also be referred to as the coloring step or colorization, in which each feature on the reticle layer is colored to reflect the assignment to a specific mask layer of the feature. The colorization step 706 may be performed on the physical design pattern prior to optical proximity effect correction (OPC). In step 708, OPC is performed on the physical design pattern to generate a plurality of possible mask designs 710. Each mask design in the plurality of mask designs corresponds to one of the plurality of possible proximities generated in step 704. By combining the plurality of possible mask designs 710, a nominal mask design with variations can be formed. Conventionally, the nominal mask design is determined by calculating the nominal profile of the mask design using a nominal dose, such as 1.0, and a threshold, such as 0.5. In one embodiment, the nominal profile of the mask design is calculated from the plurality of possible mask designs 710. The variations are calculated for all possible proximities generated in step 704.
[0037] In one embodiment of the present disclosure, the OPC step 708 includes inverse lithography technology (ILT) that forms an ideal curve ILT pattern. In other embodiments, ILT with linearization of the curved pattern is used.
[0038] The OPC features or ILT patterns for the same physical design pattern vary from proximity to proximity. A plurality of possible mask images can be calculated from the plurality of possible mask designs in each of the many possible proximities. In one embodiment, the nominal mask design is calculated from the calculated OPC features or ILT patterns in the many possible proximities. In some embodiments, the plurality of possible mask designs are stored within a file system 726 on a disk, memory, or any other storage device.
[0039] In some embodiments, the mask process simulation step 716 includes mask data preparation (MDP) that prepares a mask design for mask writing. This step includes "fracturing" the data into trapezoids, rectangles, or triangles. Mask process correction (MPC) can also be included in step 716. MPC geometrically modifies the shape and / or assigns a dose to the shape to bring the resulting shape on the mask closer to the desired shape. MDP can use the possible mask design 710 or the result of MPC as input. MPC is performed as part of fracturing or other MDP operations. Other corrections are also performed as part of fracturing or other MDP operations. Possible corrections include forward scattering, resist diffusion, Coulomb effects, etching, backscattering, fogging, loading, resist charging, and EUV mid-field scattering, among others. In step 716, pixel-level dose correction (PLDC) can also be applied. In other embodiments, a multi-beam VSB shot list or exposure information can be generated to generate a plurality of possible mask images 718 from the possible mask design 710. In some embodiments, a set of VSB shots is generated for the calculated mask patterns in a plurality of calculated mask patterns. In some embodiments, MPC and / or MDP are performed on the possible mask design 710.
[0040] In step 716, calculating the plurality of possible mask images 718 includes charged particle beam simulations. In some embodiments, the plurality of possible mask images are stored on the file system 726. The effects to be simulated include forward scattering, backscattering, resist diffusion, Coulomb effects, fogging, loading, and resist charging, etc. Also, step 716 includes mask process simulations in which the effects of various post-exposure processes are calculated. These post-exposure processes include resist baking, resist development, and etching, etc. When the charged particle beam simulation is performed on a mask on any layer, the simulation is performed over a range of process variations in order to establish a manufacturable contour for the mask itself. The contour may be an extension from the nominal contour. The nominal contour in this case may be based on a pattern generated with a specific resist threshold, for example, a threshold of 0.5. In some embodiments, by calculating the percentage difference in exposure dose, for example, a dose variation of + / - 10%, a mask image with variations is created for display in a viewport 728 that includes the upper and lower limits of the process variation band surrounding the nominal contour. In some embodiments, the positive variation and the negative variation may be different from each other, for example, +10% and -8%. The charged particle beam simulation and the mask process simulation may be performed separately from each other in step 716.
[0041] In substrate simulation step 720, calculating possible substrate patterns 722 includes lithography simulation using the calculated mask image 718. A plurality of possible patterns on the substrate are calculated from a plurality of mask images. Each pattern among the plurality of possible patterns on the substrate corresponds to a set of manufacturing variation parameters. Calculating the substrate pattern from the calculated mask image is described in U.S. Patent No. 8,719,739, owned by the assignee of this patent application, entitled "Method and System for Forming High Accuracy Patterns Using Charged Particle Beam Lithography" "Charged Particle Beam Lithography". A plurality of possible patterns on the substrate 722 may be combined to form a nominal substrate pattern with variations. In some embodiments, the sources of substrate pattern variations include some variations in depth of focus, such as + / - 10% in exposure, and some variations in exposure (dose) in combination with + / - 30 nm in depth of focus. In some embodiments, the plus and minus variations may be different from each other, for example, +5% / -7% and 30 nm / -28 nm. Conventional statistical methods are used to generate 3-sigma variations from the nominal profile. The variations include a lower 3-sigma that is smaller than the nominal profile for the minimum value and an upper 3-sigma that is larger than the nominal profile for the maximum value. In some embodiments, instead of calculating the 3-sigma variations extending from the nominal profile, a mask image with variations is created by combining a plurality of mask images 718 that include a process variation band with a lower and an upper limit. In some embodiments, the substrate pattern is formed on the wafer using an optical lithography process that uses a mask image with variations. In some embodiments, a plurality of possible patterns on the substrate are stored in the file system 726. In some embodiments, a wafer process simulation is performed on the substrate pattern.Examples of wafer process simulations include resist baking, resist development, and etching simulations. Lithography simulation 720 and wafer process simulation may be separate steps, and optionally, each step has process variations. In other embodiments, lithography simulation 720 includes flat panel display (FPD) simulation, microelectromechanical systems (MEMS) simulation, other process simulations, or other things manufactured on a substrate.
[0042] In each step of FIG. 7, the variations are statistically accumulated and take into account the variations from previous steps such that not only the variations at the final step which determine the possible patterns on the substrate 722, but also the variations in the mask process 716 and the mask design 710 are incorporated. In step 724, a process variation band is calculated from the possible substrate patterns. To more efficiently calculate many possible combinations of variations, those variations are accumulated using insights into how specific variations and pattern parameters can affect one another. For example, instead of simply feeding the minimum and maximum 3-sigma values from one step to the next, the worst-case variation fed to the next step can take into account the distance from one pattern to another pattern. This is because features that are closer to each other affect each other more than features that are far apart. Due to the impact of these variations on design performance and manufacturing reliability, it is desirable to enable designers to visualize the impact of different variations in the context of an actual circuit design. Visualizing the impact of the statistically accumulated variations predicted on the substrate can be shown by calculating the variation band in step 724 or by visualizing the impact of different variations at each step. If the variations are unacceptable in step 725, the designer can modify the physical design 702 to produce an improved physical design so as to ensure that the improved physical design is more robust to manufacturing variations. Modifications to the physical design can include modifying the possible neighborhood of the physical design or modifying the coloring, for example, modifying the shape assignment to any particular layer. In a design environment where curvilinear designs are allowed, providing the calculated nominal profile as a new manufacturable physical design has the advantage that manufacturing variations are reduced. This is because designs that can be manufactured have less variation than designs that cannot be manufactured (e.g., shapes with 90-degree angles that are essentially impossible to manufacture). It should be noted that the manufacturing variations predicted in the above steps need to be repeated with the improved physical design in order to estimate the manufacturing variations of the modified physical design.In some embodiments, the variations at each step may be shown simultaneously in a single viewport 728, along with a nominal contour with variations superimposed on a corresponding design, image, or pattern, or the variations may be shown in a plurality of viewports 728.
[0043] Calculating the pattern to be manufactured on a substrate includes calculating a plurality of substrate patterns from a plurality of mask images calculated from a plurality of mask designs. These calculations can be quite time-consuming and may take time to recover even if pre-calculated and stored. In one embodiment, the calculation of the pattern to be manufactured on a substrate may be learned in a neural network. A neural network is a framework of machine learning algorithms that cooperate to predict patterns based on previous training processes. Embodiments include training a neural network to calculate the pattern to be manufactured on a substrate using an input physical design 702 and any combination of one or more outputs shown in FIG. 7, such as a possible mask design 710, a possible mask image 718, a possible substrate pattern 722, etc. Also, step 725 includes adjusting a set of neural network parameters to reduce manufacturing variations of the calculated patterns as part of the process of training the neural network. The training of the neural network is performed using a computing hardware processor. Such training achieves similar goals as the previous embodiments, but once trained, the conversion in the trained neural network can be much faster, for example 10 times faster, than with simulation alone. In one embodiment, a trained neural network or a group of trained neural networks can convert a physical design pattern into a pattern to be manufactured on a substrate. That is, in some embodiments, calculating the pattern on a substrate includes a neural network having a physical design as an input.
[0044] In one embodiment, each of the outputs 710, 718, and 722 is generated by a trained neural network. The digital twin replicates the physical entity. Conventionally, digital twins model the characteristics, conditions, and attributes of their real-world counterparts. This is achieved through exact simulation. In this application, the simulation results can be used to train a neural network, resulting in a neural network digital twin that runs much faster than simulation alone. At any stage or combination of stages, the neural network digital twin trained using the simulated data is used to perform the image-to-image conversion. In one embodiment, for example, a deep convolutional neural network (CNN) architecture such as a fully convolutional network (FCN) is trained using a pair of image data representing the input and output of any of the computational steps in FIG. 7, respectively. In FIG. 8, an image 800 representing the physical design or CAD data is provided as an input to a CNN 810 such as an FCN, and an image 820 representing the manufactured output shape is generated by the CNN 810. Other neural network architectures such as U-Net, a type of FCN, or a generative adversarial network (GAN) are also used. In other embodiments, the neural network may be trained to generate OPC / ILT features or shapes for various proximities, generate an image optimized for mask process correction or data preparation, calculate the pattern on the substrate, or perform any combination of steps. In one embodiment, any one or more of the steps in FIG. 7 may be replaced in combination with a digital twin, a neural network, a group of digital twins, or a neural network.
[0045] In an embodiment, a method for calculating a pattern manufactured on a substrate includes inputting a physical design pattern 702, determining a plurality of possible neighborhoods of the physical design pattern (step 704), and generating a plurality of possible mask designs 710 of the physical design pattern, where the plurality of possible mask designs correspond to the plurality of possible neighborhoods. The method also includes calculating a plurality of possible patterns on the substrate corresponding to the plurality of possible mask designs (722), calculating a variation band from the plurality of possible patterns on the substrate (step 724), and modifying the physical design pattern to reduce the variation band (loop from step 725 to physical design 702).
[0046] In some embodiments, the method includes calculating a plurality of calculated mask images from the plurality of possible mask designs (step 718). In some embodiments, calculating the plurality of possible mask images includes charged particle beam simulation (step 716). In some embodiments, modifying the physical design pattern includes modifying a plurality of possible neighborhoods of the physical design pattern (step 704). In some embodiments, the variation band of step 724 corresponds to a set of manufacturing variation parameters. In some embodiments, the variation band of step 724 includes process variations having a lower limit and an upper limit surrounding a nominal substrate pattern. In some embodiments, the method includes performing a coloring step 706 that separates the shape of the physical design pattern into layers, and in a further embodiment, modifying the physical design pattern includes modifying the coloring step.
[0047] In some embodiments, physical design 702 includes optical proximity correction of physical design patterns (step 708). In some embodiments, determining a plurality of possible proximities 704, generating a plurality of possible mask designs 710, or calculating a plurality of possible patterns on substrate 722 includes using a neural network. In some embodiments, calculating a plurality of possible patterns on a substrate includes lithography simulation (step 720).
[0048] In some embodiments, physical design pattern 702 includes a part of the overall design, and the method further includes determining an actual set of proximities in which the physical design pattern is used in the overall design (step 704). The part of the overall design may be an instance of the physical design pattern, and the plurality of possible proximities includes all proximities of each instantiation.
[0049] U-Net applications such as FCN are used for predicting process variability bands related to semiconductor manufacturing. The first U-Net architecture was developed for biomedical image segmentation problems. In the first U-Net model architecture, each layer is characterized by a multi-channel feature map with multi-channels that vary across the layers. In the final layer, 1×1 convolutions are used to map each 64-component feature vector to the desired number of classes. In total, a typical network has 23 convolutional layers.
[0050] In one embodiment, the main neural network architecture for the FCN is essentially an encoder-decoder network as shown in FIG. 9, where the left encoding side and bottleneck layer 910 guide the model to learn the low-dimensional encoding of the input image 900. Next, a decoder network including layers 912, 914, 916, and 918 decodes the low-dimensional representation of the image back to full output resolution, and both sides cooperate to learn the conversion from the input image 900 to the output image 920 during training. The copy and crop operations indicated by the horizontal arrows from the encoder layers to their corresponding decoder layers act as skip connections that provide additional information from the encoder side of the network and are concatenated with the information on the decoder side to assist in the localization of information in the x, y space.
[0051] If the input image is too large to be processed at once, the input image is split into a set of image tiles. The image tiles may overlap with each other. Then, each of the smaller tiles may be processed by the network, and the output tiles may be collected and reconstructed into the final output image. To reduce artifacts at tile boundaries, the FCN includes a halo of adjacent pixels. The halo may overlap with adjacent tiles and may be used to reconstruct the large input image.
[0052] In semiconductor manufacturing applications, the input 900 to the neural network represents an input image, or tiles from an input image representing a design intent, i.e., assuming "ideal" rather than a realistic manufacturing process. In one embodiment, the output image 920 represents what would actually be manufactured by a realistic manufacturing process where sharp corners are rounded and small squares are made as circles or ellipses, etc. After training the FCN on semiconductor manufacturing image data to determine a set of model weights, the model weights are significantly different from the model weights used in other applications.
[0053] In one embodiment, the FCN architecture shown in FIG. 9 is a multi-resolution U-Net with an initial number of filters reduced from 64 to 8 in the first layer 902, where the filter doubling continues after each max pooling operation in each encoder layer 902, 904, 906, 908 and bottleneck layer 910. This has the effect of significantly reducing the overall number of trainable parameters for the network while maintaining a sufficient level of accuracy for semiconductor manufacturing applications. In another embodiment, there may be 16 filters in the first layer 902. The last encoder layer 908 and bottleneck layer 910 may each employ dropout regularization. In one embodiment, the input tile size and output tile size may be, for example, 256×256 pixels (along with a 128×128 inner core tile surrounded by a 64-pixel-wide halo). In another embodiment, the network may be further modified by removing some of the layers (shorter U depth) or adding additional layers (deeper U) as needed for accuracy. In another embodiment, instead of doubling the number of filters after each downsampling (max pooling) or upsampling convolution, different ratios are used. In one embodiment, a fixed ratio (e.g., 2.0) may be used for each layer, and in an alternative embodiment, different layer-specific ratios may be used for each layer. For example, the ratio may be lower for the U-shaped bottom bottleneck layer and gradually increase as it approaches and then decrease again as it moves away from the bottleneck layer towards the output. These ratios and other network parameters are adjusted during the training phase. That is, an initial set of parameters is input for the neural network, and the set of parameters is adjusted as the neural network is trained. In one embodiment, the adjustment may be repeated for each different manufacturing process and / or for each different layer during the manufacturing process.
[0054] In one embodiment, the network has a single input and a single output representing a manufactured output image corresponding to a single set of process conditions, such as a process corner. The input to the network consists of an image corresponding to computer-aided design (CAD) data (tiles from a physical design drawn by a circuit designer), and the output consists of an image corresponding to silicon manufactured for a particular set of process conditions.
[0055] In another embodiment, multiple sets of process conditions are represented via multiple copies of a single-output network as shown in FIG. 10, with one network for each unique set of process conditions. Each of these single-output networks 1001, 1002 - 1010 is trained in parallel. After training, each of these networks can be used to infer the output for a given CAD data input image 1000 for a unique set of process conditions 1021, 1022 - 1030, i.e., for a particular process corner.
[0056] An example of the inferred output is shown in FIG. 11. A recombined tile is shown representing the shape of the manufactured D-type flip-flop (DFF) design image 1101 under three different unique process conditions. At first glance they appear similar, but upon closer inspection it is clear that the three images are different, for example, with different amounts of corner rounding being evident in each. The shape of image 1102 is closest to the straight-line CAD shape drawn from image 1101. The shape of image 1104 is perhaps the furthest away, with a greater degree of corner rounding and narrowing of the shape. Image 1103 is somewhere between these two extremes. In this example, for simplicity, only three examples are shown as representative of semiconductor manufacturing process conditions. More comprehensive sets include dozens, representing different extremes of dose variation in mask manufacturing, and different extremes of both dose variation and depth of focus variation in semiconductor manufacturing.
[0057] In another embodiment, FIG. 12 shows that after inferring the output manufacturing images of each process corner 1211, 1212 - 1220 using one output network 1201, 1202 - 1210, the images for each corner are combined and aggregated using post - processing 1230 to generate an average image 1233 representing a typical set of manufacturing conditions, a maximum image 1231 representing the most extreme result where the most material is deposited on silicon, and a minimum image 1232 representing the most extreme result where the least amount of material is deposited on silicon.
[0058] The details of the output images generated by combining the image tiles for each corner are shown in FIG. 13. The maximum image 1301 is calculated by taking the maximum value for each pixel across all output images for each corner. The minimum image 1302 is calculated by taking the minimum value for each pixel across all output images for each corner. Comparing the minimum elliptical shapes 1311 and 1312 near the center of both images, it is clearly shown that the ellipse 1312 in the case of the minimum image 1302 is much smaller than the ellipse 1311 in the case of the maximum image 1301.
[0059] The average image 1303 is calculated by taking the sum for each pixel divided by the number of process corners, or the average for each pixel across all output images for each corner. The process variation band or PV - band image 1241 shown in FIG. 12 is calculated by a post - processing step by subtracting the minimum image from the maximum image. The PV - band image 1304 in FIG. 13 is shown in detail. The white pixels indicate locations where metal may or may not be deposited on silicon during manufacturing. That is, each white pixel represents an area of uncertainty due to process variation. The more white pixels there are, the more susceptible the design is to manufacturing process variation.
[0060] Image thresholding compares each pixel value with a predetermined threshold value (e.g., 0.5). Pixel values above the threshold are converted to white (1.0), while pixel values below the threshold are converted to black (0.0). In another embodiment, image thresholding is performed before calculating the maximum value, minimum value, or average value. This refers to, for example, in a metal manufacturing step, determining a single binary value (1 or 0 corresponding to white or black respectively) for each pixel, regardless of whether metal is present at each pixel position. In a further embodiment, the maximum value, minimum value, and average value for each pixel may be calculated first, and then image thresholding may be performed.
[0061] As shown in FIG. 12, it may be desirable to generate two additional images to calculate a metric that represents the sensitivity or tolerance of the design / process combination to process variations. These are false positives 1242, i.e., image pixel positions where material is deposited on silicon but not set in the original CAD data (i.e., unintended material) is manufactured, and false negatives 1243, i.e., output image pixel positions that were set as intended material in the original CAD data image but not deposited during manufacturing.
[0062] An example of a false negative occurs at a 90-degree corner of a drawn CAD polygon. Here, a sharp corner is drawn by the circuit designer, but during manufacturing, corner rounding and / or line retraction occurs, effectively shaving or shortening the corner of the deposited material. Examples of false positives are, for example, extra material generated at a 270° corner, or extra material generated via pinching. In one embodiment, the false positive image and the false negative image are shown in FIG. 14 and generated as post-processing steps. The false positive image 1401 is calculated by subtracting the original CAD data image from the maximum image. The false negative image 1402 is calculated by taking the product (logical AND) of the minimum image and the original CAD data image, thresholding the result, and then subtracting the thresholded result from the thresholded original CAD data image.
[0063] To reduce the post - processing burden, in one embodiment, a CNN architecture with an output consisting of multiple channels is shown in FIG. 15. In one embodiment, the first N channels may be reserved for each of the N process - corner conditions 1502. This is achieved by forming an output layer consisting of 1×1 convolution operations with a filter depth of N. Here, N is the number of process corners. The intention is to generate an image corresponding to the output produced for each of the individual process corners 1502 for a single trained multi - output network 1501.
[0064] Also, as shown in FIG. 15, the calculation of the maximum image 1504, minimum image 1505, average image 1506, PV - band image 1507, false - positive image 1508, and false - negative image 1509 is achieved as the post - processing step 1503 described in relation to FIG. 12, after a single trained CNN 1501 is used to generate the outputs 1502 for each of the multiple process conditions. It will be understood that any one or any combination of these set - output images is obtained via post - processing rather than being directly inferred by the network.
[0065] As described above, these images may be calculated via post - processing of the minimum and maximum images as well as the input CAD - data image. In one embodiment shown in FIG. 16, these images may be directly generated by a deep - neural network 1601 via the introduction of process corners 1602 and additional output - image channels such as minimum, maximum, and average values 1610. The PV - band, false - negative, and false - positive 1620 may be generated directly.
[0066] When maximum, minimum, and average images, etc. are directly generated by a trained network, the corner image 1602 for each process may not need to be learned / inferred by the network. In this case, the network is trained to directly output the aggregated images 1610 (maximum, minimum, average) and 1620 (PV band, false positive, false negative) without outputting the corner image 1602 for each process. When there are a large number of process corners to be considered, it is desirable (to reduce computing and / or GPU resources such as memory) to not output the image 1602 for each process corner and instead output only the remaining aggregated images. In this case, the filter for each corner is removed from the CNN output layer, and their corresponding images are removed during training. In one embodiment, the user can select, prior to training, whether to have the network output all, some, or none of the per-corner images. Accordingly, the network architecture and parameters of the neural network are adjusted.
[0067] The shape manufactured on silicon depends greatly on the location or vicinity of the input shape, but there are also long-range effects such as local pattern density. Simply put, the shape manufactured for the CDA data image tile of an image contains some differences when the tile is from a densely populated part of a larger design compared to when it is from a relatively isolated part of the design. To enable the CNN model to learn these density effects, embodiments expand the input to include multiple channels. In such embodiments, the local pattern density is encoded as a single number from 0.0 (fully separated) to 1.0 (completely surrounded by metal), and a grayscale image is generated such that all pixels are set to the same number. The grayscale image dimensions are set to be the same as the CAD data tile dimensions and are represented as additional channels within the input image such that a color image is represented as the R, G, B channels for normal image processing. Next, the CNN architecture is expanded to handle a two-channel input instead of a single-channel input. During the training process, the network parameters learn the relationship between the grayscale color levels and the effects corresponding to the output manufactured image.
[0068] In one embodiment, the input image may be composed of two channels, where each channel itself can be represented as a grayscale image, one for the CAD data and one for a lower-resolution image of a larger region from which the patch tile representing the local density information was obtained. In some embodiments, the output image may include multiple channels along with different grayscale images for each channel (e.g., channels representing the maximum image, minimum image, average image, PV band image, false positive image, or false negative image). In additional embodiments, the output image can also include additional channels, for example, one channel for each process corner. Here, the image for each corner represents the manufactured shape expected for a particular combination of process variables specific to that process corner.
[0069] In an embodiment, a method for calculating patterns manufactured on a substrate includes inputting a physical design 900, inputting a set of parameters for a neural network to calculate patterns manufactured on the substrate, generating a plurality of possible neighborhoods for the physical design (step 704 in FIG. 7), and calculating a plurality of patterns manufactured on the substrate for the physical design in each of the plurality of possible neighborhoods (step 722). The method also includes training a neural network using the plurality of calculated patterns (e.g., in the loop from step 725 to physical design 702), where the training is executed using a computing hardware processor, and adjusting the set of parameters to reduce manufacturing variations of the plurality of calculated patterns manufactured on the substrate (e.g., at step 725).
[0070] In some embodiments, the neural network includes using post-processing to aggregate variations within a variation band. The neural network includes multi-output channels to aggregate variations within a variation band. In some embodiments, the method includes calculating false negatives and false positives for patterns on the substrate.
[0071] In some embodiments, the neural network includes a single fully convolutional network (FCN) architecture (e.g., FIG. 9). The FCN includes a first encoding layer, a second encoding layer, a last encoding layer, and a bottleneck layer, and the last encoding layer and the bottleneck layer each employ dropout regularization. In some embodiments, the FCN includes a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer, and each decoding layer employs concatenation with additional information from the fourth encoding layer, the third encoding layer, the second encoding layer, and the first encoding layer, respectively.
[0072] In some embodiments, the physical design and the plurality of patterns to be calculated are each divided into tiles. For example, each of the tiles comprises a 256×256 pixel tile with a 128×128 pixel inner core and a 64 pixel wide halo. In some embodiments, calculating the patterns fabricated on a substrate includes charged particle beam simulation. In some embodiments, calculating the patterns fabricated on a substrate includes lithography simulation 720. In some embodiments, the method includes inputting a local pattern density of the physical design 702. Design variability metric Using various aggregated images over the variations, a scalar design variability metric can be generated.
[0073] Let TP (true positive) be the number of white pixels in the CAD design representing the locations where metal is ideally deposited in silicon manufacturing, and TN (true negative) be the number of black pixels in the same image. Let VB (variation band) be the number of white pixels in the variation band plot that serves as an upper bound on the uncertainty associated with metal deposits due to process variations.
[0074] Let FN (false negative) be the number of white pixels in the false negative design image representing the amount of metal that was intended to be deposited ideally during silicon manufacturing but was found not to be deposited due to, for example, corner rounding, line end pullback, etc. FN is a metric that functions as an upper bound on the amount of missing metal found after manufacturing.
[0075] Let FP (false positive) be the number of white pixels in the false positive design image representing how much metal was inadvertently deposited at locations where metal was not intended to be deposited ideally during silicon manufacturing. FP is a metric that functions as an upper bound on the measurement of unwanted material deposited during manufacturing.
[0076] The Matthews correlation coefficient (MCC) is defined as follows and is often used as a single metric by which a classification algorithm is measured when using the TP, FP, TN, FN measurements from the confusion matrix.
[0077]
Number
[0078] In this semiconductor manufacturing scenario, since the four variables TP, FP, TN, FN for the semiconductor manufacturing applications of the present disclosure have different meanings, the formula for MCC has a different meaning from conventional use. In this case, MCC is a function of the amount of intended metal (TP), unintended metal (TN), the upper limit of metal inadvertently removed at a location that was initially intended (FN), and the upper limit of metal inadvertently deposited at a location that was not initially intended (FP). The MCC score calculated through this formula functions as a single scalar value measure of how process variability tends to produce an image on silicon that is different from the intended image (for each initially drawn CAD data). A large value of MCC (close to 1.0) indicates a good correlation between the CAD data image and the silicon image being manufactured, i.e., high resistance to process variations. A very small value of MCC (close to 0) indicates a very small correlation between the intended image and the manufactured image. The MCC value can be improved by changing the manufacturing process to reduce difficult and costly variations, either by changing the manufacturing process by design or a combination of both techniques. Integrated device manufacturers (IDMs) are in a position to modify the process for important designs where a high level of reproducibility (yield) is required.
[0079] Precision is defined as TP divided by the sum of TP and FP, and recall is defined as TP divided by the sum of TP and FN. Precision and recall are used to calculate another metric, F1.
[0080]
Number
[0081] The formula for calculating the F1 score from these images is shown below.
[0082]
Equation
[0083] The MCC value takes into account all four quantities (TP, TN, FP, FN) and is considered a more useful metric than the F1 formula that does not include the TN quantity. Therefore, it should be noted that it can be asymmetric for imbalanced class problems (where TP and TN are significantly different). This is one of the reasons why MCC has been preferred for use in classification algorithms.
[0084] Let TP2 be the number of white pixels in the average image after the grayscale average image is thresholded. This represents the number of pixels that a designer can realistically expect for the metal deposited by a realistic process. The designer is aware that the process is not ideal and that effects such as rounding of corners occur during manufacturing. However, for the convenience of drawing, the designer continues to draw images composed of straight lines with right-angled corners during the circuit design process. By setting the number of (ideal) white pixels drawn first as TP and the number of white pixels that the designer can more realistically expect as TP2, two more quantities are defined.
[0085] Let VBI = VB / TP2. This is the ratio of the number of variable band white pixels to the average (realistically expected) image white pixels. Since the numerator VB of the fraction still contains uncertain terms and includes the amount of pixels with uncertain manufacturing output, it provides a more realistic assessment of how susceptible the design is to a given process variation. Next, VB is normalized by the denominator TP2, which is the amount of pixels where the metal can realistically be expected on average across process variations.
[0086] The second quantity VBI′ = VB / TP functions as the ratio of manufacturing uncertainty to the number of white pixels drawn first (although unrealistic, it is the predicted result in an ideal manufacturing scenario).
[0087] Using these definitions as appropriate, different designs / cells or design candidates generated by a designer are processed by a trained neural network according to an embodiment, various collected manufacturing output images are generated, and the white pixels of those images are counted as described above. And the design is later scored using metrics regarding their tolerance (MCC) to process variations or their sensitivity (VBI) to process variations. Deep learning issues In deep convolutional neural networks, i.e., in deep learning, a computer model learns to directly classify or perform regression tasks from images, text, or sound. Deep learning models can achieve state-of-the-art accuracy and, in some cognitive applications, can even exceed human-level performance. The model is trained by using a large set of labeled data and a neural network architecture with many layers. Most deep learning methods use a network architecture. This is why deep learning is referred to as deep neural networks. The term "deep" usually refers to the number of hidden layers within a neural network. Conventional neural networks typically contain only 2 to 3 hidden layers, while deep networks can have as many as 150 layers.
[0088] A deep learning model is trained by using a large set of labeled data and a neural network architecture. One of the most common types of deep neural networks is the CNN architecture. The CNN architecture convolves learned features with input data, typically using 2D convolutional layers, making this architecture well-suited for processing 2D data such as images.
[0089] The CNN eliminates the need for manual feature extraction. That is, it removes the need to pre-identify the features used to classify or predict an image. The CNN functions by directly extracting features from the image. The relevant features are not pre-trained. The relevant features are learned while the network is trained on a sufficiently large set of images. Such automated feature extraction makes deep learning models very accurate for general computer vision tasks such as object classification and semiconductor manufacturing image-image conversion tasks such as those of the present invention.
[0090] There are several main reasons why deep learning has only recently become useful. Deep learning requires large amounts of labeled data. Deep learning requires significant computing power.
[0091] Deep learning is an iterative process. Deep learning requires large amounts of labeled data. For example, the development of self-driving cars requires millions of images and thousands of hours of video. In the case of the present disclosure, obtaining labeled data refers to gathering a large collection of thousands to millions of images representing the physical designs to be manufactured and refers to the image-based outputs of the various computational steps of FIG. 7, such as OPC / ILT, mask process simulation, substrate simulation, etc.
[0092] Some of the data can be collected by the actual manufacture of dedicated test chips, but considering the mask set production and manufacturing costs for today's high-density processes, such manufacturing-based data collection methods are very expensive. As an alternative, computational simulations may be used instead of manufacturing, but significant computational costs are involved for any of the steps in Figure 7. This option has, until recently, been prohibitively expensive in terms of the computational power required. In particular, the OPC / ILT calculations required to compute the mask shapes that result in manufacturable designs have required enormous computational power. ILT / OPC calculation tools combined with highly parallel and GPU-accelerated computing design platforms now finally make it possible to use computational software to determine mask shapes for full-reticle size IC designs within a time frame that no longer prohibits the generation of sufficient amounts of labeled data required for deep learning.
[0093] Deep learning requires significant computational power. High-performance GPUs have a parallel architecture that is efficient for deep learning. When combined with cluster or cloud computing, this may enable the development team to reduce the training time for deep learning networks from weeks to less than a few hours, depending on the problems and complexities of the deep learning neural network architecture. According to dedicated architectures for high-performance computing by computers as described in Figures 17 and 18, the application of deep learning to problems that were previously difficult to solve can be facilitated.
[0094] The sequential process for deep learning layer learning includes loading / preprocessing data and fitting the model for prediction. This sequential approach is certainly reasonable and useful for understanding, but in reality, deep learning is not so linear. Instead, the actual deep learning for generating the learned image as shown in Figure 9 has a distinct periodic nature that requires a certain number of iterations, adjustments, and improvements. The cycle starts the iteration from the input physical design 900, calculates the mask image, and then compares that image with the output image 920 generated by deep learning. When each process is completed, the impact on how the model performs is measured and adjustments are made to improve performance in the next cycle.
[0095] Deep learning practitioners must handle the following iterative process. Model level: Fitting of model parameters Micro level: Adjustment of hyperparameters Macro level: Problem solving Meta level: Improvement of training / test data Model level: Fitting parameters The first level at which iteration plays a major role is the model level. Any model, whether it is a regression model, a decision tree, or a neural network, is defined by a large number (sometimes millions) of model parameters. For example, a regression model is defined by its feature coefficients, a decision tree is defined by its branching locations, and a neural network is defined by the weights connecting its layers. In deep learning, the model parameters are learned via iterative methods such as gradient descent, which is an iterative method for finding the minimum of a function. In deep learning, that function is typically a loss (or cost) function. "Loss" is a metric that quantifies the cost of an incorrect prediction, such as mean squared error, mean absolute error, cross entropy, etc. Gradient descent calculates the loss achieved by a model with a given set of parameters and then adjusts the parameters to reduce the loss. This process is repeated until the loss can no longer be substantially reduced.
[0096] Micro level: Hyperparameter tuning Hyperparameters are "high-level" parameters that cannot be directly learned from data using gradient descent or other optimization algorithms. For example, dropout is a regularization method that is similar to training multiple neural networks in parallel with different architectures. During training, some layer outputs are randomly ignored or "dropped out". This has the effect of making the layer appear and be treated as if it had a different number of nodes and connectivity to the previous layer. In fact, each update to a layer during training is performed using a different "view" of the constructed layer. Conceptually, dropout decomposes the situation where network layers adapt to each other to correct errors from the previous layer, and then makes the model more robust. Hyperparameters describe the structural information about the model that must be determined before fitting the model parameters, such as whether dropout or other forms of regularization should be included in the model, whether batch normalization should be performed before calculating the layer output, the number of epochs (external iterations used in the model parameter fitting process), the specific optimization algorithm used during model parameter fitting, and whether cross-validation should be used to validate the model during fitting. Determining appropriate values for each of these various parameters / decisions is an iterative process and requires many iterations of the model parameter fitting process described above.
[0097] Macro level: Problem solving There is no model architecture / family that works best for all problems. Different model families work better than other model families depending on various factors such as the type of data, the problem domain, the sparsity of the data, and even the amount of data collected.
[0098] Thus, one way to improve candidate solutions to a given problem is to try several different model families or model architectures, e.g., the shape of the network itself, the number of filter layers, and the size of the convolutional kernels used in convolutional layers, as well as whether to use skip layer techniques. Determining appropriate values for each of these various parameters / decisions is an iterative process and requires many iterations of the model parameter fitting and hyperparameter tuning processes described above.
[0099] Another way to improve deep - learned solutions is by combining multiple deep - learned models into an ensemble. This is a direct extension from the iterative process required to fit multiple models. A common form of forming an ensemble is to average the predictions from multiple trained models. There are more sophisticated ways to combine multiple models, but the iterations required to fit multiple models are the same. Determining the appropriate combination / ensemble for each of the various deep - learned models is an iterative process.
[0100] Meta - level: Improvement of training / test data In terms of machine learning, better data generally confers more value than better algorithms. However, better data is not the same as more data. Better data means having less missing data and lower measurement error (e.g., more accurate data). Also, the data needs to be representative and avoid problems known to those skilled in the art such as data imbalance. The overall process of obtaining a sufficient set of such clean and accurately labeled data is often an iterative process in itself. In the case of the present disclosure, the iterative process includes performing more simulations of various steps of FIG. 7 under different conditions to ensure that a sufficiently wide sampling of the possible neighborhood is performed. The learning process includes determining the various types of combinations of shapes that will exist in the cell design, such as the physical design input to FIG. 7, in order to make the deep learning model function well for shape combinations that have not been previously manifested.
[0101] The various iterations of the above process to actually train the deep learning model to a sufficient level of accuracy require a vast amount of computing power and a vast amount of data related to the semiconductor manufacturing process. Prior to recent developments in the semiconductor manufacturing computing industry and the recent ability to run computational software simulators on dedicated GPU-accelerated hardware, such deeply nested iterative processes for deeply learning the patterns fabricated on substrates such as silicon wafers were not easy to handle and there was even no motivation to attempt such an approach.
[0102] Neural networks such as CNNs need to be trained using training data. Typically, the higher the capacity of the network (the amount of the ability to generalize to unseen data), and in addition, the more parameters the network has, the larger the number of data samples required to train the network without overfitting. The networks considered here typically contain hundreds of thousands of trainable parameters and require extremely large amounts of training data samples.
[0103] The supervised training paradigm supplies the CNN with a large number (hundreds of thousands to millions) of input-output pairs. In the case of the present invention, the input items consist of patches of CAD data, i.e., a set of CAD data representing the physical design drawn by a circuit designer. The set of CAD data is rasterized and divided into patches or tiles. For each layer to be manufactured, the input image is a single-channel image having a specific width and height. The input image may be a binary image (each pixel is either black or white) or a grayscale image, and each pixel takes a continuous value from 0.0 (black) to 1.0 (white). The output item of each pair consists of the corresponding predicted image after manufacturing at a certain specific process corner. In one embodiment, the output item is a single-channel output image and includes either binary-valued pixels or grayscale (continuous-valued) pixels. The intention is to train the network so that it can infer or predict the output image when only the input image is given. Also, the intention is to train the network so that it can infer or predict the output image for a design input image that has never been seen before.
[0104] The creation of a sufficient amount of input image data (CAD data) can be done relatively quickly, but it should be noted that the creation of output image data, which is expected to correspond to the manufacturing results, is a very long-term problem for actual semiconductor manufacturing processes involving state-of-the-art process nodes such as a huge amount of computing hardware resources. The CAD data needs to be simulated using various computationally intensive algorithms including, but not limited to, wafer manufacturing simulations using OPC and ILT as well as calibrated mask models. Such simulation tools and models are used together with dedicated GPU-based hardware in the form of high-performance computing clusters (HPC) or computing data platforms (CDP) to accelerate the simulation. Only after a series of such tools have been executed can an output image be obtained. Furthermore, when process variations are taken into account, the creation of corresponding corner images for each process involves a significant additional cost. Only recently have semiconductor process manufacturing simulation tools, especially ILT, become fast enough to generate the required amount of data within a realistic time frame.
[0105] FIG. 17 shows an example of a computing hardware device 1700 used to execute the calculations described in the present disclosure. The computing hardware device 1700 includes a central processing unit (CPU) 1702 to which a main memory 1704 is attached. The CPU includes, for example, eight processing cores, and as a result, improves the performance of any part of the computer software that is multithreaded. The size of the main memory 1704 is, for example, 64 gigabytes. The CPU 1702 is connected to a PCIe (Peripheral Component Interconnect Express) bus 1720. A graphics processing unit (GPU) 1714 is also connected to the PCIe bus. In the computing hardware device 1700, the GPU 1714 may or may not be connected to a graphics output device such as a video monitor. When not connected to a graphics output device, the GPU 1714 is simply used as a high-speed parallel computing engine. The computing software can obtain significantly higher performance by using the GPU for part of the calculations as compared to using the CPU 1702 for all calculations. The CPU 1702 communicates with the GPU 1714 via the PCIe bus 1720. In other embodiments (not shown), the GPU 1714 may be integrated into the CPU 1702 instead of being connected to the PCIe bus 1720. Also, a disk controller 1708 may be attached to the PCIe bus, for example, with two disks 1710 connected to the disk controller 1708. Finally, a local area network (LAN) controller 1712 that provides gigabit Ethernet (GbE) connectivity to other computers may be connected to the PCIe bus. In some embodiments, the computer software and / or design data are stored on the disk 1710. In other embodiments, either or both of the computer program and design data may be accessed from another computer or file serving hardware via GbE Ethernet.
[0106] FIG. 18 is another embodiment of a system that executes the calculations of this embodiment. The system 1800, which may also be referred to as a CDP, includes a master node 1810, an optional viewing node 1820, an optional network file system 1830, and a GPU-enabled computing node 1840. The viewing node 1820 may be absent, may have only one node, or may have any other number of nodes. The GPU-enabled computing node 1840 can include one or more GPU-enabled nodes that form a cluster. Each GPU-enabled computing node 1840 includes, for example, a GPU, a CPU, a pair of GPU and CPU, multiple GPUs for a CPU, or some other combination of GPU and CPU. The GPU and / or CPU may be on a single chip, such as a GPU chip having a CPU accelerated by a GPU on the chip, or a CPU chip having a GPU that accelerates the CPU. The GPU may be replaced by another coprocessor.
[0107] The master node 1810 and the viewing node 1820 are connected to the network file system 1830 and the GPU-enabled computing node 1840 via switches and high-speed networks such as networks 1850, 1852, and 1854. In one embodiment, network 1850 may be a 56 Gbps network, 1852 may be a 1 Gbps network, and 1854 may be a management network. In various embodiments, the number of networks may be less than or more than this, and there may be various combinations of different types of networks such as high-speed and low-speed. The master node 1810 controls the CDP 1800. An external system can connect to the master node 1810 from the external network 1860. In some embodiments, a job may be started from an external system. Data for the job is loaded onto the network file system 1830 before starting the job, and the program is used to dispatch and monitor tasks on the GPU-enabled computing node 1840. The progress of the job can be viewed by a user on the master node 1810 via a graphical interface such as the viewing node 1820. The task is executed on the CPU using a script that launches an appropriate executable file on the CPU. The executable file connects to the GPU and disconnects from the GPU after executing various computing tasks. Also, the master node 1810 is used to operate as if the failed GPU-enabled computing node 1840 did not exist after disabling it.
[0108] Although this specification has been described in detail with respect to particular embodiments, those of ordinary skill in the art will understand that, upon understanding the foregoing, alternative, modified, and equivalent forms of these embodiments can be readily conceived. These and other changes and modifications to the method can be implemented by those of ordinary skill in the art without departing from the scope of the subject matter described in more detail by the appended claims. Furthermore, those of ordinary skill in the art will understand that the foregoing description is merely exemplary and not intended to be limiting. Steps may be added, removed, or modified to the steps described herein without departing from the scope of the present invention. Generally, the presented flowcharts are only intended to show one possible sequence of the basic operations for achieving the functions, and many variations are possible. Accordingly, the subject matter is intended to embrace changes and modifications that fall within the scope of the appended claims and their equivalents.
Claims
1. A method for adjusting a physical design manufactured on a substrate, comprising: receiving a physical design including a specific design pattern; generating a plurality of prediction patterns representing different predictions for how the specific design pattern is manufactured on the substrate for the physical design; analyzing the plurality of prediction patterns to determine that the physical design needs to be adjusted; and adjusting the physical design based on the determination. A method comprising the above steps.
2. The method according to claim 1, wherein generating the plurality of prediction patterns comprises generating the plurality of prediction patterns using a set of one or more neural networks.
3. The method according to claim 2, wherein the set of neural networks includes a convolutional neural network.
4. The method according to claim 1, wherein the set of neural networks includes a set of two or more neural networks for generating two or more prediction patterns.
5. The method according to claim 1, wherein the specific design pattern is a first pattern, and the plurality of generated prediction patterns includes two or more prediction patterns representing two or more predictions for how the first pattern in the physical design is manufactured on the substrate based on a set of two or more patterns that are possible patterns for the vicinity of the first pattern in the physical design.
6. The method according to claim 1, wherein the plurality of generated prediction patterns comprises two or more prediction patterns, and the two or more prediction patterns represent two or more predictions for how the specific design pattern is manufactured on the substrate based on two or more variations in manufacturing process parameters.
7. The method according to claim 6, wherein the manufacturing process parameters include the depth of focus of a lens used during lithography for manufacturing the physical design on the substrate.
8. The method according to claim 6, wherein the manufacturing process parameters include the exposure dose used during lithography for manufacturing the physical design on the substrate.
9. The method according to claim 6, wherein The method wherein the two or more prediction patterns include a maximum variation profile, a nominal variation profile, and a minimum variation profile, and they respectively correspond to the maximum variation, the nominal variation, and the minimum variation in the manufacturing process parameters.
10. In the method according to claim 1, The generation and analysis are performed during a mask design process that is performed after a physical design process.
11. In the method according to claim 10, The adjustment is performed by returning to the physical design process to correct the physical design.
12. A machine-readable medium storing a program executed by at least one processing unit that executes the method according to any one of claims 1 to 11.
13. An electronic device, A set of processing units, and A machine-readable medium storing a program executed by at least one processing unit that executes the method according to any one of claims 1 to 11 An electronic device comprising the same.
14. A system, A system comprising means for executing the method according to any one of claims 1 to 11.
15. A computer program product, A computer program product that includes instructions for causing a computer to execute the method according to any one of claims 1 to 11 when executed by the computer.
Citation Information
Patent Citations
Integrated circuit layout design method using process variation band
JP2007536581A
Photomask pattern verifying method, photomask pattern verifying device, method of manufacturing semiconductor integrated circuit, photomask pattern verification control program and readable storage medium
JP2009014790A
Pattern verifying method and method for manufacturing semiconductor device
JP2010177374A
Method for creating mask layout, apparatus for creating mask layout, method for manufacturing mask for lithography, method for manufacturing semiconductor device, and program operable in computer
JP2011175111A
Pattern determination method, pattern determination device, and program
JP2013120290A