Method and apparatus for layout pattern selection

By generating features from the pattern set in the lithography projection device and grouping them, and selecting representative patterns as training patterns, the problem of non-representative patterns in the prior art is solved, and a more efficient lithography process is achieved.

CN120181019APending Publication Date: 2025-06-20ASML NETHERLANDS BV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510165403.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-01-29
Filing Date
2020-01-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively select representative patterns in lithography projection equipment for training machine learning models, resulting in lithography hot spots and process window limitations, affecting lithography performance.

Method used

By generating multiple features from the pattern set, the patterns are grouped based on the similarity of these features, and selecting representative patterns from each group as training patterns, provided to a computational lithography application.

Benefits of technology

An automated pattern selection process is realized, which improves the representativeness of training patterns, reduces lithography hot spots and process window restrictions, and improves lithography performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181019A_ABST
    Figure CN120181019A_ABST
Patent Text Reader

Abstract

A method for determining a training pattern in a layout patterning process is described herein. The method includes: generating a plurality of features from a pattern in a set of patterns; grouping the patterns in the set of patterns into separate groups based on similarities of a plurality of generated features; and selecting a representative pattern from the individual groups to determine the training pattern. In some embodiments, the method is a method for training a machine learning model in a layout patterning process. For example, the method may include providing representative patterns from the individual sets to the machine learning model to train the machine learning model to predict a continuous transmission mask (CTM) map for optical proximity effect correction (OPC) in the layout patterning process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the invention name "Method and apparatus for layout pattern selection" (International Application No. PCT / EP2020 / 050494, International Filing Date: 2020-01-10) and application number 202080011671.1, which entered the Chinese national phase on July 29, 2021. Technical Field

[0002] The description herein generally relates to mask manufacturing and patterning processes. More particularly, the description relates to methods and apparatuses for layout pattern selection to train machine learning models. Background Art

[0003] A lithographic projection apparatus can be used, for example, in the manufacture of integrated circuits (ICs). In such a case, a patterning device (e.g., a mask) can comprise or provide a pattern corresponding to a single layer of the IC ("design layout"), and this pattern can be transferred onto a target portion (e.g., including one or more dies) of a substrate (e.g., a silicon wafer) that has been coated with a radiation-sensitive material ("resist") layer by radiation passing through the pattern on the patterning device. Typically, a single substrate comprises a plurality of adjacent target portions onto which the pattern is transferred successively, one target portion at a time, by the lithographic projection apparatus. In one type of lithographic projection apparatus, the pattern over the entire patterning device is transferred onto one target portion in a single operation; such an apparatus is commonly referred to as a stepper. In an alternative apparatus (commonly referred to as a step-and-scan apparatus), the projection beam scans over the patterning device in a given reference direction ("scan" direction) while the substrate is synchronously moved in a direction parallel or anti-parallel to the reference direction. Different portions of the pattern on the patterning device are transferred gradually onto one target portion. Since typically the lithographic projection apparatus will have a reduction ratio M (e.g., 4), the rate F at which the substrate is moved will be 1 / M times the rate at which the projection beam scans the patterning device. More information about lithographic apparatuses as described herein can be gathered, for example, from US 6,046,792, which is incorporated herein by reference.

[0004] Before transferring the pattern from the pattern forming device to the substrate, the substrate may undergo various processes such as underlayer coating, resist coating, and soft baking. After exposure, the substrate may undergo other processes ("post-exposure processes") such as post-exposure bake (PEB), development, hard bake, and measurement / inspection of the transferred pattern. This series of processes serves as the basis for fabricating a single layer of a device (e.g., an IC). Subsequently, the substrate may undergo various processes such as etching, ion implantation (doping), metallization, oxidation, chemical mechanical polishing, etc., all of which are aimed at finally completing a single layer of the device. If the device requires multiple layers, the entire process or its variations are repeated for each layer. Eventually, the devices will be provided in each target portion on the substrate. Thereafter, the devices are separated from each other by techniques such as dicing or cutting so that individual devices can be mounted on carriers, connected to pins, and so on.

[0005] Fabricating a device (such as a semiconductor device) typically involves using multiple manufacturing processes to process a substrate (e.g., a semiconductor wafer) to form various features and multiple layers of the device. These layers and features are typically fabricated and processed using, for example, deposition, lithography, etching, chemical mechanical polishing, and ion implantation. Multiple devices can be fabricated on multiple die on the substrate and then separated into individual devices. This device manufacturing process can be considered a patterning process. The patterning process involves a pattern formation step, such as optical and / or nanoimprint lithography using a pattern forming device in a lithography apparatus, to transfer the pattern on the pattern forming device to the substrate, and typically but optionally involves one or more associated pattern processing steps, such as resist development by a development apparatus, substrate baking using a baking tool, etching using a pattern with an etching apparatus, etc. Additionally, typically one or more metrology processes are involved in the patterning process.

[0006] As mentioned, lithography is a core step in fabricating devices (such as ICs), where the patterns formed on the substrate define the functional elements of the device, such as microprocessors, memory chips, etc. Similar lithography techniques are also used to form flat panel displays, microelectromechanical systems (MEMS), and other devices.

[0007] As the semiconductor manufacturing process continues to progress, for decades, the dimensions of the functional elements have been continuously decreasing while the number of functional elements (such as transistors) per device has been steadily increasing, following a trend commonly known as "Moore's Law". In the current state of the art, lithographic projection equipment is used to fabricate multiple layers of a device, and the lithographic projection equipment projects a design layout onto the substrate using irradiation from a deep ultraviolet light source, thereby forming individual functional elements with dimensions far less than 100 nm (i.e., less than half of the wavelength of the radiation from the irradiation source (e.g., a 193 nm irradiation source)).

[0008] The process in which features having dimensions smaller than the classical resolution limit of a lithographic projection apparatus are printed is generally referred to as low-k1 lithography. It is based on the resolution formula CD = k1×λ / NA, where λ is the wavelength of the radiation used (currently most commonly 248 nm or 193 nm), NA is the numerical aperture of the projection optics in the lithographic projection apparatus, CD is the "critical dimension" (usually the smallest feature size printed), and k1 is an empirical resolution factor. Generally, the smaller k1 is, the more difficult it becomes to reproduce on the substrate a pattern similar in shape and size to that planned by the designer to achieve a particular electrical functionality and performance. To overcome these difficulties, complex fine-tuning steps are applied to the lithographic projection apparatus, the design layout, or the patterning device. These steps include, for example, but are not limited to: optimization of NA and optical coherence settings, custom illumination schemes, use of phase-shifting patterning devices, optical proximity correction (OPC, sometimes also referred to as "optical and process correction") in the design layout, or other methods generally defined as "resolution enhancement techniques" (RET). As used herein, the term "projection optics" should be broadly interpreted to cover various types of optical systems, including, for example, refractive optical devices, reflective optical devices, apertures, and catadioptric optical devices. The term "projection optics" may also include components that operate according to any of these design types for jointly or individually guiding, shaping, or controlling a projection radiation beam. The term "projection optics" may include any optical component in a lithographic projection apparatus, regardless of where the optical component is located in the optical path of the lithographic projection apparatus. The projection optics may include optical components for shaping, adjusting, and / or projecting the radiation from the source before it passes through the patterning device, or for shaping, adjusting, and / or projecting the radiation after it passes through the patterning device. The projection optics generally do not include the source and the patterning device. SUMMARY OF THE INVENTION

[0009] According to an embodiment, a method for training a machine learning model for a wafer patterning process is provided. The method includes: generating a plurality of features for each pattern in a pattern set; grouping the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features; and providing representative patterns from the plurality of separate groups to a computational lithography application of the wafer patterning process. The application may include source mask optimization (SMO), optical proximity effect correction (OPC), lithographic manufacturability checking (LMC), etc.

[0010] In an embodiment, the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information already included in the pattern set. In an embodiment, SMO includes source and mask co-optimization for the full-chip layout of a wafer during the wafer patterning process. In an embodiment, the OPC includes full-chip OPC for the wafer during the layout patterning process. In an embodiment, LMC includes lithographic manufacturability and lithographic performance inspection for the full-chip layout of a wafer during the wafer patterning process. In an embodiment, the plurality of generated features include geometric features and lithography-aware features. In an embodiment, grouping the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features includes using a machine learning clustering method to cluster unique patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features.

[0011] According to another embodiment, a method for determining training patterns for a wafer patterning process is provided. The method includes: generating a plurality of features from patterns in a pattern set; grouping the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features; and selecting representative patterns from the multiple separate groups to determine the training patterns.

[0012] In an embodiment, the plurality of generated features include geometric features and lithography-aware features. The geometric features include one or more of the following: target mask image, frequency map, pattern density map, or pattern occurrence rate of unique patterns in the pattern set. The lithography-aware features include one or more of the following: sub-resolution assist feature guidance map (SGM), diffraction order, or diffraction pattern of unique patterns in the pattern set. In an embodiment, the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information already included in the pattern set.

[0013] In an embodiment, grouping the patterns in the pattern set into multiple groups based on the plurality of generated features is performed using unsupervised machine learning. In an embodiment, grouping the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features includes clustering unique patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features. The clustering includes a consecutive series of clustering steps performed using different features among the plurality of generated features for different clustering steps, and the consecutive series of clustering steps form subgroups of the patterns in the pattern set such that the representative patterns are selected from the subgroups to determine the training patterns. In an embodiment, the clustering includes a machine learning clustering method (e.g., k-means clustering).

[0014] In an embodiment, the successive series of clustering steps includes a cross-validation step performed for a given feature of a given step. The cross-validation step includes adjusting which patterns are included in a given subgroup.

[0015] In an embodiment, selecting representative patterns from the separate groups to determine the training patterns includes selecting a target number of representative patterns. The target number of representative patterns is determined based on a stopping criterion. The stopping criterion is configured to contribute to a change in the training patterns. In an embodiment, the method further includes determining a change amount of the training patterns. In an embodiment, the stopping criterion is further configured to ensure that the change amount of the training patterns breaks through a change amount threshold.

[0016] In an embodiment, the target number of representative patterns is randomly selected from the plurality of separate groups. In an embodiment, the target number of representative patterns is re-randomly selected in response to the change amount of the training patterns not breaking through the change amount threshold.

[0017] In an embodiment, selecting representative patterns from the plurality of separate groups to determine the training patterns includes selecting the most central pattern from each separate group. The most central pattern is closest to the centroid of a designated feature space for the separate group relative to other patterns in the plurality of separate groups. In an embodiment, the designated feature space is a target mask image feature space, a frequency mapping feature space, a pattern density mapping feature space, a pattern occurrence feature space, an SGM feature space, a diffraction order feature space, or a diffraction pattern feature space.

[0018] In an embodiment, the method further includes providing the training patterns to a deep convolutional neural network to train the deep convolutional neural network. In an embodiment, the method further includes performing optical proximity effect correction using the trained deep convolutional neural network as part of a wafer patterning process.

[0019] According to another embodiment, there is provided a computer program product including a non-transitory computer-readable medium having instructions recorded thereon, the instructions when executed by a computer implementing the method described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings incorporated in and constituting a part of this specification illustrate one or more embodiments and, together with the description, explain these embodiments. Embodiments of the present invention will now be described by way of example only with reference to the accompanying schematic drawings, in which corresponding reference numerals indicate corresponding parts, and in which:

[0021] Figure 1A block diagram showing the various subsystems of a lithography system according to an embodiment.

[0022] Figure 2 A flowchart of a method for determining a patterning device pattern (or mask pattern) from an image (e.g., a continuous transmission mask image, a binary mask image, a curved mask image, etc.) according to an embodiment, the image corresponding to a target pattern to be printed on a substrate via a patterning process involving a lithography process.

[0023] Figure 3 An exemplary flowchart for simulating lithography in a lithographic projection apparatus according to an embodiment is illustrated.

[0024] Figure 4 An overview of the operations of the present method for determining a training pattern during a wafer patterning process according to an embodiment is illustrated.

[0025] Figure 5 Aspects of feature generation according to an embodiment are illustrated.

[0026] Figure 6 Clustering patterns in a pattern set into multiple separate groups based on the similarity of multiple features generated for unique patterns according to an embodiment is illustrated.

[0027] Figure 7 An example graph of the sum of squared errors as part of a clustering validation operation versus the k value according to an embodiment is illustrated.

[0028] Figure 8 k-fold cross-validation according to an embodiment is illustrated.

[0029] Figure 9 Pattern selection for forming a representative set of example patterns for use as training patterns according to an embodiment is illustrated.

[0030] Figure 10 A block diagram of an example computer system according to an embodiment.

[0031] Figure 11 A schematic diagram of a lithographic projection apparatus according to an embodiment.

[0032] Figure 12 A schematic diagram of another lithographic projection apparatus according to an embodiment.

[0033] Figure 13 According to an embodiment Figure 12 A more detailed view of the device in

[0034] Figure 14 According to an embodiment Figure 12 And Figure 13 A more detailed view of the source collector module SO of the device of Detailed implementation mode

[0035] Performing pattern selection from a full-chip graphic database system (GDS) file (e.g., a GDSII file) is a challenging task. Training patterns are used to train a deep convolutional neural network (DCNN) and / or other machine learning models, which are part of full-chip optical proximity correction (OPC) applications, source mask optimization (SMO) applications, lithography manufacturability inspection (LMC) applications, and / or for other purposes. If a user generates a set of training patterns for layout patterns that are not entirely representative or otherwise insufficient patterns, and provides such training data to a machine learning model (e.g., a DCNN) for training, an accurate CTM mapping will not be predicted by the machine learning model (e.g., a DCNN) for full-chip OPC applications. Inaccurate CTM mapping results in lithography hotspots and process window limitations, and / or difficulties during subsequent mask correction operations when attempting to meet lithography performance specifications. Currently, users manually select training patterns from the full-chip GDS. Manual selection requires the effective work of the user, and the layout pattern selection (e.g., the representative coverage of various patterns) depends on the user's experience and prior knowledge of the full-chip GDS design. Advantageously, the present method and device systematically analyze the full-chip GDS patterns and select representative patterns to construct a machine learning model training set, where the pattern coverage sufficiently represents the target layout patterns across the full-chip GDS file.

[0036] As a brief introduction, although the manufacture of ICs may be specifically referenced herein, it should be clearly understood that the description herein has many other possible applications. For example, the description herein can be used for manufacturing integrated optical systems, guiding and detecting patterns for magnetic domain memories, liquid crystal display panels, thin film magnetic heads, etc. In these alternative applications, those skilled in the art should understand that in the context of such alternative applications, any use of the terms "reticle", "wafer" or "die" herein should be recognized as being interchangeable with the more general terms "mask", "substrate" and "target portion" respectively. Additionally, it should be noted that the methods described herein can have many other possible applications in diverse fields such as language processing systems, autonomous vehicles, medical imaging and diagnosis, semantic segmentation, denoising, chip design, electronic design automation, etc. The method can be applied to any field where it is advantageous to quantify the uncertainty in machine learning model predictions.

[0037] In this document, the terms "radiation" and "beam" are used to cover all types of electromagnetic radiation, including ultraviolet radiation (e.g., having wavelengths of 365, 248, 193, 157, or 126 nm) and EUV (i.e., extreme ultraviolet radiation, e.g., having wavelengths in the range of about 5 to 100 nm).

[0038] A pattern forming apparatus may include or form one or more design layouts. The design layouts may be generated using a CAD (i.e., computer-aided design) process. Such a process is often referred to as EDA (i.e., electronic design automation). Most CAD programs follow a set of predefined design rules in order to generate a functional design layout / pattern forming apparatus. These rules are set based on processing and design limitations. For example, the design rules regulate the space tolerances between devices (such as gates, capacitors, etc.) or interconnect lines in order to ensure that the devices or lines do not interact with each other in an undesirable manner. One or more of the design rule limitations may be referred to as "critical dimension" (CD). The critical dimension of a device may be defined as the minimum width of a line or a hole, or the minimum space / gap between two lines or two holes. Thus, the CD determines the overall size and density of the designed device. One of the goals in device manufacturing is to faithfully reproduce the original design intent (via the pattern forming apparatus) on a substrate.

[0039] As used herein, the term "mask" or "pattern forming device" may be broadly interpreted to refer to a general pattern forming device that can be used to impart a patterned cross-section to an incident radiation beam, the patterned cross-section corresponding to the pattern to be produced in a target portion of a substrate. The term "light valve" may also be used in this context. In addition to classical masks (transmission or reflection, binary, phase-shifting, hybrid, etc.), examples of other such pattern forming devices include programmable mirror arrays. An example of such a device may be a matrix-addressable surface having a viscoelastic control layer and a reflective surface. The basic principle on which such a device is based is that (for example) the addressed regions of the reflective surface reflect the incident radiation as diffracted radiation, while the non-addressed regions reflect the incident radiation as non-diffracted radiation. Using a suitable filter, the non-diffracted radiation can be filtered out of the reflected beam, so that only the diffracted radiation remains thereafter; thus, the beam is patterned according to the addressing pattern of the matrix-addressable surface. The required matrix addressing can be carried out using suitable electronics. Examples of other such pattern forming devices also include programmable LCD arrays. An example of such a configuration is given in U.S. Patent No. 5,229,872, which is incorporated herein by reference.

[0040] Figure 1FIG. shows an exemplary lithographic projection apparatus 10A. The main components are: a radiation source 12A, which may be a deep ultraviolet (DUV) excimer laser source or other type of source, including an extreme ultraviolet (EUV) source (as discussed above, the lithographic projection apparatus itself need not have a radiation source); illumination optics, which for example define partial coherence (expressed as a standard deviation) and may include optics 14A, 16Aa and 16Ab for shaping the radiation from source 12A; a patterning device 18A; and projection optics 16Ac, which projects an image of the patterning device pattern onto a substrate plane 22A. An adjustable filter or aperture 20A at the pupil plane of the projection optics may restrict the range of beam angles incident on the substrate plane 22A, where the maximum possible angle defines the numerical aperture NA = n sin(Θ max ), where n is the refractive index of the medium between the substrate and the last element of the projection optics, and Θ max is the maximum angle of the beam emerging from the projection optics that can still be incident on the substrate plane 22A.

[0041] In a lithographic projection apparatus, a source that provides illumination (i.e., radiation) to a patterning device and a projection optical system directs and shapes the illumination via the patterning device onto a substrate. The projection optical system may include at least some of components 14A, 16Aa, 16Ab, and 16Ac. A spatial image (AI) is the radiation intensity distribution at the substrate level. A resist model may be used to calculate a resist image based on the spatial image, and an example of such a situation can be found in U.S. Patent Application Publication No. US 2009-0157630, the entire disclosure of which is hereby incorporated by reference herein. The resist model is related to the properties of the resist layer (e.g., the effects of chemical processes that occur during exposure, post-exposure bake (PEB), and development). The optical properties of the lithographic projection apparatus (e.g., the properties of the illumination, the patterning device, and the projection optical system) define the spatial image and may be defined in an optical model. Since the patterning device used in the lithographic projection apparatus can be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the remainder of the lithographic projection apparatus, which at least includes the source and the projection optical system. The techniques and models for transforming a design layout into various lithographic images (e.g., spatial images, resist images, etc.), the application of optical proximity correction (OPC) using those techniques and models, and the evaluation of performance (e.g., according to a process window) are described in U.S. Patent Application Publication Nos. US 2008-0301620, 2007-0050749, 2007-0031745, 2008-0309897, 2010-0162197, and 2010-0180251, the entire disclosure of each of which is hereby incorporated by reference in its entirety.

[0042] Optical proximity correction (OPC) enhances the integrated circuit patterning process by compensating for distortions that occur during processing. The distortions occur during processing because the features printed on the wafer are smaller than the wavelength of the light used in the patterning and printing processes. OPC verification identifies OPC errors or weaknesses in the post-OPC wafer design that may potentially lead to patterning defects on the wafer. For example, the ASML Tachyon Lithography Manufacturability Check (LMC) is an OPC verification product.

[0043] OPC addresses the fact that the final size and placement of an image of a design layout projected onto a substrate will be different from or only depend on the size and placement of the design layout on the patterning device. In the context of resolution enhancement techniques (RET) such as OPC, it is not necessary to use a physical patterning device, but the design layout can be used to represent the physical patterning device. For small feature sizes and high feature densities present in some design layouts, the position of a particular edge of a given feature will be affected to some extent by the presence or absence of other adjacent features. These proximity effects originate from minute amounts of radiation coupled from one feature to another and / or non-geometric optical effects such as diffraction and interference. Similarly, proximity effects can originate from diffusion and other chemical effects during post-exposure bake (PEB), resist development, and etching, which typically occur after lithography.

[0044] To increase the chance that the projected image of the design layout meets the requirements of a given target circuit design, complex numerical models, corrections, or pre-distortions of the design layout can be used to predict and compensate for proximity effects. The paper “Full-Chip Lithography Simulation and Design Analysis - how OPC Is Changing IC Design” (C. Spence, Proc. SPIE, Vol. 5751, pp. 1-14 (2005)) provides a review of current “model-based” optical proximity correction processes. In typical high-end designs, almost every feature of the design layout has some modification in order to achieve a high-fidelity projection of the image to the target design. These OPC modifications can include offsets or biases in edge position or line width and / or the application of “assist” features that are expected to aid in the projection of other features.

[0045] One of the simplest forms of OPC is selective biasing. Given a CD vs. pitch curve, by changing the CD at the patterning device level, all different pitches can be forced to have the same CD, at least under best focus and exposure conditions. Thus, if a feature is printed too small at the substrate level, the patterning device level feature will be biased to be slightly larger than the nominal feature, and vice versa. Since the pattern transfer process from the patterning device level to the substrate level is non-linear, the biasing amount is not simply the CD error measured under best focus and exposure conditions multiplied by the reduction ratio, but rather an appropriate bias can be determined using modeling and experimentation. Selective biasing is an incomplete solution to the problem of proximity effects, especially when it is applied only under nominal process conditions. Although in principle this bias can be applied to give a uniform CD vs. pitch curve under best focus and exposure conditions, once the exposure process changes from the nominal conditions, each biased pitch curve will respond differently, resulting in different process windows for different features. A process window is the range of values of two or more process parameters (e.g., focus and radiation dose in a lithography tool) within which a feature is produced adequately (e.g., the CD of the feature is within a certain range, such as ±10% or ±5%). Thus, the "optimal" bias for giving the same CD vs. pitch can even have a negative impact on the overall process window, shrinking rather than enlarging the focus and exposure ranges within which all target features can be printed on the substrate within the desired process tolerance.

[0046] Other, more complex OPC techniques have been developed for applications beyond the one-dimensional biasing examples above. Two-dimensional proximity effects are line-end shortening. Line ends have a tendency to "pull back" from their desired end positions depending on exposure and focus. In many cases, the degree of end shortening of long line ends can be several times greater than the corresponding line narrowing. This type of line-end pullback can lead to severe failure of the fabricated device if the line end cannot completely span the underlying layer it is intended to cover (such as a polysilicon gate layer over a source-drain region). Since this type of pattern is extremely sensitive to focus and exposure, simply biasing the line end to be longer than the designed length is not sufficient, because the line under best focus and exposure conditions or under under-exposure conditions will be too long, resulting in a short circuit when the extended line end touches an adjacent structure, or an unnecessarily large circuit size if more space is added between individual features in the circuit. Since one of the goals in integrated circuit design and manufacturing is to minimize the area required per chip while maximizing the number of functional elements, adding excessive spacing is not a desirable solution.

[0047] The two-dimensional OPC method can help solve the line-end pullback problem. Additional structures such as "hammerheads" or "wiring" (also known as "auxiliary features") can be added to the line ends to effectively anchor the line ends in place and provide reduced pullback throughout the process window. Even under the best focus and exposure conditions, these additional structures are not resolved, but they change the appearance of the main feature without being fully resolved individually. As used herein, "main feature" means a feature expected to be printed on a substrate under some or all conditions within the process window. The auxiliary features can take forms that are much more aggressive than simple hammerheads added to the line ends, to the extent that the pattern on the patterning device is no longer just the desired substrate pattern with its size increased according to the reduction ratio. Compared to merely reducing line-end pullback, auxiliary features such as wiring can be applied to more situations. Inner wiring or outer wiring can be applied to any edge, especially two-dimensional edges, to reduce corner rounding (i.e., chamfering) or edge squeezing. With sufficient selective biasing and auxiliary features of all sizes and polarities, the features on the patterning device bear less and less resemblance to the final pattern desired at the substrate level. Generally, the patterning device pattern becomes a pre-distorted version of the substrate-level pattern, where the distortion is expected to cancel or reverse the pattern deformation that will occur during the manufacturing process to produce a pattern on the substrate as close as possible to what the designer expects.

[0048] Instead of or in addition to those auxiliary features (e.g., wiring) that are connected to the main feature, another OPC technique also involves using completely independent and unresolved auxiliary features. The term "independent" here means that the edges of these auxiliary features are not connected to the edges of the main feature. These independent auxiliary features are not expected or required to be printed as features on the substrate, but are expected to modify the spatial image of the nearby main feature to enhance the printability and process tolerance of the main feature. These auxiliary features (often referred to as "scattering bars" or "SBAR") can include: sub-resolution auxiliary features (SRAF), which are features outside the edges of the main feature; and sub-resolution inverse features (SRIF), which are features dug out from inside the edges of the main feature. The presence of SBAR adds another level of complexity to the patterning device pattern. A simple example of the use of scattering bars is where a regular array of unresolved scattering bars is dragged on both sides of an isolated line feature, which has the effect of making the isolated line appear, from the perspective of the spatial image, as more of a single line within an array of dense lines, resulting in the process window being closer to the focus and exposure tolerances of the dense pattern in terms of focus and exposure tolerances. Compared to the case of a feature dragged in isolation at the patterning device level, the common process window between this decorated isolated feature and the dense pattern will have a greater common tolerance for focus and exposure variations.

[0049] Auxiliary features can be regarded as the differences between the features on the patterning device and the features in the design layout. The terms "main feature" and "auxiliary feature" do not imply that a particular feature on the patterning device must be labeled as a main feature or an auxiliary feature.

[0050] Another aspect of understanding the lithography process is understanding the interaction of the radiation with the patterning device. The electromagnetic field of the radiation after it has passed through the patterning device can be determined based on the electromagnetic field of the radiation before it reaches the patterning device and a function characterizing the interaction. Such a function can be called a mask transmission function (which can be used to describe the interaction of a transmissive patterning device and / or a reflective patterning device).

[0051] The mask transmission function can have a variety of different forms. One form is binary. The binary mask transmission function has one of two values (e.g., zero and a positive constant) at any given location on the patterning device. A mask transmission function in binary form can be called a binary mask. Another form is continuous. That is, the modulus of the transmittance (or reflectance) of the patterning device is a continuous function of the location on the patterning device. The phase of the transmittance (or reflectance) can also be a continuous function of the location on the patterning device. A mask transmission function in continuous form can be called a continuous tone mask or a continuous transmission mask (CTM). For example, a CTM can be represented as a pixelated image, where a value between 0 and 1 (e.g., 0.1, 0.2, 0.3, etc.) can be assigned to each pixel instead of a binary value of 0 or 1. In an embodiment, the CTM can be a pixelated grayscale image, where each pixel has a plurality of values (e.g., within the range [-255, 255], a normalized value within the range [0, 1] or [-1, 1] or other appropriate ranges).

[0052] The thin mask approximation, also known as the Kirchhoff boundary condition, is widely used to simplify the determination of the interaction between the radiation and the patterning device. The thin mask approximation assumes that the thickness of the structures on the patterning device is very small compared to the wavelength, and the width of the structures on the mask is very large compared to the wavelength. Thus, the thin mask approximation assumes that the electromagnetic field after the patterning device is the product of the incident electromagnetic field and the mask transmission function. However, as the lithography process uses radiation with an increasingly short wavelength, and the structures on the patterning device become smaller and smaller, the assumptions of the thin mask approximation break down. For example, the interaction of the radiation with the structures (such as the edges between the top surface and the sidewalls) can become significant due to the finite thickness of the structures ("mask 3D effect" or "M3D"). Incorporating such scattering in the mask transmission function may enable the mask transmission function to better capture the interaction between the radiation and the patterning device. The mask transmission function under the thin mask approximation may be referred to as the thin mask transmission function. The mask transmission function incorporating M3D may be referred to as the M3D mask transmission function.

[0053] Figure 2 is a flowchart of a method 200 for determining a patterning device pattern (or mask pattern hereinafter) from an image corresponding to a target pattern (e.g., a continuous transmission mask image, a binary mask image, a curve mask image, etc.) to be printed on a substrate via a patterning process involving a lithography process. In an embodiment, the design layout or the target pattern may be a binary design layout, a continuous tone design layout, or another suitable form of design layout.

[0054] Method 200 is an iterative process in which an initial image (e.g., an enhanced image, a mask variable initialized from a CTM image, etc.) is progressively modified to generate different types of images according to different processes of the present disclosure to ultimately generate information including a mask pattern or an image (e.g., a mask variable corresponding to a final curve mask) further used to fabricate / manufacture a mask. The iterative modification of the initial image may be based on a cost function, where during the iteration, the initial image may be modified such that the cost function is reduced, minimized in an embodiment. In an embodiment, method 200 may also be referred to as a binary CTM process, where the initial image is an optimized CTM image, and the optimized CTM image is further processed according to the present disclosure to generate a curve mask pattern (e.g., the geometry or polygon representation shape of a curve mask or a curve pattern). In an embodiment, the initial image may be an enhanced image of a CTM image). The curve mask pattern may be in the form of a vector, a table, a mathematical equation, or other forms representing a geometric / polygon shape.

[0055] In an embodiment, process P201 may involve obtaining an initial image (e.g., a CTM image or an optimized CTM image, or a binary mask image). In an embodiment, the initial image 201 may be a CTM image generated by a CTM generation process based on a target pattern to be printed on a substrate. The CTM image may then be received by the process P201. In an embodiment, the process P201 may be configured to generate a CTM image. For example, in CTM generation techniques, the inverse lithography problem is formulated as an optimization problem. The variables are related to the values of the pixels in the mask image, and lithography metrics such as EPE or sidelobe printing are used as cost functions. In the iterations of the optimization, the mask image is constructed from the variables and then a process model (e.g., the Tachyon model) is applied to obtain an optical or resist image and the cost function is calculated. The cost calculation then gives a gradient value that is used in the optimization solver to update the variables (e.g., pixel intensities). After several iterations during the optimization, a final mask image is generated, which is additionally used as a guiding map for pattern extraction (e.g., as implemented in the Tachyon SMO software). Such an initial image (e.g., a CTM image) may include one or more features (e.g., features of the target pattern, SRAF, SRIF, etc.) corresponding to the target pattern to be printed on the substrate via the patterning process.

[0056] In an embodiment, a CTM image (or an enhanced version of the CTM image) may be used to initialize mask variables that may be used as the initial image 201, which is iteratively modified as discussed below.

[0057] Process P201 may involve generating an enhanced image 202 based on the initial image 201. The enhanced image 202 may be an image in which certain selected pixels within the initial image 201 are magnified. The selected pixels may be pixels within the initial image 201 that have a relatively low value (or weak signal). In an embodiment, the selected pixels are pixels having a signal value lower than, for example, the average intensity of the pixels throughout the initial image or a given threshold. In other words, the pixels within the initial image 201 having a weak signal are magnified, thus enhancing one or more features within the initial image 201. For example, the second-order SRAF around a target feature may have a weak signal that can be magnified. Thus, the enhanced image 202 may highlight or identify additional features (or structures) that may be included in the mask image (generated later in the method). In conventional methods of determining the mask image (e.g., the CTM method), weak signals within the initial image may be ignored, and as such, the mask image may not include features that may be formed by the weak signals in the initial image 201.

[0058] The generation of the enhanced image 202 involves applying image processing operations such as filters (e.g., edge detection filters) to amplify weak signals within the initial image 201. Alternatively or additionally, the image processing operation can be deblurring, averaging, and / or feature extraction or other similar operations. Examples of the edge detection filters include the Prewitt operator, the Laplacian operator, the Laplacian of Gaussian (LoG) filter, etc. The generation step can also involve combining the amplified signal of the initial image 201 with the original signal of the initial image 201 with or without modifying the original strong signal of the initial image 201. For example, in an embodiment, for one or more pixel values at one or more locations across the initial image 201 (e.g., at contact holes), if the original signal is relatively strong (e.g., higher than a certain threshold such as 150 or lower than -50), the original signal at the one or more locations (e.g., at contact holes) may not be modified or combined with the amplified signal for that location.

[0059] In an embodiment, the noise (e.g., random variations in brightness or color or pixel values) in the initial image 201 can also be amplified. Thus, alternatively or additionally, a smoothing process can be applied to reduce the noise (e.g., random variations in brightness or color or pixel values) in the combined image. Examples of image smoothing methods include Gaussian blur, running average, low-pass filter, etc.

[0060] In an embodiment, an edge detection filter may be used to generate the enhanced image 202. For example, an edge detection filter may be applied to the initial image 201 to generate a filtered image that highlights the edges of one or more features within the initial image 201. The resulting filtered image may be further combined with the original image (i.e., the initial image 201) to generate the enhanced image 202. In an embodiment, the combination of the initial image 201 and the image obtained after edge filtering may involve modifying only those portions of the initial image 201 that have weak signals without modifying regions with strong signals, and the combination process may be weighted based on signal strength. In an embodiment, amplifying the weak signals may also amplify the noise within the filtered image. Thus, according to an embodiment, a smoothing process may be performed on the combined image. Smoothing of an image may refer to an approximation function that attempts to capture important patterns (e.g., target patterns, SRAFs) in the image while omitting noise or other fine-scale structures / rapid phenomena. In smoothing, the data points of the signal may be modified such that individual points (roughly due to noise) may be reduced and points that may be lower than neighboring points may be increased, resulting in a smoother signal or a smoother image. Consequently, after the smoothing operation, according to an embodiment of the present disclosure, a further smoothed version of the enhanced image 202 with reduced noise may be obtained.

[0061] In process P203, the method may involve generating a mask variable 203 based on the enhanced image 202. In a first iteration, the enhanced image 202 may be used to initialize the mask variable 203. In later iterations, the mask variable 203 may be iteratively updated.

[0062] The contour extraction of a real-valued function f of n real variables is a set of the following form:

[0063]

[0064] In two-dimensional space, the set defines the points on the surface where the function f equals a given value c. In two-dimensional space, the function f is capable of extracting a closed contour to be rendered to the mask image.

[0065] In the above equation, x1, x2,... x n refers to a mask variable such as the intensity of an individual pixel, and the mask variable determines the locations where the curve mask edge exists at a given constant value c (e.g., in a threshold plane as discussed in process P205 below).

[0066] In an embodiment, at an iteration, generation of the mask variable 203 can involve modifying one or more values (e.g., pixel values at one or more locations) of a variable within the enhanced image 202 based on, for example, an initialization condition or a gradient map (which can be generated subsequently in the method). For example, the one or more pixel values can be increased or decreased. In other words, the amplitude of one or more signals within the enhanced image 202 can be increased or decreased. Such a modified amplitude of the signal enables different curve patterns to be generated depending on the amount of change in the amplitude of the signal. Thus, the curve pattern gradually evolves until the cost function is decreased (minimized in an embodiment). In an embodiment, further smoothing can be performed on the horizontal mask variable 203.

[0067] In addition, process P205 involves generating a curve mask pattern 205 (e.g., having a polygonal shape represented in vector form) based on the mask variable 203. Generation of the curve mask pattern 205 can involve thresholding the mask variable 203 to trace or generate a curve (or curved) pattern from the mask variable 203. For example, thresholding can be performed using a threshold plane (e.g., the x - y plane) having a fixed value that intersects the signal of the mask variable 203. The intersection of the threshold plane with the signal of the mask variable 203 produces a trace or contour (i.e., a curved polygonal shape) that forms the polygonal shape serving as the curve pattern for the curve mask pattern 205. For example, the mask variable 203 can intersect a zero plane parallel to the (x,y) plane. Thus, the curve mask pattern 205 can be any curve pattern generated as described above. In an embodiment, the curve pattern traced or generated from the mask variable 203 depends on the signal of the enhanced image 202. As such, the image enhancement process P203 contributes to the improvement of the pattern generated for the final curve mask pattern. The final curve mask pattern can be further used by a mask manufacturer to fabricate a mask for use in a lithography process.

[0068] Process P207 can involve rendering the curve mask pattern 205 to produce a mask image 207. Rendering is an operation performed on the curve mask pattern and is a process similar to converting a rectangular mask polygon into a discrete grayscale image representation. Such a process can generally be understood as sampling a box function of continuous coordinates (the polygon) into values at each point of the image pixels.

[0069] The method also involves forward simulation of the patterning process using a process model that generates or predicts a pattern 209 that can be printed on the substrate based on the mask image 207. For example, process P209 can involve using the mask image 207 as an input to execute and / or simulate the process model and generating a process image 209 (e.g., a spatial image, a resist image, an etch image, etc.) on the substrate. In an embodiment, the process model can include a mask transmission model coupled to an optical device model, which is further coupled to a resist model and / or an etch model (e.g., as described below). The output of the process model can be a process image 209 that takes into account different process variations as factors during the simulation process. The process image can be further used to determine parameters of the patterning process (e.g., EPE, CD, overlay, sidelobes, etc.) by, for example, tracing the contours of the patterns within the process image. The parameters can also be used to define a cost function that is further used to optimize the mask image 207 such that the cost function is reduced or, in an embodiment, minimized.

[0070] In process P211, a cost function can be evaluated based on the process model image 209 (also referred to as a simulated substrate image or a substrate image or a wafer image). Thus, the cost function can be considered process-aware in the case of variations in the patterning process, enabling the generation of a curved mask pattern that takes into account the variations in the patterning process. For example, the cost function can be an edge placement error (EPE), a sidelobe, a mean squared error (MSE), a pattern placement error (PPE), a normalized image logarithm, or other suitable variables defined based on the pattern contours in the process image. The EPE can be an edge placement error associated with one or more patterns and / or the sum of all edge placement errors associated with all patterns in the process model image 209 and the corresponding target patterns. In an embodiment, the cost function can include more than one condition that can be simultaneously reduced or minimized. For example, in addition to the MRC violation probability, the number of defects, EPE, overlay, CD, or other parameters can also be included, and all conditions can be simultaneously reduced (or minimized).

[0071] In addition, one or more gradient maps can be generated based on the cost function (e.g., EPE), and mask variables can be modified based on such gradient maps. Mask variables (MV) refer The strength. Thus, the gradient calculation can be expressed as dEPE / d∅, and the gradient value is updated by capturing the inverse mathematical relationship from the mask image (MI) to the curved mask polygon to the mask variable. Thus, the derivative chain of the cost function with respect to the mask image can be calculated from the mask image to the curved mask polygon and from the curved mask polygon to the mask variable, which allows modifying the value of the mask variable at the mask variable.

[0072] In an embodiment, image regularization can be added to reduce the complexity of the mask pattern that can be generated. Such image regularization can be mask rule checking (MRC). MRC refers to the constraints of the mask manufacturing process or equipment. Thus, the cost function can include different components, for example, based on EPE and MRC violation penalties. The penalty can be a term of the cost function that depends on the amount of violation, such as the difference between the mask measurement and a given MRC or mask parameter (e.g., the mask pattern width and the allowed (e.g., minimum or maximum) mask pattern width). Thus, according to an embodiment of the present disclosure, a mask pattern can be designed and a corresponding mask can be fabricated not only based on the forward simulation of the patterning process but also additionally based on the manufacturing constraints of the mask manufacturing equipment / process. Thus, a manufacturable curved mask with high yield (i.e., minimum defects) and high accuracy can be obtained according to, for example, EPE or overlap on the printed pattern.

[0073] The pattern corresponding to the process image should be exactly the same as the target pattern. However, such an exact target pattern may be infeasible (e.g., usually sharp corners), and some conflicts are introduced due to variations in the patterning process itself and / or approximations in the model of the patterning process. In the first iteration of the method, the mask image 207 may not produce a pattern similar to the target pattern (in the resist image). The determination of the accuracy or acceptability of the printed pattern in the resist image (or etched image) can be based on a cost function such as EPE. For example, if the EPE of the resist pattern is high, it indicates that the printed pattern using the mask image 207 is unacceptable and the pattern in the mask variable 203 must be modified.

[0074] To determine whether the mask image 207 is acceptable, process P213 can involve determining whether the cost function is reduced or minimized, or whether a given number of iterations is reached. For example, the EPE value of the previous iteration can be compared with the EPE value of the current iteration to determine whether the EPE has been reduced, minimized, or converged (i.e., no significant improvement in the printed pattern is observed). When the cost function is minimized, the method can stop, and the generated curved mask pattern information is regarded as the optimization result.

[0075] However, if the cost function is not reduced or minimized, the mask-related variables or enhanced image-related variables (e.g., pixel values) may be updated. In an embodiment, the update may be according to a gradient-based method. For example, if the cost function is not reduced, the method 200 proceeds to the next iteration of generating the mask image after performing processes P215 and P217 that indicate how to further modify the mask variable 203.

[0076] Process P215 may involve generating a gradient map 215 based on the cost function. The gradient map may be the derivative and / or partial derivative of the cost function. In an embodiment, the partial derivative of the cost function may be determined with respect to the pixels of the mask image, and the derivatives may be further chained to determine the partial derivative with respect to the mask variable 203. Such gradient calculation may involve determining the inverse relationship between the mask image 207 and the mask variable 203. In addition, the inverse relationship of any smoothing operation (or function) performed in processes P205 and P203 must be considered.

[0077] The gradient map 215 may provide a suggestion on how to increase or decrease the value of the mask variable in a way that reduces (minimized in an embodiment) the value of the cost function. In an embodiment, an optimization algorithm may be applied to the gradient map 215 to determine the mask variable value. In an embodiment, an optimization solving process may be used to perform the gradient-based calculation (in process P217).

[0078] In an embodiment, for an iteration, the mask variable may be changed while the threshold plane may be kept fixed or unchanged to gradually reduce or minimize the cost function. Thus, the resulting curve pattern may gradually evolve during the iteration such that the cost function is reduced, or minimized in an embodiment. In another embodiment, both the mask variable and the threshold plane may be changed to achieve a faster convergence of the optimization process. A final set of binary CTM results (i.e., a modified version of the enhanced image, mask image, or curve mask) may be produced after several iterations and / or minimization of the cost function.

[0079] In embodiments of the present disclosure, the transition from CTM optimization using a grayscale image to binary CTM optimization using a curve mask can be simplified by replacing the threshold setting processes (i.e., P203 and P205) with different processes, at which an S-shaped transformation is applied to the enhanced image 202 and corresponding changes in gradient calculation are performed. The S-shaped transformation of the enhanced image 202 produces a transformed image that gradually evolves into a curve pattern during an optimization process (e.g., minimizing a cost function). During an iterative or optimization step, variables related to the S-shaped function (e.g., steepness and / or threshold) can be modified based on gradient calculation. Since the S-shaped transformation becomes steeper during successive iterations (e.g., the steepness of the slope of the S-shaped transformation increases), a gradual transition from the CTM image to the final binary CTM image can be achieved, allowing for improved results in the final binary CTM optimization using a curve mask pattern.

[0080] In embodiments of the present disclosure, additional steps / processes can be inserted into the iterative loop of the optimization to enhance the result to have selected or desired properties. For example, smoothness can be ensured by adding a smoothing step, or other filters can be used to enhance the image for horizontal / vertical structures.

[0081] The method has several features or aspects. For example, a CTM mask image optimized using an image enhancement method is used to improve the signal, which can also be used as seeding in the optimization process. In another aspect, a threshold setting method using CTM technology (referred to as binary CTM) enables the generation of a curve mask pattern. In yet another aspect, the complete formulation of gradient calculation (i.e., closed-loop formulation) also allows the use of gradient-based solution processes for mask variable optimization. The binary CTM result can be used as a local solution (as a hot spot fix) or as a full-chip solution. The binary CTM result can be used as an input together with machine learning. This can allow the use of machine learning to accelerate binary CTM. In yet another aspect, the method includes an image regularization method to improve the result. In another aspect, the method involves successive optimization stages to achieve a smoother transition from grayscale image CTM to binary curve mask binary CTM. The method allows tuning of the optimization threshold to improve the result. The method includes additional transformations to the iterative optimization to enhance the desirable properties of the result (requiring smoothness in the binary CTM image).

[0082] As the lithography nodes continue to shrink, increasingly complex masks are required. The method described herein can be used in critical layers with DUV scanners, EUV scanners, and / or other scanners. The method according to the present disclosure can be included in different aspects of the mask optimization process (including source mask optimization (SMO), mask optimization, and / or OPC).

[0083] As described above, it is often desirable to be able to computationally determine how a patterning process will produce a desired pattern on a substrate. Thus, simulations can be provided to simulate one or more parts of the process. For example, it is desirable to be able to simulate the lithography process of transferring a patterning device pattern to a resist layer on a substrate and the pattern generated in the resist layer after development of the resist.

[0084] Figure 3 An exemplary flowchart for simulating lithography in a lithographic projection apparatus is illustrated. The illumination model 331 represents the optical characteristics of the illumination (including the radiation intensity distribution and / or phase distribution). The projection optics model 332 represents the optical characteristics of the projection optics (including the change in the radiation intensity distribution and / or phase distribution caused by the projection optics). The design layout model 335 represents the optical characteristics of the design layout (including the change in the radiation intensity distribution and / or phase distribution caused by a given design layout), where the design layout is a representation of the arrangement of features on or formed by a patterning device. The illumination model 331, the projection optics model 332, and the design layout model 335 can be used to simulate the aerial image 336. The resist image 338 can be simulated from the aerial image 336 using the resist model 337. The simulation of lithography can, for example, predict the profile and / or CD in the resist image.

[0085] More specifically, the illumination model 331 may represent the optical characteristics of the illumination, which may include but are not limited to NA - root mean square deviation (σ) settings, and any particular illumination shape (e.g., off - axis illumination such as annular, quadrupole, dipole, etc.). The projection optics model 332 may represent the optical characteristics of the projection optics, including for example aberrations, distortions, refractive index, physical size or dimensions, etc. The design layout model 335 may also represent one or more physical properties of the patterning device, as described for example in U.S. Patent No. 7,587,704, which is hereby incorporated by reference in its entirety. The optical properties associated with the lithographic projection apparatus (e.g., the properties of the illumination, the patterning device, and the projection optics) define the aerial image. Since the patterning device used in the lithographic projection apparatus may be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the remainder of the lithographic projection apparatus, which at least includes the illumination and the projection optics (and thus the design layout model 335).

[0086] The resist model 337 may be used to calculate the resist image based on the aerial image, an example of which can be found in U.S. Patent No. 8,200,468, which is hereby incorporated by reference in its entirety. The resist model is typically related to the properties of the resist layer (e.g., the effects of chemical processes occurring during exposure, post - exposure bake, and / or development).

[0087] The goal of the simulation is to accurately predict (e.g.) edge placement, aerial image intensity slope, and / or CD, which can then be compared to the expected design. The expected design is typically defined as a pre - OPC design layout, which may be provided in a standardized digital file format such as GDSII, OASIS, or another file format.

[0088] From the design layout, one or more portions referred to as "chips" can be identified. In an embodiment, a set of chips is extracted, which represents complex patterns (e.g., a set of patterns) in the design layout (typically on the order of 50 to 1000 chips, but any number of chips can be used). As would be understood by one of ordinary skill in the art, these patterns or chips represent small portions of the design (e.g., circuits, cells, etc.) for which specific attention and / or verification (e.g., a set of patterns) is required. In other words, a chip can be a portion of the design layout, or can be similar, or have similar behavior to a portion of the design layout whose critical features are identified empirically (including chips provided by a customer), identified by trial and error, or identified by performing full-chip simulations. A chip typically contains one or more test patterns or gauge patterns. An initial larger set of chips can be provided a priori by a customer based on known critical feature regions in the design layout that require specific image optimization. Alternatively, in another embodiment, an initial larger set of chips can be extracted from the entire design layout by using the process for identifying the critical feature regions described below.

[0089] In some examples, the simulation and modeling can be used to configure one or more features of the patterning device pattern (e.g., perform optical proximity correction), one or more features of the illumination (e.g., change one or more characteristics of the spatial / angular intensity distribution of the illumination, such as shape), and / or one or more features of the projection optics (e.g., numerical aperture, etc.). Such configuration can generally be referred to as mask optimization, source optimization, and projection optimization, respectively. Such optimizations can be performed independently or in various combinations. One such example is source-mask optimization (SMO), which involves configuring one or more features of the patterning device pattern together with one or more features of the illumination. The optimization techniques can focus on one or more of the chips in the set. The optimization can use the machine learning models described herein to predict values of various parameters (including images, etc.).

[0090] In some embodiments, the illumination model 331, the projection optics model 332, the design layout model 335, the resist model 337, the SMO model, and / or other models associated with and / or included in the integrated circuit manufacturing process can be empirical models that perform operations of the methods described herein. The empirical models can predict outputs based on correlations between various inputs (e.g., one or more characteristics of a mask or wafer image, one or more characteristics of the design layout, one or more characteristics of the patterning device, one or more characteristics of the illumination used in the lithography process (such as wavelength), etc.).

[0091] As an example, the empirical model can be a machine learning model. In some embodiments, the machine learning model can be and / or include mathematical equations, algorithms, plots, charts, networks (such as neural networks), and / or other tools and machine learning model components. For example, the machine learning model can be and / or include one or more neural networks having an input layer, an output layer, and one or more intermediate or hidden layers. In some embodiments, the one or more neural networks can be and / or include deep neural networks (e.g., neural networks having one or more intermediate or hidden layers between the input layer and the output layer).

[0092] As an example, the one or more neural networks can be based on a large collection of nerve units (or artificial neurons). The one or more neural networks can loosely, i.e., not strictly, mimic the way a biological brain works (e.g., via large clusters of biological neurons connected by axons). Each nerve unit of a neural network can be connected to many other nerve units of the neural network. Such connections can strengthen or inhibit their influence on the activation state of the connected nerve units. In some embodiments, each individual nerve unit can have a summation function that combines the values of all its inputs. In some embodiments, each connection (or the nerve unit itself) can have a threshold function such that a signal must exceed the threshold before it is allowed to propagate to other nerve units. These neural network systems can be self-learning and trained, rather than explicitly programmed, and can perform significantly better in certain problem-solving areas compared to traditional computer programs. In some embodiments, one or more neural networks can include multiple layers (e.g., where the signal path traverses from a previous layer to a subsequent layer). In some embodiments, the backpropagation technique can be utilized by the neural network, where forward stimuli are used to reset the weights of the "front-end" nerve units. In some embodiments, the stimuli and inhibitions for one or more neural networks may flow more freely, where the connections interact in a more chaotic and complex manner. In some embodiments, the intermediate layer of one or more neural networks includes one or more convolutional layers, one or more recurrent layers, and / or other layers.

[0093] A set of training data can be used to train the one or more neural networks (i.e., determine their parameters). The training data can include a set of training samples. Each sample can be a pair including an input object (usually a vector, which can be referred to as a feature vector) and a desired output value (also referred to as a governing signal). The training algorithm analyzes the training data and adjusts the behavior of the neural network by adjusting the parameters of the neural network (e.g., the weights of one or more layers) based on the training data. For example, given a set of N training samples of the form such that xi is the feature vector for the i-th example and y i In the case of being its management signal, the training algorithm searches for a neural network g: X → Y, where X is the input space and Y is the output space. A feature vector is an n-dimensional vector representing the numerical features of a certain object (e.g., a wafer design as in the above example). The vector space associated with these vectors is often referred to as the feature space. After training, the neural network can be used to make predictions using new samples.

[0094] For example, as described above, in order to train a DCNN and / or other machine learning models to predict CTM maps, it is necessary for the user to generate CTM images as training data (e.g., via ASML's Tachyon product). However, it is difficult to select an appropriate representative portion of the patterns in the full-chip GDS as the training patterns that are set to train the DCNN to predict CTM maps. Manual selection requires the user to have significant prior knowledge, i.e., preparatory knowledge, about the full-chip GDS design and invest several hours in the selection process. Although random selection is not time-consuming compared to manual selection, the unclear pattern coverage and low stability associated with random selection make random selection infeasible or non-persistent in real-world applications.

[0095] To address these and other disadvantages of prior systems, the present method and device provide an effective tool for the user to automatically perform training pattern selection from a full-chip GDS file. The present method and device are configured to automatically select representative patterns from a set of patterns in less time compared to prior art systems.

[0096] Figure 4 Illustrates an overview of the operations of the present method 400 for determining training patterns during the wafer patterning process. Figure 4 The method shown is a method or part of a method for training a machine learning model during a wafer patterning process (e.g., as described herein). The operations include: generating 441 (e.g., using a feature generation engine) a plurality of features from unique patterns in a set of patterns 443 to form an enhanced set of patterns 445 with additional features; grouping 442 (e.g., using a pattern clustering engine) the patterns in the enhanced set of patterns 445 into a plurality of individual groups 447 based on the similarity of the plurality of generated features; and selecting 444 (e.g., using a pattern selection engine) representative patterns 448 from the plurality of individual groups 447 to determine the training patterns 449. The selected representative patterns 448 from the plurality of individual groups 447 (e.g., which form the training patterns 449) can be provided to a machine learning model (e.g., DCNN, Figure 4(not shown in the figure) to train the machine learning model to predict a continuous transmission mask (CTM) map for optical proximity correction (OPC) in the wafer patterning process, and / or can be provided for other applications. For example, the patterns in pattern set 443 can be, for example, a unique pattern library 446 generated by the pattern collection from layout (PCL) function 451 in, for example, the Tachyon PRO (pattern recognition and optimization) product (and / or other similar products) based on the full-chip layout GDS file 453, and / or be part of the unique pattern library 446.

[0097] In operation 441, the multiple generated features include geometric features, lithography-aware features, and / or other features. The geometric features include one or more of the following: target mask image, frequency map, pattern density map, pattern occurrence rate of unique patterns in the pattern set, and / or other geometric features. The lithography-aware features include one or more of the following: sub-resolution assist feature guide map (SGM), diffraction order, diffraction pattern of unique patterns in the pattern set, and / or other lithography-aware features. Geometric features and lithography-aware features are generated for each of the multiple patterns in the pattern set. In an embodiment, a feature set (including geometric and / or lithography-aware features) is generated for each pattern (in set 443). The multiple features generated from the multiple patterns in the pattern set are features in addition to the geometric information and / or vertex information already included in the pattern set.

[0098] For example, the pattern set 446 (e.g., pattern library (PLIB)) generated by PCL 451 can include geometric information associated with multiple individual unique patterns. However, only the vertex information for a single pattern can be stored as part of the pattern set 446 to represent a specific pattern design. For example, vertex information is easy to use for exact pattern matching, but vertex information is not sufficient to guide a pattern grouping method to group similar patterns (which are not exactly the same) adequately because this process is a fuzzy matching process. To improve the robustness of pattern grouping in operation 442, operation 441 includes generating additional features (e.g., in addition to vertex geometric information) for the multiple individual patterns, including both geometric and lithography-aware features. As described above, the geometric features include one or more of the following: target mask image, frequency map, pattern density map, pattern occurrence rate of unique patterns in the pattern set, and / or other geometric features. The lithography-aware features include one or more of the following: sub-resolution assist feature guide map (SGM), diffraction order, diffraction pattern of unique patterns in the pattern set, and / or other lithography-aware features. However, this specification is not intended to be restrictive. Operation 441 can be customized by the user to extend to additional and / or different features generated according to the user's specifications.

[0099] Figure 5 illustrates Figure 4 the features shown in generate aspects of operation 441. As Figure 5 shown, operation 441 includes generating geometric feature 502 and lithography-aware feature 504. Geometric feature 502 and lithography-aware feature 504 are generated for each pattern 501 in pattern set 443 ( Figure 4 ) such that a set of features 502, 504 is generated for each pattern 501. Examples of geometric feature 502 generated as part of operation 441 (e.g., by a feature generation engine) include target mask image 506, frequency map 508, pattern density map 510, and pattern occurrence rate (count) 512 of the pattern in the pattern set. Examples of lithography-aware feature 504 generated as part of operation 441 (e.g., by a feature generation engine) include SGM map 514, diffraction order 516, and diffraction pattern 518. These features are only examples and are not intended to be limiting.

[0100] Returning to Figure 4 , in operation 442, the patterns in enhanced pattern set 445 are grouped into multiple individual groups based on the similarity of multiple generated features (e.g., generated at operation 441). Operation 442 is performed using unsupervised machine learning. In an embodiment, the grouping includes clustering and / or other grouping operations. The clustering includes a series of consecutive clustering steps 455, 457, 459 that use different features among the multiple generated features for different clustering steps 455, 457, 459. In the example shown in Figure 4 , clustering step 455 can be performed based on a first feature, clustering step 457 can be performed based on a second feature, and clustering step 459 (and / or any other intermediate clustering step) can be performed based on an nth feature. The series of consecutive clustering steps 455, 457, 459 form corresponding subgroups 461, 463, 465 of the unique patterns in the pattern set such that representative pattern 448 (operation 444) is selected from subgroups 461, 463, 465 to determine training pattern 449.

[0101] Figure 6 illustrates additional details related to Figure 4 operation 442 shown in. For example, Figure 6 illustrates based on multiple generated features (e.g., at Figure 4The similarity generated at operation 441 shown (e.g., by a feature generation engine) clusters the patterns in the enhanced pattern set 445 into multiple separate groups. Since the patterns exhibit multi-dimensional features, clustering can be performed based on multiple individual features. Clustering can be performed on each pattern through unsupervised machine learning Figure 6 The clustering shown, considering multiple features for each pattern in a step-by-step manner to group together patterns sharing similarities in multi-dimensional features (e.g., as shown by Figure 6 The branches and sub-branches shown for forming groups). For example, the user can define which features should be used for clustering the patterns and the sequence of features for clustering. In the Figure 6 example shown, the clustering includes a continuous series of clustering steps 455, 457, 459 that use different features among the multiple generated features of the different clustering steps 455, 457, 459 (e.g., the identification and order of these clustering steps are defined by the user). For example, clustering step 455 can be performed based on a first feature, clustering step 457 can be performed based on a second feature, and clustering step 459 (and / or any other intermediate clustering steps) can be performed based on an nth feature. The continuous series of clustering steps 455, 457, 459 form corresponding subgroups 461 (e.g., groups A and B), 463 (e.g., groups A-a, A-b, B-a, B-b, B-c), and 465 (e.g., A-a-n, A-b-n, B-a-n, B-b-n, B-c-n, etc.) of unique patterns. For example, Figure 6 The patterns in the group in the nth layer of the hierarchy shown (clustering step 459) share similarities in n features.

[0102] In an embodiment, the continuous series of clustering steps 455, 457, 459 includes a cross-validation step performed using a given feature for a given step. The cross-validation step includes: adjusting the number of groups clustered at the multiple individual clustering steps 455, 457, 459, determining and / or adjusting which unique patterns to include in a given subgroup, and / or other cross-validation operations. In an embodiment, the cross-validation step includes: optimizing the number of groups clustered at the multiple individual clustering steps 455, 457, 459, and optimizing which unique patterns to include in a given subgroup, and / or other cross-validation operations.

[0103] In an embodiment, clustering and / or cross-validation includes machine learning clustering methods such as k-means clustering and / or other clustering methods. K-means clustering is configured to partition n observations into k clusters, where each observation belongs to the cluster with the closest mean, which serves as a representative example of the cluster. This results in the data space being partitioned into Voronoi cells. Given a set of patterns (x1, x2, ……, xn ), or one of its corresponding features, then the feature can be regarded as a d-dimensional real vector, such that the k-means clustering is configured to partition n observations into k (≤n) sets S={S1, S2, ……, S k} to minimize the within-cluster sum of squares. Formally, the goal is to find:

[0104]

[0105] where μ i is the mean of the points in S i .

[0106] Although k-means clustering is an efficient clustering method, one of its drawbacks is that the number of clusters k needs to be specified before the clustering occurs. To address the need for pre-selecting a certain number of clusters for each k-means clustering process, the present device and method are configured to perform k-fold cross-validation tests using the elbow method to determine the value of k (e.g., the optimal value). The elbow method includes running k-means clustering on a dataset for a range of k values (e.g., k ranges from 1 to 14, which is not intended to be restrictive), and for each k value, determining the sum of squared errors (SSE). After determining the SSE, a line graph of the SSE is plotted for each k value.

[0107] Figure 7 Illustrates an example line graph 700 of SSE 702 versus k value 704. If the line graph 700 looks similar to an arm, then the "elbow" 706 on the arm corresponds to the optimal value of k. In principle, a relatively small SSE is preferred, but the SSE tends to decrease towards zero as k increases (the SSE is 0 when k is equal to the number of data points in the dataset because each data point is its own cluster and there is no error between the data point and the center of its cluster). The present method and device are configured to determine a smaller k value that still has a low SSE, and the "elbow" 706 in the line graph 700 corresponds to the point where increasing k in the line graph 700 results in diminishing returns. For example, in Figure 7 , the "elbow" 706 is located at k = 5, indicating that for this example dataset, the optimal k is 5. As k increases above 5, the SSE 702 changes less and less as it approaches zero.

[0108] The method and apparatus are configured such that for the determination of the sum of squared errors for each k value, a k-fold cross-validation test is performed. In k-fold cross-validation, the original sample is randomly divided into p equal-sized subsamples. A single subsample among the p subsamples is retained as validation data for determining the sum of squared errors, and the remaining p - 1 subsamples are used as training data for k-means clustering. Subsequently, the cross-validation process is repeated p times, where each of the p subsamples is used as validation data exactly once. The p results can then be averaged to produce a single estimate of the sum of squared errors. The advantage of this method over repeated random subsampling is that multiple individual observations are used for both training and validation, and multiple individual observations are used for validation exactly once.

[0109] By way of non-limiting example, Figure 8 illustrates k-fold cross-validation, where k = 4. As Figure 8 shown, the original sample 800 is randomly divided into p equal-sized subsamples 802 (e.g., five points in each subsample 802). In Figure 8 , there are p = 4 subsamples in the sample 800. A single subsample 804 among the p = 4 subsamples is retained as validation data 806 for determining the sum of squared errors, and the remaining p - 1 subsamples 802 are used as training data 808 for k-means clustering. Subsequently, the cross-validation process is repeated p times, where each of the p subsamples is used as validation data (e.g., 810, 812, 814) exactly once. The p results can then be averaged to produce a single estimate of the sum of squared errors.

[0110] Returning to Figure 4 , 444 representative patterns 448 are selected from multiple individual groups 447 to determine that the training pattern 105 includes selecting a target number 467 ("M") of representative patterns 448 from a corresponding number of groups 447 (e.g., groups G(1), G(2),..., G(M)). In an embodiment, selecting representative patterns 448 from a target number 467 of multiple individual groups 447 to determine the training pattern 449 includes selecting the most central patterns 469 (such as CP(1), CP(2),..., CP(M) as shown in Figure 4 ) from the target number 467 of multiple individual groups 447 (G(1), G(2),..., G(M)).

[0111] In an embodiment, the most central pattern 469 from a given set (CP(1), CP(2), …, CP(M)) is the pattern that is closest to the centroid of the specified feature space for the multiple individual sets relative to the other patterns in the multiple individual sets. For example, the specified feature space can be a target mask image feature space, a frequency mapping feature space, a pattern density mapping feature space, a pattern occurrence feature space, an SGM feature space, a diffraction order feature space, a diffraction pattern feature space, and / or other feature spaces. In an embodiment, the method and apparatus can be configured such that by default the target mask image feature space is used to select the central sample. However, because each sample in the enhanced pattern set (e.g., Figure 4 445 therein) includes multi-dimensional features (e.g., frequency mapping, pattern density mapping, etc.), the method and apparatus can be configured such that the user can thus specify the feature space and / or select the central sample based on the target feature.

[0112] The method and apparatus are configured such that: assuming there are a total of (N) multiple individual groups 447, and in response to N being greater than the target number M 467 of representative patterns 448, then M patterns are selected from the N groups as representative example patterns that are set to be used as training patterns 449. The M selected patterns are configured to maximize the variation of the training patterns 449 in a given (e.g., user-selected) feature space (e.g., the default target mask image feature space).

[0113] Figure 9 Illustrated is the pattern selection from n groups 900 (e.g., Figure 4 the “n” therein is similar to the “N” described above) to form a set of representative example patterns for use as training patterns (e.g., Figure 9 449 therein). The groups 900 are formed by images 901 (e.g., clips) C1 902, C2 904, …, Cn 906. As Figure 9 shown, for a pattern set including n samples (e.g., n images (or clips) of different patterns), with a sample image size of H pixels by W pixels (e.g., H×W), C j P i indicates the i-th pixel in the j-th image (clip), where j ϵ [1,n] and i ϵ [1,H*W]. For pixel P i , P i 's variation along different images is given as follows:

[0114]

[0115] For each term (P ij - P i ) in the above formula 2, the change of the j-th pixel among n images at the same pixel position is determined, and the pixel changes of all H×W pixels are summed to represent a criterion for selecting M patterns from N groups as a set of representative example patterns, where the selection criterion is configured to maximize the change in the training patterns.

[0116] Return to Figure 4 , as described above, the method and apparatus are configured to select M samples from N candidates (M < N) to maximize the variation of the training pattern 449 (e.g., in the target mask image feature space). In an embodiment, determining the training pattern 449 may include performing a global search method across all N candidates such that the training pattern 449 has the maximum amount of variation. The global search method includes traversing all permutations of multiple candidate patterns and multiple groups, since this is a gradient-free search. The total number of permutations is:

[0117]

[0118] Typically, both M and N are usually greater than hundreds or even thousands. Thus, the total number of permutations is an extremely large number. This makes it difficult to traverse all permutations within a limited and feasible amount of time.

[0119] In an embodiment, different methods may be used to select M samples from N candidates. In an embodiment, the target number 467 of the representative pattern 448 is determined based on a stopping criterion and / or other information. The stopping criterion is configured to contribute to the variation in the training pattern 449. The stopping criterion may be determined based on information from previous training patterns, entered and / or selected by a user, determined at the manufacturing site of the apparatus, and / or determined in other ways. In an embodiment, method 400 includes determining the amount of variation of the training pattern 449. In an embodiment, the stopping criterion is configured to ensure that the amount of variation of the training pattern 449 exceeds a variation threshold. The variation threshold may be determined based on information from previous training patterns, entered and / or selected by a user, determined at the manufacturing site of the apparatus, and / or determined in other ways.

[0120] In an embodiment, a target number 467 of representative patterns 448 are randomly selected from a total number (N) of individual groups 447 (e.g., patterns are randomly selected from M groups). In an embodiment, in response to the amount of change in the training pattern not exceeding a change threshold, in response to the amount of change in a subsequently selected training pattern 449 increasing relative to the immediately preceding iteration in the training pattern 449, and / or for other reasons, the target number of representative patterns are re-randomly selected. For example, the amount of change in the training pattern 449 may be determined after an iteration of randomly selecting patterns from M groups out of a total number N of groups. If the change in the current iteration is greater than the previous change (e.g., because an increased change is needed), the training pattern 449 may be updated (e.g., the randomly selected patterns that make up the training pattern 449 may be randomly re-selected). The stopping criterion may be and / or include a maximum number of iterations (e.g., N_iter > max_iter), a change amount that breaks through a threshold (e.g., change >= threshold), a breakthrough of a maximum iteration time, and / or other stopping criteria. Such a search method is time-feasible and can be used to utilize (e.g., good enough) changes to drive the training set.

[0121] In an embodiment, method 400 further includes providing the training pattern 449 ( Figure 4 not shown therein) to a deep convolutional neural network and / or other machine learning model to train the deep convolutional neural network. In an embodiment, method 400 further includes using the trained deep convolutional neural network to perform ( Figure 4 also not shown therein) optical proximity effect correction as part of a wafer patterning process. For example, the training pattern 449 (e.g., generated based on a full-chip GDS as described herein) may be provided to train a deep convolutional neural network and / or other machine learning model to predict a CTM map. For example, the training pattern 449 (e.g., a continuous transmission mask (CTM) image) is a set of representative patterns that are provided to a machine learning model (e.g., a DCNN) for training such that an accurate CTM map will be predicted by the machine learning model (e.g., a DCNN) for full-chip OPC applications. Such a more accurate CTM map results in fewer lithography hotspots and fewer process window limitations, and / or reduces difficulties during subsequent mask correction operations when attempting to meet lithography performance specifications (e.g., related to aerial images, resist images, etc.). Other applications of method 400 are contemplated.

[0122] Figure 10FIG. is a block diagram of a computer system 100 that can assist in implementing the methods, processes, or apparatuses disclosed herein. Computer system 100 includes a bus 102 or other communication mechanism for communicating information, and a processor 104 (or processors 104 and 105) coupled to bus 102 for processing information. Computer system 100 also includes a main memory 106 coupled to bus 102 for storing information and instructions to be executed by processor 104, such as random access memory (RAM) or other dynamic storage device. Main memory 106 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 104. Computer system 100 further includes a read-only memory (ROM) 108 or other static storage device coupled to bus 102 for storing static information and instructions for processor 104. A storage device 110, such as a magnetic disk or optical disk, is provided and coupled to bus 102 for storing information and instructions.

[0123] Computer system 100 may be coupled via bus 102 to a display 112 for displaying information to a computer user, such as a cathode ray tube (CRT), flat panel display, or touch panel display. An input device 114 including alphanumeric keys and other keys is coupled to bus 102 for communicating information and command selections to processor 104. Another type of user input device is a cursor control 116 for communicating direction information and command selections to processor 104 and for controlling cursor movement on display 112, such as a mouse, trackball, or cursor direction keys. Such input devices typically have two degrees of freedom in two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a position in a plane. A touch panel (screen) display may also be used as an input device.

[0124] According to one embodiment, according to one embodiment, portions of one or more of the methods described herein may be performed by computer system 100 in response to processor 104 executing one or more sequences of one or more instructions contained in main memory 106. These instructions may be read into main memory 106 from another computer-readable medium, such as storage device 110. Execution of the instruction sequence contained in main memory 106 causes processor 104 to perform the process steps described herein. One or more processors in a multiprocessing arrangement may also be used to execute the instruction sequence contained in main memory 106. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Accordingly, the description herein is not limited to any specific combination of hardware circuitry and software.

[0125] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 104 for execution. Such a medium may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 110. Volatile media includes volatile memory, such as main memory 106. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and in frared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, DVD, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave as described hereinafter, or any other medium from which a computer can read.

[0126] Various forms of computer-readable media are involved in carrying one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions can initially be carried on a magnetic disk of a remote computer. The remote computer can load the instructions into its volatile memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 102 can receive the data carried in the infrared signal and place the data on bus 102. Bus 102 carries the data to main memory 106, from which processor 104 retrieves and executes the instructions. The instructions received by main memory 106 can optionally be stored on storage device 110 either before or after execution by processor 104.

[0127] Computer system 100 can also include a communication interface 118 coupled to bus 102. Communication interface 118 provides a two-way data communication coupling to network link 120 that is connected to a local area network 122. For example, communication interface 118 can be an integrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 118 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links can also be implemented. In any such implementation, communication interface 118 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0128] Network link 120 generally provides data communication to other data devices via one or more networks. For example, network link 120 can provide a connection to a host computer 124 or to data equipment operated by an Internet service provider (ISP) 126 via a local area network 122. The ISP 126 in turn provides data communication services via a global packet data communication network (now commonly referred to as the "Internet" 128). Both the local area network 122 and the Internet 128 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals via the various networks and signals on network link 120 and via communication interface 118 (which carry digital data to and from the computer system 100) are example forms of carriers that convey information.

[0129] Computer system 100 can send messages and receive data including process code via a network, network link 120, and communication interface 118. In the Internet example, server 130 can transmit the requested process code for an application via the Internet 128, ISP 126, local area network 122, and communication interface 118. For example, one such download application can provide all or part of the methods described herein. The received process code can be executed by processor 104 upon receipt, and / or stored in storage device 110 or other non-volatile memory for later execution. In this way, computer system 100 can obtain application program code in the form of a carrier wave.

[0130] Figure 11 Schematically depicts an exemplary lithographic projection apparatus that can be utilized in conjunction with the techniques described herein. The apparatus includes:

[0131] - An illumination system IL, which is used to condition a radiation beam B. In such a particular case, the illumination system also includes a radiation source SO;

[0132] - A first object table (e.g., a patterning device table) MT, having a patterning device holder for holding a patterning device MA (e.g., a mask table) and connected to a first positioner for accurately positioning the patterning device relative to an item PS;

[0133] - A second object table (substrate table) WT, having a substrate holder for holding a substrate W (e.g., a silicon wafer coated with resist) and connected to a second positioner for accurately positioning the substrate relative to an item PS; and

[0134] - A projection system ("lens") PS (e.g., a refractive, reflective, or catadioptric optical system), which can image a radiated portion of the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W.

[0135] As depicted herein, the device may belong to the transmissive type (e.g., employing a transmissive patterning device). However, in general, it may belong to the reflective type (e.g., employing a reflective patterning device). The device may employ a different kind of patterning device than a classical mask; examples include a programmable mirror array or an LCD matrix.

[0136] A source SO (e.g., a mercury lamp or an excimer laser, a laser-produced plasma (LPP) EUV source) generates a radiation beam. For example, this beam is fed directly or after traversing an adjusting device such as a beam expander Ex into an illumination system (illuminator) IL. The illuminator IL may include an adjusting device AD for setting an outer radial range and / or an inner radial range of the intensity distribution in the beam (commonly referred to as σ - outer and σ - inner, respectively). Additionally, the illuminator IL typically includes various other components, such as a collector IN and a condenser CO. Thus, the beam B incident on the patterning device MA has a desired uniformity and intensity distribution in its cross-section.

[0137] Regarding Figure 10 It should be noted that the source SO may be within the housing of the lithographic projection apparatus (which is often the case when the source SO is, for example, a mercury lamp), but it may also be remote from the lithographic projection apparatus and the radiation beam it generates is guided into the apparatus (e.g., by means of a suitable directing mirror); the latter case is often when the source SO is an excimer laser (e.g., based on KrF, ArF, or F2 laser action).

[0138] The beam PB then intercepts the patterning device MA held on the patterning device table MT. In the case of traversing the patterning device MA, the beam PB may pass through a lens PL, which focuses the beam B onto a target portion C of the substrate W. By means of a second positioning device PW2 (and an interferometric device IF), the substrate table WT can be accurately moved, for example, to position different target portions C in the path of the beam PB. Similarly, a first positioning device can be used to accurately position the patterning device MA, for example, after mechanically retrieving the patterning device MA from a patterning device library or during scanning relative to the path of the beam B. Generally, the movement of the object tables MT, WT can be achieved by means of long-stroke modules (coarse positioning) and short-stroke modules (fine positioning) not explicitly depicted in Figure 11 However, in the case of a stepper (as opposed to a step-and-scan tool), the patterning device table MT may be connected only to a short-stroke actuator or may be fixed.

[0139] The depicted tool can be used in two different modes:

[0140] - In the step mode, the patterning device table MT is kept substantially stationary and the entire patterning device image is projected onto the target portion C in one go (i.e., a single “flash”). Subsequently, the substrate table WT is displaced in the x and / or y direction so that different target portions C can be irradiated by the beam PB;

[0141] - In the scan mode, substantially the same situation applies, except that a given target portion C is not exposed in a single “flash”. Alternatively, the patterning device table MT can be moved at a rate v in a given direction (the so-called “scan direction”, e.g., the y direction) such that the projection beam B scans across the patterning device image; simultaneously, the substrate table WT is moved at a rate V = Mv in the same or opposite direction, where M is the magnification of the lens PL (typically M = 1 / 4 or 1 / 5). In this way, a relatively large target portion C can be exposed without sacrificing resolution.

[0142] Figure 12 Schematically depicts another exemplary lithographic projection apparatus 1000 that can be utilized in conjunction with the techniques described herein.

[0143] The lithographic projection apparatus 1000 includes:

[0144] - A source collector module SO;

[0145] - An illumination system (illuminator) IL configured to condition a radiation beam B (e.g., EUV radiation);

[0146] - A support structure (e.g., patterning device table) MT configured to support a patterning device (e.g., a mask or reticle) MA and connected to a first positioning device PM configured to accurately position the patterning device;

[0147] - A substrate table (e.g., wafer table) WT configured to hold a substrate (e.g., a wafer coated with resist) W and connected to a second positioning device PW configured to accurately position the substrate; and

[0148] - A projection system (e.g., a reflective projection system) PS configured to project the pattern imparted to the radiation beam B by the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W.

[0149] As Figure 12As depicted, the apparatus 1000 is of the reflective type (e.g., using a reflective patterning device). It should be noted that since most materials are absorptive in the EUV wavelength range, the patterning device may have a multilayer reflector including, for example, multiple stacks of molybdenum and silicon. In one example, the multilayer reflector has 40 layer pairs of molybdenum and silicon, with each layer having a thickness of a quarter wavelength. X-ray lithography can be utilized to generate smaller wavelengths. Since most materials are absorptive at EUV and x-ray wavelengths, a thin sheet of patterned absorptive material on the patterning device topography (e.g., a TaN absorber on top of the multilayer reflector) defines where features will be printed (positive resist) or where features will not be printed (negative resist).

[0150] The illuminator IL receives an extreme ultraviolet (EUV) radiation beam from the source collector module SO. Methods for generating EUV radiation include, but are not limited to, converting a material into a plasma state having at least one element such as xenon, lithium, or tin using one or more emission spectral lines in the EUV range. In one such method, often referred to as laser-produced plasma ("LPP"), a plasma can be generated by irradiating a fuel (such as a droplet, stream, or cluster of material having a spectral emission element) with a laser beam. The source collector module SO can be part of an EUV radiation system including a laser ( Figure 12 (not shown in the figure) that provides the laser beam for exciting the fuel. The resulting plasma emits output radiation, such as EUV radiation, which is collected using a radiation collector disposed in the source collector module. For example, when a CO2 laser is used to provide the laser beam for fuel excitation, the laser and the source collector module can be separate entities.

[0151] In these cases, the laser is not considered part of the lithographic apparatus, and the radiation beam is transmitted from the laser to the source collector module by means of a beam delivery system including, for example, suitable steering mirrors and / or beam expanders. In other cases, such as when the radiation source is a discharge-produced plasma EUV generator, often referred to as a DPP source, the source can be an integral part of the source collector module. In an embodiment, a DUV laser source can be used.

[0152] The illuminator IL may include an adjuster for adjusting the angular intensity distribution of the radiation beam. Generally, at least the outer radial range and / or the inner radial range of the intensity distribution in the pupil plane of the illuminator can be adjusted (commonly referred to as σ - outer and σ - inner, respectively). In addition, the illuminator IL may include various other components, such as faceted field mirror devices and faceted pupil mirror devices. The illuminator can be used to adjust the radiation beam to have a desired uniformity and intensity distribution in its cross-section.

[0153] The radiation beam B is incident on a patterning device (e.g., a mask) MA held on a support structure (e.g., a patterning device table) MT and is patterned by the patterning device. After reflection from the patterning device (e.g., a mask) MA, the radiation beam B passes through a projection system PS which focuses the beam onto a target portion C of a substrate W. By means of a second positioning device PW and a position sensor PS2 (e.g., an interferometric device, a linear encoder or a capacitive sensor), the substrate table WT can be accurately moved, for example in order to position different target portions C in the path of the radiation beam B. Similarly, a first positioning device PM and another position sensor PS1 can be used to accurately position the patterning device (e.g., a mask) MA relative to the path of the radiation beam B. Patterning device alignment marks M1, M2 and substrate alignment marks P1, P2 can be used to align the patterning device (e.g., a mask) MA with the substrate W.

[0154] The depicted apparatus 1000 can be used in at least one of the following modes:

[0155] In a step mode, the entire pattern imparted to the radiation beam is projected onto the target portion C in one go (i.e., a single static exposure) while keeping the support structure (e.g., the patterning device table) MT and the substrate table WT substantially stationary. Subsequently, the substrate table WT is offset in the X and / or Y direction so that different target portions C can be exposed.

[0156] In a scan mode, the support structure (e.g., the patterning device table) MT and the substrate table WT are scanned synchronously while the pattern imparted to the radiation beam is projected onto the target portion C (i.e., a single dynamic exposure). The speed and direction of the substrate table WT relative to the support structure (e.g., the patterning device table) MT can be determined by the (reduction) magnification and image inversion characteristics of the projection system PS.

[0157] In another mode, the support structure (e.g., the patterning device table) MT is kept substantially stationary so as to hold a programmable patterning device, and the substrate table WT is moved or scanned while the pattern imparted to the radiation beam is projected onto the target portion C. In this mode, a pulsed radiation source is typically used and the programmable patterning device is updated as required after each movement of the substrate table WT or between successive radiation pulses during the scan. This mode of operation can be readily applied to maskless lithography using a programmable patterning device such as a programmable mirror array of the type mentioned above.

[0158] Figure 13Device 1000 is shown in more detail, the device including a source collector module SO, an illumination system IL, and a projection system PS. The source collector module SO is constructed and arranged such that a vacuum environment can be maintained within the enclosure structure 220 of the source collector module SO. A plasma source for EUV radiation emission 210 can be formed by a discharge to generate a plasma. EUV radiation can be generated by a gas or vapor (e.g., Xe gas, Li vapor, or Sn vapor), where a very hot plasma 210 is generated to emit radiation in the EUV range of the electromagnetic spectrum. The very hot plasma 210 is generated by, for example, causing a discharge that at least partially ionizes the plasma. For efficient generation of the radiation, a partial pressure of, for example, 10 Pa of Xe, Li, Sn vapor, or any other suitable gas or vapor may be required. In an embodiment, an excited tin (Sn) plasma is provided to generate EUV radiation.

[0159] The radiation emitted by the hot plasma 210 is transferred from the source chamber 211 to the collector chamber 212 through an optional gas barrier or contaminant trap 230 (also referred to in some cases as a contaminant barrier or foil trap) located in or behind an opening in the source chamber 211. The contaminant trap 230 may include a channel structure. The contaminant trap 230 may also include a gas barrier, or a combination of a gas barrier and a channel structure. As is known in the art, the contaminant trap or contaminant barrier 230 further indicated herein includes at least a channel structure.

[0160] The collector chamber 211 may include a radiation collector CO that may be a so-called grazing incidence collector. The radiation collector CO has an upstream radiation collector side 251 and a downstream radiation collector side 252. The radiation traversing the collector CO can be reflected from the grating spectral filter 240 to be focused at a virtual source point IF along the optical axis indicated by the dotted line "O". The virtual source point IF is generally referred to as an intermediate focus, and the source collector module is arranged such that the intermediate focus IF is located at or near the opening 221 in the enclosure structure 220. The virtual source point IF is an image of the radiation emission plasma 210.

[0161] Subsequently, the radiation traverses the illumination system IL, which may include a faceted field mirror device 22 and a faceted pupil mirror device 24. The faceted field mirror device and the faceted pupil mirror device are arranged to provide a desired angular distribution of the radiation beam 21 at the patterning device MA, and a desired uniformity of the radiation intensity at the patterning device MA. After reflection of the radiation beam 21 at the patterning device MA held by the support structure MT, a patterned beam 26 is formed and imaged onto a substrate W held by a substrate stage WT through the projection system PS via reflection elements 28, 30.

[0162] Typically, there may be more elements in the illumination optical device unit IL and the projection system PS than those shown. Depending on the type of lithographic apparatus, a grating spectral filter 240 may optionally be present. Additionally, there may be more mirrors than the mirrors shown in the figure. For example, in the projection system PS, there may be 1 to 6 additional reflective elements more than the Figure 13 reflective elements shown.

[0163] As Figure 14 illustrated, the collector optical device CO is depicted as a nested collector having grazing-incidence reflectors 253, 254, and 255, merely as an example of a collector (or collector mirror). The grazing-incidence reflectors 253, 254, and 255 are arranged axially symmetrically about the optical axis O, and a collector optical device CO of this type can be used in combination with a discharge-produced plasma source often referred to as a DPP source.

[0164] Alternatively, the source collector module SO can be part of an LPP radiation system as shown in Figure 14 . A laser LA is arranged to deposit laser energy into a fuel such as xenon (Xe), tin (Sn), or lithium (Li), thereby generating a highly ionized plasma 210 having an electron temperature of several tens of eV. The high-energy radiation generated during the de-excitation and recombination of these ions is emitted from the plasma, collected by a near-normal-incidence collector optical device CO, and focused onto an opening 221 in the enclosure structure 220.

[0165] The embodiments can be further described using the following aspects:

[0166] 1. A method for training a machine learning model for a layout patterning process, the method comprising:

[0167] Generating a plurality of features from patterns in a pattern set;

[0168] Grouping the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features; and

[0169] Providing representative patterns from the plurality of separate groups to the machine learning model to train the machine learning model to predict a continuous transmission mask (CTM) map for optical proximity correction (OPC) for the layout patterning process.

[0170] 2. The method according to aspect 1, wherein the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information already included in the pattern set.

[0171] 3. The method according to aspect 1 or 2, wherein the OPC includes full-chip OPC for the wafer during the layout patterning process.

[0172] 4. The method according to any one of aspects 1 to 3, wherein the plurality of generated features include geometric features and lithography-aware features.

[0173] 5. The method according to any one of aspects 1 to 4, wherein grouping the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features includes using a machine learning clustering method to cluster unique patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features.

[0174] 6. A method for determining training patterns for a layout patterning process, the method comprising:

[0175] generating a plurality of features from patterns in a pattern set;

[0176] grouping the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features; and

[0177] selecting representative patterns from the plurality of separate groups to determine the training patterns.

[0178] 7. The method according to aspect 6, wherein the plurality of generated features include geometric features and lithography-aware features.

[0179] 8. The method according to aspect 7, wherein the geometric features include one or more of the following: target mask image, frequency map, pattern density map, or pattern occurrence rate of unique patterns in the pattern set.

[0180] 9. The method according to any one of aspects 7 or 8, wherein the lithography-aware features include one or more of the following: sub-resolution assist feature guidance map (SGM), diffraction order, or diffraction pattern of unique patterns in the pattern set.

[0181] 10. The method according to any one of aspects 6 to 9, wherein the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information that has been included in the pattern set.

[0182] 11. The method according to any one of aspects 6 to 10, wherein grouping the patterns in the pattern set into multiple groups based on the plurality of generated features is performed using unsupervised machine learning.

[0183] 12. The method according to any one of aspects 6 to 11, wherein grouping the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features includes clustering the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features.

[0184] 13. The method according to aspect 12, wherein the clustering includes a successive series of clustering steps performed using different features among the plurality of generated features for different clustering steps, the successive series of clustering steps forming subgroups of distinct patterns in the pattern set such that the representative pattern is selected from the subgroups to determine the training pattern.

[0185] 14. The method according to aspect 12 or 13, wherein the clustering includes a machine learning clustering method.

[0186] 15. The method according to aspect 13 or 14, wherein the successive series of clustering steps includes a cross-validation step performed using a given feature for a given step, the cross-validation including adjusting which patterns are included in a given subgroup.

[0187] 16. The method according to any one of aspects 6 to 15, wherein selecting a representative pattern from the separate groups to determine the training pattern includes selecting a target number of representative patterns.

[0188] 17. The method according to aspect 16, wherein the target number of representative patterns is determined based on a stopping criterion configured to contribute to a change in the training pattern.

[0189] 18. The method according to aspect 17, the method further comprising: determining a change amount of the training pattern.

[0190] 19. The method according to aspect 18, wherein the stopping criterion is further configured to ensure that the change amount of the training pattern breaks through a change amount threshold.

[0191] 20. The method according to aspect 19, wherein the target number of representative patterns is randomly selected from the plurality of separate groups.

[0192] 21. The method according to aspect 20, wherein the target number of representative patterns is re-randomly selected in response to the change amount of the training pattern not breaking through the change amount threshold.

[0193] 22. The method according to any one of aspects 6 to 21, wherein selecting a representative pattern from the plurality of individual groups to determine the training pattern includes selecting the most central pattern from each individual group, the most central pattern being closest to the centroid of the specified feature space for the individual group relative to other patterns in the plurality of individual groups.

[0194] 23. The method according to aspect 22, wherein the specified feature space is a target mask image feature space, a frequency mapping feature space, a pattern density mapping feature space, a pattern occurrence feature space, an SGM feature space, a diffraction order feature space, or a diffraction pattern feature space.

[0195] 24. The method according to any one of aspects 6 to 23, further comprising: providing the training pattern to a deep convolutional neural network to train the deep convolutional neural network.

[0196] 25. The method according to aspect 24, further comprising: performing optical proximity effect correction using the trained deep convolutional neural network as part of a wafer patterning process.

[0197] 26. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions when executed by a computer implementing the method according to any one of aspects 1 to 25.

[0198] The concepts disclosed herein can model any general imaging system for imaging sub-wavelength features either analogously or mathematically and can be used in particular with emerging imaging technologies capable of generating increasingly shorter wavelengths. Emerging technologies already in use include extreme ultraviolet (EUV), DUV lithography capable of generating a 193 nm wavelength by using an ArF laser and even capable of generating a 157 nm wavelength by using a fluorine laser. In addition, EUV lithography can generate wavelengths in the range of 5 nm to 20 nm by using a synchrotron or by using high-energy electrons to strike a material (solid or plasma) in order to generate photons in this range.

[0199] Although the concepts disclosed herein can be used to image on a substrate such as a silicon wafer, it should be understood that the disclosed concepts can be applicable with any type of lithographic imaging system, e.g., a lithographic imaging system for imaging on a substrate other than a silicon wafer.

[0200] The foregoing description is intended to be exemplary and not restrictive. Accordingly, those skilled in the art will appreciate that modifications can be made to the described invention without departing from the scope of the claims set forth below.

Claims

1. A non - transitory computer - readable medium having instructions recorded thereon, which when executed by a computer, implement training a machine - learning model for a layout patterning process, including: Generate a plurality of features from patterns in a pattern set; Group the patterns in the pattern set into a plurality of separate groups based on the similarity of the plurality of generated features; And Provide representative patterns from the plurality of separate groups to the machine learning model to train the machine learning model to predict a continuous transmission mask (CTM) map for optical proximity correction (OPC) for the layout patterning process.

2. The computer - readable medium according to claim 1, wherein the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information already included in the pattern set.

3. The computer - readable medium according to claim 1 or 2, wherein the OPC includes full - chip OPC for a wafer in the layout patterning process.

4. The computer - readable medium according to claim 1, wherein the plurality of generated features include geometric features and lithography - aware features.

5. The computer - readable medium according to claim 1, wherein grouping the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features includes using a machine - learning clustering method to cluster the unique patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features.

6. The computer - readable medium according to claim 4, wherein the geometric features include one or more of the following: a target mask image, a frequency map, a pattern density map, or a pattern occurrence rate of a unique pattern in the pattern set.

7. The computer - readable medium according to claim 4, wherein the lithography - aware features include one or more of the following: a sub - resolution assist feature guidance map (SGM), a diffraction order, or a diffraction pattern of a unique pattern in the pattern set.

8. The computer - readable medium according to claim 1, wherein the plurality of features generated from the patterns in the pattern set are information other than geometric information and / or vertex information already included in the pattern set.

9. The computer - readable medium according to claim 1, wherein grouping the patterns in the pattern set into multiple groups based on the plurality of generated features is performed using unsupervised machine learning.

10. The computer - readable medium according to claim 1, wherein grouping the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features includes clustering the patterns in the pattern set into multiple separate groups based on the similarity of the plurality of generated features.

11. The computer-readable medium according to claim 10, wherein the clustering includes a consecutive series of clustering steps performed using different features among a plurality of generated features for different clustering steps, the consecutive series of clustering steps forming a subgroup of unique patterns in the set of patterns, such that the representative pattern is selected from the subgroup to determine the training pattern.

12. The computer-readable medium according to claim 10 or 11, wherein the clustering includes a machine learning clustering method.

13. The computer-readable medium according to claim 11, wherein the consecutive series of clustering steps includes a cross-validation step performed using a given feature for a given step, the cross-validation including adjusting which patterns are included in a given subgroup.

14. The computer-readable medium according to claim 11, the instructions further implementing, when executed by a computer, selecting a representative pattern from the separate groups to determine the training pattern, selecting a representative pattern from the separate groups to determine the training pattern including selecting a target number of representative patterns, wherein the target number of representative patterns is determined based on a stopping criterion configured to contribute to a change in the training pattern, and the instructions further implementing, when executed by a computer, determining a change amount of the training pattern; and wherein the target number of representative patterns is re-randomly selected in response to the change amount of the training pattern not exceeding a change amount threshold.

15. The computer-readable medium according to claim 14, wherein selecting a representative pattern from the plurality of separate groups to determine the training pattern includes selecting a central pattern that is closer to the centroid of a specified feature space for a separate group relative to other patterns in the plurality of separate groups, wherein the specified feature space is a target mask image feature space, a frequency mapping feature space, a pattern density mapping feature space, a pattern occurrence feature space, an SGM feature space, a diffraction order feature space, or a diffraction pattern feature space.

Citation Information

Patent Citations

  • System and method for creating a focus-exposure model of a lithography process

    US20070031745A1

  • Method for identifying and using process window signature patterns for lithography process control

    US20070050749A1

  • System and method for model-based sub-resolution assist feature generation

    US20080301620A1

  • Multivariable solver for optical proximity correction

    US20080309897A1

  • Method of extracting data and recommending and generating visual displays

    US20090157630A1