Method for determining a pattern in a patterning process

By using the deep learning convolutional neural network framework to train the patterning process model in lithography technology, the problem of difficult pattern reproduction in low k1 lithography is solved, and the accuracy and efficiency of the lithography process are achieved is improved.

CN113892059BActive Publication Date: 2025-05-06ASML NETHERLANDS BV
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202080024772.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-20
Filing Date
2020-03-05
Publication Date
2025-05-06
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

When existing lithography technology manufactures functional components with a size of less than 100nm, it is difficult to effectively solve the complexity and fine adjustment steps requirements in low k1 lithography, resulting in difficulty in reproducing patterns.

Method used

The patterning process model is trained using the deep learning convolutional neural network framework. Through the first model and machine learning model working collaboratively, the pattern characteristics during the patterning process are predicted and used to determine optical proximity effect correction and etch deviation.

Benefits of technology

It improves the accuracy of measurement and prediction of pattern characteristics, saves measurement time and resources, and improves the accuracy and efficiency of the lithography process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113892059B_ABST
    Figure CN113892059B_ABST
Patent Text Reader

Abstract

A method for training a patterning process model, the patterning process model being configured to predict a pattern to be formed on a patterning process. The method involves: obtaining image data associated with a desired pattern, a measured pattern of a substrate, a first model including a first set of parameters, and a machine learning model including a second set of parameters; and iteratively determining values ​​of the first set of parameters and the second set of parameters to train the patterning process model. Iterations involve: executing the first model and the machine learning model using the image data to collaboratively predict a printed pattern of the substrate; and modifying the values ​​of the first set of parameters and the second set of parameters so that a difference between the measured pattern and a predicted pattern of the patterning process model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Application No. 62 / 823,029, filed on March 25, 2019, and U.S. Application No. 62 / 951,097, filed on December 20, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The description herein relates to lithographic apparatus and processes, and more particularly, to a tool for training a patterning process model and using the trained model to determine a pattern to be printed on a substrate during the patterning process. Background Art

[0004] A lithographic projection apparatus may be used, for example, in the manufacture of integrated circuits (ICs). In this case, a patterning device (e.g., a mask) may contain or provide a circuit pattern (a "design layout") corresponding to a single layer of the IC, and this circuit pattern may be transferred to a target portion (e.g., comprising one or more dies) on a substrate (e.g., a silicon wafer) that has been coated with a layer of radiation-sensitive material ("resist"), for example, by irradiating the target portion through the circuit pattern on the patterning device. Typically, a single substrate comprises a plurality of adjacent target portions to which the circuit pattern is transferred successively, one target portion at a time, via the lithographic projection apparatus. In one type of lithographic projection apparatus, the circuit pattern on the entire patterning device is transferred to one target portion at once, and such an apparatus is typically referred to as a wafer stepper. In an alternative apparatus, typically referred to as a step-and-scan apparatus, the projection beam is scanned over the patterning device in a given reference direction (the "scanning" direction) while the substrate is synchronously moved in a direction parallel or antiparallel to the reference direction. Different portions of the circuit pattern on the patterning device are progressively transferred to one target portion. Since typically the lithographic projection apparatus will have a magnification factor M (typically <1), the speed F at which the substrate is moved will be M times the speed at which the projection beam scans the patterning device. More information on the lithographic apparatus described herein can be found, for example, in U.S. Patent No. 6,046,792, which is incorporated herein by reference.

[0005] Before the circuit pattern is transferred from the pattern forming device to the substrate, the substrate may undergo various processes, such as priming, resist coating and soft baking. After exposure, the substrate may undergo other processes, such as post-exposure baking (PEB), development, hard baking, and measurement / inspection of the circuit pattern after transfer. This array of processes is used as the basis for manufacturing a single layer of a device such as an IC. The substrate may then undergo various processes, such as etching, ion implantation (doping), metallization, oxidation, chemical mechanical polishing, etc., all of which are expected to eventually complete a single layer of the device. If several layers are needed in the device, the entire process or its variation is repeated for each layer. Eventually, there will be a device in each target portion on the substrate. These devices are then separated from each other by techniques such as dicing or sawing, so that a separate device can be mounted on a carrier, connected to a pin, etc.

[0006] As noted, microlithography is a core step in the fabrication of integrated circuits, where patterns formed on a substrate define the functional elements of the IC, such as microprocessors, memory chips, etc. Similar lithographic techniques are also used to form flat panel displays, microelectromechanical systems (MEMS), and other devices.

[0007] As semiconductor manufacturing processes continue to advance, the size of functional elements has been continuously reduced over the decades, while the number of functional elements (such as transistors) per device has been steadily increasing following a trend generally referred to as "Moore's Law." In the case of the prior art, the layers of the device are manufactured by using a lithographic projection device that projects a design layout onto a substrate using illumination from a deep ultraviolet illumination source, thereby producing individual functional elements having a size well below 100 nm, i.e., the size of the functional element is less than half the wavelength of the radiation of the illumination source (e.g., a 193 nm illumination source).

[0008] The process of printing features with dimensions smaller than the classical resolution limit of a lithographic projection apparatus is often referred to as low-k1 lithography, and is based on the resolution formula CD=k1×λ / NA, where λ is the wavelength of radiation employed (currently 248 nm or 193 nm in most cases), NA is the numerical aperture of the projection optics in the lithographic projection apparatus, CD is the "critical dimension" (usually the smallest feature size printed), and k1 is an empirical resolution factor. In general, the smaller k1 is, the more difficult it becomes to replicate a pattern on a wafer (similar to the shape and dimensions designed by a circuit designer to achieve specific electrical functions and performance). To overcome these difficulties, complex fine-tuning steps are applied to the lithographic projection apparatus and the design layout. These include, for example, but are not limited to, optimization of NA and optical coherence settings, customized illumination schemes, use of phase-shifted patterning devices, optical proximity correction (OPC, sometimes referred to as "optical and process correction") in the design layout, or other methods often defined as "resolution enhancement techniques (RET)". The term "projection optical element" as used herein should be broadly interpreted to include various types of optical systems, including, for example, refractive optical devices, reflective optical devices, apertures, and refractive-reflective optical devices. The term "projection optical element" may also collectively or individually include components that operate according to any of these design types for guiding, shaping, or controlling a radiation projection beam. The term "projection optical element" may include any optical component in a lithographic projection device, regardless of where the optical component is located on the optical path of the lithographic projection device. Projection optical elements may include optical components for shaping, adjusting, and / or projecting radiation from a source before the radiation passes through a pattern forming device, and / or optical components for shaping, adjusting, and / or projecting radiation after the radiation passes through a pattern forming device. Projection optical elements typically do not include a source and a pattern forming device. Summary of the invention

[0009] The present disclosure provides multiple improvements in the area of ​​computational lithography. In particular, a training patterning process model includes a first model and a machine learning model in a framework such as a deep learning convolutional neural network. The trained model can also be used to determine the pattern to be printed on the substrate during the patterning process. Advantages of the present disclosure include, but are not limited to, providing an improved way to measure the characteristics of the pattern to be printed on the substrate, and making accurate predictions of the measurement images, thereby saving measurement time and resources.

[0010] According to an embodiment, a method for training a patterning process model is provided, the patterning process model being configured to predict a pattern to be formed on a patterning process. The method involves: obtaining (i) image data associated with a desired pattern, (ii) a measured pattern of a substrate, the measured pattern being associated with the desired pattern, (iii) a first model associated with an aspect of the patterning process, the first model comprising a first parameter set, and (iv) a machine learning model associated with another aspect of the patterning process, the machine learning model comprising a second parameter set; and iteratively determining values ​​of the first parameter set and the second parameter set to train the patterning process model. Iteration involves: using the image data to execute the first model and the machine learning model to collaboratively predict a printed pattern of the substrate; and modifying the values ​​of the first parameter set and the second parameter set so that the difference between the measured pattern and the predicted pattern of the patterning process model is reduced.

[0011] In an embodiment, the first model and the machine learning model are configured and trained in a deep convolutional neural network framework.

[0012] In an embodiment, the training involves: predicting the printed pattern by forward propagation of the output of the first model and the machine learning model; determining the difference between the measured pattern and the predicted pattern of the patterning process model; determining the differential of the difference relative to a first parameter set and a second parameter set; and determining the values ​​of the first parameter set and the second parameter set by backward propagation of the output of the first model and the machine learning model based on the differential of the difference.

[0013] In addition, according to an embodiment, a method for determining optical proximity effect correction for a patterning process is provided, the method comprising: obtaining image data associated with a desired pattern; using the image data to execute a trained patterning process model to predict a pattern to be printed on a substrate; and using the predicted pattern to be printed on the substrate undergoing the patterning process to determine optical proximity effect correction and / or defects.

[0014] In addition, according to an embodiment, a method for training a machine learning model is provided, wherein the machine learning model is configured to determine an etching deviation associated with an etching process, the method comprising: obtaining (i) resist pattern data associated with a target pattern to be printed on a substrate, (ii) physical effect data characterizing the effect of the etching process on the target pattern, and (iii) a measured deviation between the resist pattern and an etching pattern formed on the printed substrate; and training the machine learning model based on the resist pattern data, the physical effect data, and the measured deviation to reduce the difference between the measured deviation and the predicted etching deviation.

[0015] In addition, according to an embodiment, a system for determining an etching deviation related to an etching process is provided. The system includes: a semiconductor process device; and a processor. The processor is configured to: determine physical effect data characterizing the effect of the etching process on a substrate by executing a physical effect model; use the resist pattern and the physical effect data as input to execute a trained machine learning model to determine the etching deviation; and control the semiconductor device or the etching process based on the etching deviation.

[0016] In addition, according to an embodiment, a method for calibrating a process model is provided, the process model being configured to generate a simulated profile. The method comprises: obtaining (i) measurement data at a plurality of measurement locations on a pattern, and (ii) a profile constraint specified based on the measurement data; and calibrating the process model by adjusting values ​​of model parameters of the process model until the simulated profile satisfies the profile constraint.

[0017] In addition, according to an embodiment, a method for calibrating a process model is provided, the process model being configured to predict an image of a target pattern. The method includes: obtaining (i) a reference image associated with the target pattern, and (ii) a gradient constraint specified relative to the reference image; and calibrating the process model so that the process model generates a simulated image, the simulated image (i) minimizing an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfying the gradient constraint.

[0018] In addition, according to an embodiment, a system for calibrating a process model is provided, the process model being configured to generate a simulated profile. The system comprises: a metrology tool configured to obtain measurement data at a plurality of measurement locations on a pattern; and a processor. The processor is configured to: calibrate the process model by adjusting values ​​of model parameters of the process model until the simulated profile satisfies the profile constraint, the profile constraint being based on the measurement data.

[0019] In addition, according to an embodiment, a system for calibrating a process model is provided, wherein the process model is configured to predict an image of a target pattern. The system includes: a measurement tool, wherein the measurement tool is configured to obtain a reference image associated with the target pattern; and a processor. The processor is configured to: calibrate the process model so that the process model generates a simulated image, wherein the simulated image (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfies a gradient constraint associated with the reference image.

[0020] In addition, according to an embodiment, a non-transitory computer-readable medium comprising instructions is provided, which, when executed by one or more processors, causes operations including the following: obtaining (i) resist pattern data associated with a target pattern to be printed on a substrate, (ii) physical effect data characterizing the effect of an etching process on the target pattern, and (iii) measured deviations between the resist pattern and an etched pattern formed on the printed substrate; and training the machine learning model based on the resist pattern data, the physical effect data, and the measured deviations to reduce the difference between the measured deviations and the predicted etching deviations.

[0021] Additionally, according to an embodiment, a non-transitory computer-readable medium comprising instructions is provided, which when executed by one or more processors results in operations comprising: obtaining (i) measurement data at a plurality of measurement locations on a pattern, and (ii) contour constraints specified based on the measurement data; and calibrating a process model by adjusting values ​​of model parameters of the process model until the simulated contour satisfies the contour constraints.

[0022] In addition, according to an embodiment, a non-transitory computer-readable medium including instructions is provided, which, when executed by one or more processors, causes operations including: obtaining (i) a reference image associated with a target pattern, and (ii) a gradient constraint specified relative to the reference image; and calibrating a process model so that the process model produces a simulated image that (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfies the gradient constraint. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0024] Figure 1 is a block diagram of various subsystems of a lithography system according to an embodiment.

[0025] Figure 2 According to the embodiment Figure 1 Block diagram of the simulation model corresponding to the subsystem in .

[0026] Figure 3 is a flow chart of a method for training a patterning process model configured to predict a pattern to be formed in a patterning process according to an embodiment.

[0027] Figure 4A , Figure 4B and Figure 4C An example configuration of a patterning process model including a first model and a second model (eg, a machine learning model) according to an embodiment is illustrated.

[0028] Figure 5 According to an embodiment, Figure 3 Flow chart of a method for determining optical proximity effect correction of a patterning process based on a predicted pattern of a trained patterning process model.

[0029] Figure 6 is a flow chart of a method for training a machine learning model to determine etch bias associated with an etch process according to an embodiment.

[0030] Figure 7 is an example resist pattern according to an embodiment.

[0031] Figure 8 is example physical effect data based on a resist pattern according to an embodiment.

[0032] Fig. 9 are example acid and base concentrations in the resist according to the embodiment.

[0033] Fig.10 is example physical effect data based on acid-base distribution according to an embodiment.

[0034] Fig.11 is an example etch bias applied to the developed image (ADI) profile according to an embodiment.

[0035] Fig.12 is a flow chart of a method for calibrating a process model based on physical constraints related to a contour shape (or silhouette) of a pattern, according to an embodiment.

[0036] Fig.13A FIG. 1 illustrates the satisfaction and Fig.11 Example model output with relevant physical constraints.

[0037] Fig. 13B Illustration of the unsatisfactory Fig.11 Example model output for physical constraints.

[0038] Fig.14 is a flow chart of another method for calibrating a process model based on other physical constraints according to an embodiment.

[0039] Fig.15A FIG. 1 illustrates a reference intensity distribution of an aerial image or a resist image according to an embodiment.

[0040] Fig. 15B Schematic diagram of the embodiment according to the embodiment Fig.15A The intensity distribution associated with the model of physical constraints.

[0041] Fig. 15C FIG. 1 shows a diagram of a method according to an embodiment and a method not satisfying the embodiment. Fig.15AThe intensity distribution associated with the model of physical constraints.

[0042] Fig.16 An embodiment of a scanning electron microscope (SEM) according to an embodiment is schematically depicted.

[0043] Fig.17 Embodiments of electron beam inspection apparatus according to embodiments are schematically depicted.

[0044] Fig.18 is a flow chart illustrating aspects of an example method of joint optimization according to an embodiment.

[0045] Fig.19 An embodiment of another optimization method according to an embodiment is shown.

[0046] Fig. 20A , Fig. 20B and Fig.21 Example flow charts illustrating various optimization processes according to embodiments.

[0047] Fig. 22 is a block diagram of an example computer system according to an embodiment.

[0048] Fig.23 is a schematic diagram of a lithographic projection apparatus according to an embodiment.

[0049] Fig.24 is a schematic diagram of another lithography projection apparatus according to an embodiment.

[0050] Fig.25 According to the embodiment Fig.24 A more detailed view of the devices in .

[0051] Fig.26 According to the embodiment Fig.24 and Fig.25 A more detailed view of the source collector module SO of the device.

[0052] Embodiments will now be described in detail with reference to the accompanying drawings, which are provided as illustrative examples so that those skilled in the art can practice the embodiments. It is noteworthy that the following figures and examples are not intended to limit the scope to a single embodiment, but other embodiments are possible by means of the interchange of some or all of the elements described or illustrated. Wherever convenient, the same reference numerals will be used throughout the drawings to refer to the same or similar parts. In the case where certain elements of these embodiments can be implemented partially or completely using known components, only those parts of these known components that are necessary for understanding the embodiments will be described, and detailed descriptions of other parts of these known components will be omitted so as not to make the description of the embodiments unclear. In this specification, embodiments showing a single component should not be considered restrictive; rather, unless otherwise explicitly stated in the present invention, the scope is expected to cover other embodiments including multiple identical components, and vice versa. In addition, the applicant does not intend to attribute any term in this specification or claims to an uncommon or specific meaning unless so explicitly stated. In addition, the scope covers the current and future known equivalents of the components mentioned in the present invention by means of illustration. DETAILED DESCRIPTION

[0053] Although the embodiments are specifically referred to herein as being used to manufacture ICs, it should be clearly understood that the description herein may have many other possible applications. For example, it may be used to manufacture integrated optical systems, guidance and detection patterns of magnetic domain memories, liquid crystal display panels, thin film magnetic heads, etc. It should be understood by those skilled in the art that in the case of such alternative applications, any term "reticle", "wafer" or "die" used in this context may be considered to be interchangeable with the more general term "mask", "substrate" or "target portion".

[0054] In this document, the terms "radiation" and "beam" are used to include various types of electromagnetic radiation, including ultraviolet radiation (e.g., having a wavelength of 365, 248, 193, 157 or 126 nm) and EUV (extreme ultraviolet radiation, e.g., having a wavelength in the range of 5-20 nm).

[0055] As used herein, the terms "optimize" and "optimize" mean adjusting the lithography projection equipment so that the results and / or process of lithography have more ideal characteristics, such as higher projection accuracy of the design layout on the substrate, a larger process window, etc.

[0056] Furthermore, the lithographic projection apparatus may be of a type having two or more substrate tables (and / or two or more patterning device tables). In such a "multi-stage" apparatus, the additional tables may be used in parallel, or preparatory steps may be performed on one or more tables while one or more other tables are being used for exposure. A dual stage lithographic projection apparatus is described, for example, in U.S. Pat. No. 5,969,441, which is incorporated herein by reference.

[0057] The pattern forming device mentioned above includes or can form a design layout. The design layout can be generated using a CAD (computer-aided design) program, which is usually referred to as EDA (electronic design automation). Most CAD programs follow a set of predetermined design rules for generating functional design layout / pattern forming devices. These rules are set by processing and design constraints. For example, the design rules define the spacing tolerance between circuit devices (such as gates, capacitors, etc.) or interconnects, so as to ensure that the circuit devices or lines do not interact in an undesirable manner. The design rule restrictions are typically referred to as "critical dimensions" (CD). The critical dimension of a circuit can be defined as the minimum width of a line or hole or the minimum spacing between two lines or two holes. Therefore, CD determines the overall size and density of the designed circuit. Of course, one of the goals in integrated circuit manufacturing is to faithfully reproduce the original circuit design on a substrate (via a pattern forming device).

[0058] The terms "mask" or "patterning device" used in this context should be broadly interpreted to refer to a general patterning device that can be used to impart an incident radiation beam with a patterned cross-section corresponding to the pattern to be created in a target portion of the substrate; the term "light valve" may also be used in this context. In addition to conventional masks (transmissive or reflective; binary, phase-shift, hybrid, etc.), other examples of patterning devices include:

[0059] -Programmable mirror array. An example of such a device is a matrix-addressable surface having a viscoelastic control layer and a reflective surface. The basic principle on which such a device is based is that (for example) addressed areas of the reflective surface reflect incident radiation as diffracted radiation, while unaddressed areas reflect incident radiation as non-diffracted radiation. Using suitable filters, the non-diffracted radiation can be filtered out from the reflected beam, so that only diffracted radiation remains; in this way, the beam is patterned according to the addressing pattern of the matrix-addressable surface. The required matrix addressing can be performed by using suitable electronic devices. More information about such mirror arrays can be found, for example, in U.S. Patent Nos. 5,296,891 and 5,523,193, which are incorporated herein by reference.

[0060] - Programmable LCD array. An example of such a configuration is given in US Pat. No. 5,229,872, which is incorporated herein by reference.

[0061] As a brief introduction, Figure 1 An exemplary lithographic projection apparatus 10A is shown. The main components are: a radiation source 12A, which may be a deep ultraviolet excimer laser source or other types of sources including extreme ultraviolet (EUV) sources; a source (which, as discussed above, is not required by the lithographic projection apparatus itself); an illumination optic, which defines a partial coherence (labeled σ) and may include optics 14A, 16Aa, and 16Ab, which shapes the radiation from source 12A; a patterning device 14A; and a transmission optic 16Ac, which projects an image of the patterning device pattern onto a substrate plane 22A. An adjustable filter or aperture 20A at the pupil plane of the projection optics may limit the range of beam angles impinging on the substrate plane 22A, with the largest possible angle defining the numerical aperture NA=sin(θmax) of the projection optics.

[0062] In the optimization process of the system, the quality factor of the system can be expressed as a cost function. The optimization process is reduced to the process of finding a set of system parameters (design variables) that minimize the cost function. The cost function can have any suitable form depending on the goal of the optimization. For example, the cost function can be the weighted root mean square (RMS) of the deviation of specific characteristics (evaluation points) of the system from the expected values ​​(e.g., ideal values) of these characteristics; the cost function can also be the maximum value (i.e., the worst deviation) of these deviations. The term "evaluation point" here should be broadly interpreted to include any characteristic of the system. The design variables of the system can be limited to a limited range and / or be interdependent due to the practicality of the implementation of the system. In the case of a lithographic projection device, these constraints are generally related to the physical properties and characteristics of the hardware (such as an adjustable range) and / or the manufacturability design rules of the pattern forming device, and the evaluation points can include physical points on the resist image on the substrate and non-physical characteristics such as dose and focal length.

[0063] In a lithographic projection apparatus, a source provides illumination (i.e., light); the projection optical element guides and shapes the illumination through the pattern forming device and onto the substrate. The term "projection optical element" is broadly defined here to include any optical component that can change the wavefront of the radiation beam. For example, the projection optical element may include at least some of components 14A, 16Aa, 16Ab, and 16Ac. The aerial image (AI) is the radiation intensity distribution at the substrate level. The resist layer on the substrate is exposed, and the aerial image is transferred to the resist layer as a potential "resist image" (RI) therein. The resist image (RI) can be defined as the spatial distribution of the solubility of the resist in the resist layer. Resist models can be used to calculate resist images from aerial images, examples of which can be found in commonly assigned U.S. Patent Application No. 12 / 315,849, the disclosure of which is incorporated herein by reference in its entirety. The resist model is only related to the properties of the resist layer (e.g., the effects of chemical processes occurring during exposure, PEB, and development). The optical properties of a lithographic projection apparatus (e.g., properties of the source, patterning device, and projection optics) define the aerial image. Because the patterning device used in a lithographic projection apparatus can be varied, it is desirable to separate the optical properties of the patterning device from the optical properties of the rest of the lithographic projection apparatus, including at least the source and projection optics.

[0064] exist Figure 2 An exemplary flow chart of simulated lithography in a lithographic projection apparatus is shown in FIG. 3 . A source model 31 represents the optical properties of a source (including radiation intensity distribution and / or phase distribution). A projection optical model 32 represents the optical properties of a projection optical element (including changes in radiation intensity distribution and / or phase distribution caused by the projection optical element). A design layout model 35 represents the optical properties of a design layout 33 (including changes in radiation intensity distribution and / or phase distribution caused by a given design layout 33), which is a representation of the arrangement of features on or formed by a pattern forming device. An aerial image 36 can be simulated by the design layout model 35, the projection optical element model 32, and the design layout model 35. A resist image 38 can be simulated by the aerial image 36 using a resist model 37. The simulation of lithography can, for example, predict the profile and CD in the resist image.

[0065] More specifically, it is noted that the source model 31 can represent the optical properties of the source, including but not limited to the NA-sigma (σ) setting and any specific illumination source shape (e.g., off-axis radiation sources such as annular, quadrupole, and dipole, etc.). The projection optical element model 32 can represent the optical properties of the projection optical element, including aberrations, deformations, refractive indices, physical sizes, physical dimensions, etc. The design layout model 35 can also represent the physical properties of the physical pattern forming device, as described, for example, in U.S. Patent No. 7,587,704, the entire contents of which are incorporated herein by reference. The goal of the simulation is to accurately predict, for example, the positioning of edges, spatial image intensity slopes, and CD, which can then be compared to the desired design. The desired design is typically defined as a pre-OPC design layout, which can be provided in a standard digital file format (such as GDSII or OASIS) or other file format.

[0066] One or more portions, referred to as "snippets," may be identified from this design layout. In an embodiment, a set of snippets are extracted that represent complex patterns in the design layout (typically about 50 to 1000 snippets, although any number of snippets may be used). As will be appreciated by those skilled in the art, these patterns or snippets represent small portions of the design (i.e., circuits, cells, or patterns), and in particular snippets represent small portions that require special attention and / or verification. Alternatively, a snippet may be a portion of the design layout or may be similar to or have similar behavior to a portion of the design layout, where critical features are identified through experience (including snippets provided by customers), through trial and error, or by running full-chip simulations. A snippet typically contains one or more test patterns or metrology patterns.

[0067] An initial larger set of clips that require specific image optimization may be provided a priori by the customer based on known critical feature areas in the design layout. Alternatively, in another embodiment, the initial larger set of clips may be extracted from the entire design layout using some type of automated (such as machine vision) or manual algorithm that identifies critical feature areas.

[0068] Random variations in patterning processes (e.g., resist processes) potentially limit lithography (e.g., EUV lithography), e.g., in terms of shrink potential of features and exposure-dose specifications, which in turn affects the wafer throughput of the patterning process. In embodiments, random variations in the resist layer may manifest as / embody random failures or random faults, such as closed holes or grooves or folds / broken lines. Such resist-related random variations more affect and limit successful high-volume manufacturing, i.e., high-volume manufacturing (HVM), than, for example, random CD variations, which are traditionally of interest as an indicator for measuring and adjusting the performance of the patterning process.

[0069] In a patterning process (e.g., photolithography, e-beam lithography, etc.), an energy-sensitive material (e.g., resist) deposited on a substrate is subjected to a pattern transfer step (e.g., exposure). After the pattern transfer step, multiple post-steps such as resist baking and subtractive processes such as resist development, etching, etc. are applied. These post-exposure steps or processes exert multiple effects, resulting in the formation of structures having dimensions different from the target dimensions through the patterned layer or the etched substrate.

[0070] In computational lithography, patterning process models related to different aspects of the patterning process (e.g. Figure 2 ), such as mask models, optical models, resist models, post-exposure models, etc., can be used to predict the pattern to be printed on the substrate. The patterning process model, when properly calibrated (e.g., using measurement data associated with printed wafers), can produce accurate predictions of the pattern dimensions output from the patterning process. For example, the patterning process model of the post-exposure process is calibrated based on empirical measurements. The calibration process involves exposing a test substrate by varying different process parameters (e.g., dose, focus, etc.), measuring the resulting critical dimension printed pattern after the post-exposure process, and calibrating the patterning process model to the measurement results. In practice, fast and accurate models are used to improve device performance (e.g., yield), enhance process windows, patterning options, and / or increase the complexity of designed patterns.

[0071] The patterning process is a complex process, and not all aspects can be modeled based on the physical / chemical properties involved in the patterning process. For example, some effects of the post-exposure process are easy to understand and can be modeled using mathematical expressions of physical terms (e.g., parameters associated with the resist process) that describe the physical / chemical properties of the process. For example, acid-base diffusion after exposure can be modeled by a Gaussian filter on the aerial image. In an embodiment, some physical terms (e.g., associated with dose, focus, intensity, pupil, etc.) are associated with tunable parameters (e.g., tunable knobs) of the lithographic apparatus and can be tuned by the tunable parameters, thereby achieving real-time control of the patterning process via the tunable knob. In an embodiment, some physical terms may not be directly tunable via the tunable knob, but can explain the physical / chemical properties of the process (e.g., aerial image formation, resist image formation, etc.). For example, the resist model includes a Gaussian filter on the aerial image for modeling acid-base diffusion in the resist after the exposure. Such a mean square deviation term is typically not tunable via a tunable knob. Even so, the value of such a physical term (eg, mean square deviation) may be determined based on empirical equations, or physics-based equations that model the effects of the process (eg, resist).

[0072] However, several aspects or effects of other post-exposure effects are not well understood and therefore difficult to model using equations based on physical / chemical properties. In such cases, in the present disclosure, machine learning models such as deep convolutional neural networks (CNNs) are trained to model less understood aspects of the (e.g., post-exposure) patterning process. The process models of the present disclosure alleviate the need to understand, for example, post-exposure processes for model development and remove the reliance on the engineer's personal experience for model tuning. In embodiments of the present disclosure, the trained CNNs yield model accuracy comparable to or superior to that produced using conventional techniques.

[0073] Figure 3 It is a flow chart of a method for training a patterning process model, which is configured to predict patterns that will be formed in a patterning process. As previously mentioned, some aspects of the patterning process are well understood and can be modeled using mathematical expressions that employ physical terms that are configured to accurately describe the physical effects of the patterning process. In addition, there are some aspects that may not be accurately modeled using physical terms. The present method employs two different models: a first model that is configured to describe known aspects via physical terms (e.g., optics-related parameters, resist-related parameters of the patterning process), and a second model (i.e., a machine learning model) that is configured to describe aspects that are not yet fully understood (i.e., in terms of physical / chemical properties).

[0074] There are several advantages to using a hybrid model (i.e., a first model and a second model) to collaboratively predict the pattern of the patterning process. For example, the calculation of the physical terms is relatively simple, and models that employ such physical terms are less susceptible to overfitting. Incorporating physical terms in conjunction with, for example, a CNN model reduces the complexity of the CNN by several times, reduces the risk of overfitting, and improves the run time of the patterning process simulation. On the other hand, existing models employ CNNs to model both known and unknown effects, which can result in unnecessarily complex CNN models that tend to overfit and have slow run times.

[0075] In process P301, method 300 involves obtaining (i) image data 302 associated with a desired pattern, (ii) a measured pattern 304 of the substrate, the measured pattern 304 associated with the desired pattern, (iii) a first model associated with one aspect of the patterning process (e.g., whose effects can be accurately modeled by equations based on physical / chemical properties), the first model comprising a first set of parameters 307, and (iv) a machine learning model associated with another aspect of the patterning process (e.g., whose effects may not be accurately modeled by equations based on physical / chemical properties), the machine learning model comprising a second set of parameters 308.

[0076] In an embodiment, the image data 302 generally refers to any input to the patterning process model, which is configured to predict the effect of an aspect of the patterning process or the final pattern to be printed on the substrate. In an embodiment, the image data 302 is an aerial image, a mask image, a resist image, or other output related to one or more aspects of the patterning process. In an embodiment, the acquisition of the aerial image, mask image, resist image, etc. includes simulating the patterning process, such as Figure 2 As discussed in .

[0077] In an embodiment, the patterning process model includes a first model coupled to a second model (e.g., a machine learning model). The first model can be connected to the machine learning model in series or in a parallel combination (e.g., FIG. 4A to FIG. 4C In the example configuration (e.g., Figure 4A ), the serial combination of the models involves providing the output of the first model as an input to the machine learning model. In another example configuration, the serial combination of the models involves providing the output of the machine learning model as an input to the first model. In yet another example (e.g., Figure 4B ), the parallel combination of the models involves providing the same input to the first model and the machine learning model, combining the outputs of the first model and the machine learning model, and determining the predicted printed pattern based on the combined outputs of the respective models. In yet another example (e.g., Figure 4C ), the patterning process model can be configured to include both a series arrangement and a parallel arrangement of the first model, the machine learning model, and / or another physics-based model or machine learning model.

[0078] In an embodiment, the first model is an empirical model including physical terms that accurately describe the effects of the first aspect of the patterning process. In an embodiment, the first model corresponds to the first aspect related to acid-base diffusion after exposure of the substrate. In an embodiment, an example first model is a resist model. The first parameter set of the resist model corresponds to at least one of the following physical terms: initial acid distribution; acid diffusion; image contrast; long-range pattern loading effect; long-range pattern loading effect; acid concentration after neutralization; base concentration after neutralization; diffusion due to high acid concentration; diffusion due to high base concentration; resist shrinkage; resist development; or two-dimensional convex curvature effect. Example empirical models using physical terms are as follows:

[0079] R=cA×A+cMav×MAV+cAp×Ap*GAp+cBP×Bp*GBp

[0080] +cAm×A*GAm+…

[0081] In the above equation, R is a predicted resist image based on the physical terms and the coefficients associated therewith (an example of a first parameter set). In the above equation, cA is a coefficient of the initial acid distribution A, which can be represented by an aerial image, cMav is a coefficient of the long-range pattern loading effect MAV, which can be determined as an average value of the mask image, and similarly, other physical terms are associated with one or more coefficients. These coefficients are determined during the training process discussed below.

[0082] In an embodiment, the machine learning model is a neural network that models a second aspect of the patterning process that has relatively less physics-based understanding. In an embodiment, the second set of parameters includes weights and biases of one or more layers of the neural network. During a training process (e.g., involving processes P303, P307), the weights and biases are adjusted in conjunction with the first set of parameters so that the difference between the predicted pattern and the measured pattern is reduced. In an embodiment, the patterning process model corresponds to the second aspect of the post-exposure process of the patterning process.

[0083] In an embodiment, the method further involves iteratively determining values ​​of the first set of parameters 307 and the second set of parameters 308 to train the patterning process model. In an embodiment, iteration involves performing processes P303, P305 and P307.

[0084] In process P303, the method 300 involves executing the first model and the machine learning model to collaboratively predict a pattern of the substrate using the image data 302. Process P305 involves determining a difference between the measured pattern 304 and the predicted pattern 305 of the patterning process model, and further determining whether the difference is reduced or minimized. Process P307 involves modifying the values ​​of the first parameter set 307 and the second parameter set 308 so that the difference between the measured pattern 304 and the predicted pattern 305 of the patterning process model is reduced.

[0085] In an embodiment, the modified values ​​are determined based on an optimization technique such as a gradient descent method, which guides how to modify the values ​​of the first set of parameters and the second set of parameters so that the gradient of the difference with respect to the corresponding parameters decreases. After several iterations, a global optimum or a local optimum of the parameters is obtained, minimizing the difference between the prediction and the measurement. Thus, the patterning process model is calibrated (or trained) and can also be employed for improving the patterning process via OPC, defect detection, hot spot sorting, or other known applications of the patterning process model.

[0086] Figure 4A , Figure 4B and Figure 4C An example configuration of the patterning process model including the first model and the second model is illustrated. The first model and the second model (CNN) are trained together, during which the first set of parameters of the first model and the second set of parameters of the CNN are determined.

[0087] Figure 4A The first model and the second model are combined in series, wherein the first model is represented as a function of a physical parameter and the second model is represented as a CNN. The training process is iterative, wherein an iteration (or each iteration) involves determining a parameter such as c i ) (which are coefficients associated with the physical terms), and weights (such as w of the CNN). i and u i ) or the like. Coefficient c i Directly using the physical term i The values ​​of the variables (e.g., such as dose, focus, acid concentration, etc. associated with the patterning process) are operated (e.g., via multiplication, squaring, addition, or other mathematical operations) to determine a first output of the first model. In an embodiment, the physical term term i It can also be with the parameter param iAlthough the weights of the CNN do not operate directly with specific physical terms, the CNN can receive as input a first output from the first model and predict a second output associated with another (e.g., not well understood) aspect of the patterning process.

[0088] In an embodiment, initial values ​​of the first set of parameters and the second set of parameters may be assigned to start a simulation process. In an embodiment, the input to the first model is, for example, an aerial image of a desired pattern to be printed on the substrate. Based on the input (e.g., aerial image), the first model predicts a first output (e.g., a resist pattern) of aspects of the patterning process (e.g., exposing the resist using the aerial image). The first output is further input to the CNN, which further predicts a pattern to be printed on the substrate. The predicted pattern is compared with the desired output. The desired output may be a measured pattern corresponding to the desired pattern measured via, for example, a SEM tool. The comparison involves calculating the difference between the predicted pattern and the measured pattern. Based on the difference, a back propagation may be performed, in which a W of a CNN such as a W may be calculated. i and u i The value of the weight such as so that the difference is reduced. In addition, c can be calculated i and / or param i For example, a gradient-based approach may be employed, in which the difference relative to the weight is calculated to produce a gradient map. The gradient map acts as a guide to modify the weights and / or c i and / or param i After training the patterning process model, the model is able to determine patterns that take into account both well-understood physical effects of the patterning process (e.g., post-exposure processes) and less-well-understood physical effects.

[0089] Figure 4B The first model and the second model are combined in parallel. Figure 4AIn a similar manner as discussed in . In addition, the initial values ​​of the first parameter set and the second parameter set can be similar to the initialization of the training process to determine the final values ​​of the first parameter set and the second parameter set. In a parallel combination, the same input (e.g., a spatial image of a desired pattern) is provided to both the first model and the second model simultaneously. Each model prediction can be combined to form an output of a pattern to be printed on the substrate. The predicted output can be compared with the desired output (e.g., a measured pattern of the printed substrate), as discussed above. Then, back propagation can be performed, and a gradient descent method can be applied to determine the difference of the difference relative to each of the first parameter and the second parameter, as discussed previously. In addition, the values ​​of the first parameter and the second parameter are selected so that the difference is reduced (in an embodiment, minimized). After several iterations, the predicted pattern converges to the desired pattern, and model training is called completed.

[0090] Figure 4C A more general patterning process model is illustrated, wherein the process model is configured to include (i) one or more models including physical terms of the patterning process (e.g., variables of the patterning process), and (ii) one or more machine learning models (CNNs). In an embodiment, inputs and outputs can be communicated between different models as shown to collaboratively predict the pattern to be printed on the substrate. The predicted pattern can be compared to the expected pattern to determine the values ​​of the parameters of each model in the model. The values ​​can be determined, for example, using a gradient descent method as previously discussed. The values ​​are modified during the back propagation of the outputs of the different layers of the CNN and / or model, such as Figure 4C shown.

[0091] Thus, in the embodiment, as above FIG. 4A to FIG. 4C As discussed in , the first model and the machine learning model are configured and trained together in a deep convolutional neural network framework (DCNN). Training involves: outputs of the first model and the machine learning model (e.g., FIG. 4A to FIG. 4C , y, z, etc. in the image processing unit) to predict the printed pattern; determining the difference (e.g., output) between the measured pattern and the predicted pattern (e.g., output) of the patterning process model; FIG. 4A to FIG. 4C determining the difference (e.g., d(loss)) relative to the first parameter (e.g., c i 、param i 、z i 、u i 、w iand the difference between the first set of parameters and the second set of parameters; and back propagation of the outputs of the first model and the machine learning model to determine the values ​​of the first set of parameters and the second set of parameters based on the difference of the difference. FIG. 4A to FIG. 4C During back-propagation, the following differences may be computed and used to adjust the first and second sets of parameters: d(loss) / dt; d(loss) / dz; d(loss) / du; d(loss) / dw; and so on.

[0092] Figure 5 1 is a flow chart of a method for determining an optical proximity correction for a patterning process. The optical proximity correction is associated with a desired pattern to be printed on the substrate. In process P501, method 500 involves obtaining image data 502 associated with the desired pattern. In an embodiment, the image data 502 is an aerial image and / or a mask image of the desired pattern.

[0093] In addition, process P503 involves executing the trained patterning process model 310 using the image data 502 to predict the pattern to be printed on the substrate. As previously discussed in method 300, the trained patterning process model 310 includes a first model of a first aspect of the patterning process and a machine learning model of a second aspect of the patterning process, which are configured to collaboratively predict the pattern to be printed on the substrate. The first model and the machine learning model are in a series combination and / or a parallel combination, such as with respect to FIG. 4A to FIG. 4C In an embodiment, the first model is an empirical model (e.g., the resist model discussed previously) that accurately models the physics of the first aspect of the post-exposure process of the patterning process. In an embodiment, the first model corresponds to the first aspect related to acid-base diffusion after exposure of the substrate. In an embodiment, the machine learning model is a neural network that models the second aspect of the patterning process for which there is relatively little physics-based understanding.

[0094] Based on the predicted pattern, process P505 involves determining optical proximity corrections and / or defects. In an embodiment, determining optical proximity corrections involves adjusting the desired pattern and / or placing assist features around the desired pattern so that the difference between the predicted pattern and the desired pattern is reduced. Figures 18 to 21 An example OPC process is discussed.

[0095] In an embodiment, determining defects involves performing a lithography manufacturability check (LMC) on the predicted pattern. The LMC determines whether a feature of the predicted pattern meets a desired specification. If the LMC determines that the specification is not met, the feature is considered to be defective. Such defect information can be applicable to determining a yield of the patterning process. Further based on the defects (or yield), one or more variables of the patterning process can be modified to improve the yield.

[0096] As previously mentioned, some effects of the patterning process or post-exposure process are easy to understand and can be modeled using mathematical expressions including physical terms related to the pattern formed on the substrate. For example, some of the physical terms (e.g., associated with dose, focal length, intensity, pupil, etc.) are associated with tunable parameters (e.g., tunable knobs) of the lithographic apparatus and can be tuned via the tunable parameters, thereby enabling real-time control of the patterning process via the tunable knob. In an embodiment, some physical terms may not be directly tuned via the tuning knob, but the physical / chemical properties of the process (e.g., aerial image formation, resist image formation, etc.) can be explained. For example, the resist model includes a Gaussian filter (including a mean square deviation or variance term) on the aerial image, which is used to model the acid-base diffusion in the resist after the exposure. Such a mean square deviation term is generally not tunable via a tunable knob. Even so, the value of such a physical term (eg, mean square deviation) may be determined based on a physics-based equation, or an empirical equation, that models the effects of the process (eg, resist).

[0097] As discussed herein, various methods are provided for training a patterning process model based on the physical terms (e.g., FIG. 4A to FIG. 4C ). There are several advantages to using physical terms to train or calibrate patterning process models. For example, the calculation of physical terms is relatively simple, and models that employ such physical terms are less susceptible to overfitting. In an embodiment, incorporating physical terms in conjunction with, for example, a CNN model reduces the complexity of the CNN by several times, reduces the risk of overfitting, and improves the run time of the patterning process simulation. The following description discusses additional methods for training and calibrating process models based on physical terms.

[0098] Figure 6 6 is a flow chart of a method 600 for training a machine learning model to determine an etching bias associated with an etching process. In an embodiment, predicting such an etching bias may be useful for improving an etching recipe or a current lithography equipment setting. The method 600 includes several steps described in detail below.

[0099] Step P601 includes obtaining (i) resist pattern data 602 associated with a target pattern to be printed on a substrate, (ii) physical effect data 604 characterizing the effect of the etching process on the target pattern, and (iii) measured deviations 606 between the resist pattern and the etched pattern formed on the printed substrate.

[0100] In an embodiment, the measured deviation 606 data may be determined based on metrology data of a previously patterned substrate. For example, the measured deviation 606 may be the difference between a resist pattern formed on the substrate and an etch pattern formed on a printed substrate. The resist pattern may be determined via a metrology tool or a simulation of a patterning process. In an embodiment, the measured deviation 606 data may be determined based on metrology data of a previously patterned substrate. For example, the measured deviation 606 may be the difference between a resist pattern formed on the substrate and an etch pattern formed on a printed substrate. Fig.16 and Fig.17 The etched pattern formed on the printed pattern can be measured using a SEM tool as described above, or an optical metrology tool. In an embodiment, the size of the resist pattern (e.g., the CD of the features) can be increased compared to the etched pattern due to, for example, the removal of material (e.g., via descumming). The difference between the resist pattern and the etched pattern on the printed substrate includes changes due to the etching process. For example, the changes are due to changes in etching rate, changes in the amount of plasma concentration, changes in aspect ratio (e.g., height of features / width of features), or other physical aspects related to the resist pattern, the etching process, or a combination thereof.

[0101] In an embodiment, the resist pattern data 602 is represented as a resist image. The resist image may be a pixelated image, wherein the intensity of the pixels indicates the resist area and the pattern portion formed within the resist portion. For example, the pattern portion may be an edge / contour of a resist pattern. In an embodiment, obtaining the resist pattern data 602 involves executing one or more process models including a resist model of the patterning process using the target pattern to be printed on the substrate.

[0102] In an embodiment, the physical effect data 604 may be data related to etching terms characterizing etching effects, the etching terms including at least one of: a concentration of plasma within a trench of the resist pattern associated with the target pattern; a concentration of plasma on top of a resist layer of the substrate; a loading effect determined by convolving the resist pattern with a Gaussian kernel having specified model parameters; a change in the loading effect on the resist pattern during the etching process; a relative position of the resist pattern relative to an adjacent pattern on the substrate; an aspect ratio of the resist pattern; or a term related to a combined effect of two or more etching process parameters.

[0103] In an embodiment, obtaining the physical effect data 604 involves executing a physical effect model, the physical effect model including one or more of the etching items and a Gaussian kernel specified for the corresponding one or more of the etching items. In an embodiment, the physical effect data 604 is represented as a pixelated image, wherein each pixel intensity indicates a physical effect on the resist pattern associated with the target pattern. Figures 7 to 10 Some example physical effect data 604 is illustrated.

[0104] Figure 7 is an example resist image including a resist pattern formed in the resist 702. The resist pattern includes a groove region 704 formed in the resist 702. In an embodiment, the etching model can be calibrated based on a plasma concentration etching method (CEM). The CEM method uses plasma loading on the edge of the resist groove 704 to characterize the bias behavior or deviation behavior caused by etching. In an embodiment, an evaluation point such as 706 can be located at the edge of the resist groove (e.g., 704), and CEM_range is the etching proximity range currently considered by the etching model. In an embodiment, the etching physical item can be a CR image or a CT image related to the plasma loading effect generated from the resist pattern. For example, CT0 is defined as the plasma loading per unit edge length of the plasma from the groove region (e.g., 704) at the beginning of etching, and CR0 is defined as the plasma loading per unit edge length of the plasma from the resist region (e.g., 702) at the beginning of etching. At time t, CT0 and CR0 become CT and CR, respectively, where the conversion with respect to time may be based on an exponential term including parameters related to the sidewall deposition or etching reaction constant. In an embodiment, the product of the CT image and the CR image may be used to characterize the etching bias due to the proximity effect.

[0105] Figure 8 An example etching physical item (e.g., CR image) generated from a resist pattern 801 is illustrated. The resist pattern 801 includes a resist profile 802 (or a resist pattern edge), and then the plasma loading effect from the resist area (e.g., the CR discussed above) can be calculated, which is illustrated as a CR image 810. In an embodiment, a CT image can be generated, wherein the CT image is a flipped tone image of the CR. For example, a bright pixel in the CR becomes a dark pixel in the CT, and vice versa. The CR image 810 can be used to characterize an etching bias from a proximity effect. This example etching bias is an example bias applied to the resist profile 802 to compensate for this proximity effect.

[0106] Fig. 9Another example modeling of a physical term (e.g., an acid-base reaction) is shown. In an embodiment, an acid-base reaction can be modeled by truncating the acid concentration by a quencher base. For example, an acid-base reaction (e.g., illustrated as image 901) can be characterized by a linear combination coefficient of an acid concentration 910 and an alkali concentration 920. In an embodiment, a truncation term (e.g., 903) simulates the reaction and diffusion of an acid and an alkali when forming the final acid density distribution image 901. Multiple truncation terms represent different times during post-exposure baking. Depending on the truncation term 903 (e.g., the cutoff value of the truncation term), the acid concentration 910 and the alkali concentration 920 will change.

[0107] Fig.10 is another example physical item generated using the aerial image 1010. For example, the physical item may be an initial acid distribution 1020 at a particular location of the aerial image. In an embodiment, a linear transformation of the aerial image may be performed using a Gaussian filter on the aerial image. In an embodiment, the Gaussian filter includes a mean square deviation term that is not typically tunable via a tunable knob, but may be set to determine long-term effects, mid-range effects, and short-range effects associated with etching. For example, the mean square deviation value may be set based on data from previously printed and etched substrate data.

[0108] Return to reference Figure 6 , step P603 includes training the machine learning model 603 based on the resist pattern data 602, the physical effect data 604, and the measured deviation 606 to reduce the difference between the measured deviation 606 and the predicted etching deviation. After the training process is completed, the machine learning model 603 can be referred to as a trained machine learning model 603. This trained machine learning model 603 can be used in the patterning process to improve performance indicators, such as the yield of the printed substrate. For example, based on the etching deviation predicted by the trained machine learning model 603, the process parameters can be adjusted so that the number of failures of the pattern is reduced, thereby improving the yield.

[0109] In an embodiment, the machine learning model 603 is configured to receive the resist pattern data 602 at a first layer of the machine learning model 603, and the physical effect data 604 is received at a last layer of the machine learning model 603. In an embodiment, the machine learning model 603 is configured to receive the resist pattern data 602 and the physical effect data 604 at a first layer of the machine learning model 603. It will be appreciated by those skilled in the art that the present disclosure is not limited to a specific configuration of the machine learning model 603.

[0110] In an embodiment, the output of the last layer of the machine learning model 603 is a linear combination of: (i) the etching bias predicted by executing the machine learning model 603 using the resist pattern data 602 as input, and (ii) another etching bias determined based on physical effect data 604 associated with the etching process.

[0111] In an embodiment, the output of the last layer of the machine learning model 603 is an etch bias map from which the etch bias is extracted. The etch bias map is generated by: executing the machine learning model 603 using the resist pattern data 602 as input to output an etch bias map, wherein the etch bias map includes a biased resist pattern; and combining the etch bias map with the physical effect data 604.

[0112] In an embodiment, the training of the machine learning model 603 is an iterative process, the iterations involving: (a) predicting the etching bias by executing the machine learning model 603 using the resist pattern data 602 and the physical effect data 604 as inputs; (b) determining the difference between the measured bias 606 and the predicted etching bias; (c) determining the gradient of the difference relative to a model parameter (e.g., a weight associated with a layer) of the machine learning model 603; (d) using the gradient as a guide to adjust model parameter values ​​so that the difference between the measured bias 606 and the predicted etching bias is reduced; (e) determining whether the difference is minimized or exceeds a training threshold; and (f) in response to the difference not being minimized or not exceeding the training threshold, performing steps (a) to (e).

[0113] In an embodiment, the method 600 may also involve, at step P605, obtaining a resist profile of the resist pattern (e.g., 602 discussed above); and generating an etching profile 605 by applying an etching bias (e.g., determined by executing a trained machine learning model 603) to the resist profile (e.g., 602).

[0114] Fig.11An example etching bias determined via an etching model is illustrated. In an embodiment, the etching model calculates an after-etching image (AEI) profile (e.g., 1130) by directly biasing (e.g., 1120) the after-development image (ADI) profile (e.g., 1110). In an embodiment, the bias direction of 1120 is perpendicular to the ADI profile 1110. Depending on the environment (e.g., feature density) of the ADI pattern and the physical items associated with etching, the bias amount of 1120 is variable. For example, a positive bias amount moves the ADI profile outward, while a negative bias (e.g., 1120) amount moves the ADI profile 1110 inward. In other words, the etching bias can be positive, where the size of the pattern 1110 element is larger before etching than after etching, or the etching bias is negative, where the size is smaller before etching than after etching. In an embodiment, the etching model uses a calibration / check gauge, and the vertical direction of the ADI profile is relative to such a gauge. Then, the model can directly output the AEI profile for LMC / OPC application. For example, LMC can determine whether the AEI profile meets the size constraints associated with the target pattern. The AEI profile can be used to determine the OPC of the mask pattern so that the overall yield of the patterning process is improved. For example, in the OPC process, a simulated profile (e.g., based on a patterning process simulation) can be compared with the AEI profile, and the OPC can be determined based on the comparison. For example, the mask pattern is modified so that the simulated pattern closely matches the AEI profile. Therefore, accurate prediction of the AEI profile will improve the OPC of the mask pattern.

[0115] In an embodiment, the offset determined by the model (eg, 603) may be a plurality of physical terms Term i A linear combination of , where the physical term is a function of the environment at an evaluation point i.

[0116]

[0117] In an embodiment, the physical item Term i It may be, for example, a local or long range loading effect. In an example, the effect may be determined by rasterization (e.g., convolving the resist profile with a Gaussian kernel or filter having a first set of parameters (e.g., a mean square deviation between 90 and 100 nm). Another physical term may be intermediate stage loading determined using a Gaussian kernel or filter having a second set of parameters (e.g., a mean square deviation between 100 and 200 nm). Another example physical term may be aspect ratio. In an embodiment, the term may be a high order, nonlinear, or combined effect.

[0118] Modeling the etching bias mathematically can improve the generation of the feature size of the final device. The results of this modeling can be used for various purposes. For example, this result can be used to adjust the patterning process in terms of changing the design, control parameters, etc. For example, the result can be used to adjust one or more spatial properties of one or more elements in the elements provided by the pattern forming device, wherein the pattern of the pattern forming device is used to produce a device to be used for etching on the substrate pattern. Thus, once the pattern of the pattern forming device is transferred to the substrate, the device pattern on the substrate is effectively adjusted before etching to compensate for the etching changes expected to occur during the etching. As another example, one or more adjustments can be made to the lithographic equipment in terms of adjustment of dose, focal length, etc. As will be appreciated, there may be more applications. Thus, compensating for etching changes may cause the device to have more than one uniform feature size, one or more uniform electrical properties, and / or one or more improved (e.g., closer to the desired result) performance characteristics.

[0119] In addition, although etching variations are sometimes not conducive to making devices on a substrate, the etching deviation can be used to produce desired structures on the substrate. By taking into account the degree of etching deviation when making patterned devices, it is possible to make device features in the device on the substrate that are smaller than the optical resolution limit of the pattern transfer process from the patterned device to the substrate. Therefore, in this regard, the modeling results of the etching deviation can be used to adjust the patterning process in terms of changing the design, control parameters, etc. Therefore, modeling the etching deviation during the etching process can help produce more accurate device features, such as by compensating for etching variations, such as by adjusting the patterned device to (correctly) predict possible etching variations of the etching process (e.g., depending on the pattern density). The variation allows the actual features produced by the etching process after (adjusted) lithography to be closer to the desired product specifications.

[0120] In an embodiment, a system for determining an etch deviation associated with an etch process implementing the procedures discussed herein (eg, method 600) is described. For example, the system includes: a semiconductor process device (eg, Figure 1 , Fig.23 , Fig.24 , Fig.25 ), and one or more processors (e.g., Fig. 22104 / 105), which is configured to: determine physical effect data 604 characterizing the effect of the etching process on the substrate via executing a physical effect model; execute a trained machine learning model 603 using the resist pattern and the physical effect data 604 as input to determine the etching bias; and control the semiconductor device (e.g., Figure 1 ) or the etching process.

[0121] In an embodiment, the trained machine learning model 603 is trained, for example, according to the method 600. For example, the trained machine learning model 603 is trained using a plurality of resist patterns, the physical effect data 604 associated with each of the resist patterns, and the measured deviation 606 associated with each resist pattern, such that the difference between the measured deviation 606 and the determined etching deviation is minimized.

[0122] In an embodiment, the trained machine learning model 603 is a convolutional neural network (CNN) including specific weights and biases, wherein the weights and biases of the CNN are determined via a training process using a plurality of resist patterns, physical effect data 604 associated with each of the resist patterns, and measured deviations 606 associated with each resist pattern so that the difference between the measured deviations 606 and the determined etching deviations is minimized.

[0123] In an embodiment, controlling semiconductor process equipment (e.g., Figure 1 , Fig.23 , Fig.24 , Fig.25 ) includes adjusting the value of one or more parameters of the semiconductor device so that the yield of the patterning process is improved. In an embodiment, adjusting the value of one or more parameters of the semiconductor process equipment is an iterative process. The iterative process involves: (a) changing the current value of the one or more parameters via a mechanism of adjusting the semiconductor process equipment; (b) obtaining the resist pattern printed on the substrate via the semiconductor process equipment; (c) determining the etching bias by executing the trained machine learning model 603 using the resist pattern, and further determining the etching pattern by applying the etching bias to the resist pattern; (d) determining whether the yield of the patterning process is within a desired yield range based on the etching pattern; in response to not being within the yield range, performing steps (a) to (d).

[0124] In an embodiment, controlling the etching process involves: determining the etching pattern by applying the etching bias to the resist pattern; determining the yield of the patterning process based on the etching pattern; and determining an etching selection scheme for the etching process based on the etching pattern, so that the yield of the patterning process is improved. In an embodiment, the yield of the patterning process is the percentage of etching patterns that meet design specifications across the entire substrate. In an embodiment, the semiconductor process equipment is a lithography equipment (e.g., Figure 1 , Fig.23 , Fig.24 , Fig.25 ).

[0125] In today's semiconductor field, as technology nodes keep shrinking, it is desirable to have better models for lithography and etching. A good model meets both accuracy (e.g., model results match measurements of real wafers) and good wafer prediction (e.g., performs according to physical limitations). Meeting both accuracy and prediction specifications can be difficult with current complex model forms because better fit capability indicates overfitting of the model. An overfit model can produce irregularly shaped patterns that may not normally be expected to be printed on a substrate.

[0126] The current method for solving the overfitting problem or prediction related problems is to have more measurement information during model calibration. For example, more information includes data related to more pattern coverage or more evaluation points. For example, a SEM tool can be configured to generate a large number of EP gauges for a specific pattern. However, increasing the measurement will increase the cost and time of the patterning process. Typically, the pattern coverage is significantly less than 100% of the total number of patterns in the design layout. Thus, model calibration cannot be performed with all possible patterns to be printed on the wafer. Improving pattern coverage is time-consuming and cost-inefficient. Recalibration and data collection will be performed for several rounds or cycles. In addition, if the modeling is complex enough, it will still try to overfit. Therefore, a more basic solution is proposed to make the model perceive in a physical way by performing a model calibration based on physical constraints, rather than dealing with the problem by feeding more data into the model calibration.

[0127] In embodiments, the physical constraints may be associated with a contour shape obtained from a metrology tool (e.g., SEM), a solid image (e.g., resist image, aerial image) obtained from the metrology tool (e.g., SEM), or a combination thereof. Example methods for implementing physical constraints are discussed herein, for example, Fig.12 and Fig.14 method.

[0128] Fig.122 is a flow chart of a method 2000 for calibrating a process model based on physical constraints related to the contour shape (or contour) of a pattern. The method 2000 calibrates the process model to produce a simulated contour that satisfies the shape constraints. The detailed procedures of the method 2000 are discussed below.

[0129] Process P2001 involves obtaining (i) measurement data 2002 at a plurality of measurement locations on a pattern, and (ii) contour constraints 2004 specified based on the measurement data 2002. In an embodiment, the plurality of measurement locations are edge placement (EP) gauges placed on the printed pattern or on a printed contour of the printed pattern.

[0130] In an embodiment, the measurement data 2002 includes a plurality of angles, each angle being defined at each measurement location placed on the pattern or on a printed contour of the printed pattern. In an embodiment, each angle at each measurement location defines a direction in which an edge placement error between the printed contour and a target contour is determined. Fig.13A and Fig. 13B Example measurement locations EP1, EP2, and EP3 are illustrated. Next, the angle associated with each measurement location EP1 to EP3 is the angle or direction along which the EPE can be calculated. In other words, for example, at point EP1, the contour 1110 (or Fig. 13B The distance between the contour 1110 (or 1120) is measured in the direction shown by the arrow pointing away from the contour 1110 (or 1120). Depending on the shape of the contour 1110 / 1120, such measurement results may vary.

[0131] In an embodiment, each contour constraint is a function of the tangent angle between a tangent to a simulated contour (e.g., contour 1110 or 1120) at a given measurement site (e.g., EP1 to EP3) and the angle of the measurement data 2002 at the given site. Fig.13A , the contour constraint may be that the angle θ1 between the tangent of the simulated contour 1110 and the arrow at EP1 (which indicates the angle of the measurement data 2002) should be within a vertical range. In an embodiment, the vertical range is the value of the angle θ1 that should be between 88° and 92°, preferably 90°. Each point EP1, EP2, and EP3 may be associated with such a verticality constraint.

[0132] In another example, Fig. 13BThe results of a calibrated model that produces a simulated profile 1120 that does not satisfy the physical constraints are illustrated. For example, the calibrated model may be overfitted such that it makes good predictions with respect to the measured data. For example, the model is overfitted due to an excess of data, where the fitting is focused on minimizing EP errors. Such an overfitted model may produce an irregularly shaped profile such as profile 1120. Subsequently, the tangents to profile 1120 at EP1, EP2, and EP3 may not be perpendicular to the measured data (e.g., the measured angles indicated by the arrows at EP1 to EP3, respectively). For example, as Fig. 13B As seen in , the lines perpendicular to the arrows at EP1, EP2, and EP3 are not tangent to the profile 1120. Therefore, although the simulated profile 1120 fits the measured data (e.g., EPE or CD values) to minimize the sum of errors associated with the simulated profile 1120, the shape of the profile 1120 may not be physically accurate.

[0133] Return to reference Fig.12 , step P2003 involves calibrating the process model 2003 by adjusting the values ​​of the model parameters of the process model until the simulation profile (e.g., 1110) satisfies the profile constraints 2004. After the calibration process, the process model may be referred to as a calibrated process model 2003. In an example, the model 2003 that produced the simulation profile 1120 may not be considered calibrated because the model does not satisfy the profile constraints 2004 at several points EP1, EP2, and EP3. In an embodiment, the calibration of the process model may be limited to a selected number of points EP1 and EP3, or all points, e.g., EP, EP2, and EP3.

[0134] In an embodiment, adjusting the value of the model parameter is an iterative process. The iteration involves: (a) executing the process model 2003 using a given value of the model parameter to generate the simulated profile, wherein the given value is a random value at a first iteration and an adjusted value at a subsequent iteration; (c) determining a tangent to the simulated profile at each of the measurement sites; (d) determining a tangent angle between an angle of the measurement data 2002 at each of the measurement sites and the tangent; (e) determining whether the tangent angle is within a vertical range at one or more of the plurality of measurement sites; and (f) in response to the tangent angle not being within a vertical range, adjusting the value of the model parameter and performing steps (a) to (d).

[0135] In an embodiment, at each iteration, a simulated profile may be obtained and a tangent line may be drawn or calculated (e.g., via the trigonometric relationship "tan") at the measurement site. Next, the angle between the tangent line and the EPE angle of the measurement data 2002 may be determined, thereby detecting whether the tangent line angle is within a vertical range (e.g., between 88° and 92°, preferably 90°).

[0136] In an embodiment, the adjustment is based on a gradient of each tangent angle relative to the model parameter, wherein the gradient indicates how sensitive the tangent angle is to changes in the model parameter value.

[0137] In an embodiment, the process model 2003 is a data driven model including an empirical model and / or a machine learning model. For example, the machine learning model is a convolutional neural network, wherein the model parameters are weights and biases associated with multiple layers. The present disclosure is not limited to a particular type of model or a particular process of the patterning process. The method 2000 can be modified or adapted for use with any process model and any process (or combination of processes) of the patterning process.

[0138] In an embodiment, the method 2000 may also involve, at step P2005, obtaining a resist profile of the resist pattern (e.g., 602 discussed above); and generating an etching profile 2005 by applying the etching bias (e.g., determined by executing a calibrated model) to the resist profile.

[0139] Fig.14 is a flow chart of another method 3000 for calibrating a process model based on physical constraints. In an embodiment, the process model is configured to predict an image of a target pattern. Then, the calibration can be based on image-based constraints. The method 3000 involves the following steps.

[0140] Process P3001 involves obtaining (i) a reference image 3002 associated with the target pattern, and (ii) a gradient constraint 3004 specified relative to the reference image 3002. In an embodiment, the reference image 3002 is obtained by simulating a physics-based model of a patterning process using the target pattern. The reference image 3002 includes, but is not limited to, at least one of: an aerial image of the target pattern; a resist image of the target pattern; or an etched image of the target pattern. In an embodiment, the simulated gradient is determined by taking a first-order derivative of a signal along a given line through the simulated image. In an embodiment, the gradient constraint 3004 is obtained by taking a first-order derivative of a signal along a given line through the reference image 3002.

[0141] Step P3003 involves calibrating the process model 3003 so that the process model 3003 produces a simulated image that (i) minimizes the intensity difference or frequency difference between the simulated image and the reference image 3002, and (ii) satisfies the gradient constraint 3004. After calibration, the process model 3003 may be referred to as a calibrated process model 3003.

[0142] FIG. 15A to FIG. 15C An example of a gradient constraint 3004 is illustrated. Fig.15A An example reference intensity distribution 1510 at a given location in a solid image (e.g., an aerial image, a resist image, and an ADI) is illustrated. In an embodiment, the similarity between the reference image 3002 and the simulated image can be used to quantify the resist model stability to understand the overfitting risk level. For example, similarity can be evaluated as the intensity difference between the reference image 3002 and the simulated image, or the frequency difference between the reference image 3002 and the simulated image (e.g., via an FFT of the image).

[0143] Thus, in an embodiment, the gradient of the intensity difference or frequency difference may be applied as a constraint during calibration of the process model (e.g., 3003). For example, after applying the gradient constraint relative to the reference intensity distribution 1510 during the calibration process, the process model may generate a signal having an intensity distribution 1520 (at Fig. 15B 1520 has a similar shape (e.g., similar peaks and valleys) to the reference intensity distribution 1510. Thus, the calibrated process model (e.g., 3003) is considered to follow the physical terms (e.g., intensity distribution or frequency distribution) of the reference image associated with the patterning process. In embodiments, the process model is not calibrated according to the gradient associated with the physical terms (e.g., AI / RI), and such a process model may produce an intensity distribution 1530 (e.g., in FIG. 1530 ) that is unacceptable compared to the reference intensity distribution. Fig. 15C For example, the peaks at 1510 and 1530 are clearly different.

[0144] In an embodiment, calibrating the process model is an iterative process. Iteration involves (a) executing the process model using the target pattern to generate the simulated image; (b) determining the intensity difference between the intensity values ​​of the simulated image and the reference image 3002, and / or transforming the simulated image and the reference image 3002 into the frequency domain via Fourier transform and determining the frequency difference between the frequencies associated with the simulated image and the reference image 3002; (c) determining the simulated gradient of the signal in the simulated image, wherein the signal is a signal along a given line passing through the simulated image; (d) determining whether the conditions are met: (i) the intensity difference or the frequency difference is minimized, and (ii) the simulated gradient satisfies the gradient constraint 3004 associated with the reference image 3002; (e) in response to not meeting the conditions (i) and (ii), adjusting the values ​​of the model parameters of the process model, and performing steps (a) to (d) until the conditions (i) and (ii) are met.

[0145] In an embodiment, method 3000 also involves extracting a simulated contour from a simulated image, extracting a reference contour from a reference image 3002, and calibrating the process model so that the simulated contour satisfies a contour shape constraint at step P3005. For example, an edge detection algorithm or other contour extraction techniques for extracting contours associated with a target pattern can be used to extract contours from an image. In an embodiment, the simulated contour and the reference contour are associated with the target pattern, and the contour shape constraint ensures that the shape of the simulated contour matches that of the reference contour. In an embodiment, the contour shape constraint can be implemented as an invariant condition that the model output should satisfy. If the invariant condition is not satisfied, the value of the model parameter is adjusted until such invariant condition is satisfied.

[0146] In an embodiment, determining whether the contour shape constraint is satisfied involves determining that the second order derivative of the simulated contour is within an expected range of the second order derivative of the reference contour. In an embodiment, the contour shape is represented as a polygon, and therefore, the second order derivative of the polygon can be calculated using computing software.

[0147] In an embodiment, the process model may be configured to satisfy contour constraints defined relative to a printed contour of a pattern on a printed substrate (e.g., Fig.12 For example, each contour constraint is a function of a tangent angle between a tangent to a simulated contour at a given measurement location and an angle of the measurement data 2002 at the given location, wherein the simulated contour is a contour of a simulated pattern determined by executing the process model using the target pattern.

[0148] In an embodiment, the method 3000 may also involve, at step P3005, obtaining a resist profile of the resist pattern (e.g., 602 discussed above); and generating an etching profile 3005 by applying an etching bias (e.g., determined by executing a calibrated model) to the resist profile.

[0149] In an embodiment, a system for calibrating a process model implementing a process discussed herein (e.g., method 2000) is described. The process model is configured to generate a simulation profile. The system includes: a measurement tool (e.g., Fig.16 and Fig.17 ), the metrology tool being configured to obtain measurement data 2002 at a plurality of measurement locations on the pattern; and one or more processors (e.g., Fig. 22 The processor (eg, 104 / 105) may be configured to calibrate the process model by adjusting values ​​of model parameters of the process model until a simulation profile complies with profile constraints 2004, the profile constraints 2004 being based on the measurement data 2002.

[0150] In an embodiment, a metrology tool, such as a SEM, is configured to obtain measurements at a plurality of measurement locations, such as an edge placement (EP) gauge placed on a printed pattern or on a printed profile of the printed pattern. In an embodiment, the measurement data 2002 includes a plurality of angles, each angle being defined at each measurement location placed on the pattern or on a printed profile of the printed pattern. In an embodiment, each angle at each measurement location defines a direction in which an edge placement error between the printed profile and the target profile is determined. As previously mentioned, the metrology tool may be an electron beam device (e.g., Fig.16 and Fig.17 In an embodiment, the metrology tool is a scanning electron microscope configured to identify and extract contours from a captured image of a pattern on a printed substrate.

[0151] In an embodiment, the processor is configured to include each contour constraint as a function of a tangent angle between a tangent to the simulated contour at a given measurement location and an angle of the measurement data 2002 at the given location.

[0152] In an embodiment, the processor is configured to adjust the value of the model parameter in an iterative manner. For example, the iteration includes: (a) executing the process model using a given value of the model parameter to generate a simulated profile, wherein the given value is a random value at a first iteration and an adjusted value at a subsequent iteration; (c) determining a tangent of the simulated profile at each of the measurement sites; (d) determining a tangent angle between an angle of the measurement data 2002 at each of the measurement sites and the tangent; (e) determining whether the tangent angle is within a vertical range at one or more of the measurement sites; and (f) in response to the tangent angle not being within the vertical range, adjusting the value of the model parameter and performing steps (a) to (d).

[0153] In an embodiment, the vertical range is a value of an angle between 88° and 92°, preferably 90°. In an embodiment, the processor is configured to make adjustments based on a gradient of each tangent angle relative to a model parameter, wherein the gradient indicates how sensitive the tangent angle is to changes in the model parameter value. In an embodiment, the process model is a data driven model comprising an empirical model and / or a machine learning model.

[0154] In an embodiment, the machine learning model is a convolutional neural network where the model parameters are weights and biases associated with multiple layers.

[0155] Similarly, in an embodiment, a system for calibrating a process model according to a process discussed herein (e.g., method 3000) is described. The process model can be configured to predict an image of a target pattern. The system includes: a measurement tool (e.g., Fig.16 and Fig.17 The SEM tool in the embodiment of the present invention is configured to obtain a reference image 3002 associated with a target pattern; and one or more processors ( Fig. 22 The processor is configured to calibrate the process model so that the process model produces a simulated image that (i) minimizes the intensity difference or frequency difference between the simulated image and the reference image 3002, and (ii) satisfies the gradient constraint 3004 associated with the reference image 3002.

[0156] In an embodiment, the processor is configured to calibrate the process model in an iterative manner. The iteration includes (a) executing the process model using the target pattern to generate a simulated image; (b) determining the intensity difference between the intensity values ​​of the simulated image and the reference image 3002, and / or transforming the simulated image and the reference image 3002 into the frequency domain via Fourier transform and determining the frequency difference between the frequencies associated with the simulated image and the reference image 3002; (c) determining the simulated gradient of the signal in the simulated image, wherein the signal is a signal along a given line passing through the simulated image; (d) determining whether the conditions are met: (i) the intensity difference or the frequency difference is minimized, and (ii) the simulated gradient satisfies the gradient constraint 3004 associated with the reference image 3002; (e) in response to not meeting the conditions (i) and (ii), adjusting the values ​​of the model parameters of the process model, and performing steps (a) to (d) until the conditions (i) and (ii) are met.

[0157] As mentioned previously, in an embodiment, the simulated gradient is determined by taking the first derivative of the signal along a given line through the simulated image. In an embodiment, the gradient constraint 3004 is obtained by taking the first derivative of the signal along a given line through the reference image 3002.

[0158] In an embodiment, the processor is further configured to extract a simulated contour from the simulated image and a reference contour from the reference image 3002, wherein the simulated contour and the reference contour are associated with the target pattern; and calibrate the process model so that the simulated contour satisfies a contour shape constraint, wherein the contour shape constraint ensures that the simulated contour matches the shape of the reference contour.

[0159] In an embodiment, determining whether the contour shape constraint is satisfied involves determining that a second order derivative of the simulated contour is within an expected range of a second order derivative of the reference contour.

[0160] In an embodiment, the reference image 3002 is obtained by simulating a physics-based model of a patterning process using a target pattern. The simulation may be performed on a processor. In an embodiment, the reference image 3002 may be obtained from a metrology tool (e.g., a SEM). In an embodiment, the reference image 3002 includes an aerial image of the target pattern; a resist image of the target pattern; and / or an etch image of the target pattern.

[0161] In an embodiment, the process model is configured to satisfy contour constraints 2004 defined relative to a printed contour of a pattern on a printed substrate, such as Fig.12 For example, each contour constraint is a function of the tangent angle between a tangent to a simulated contour at a given measurement location and the angle of the measurement data 2002 at the given location, wherein the simulated contour is the contour of a simulated pattern determined by executing the process model using the target pattern.

[0162] In an embodiment, a non-transitory computer-readable medium including instructions is provided, which when executed by one or more processors causes operations including: obtaining (i) resist pattern data 602 associated with a target pattern to be printed on a substrate, (ii) physical effect data 604 characterizing the effect of an etching process on the target pattern, and (iii) a measured deviation 606 between the resist pattern and an etched pattern formed on the printed substrate; and training the machine learning model based on the resist pattern data 602, the physical effect data 604, and the measured deviation 606 to reduce the difference between the measured deviation 606 and the predicted etching deviation. In addition, the non-transitory computer-readable medium may include instructions for Figure 6 Additional instructions discussed (eg, related to P601, P603, and P605).

[0163] In an embodiment, a non-transitory computer-readable medium including instructions is provided, which when executed by one or more processors results in operations including: obtaining (i) measurement data 2002 at a plurality of measurement locations on a pattern, and (ii) contour constraints 2004 specified based on the measurement data 2002; and calibrating a process model by adjusting values ​​of model parameters of the process model until the simulated contour satisfies the contour constraints 2004. In addition, the non-transitory computer-readable medium may include instructions regarding Fig.12 Additional instructions discussed (eg, associated with processes P2001, P2003, and P2005).

[0164] In an embodiment, a non-transitory computer-readable medium including instructions is provided, which when executed by one or more processors causes operations including: obtaining (i) a reference image 3002 associated with a target pattern, and (ii) a gradient constraint 3004 specified relative to the reference image 3002; and calibrating a process model so that the process model produces a simulated image that (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image 3002, and (ii) satisfies the gradient constraint 3004. In addition, the non-transitory computer-readable medium may include instructions for Fig.14 Additional instructions discussed (eg, associated with processes P3001, P3003, and P3005).

[0165] According to the present disclosure, combinations and sub-combinations of the disclosed elements constitute separate embodiments. For example, a first combination includes determining an etch profile based on a trained machine learning model. In another example, a combination includes determining a simulation profile based on a model calibrated according to physical constraints.

[0166] In some embodiments, a scanning electron microscope (SEM) obtains images of structures (eg, some or all structures of a device) that are exposed or transferred onto a substrate. Fig.16 Depicted is an embodiment of a SEM 200. A primary electron beam 202 emitted from an electron source 201 is converged by a condenser lens 203 and then passes through a beam deflector 204, an EB deflector 205, and an objective lens 206 to irradiate a substrate 100 on a substrate stage 101 at a focal point.

[0167] When the substrate 100 is irradiated with the electron beam 202, secondary electrons are generated from the substrate 100. The secondary electrons are deflected by the E×B deflector 205 and detected by the secondary electron detector 207. A two-dimensional electron beam image can be obtained by detecting the electrons generated from the sample in synchronization with, for example, two-dimensionally scanning the electron beam by the beam deflector 204 or repeatedly scanning the electron beam 202 in the X direction or the Y direction by the beam deflector 204, and continuously moving the substrate 100 in the other of the X direction or the Y direction by the substrate stage 101.

[0168] The signal detected by the secondary electron detector 207 is converted into a digital signal by an analog / digital (A / D) converter 208, and the digital signal is sent to the image processing system 300. In an embodiment, the image processing system 300 may have a memory 303 for storing all or part of the digital image for processing by the processing unit 304. The processing unit 304 (e.g., specially designed hardware or a combination of hardware and software) is configured to convert or process the digital image into a data set representing the digital image. In addition, the image processing system 300 may have a storage medium 301 configured to store the digital image and the corresponding data set in a reference database. The display device 302 can be connected to the image processing system 300 so that the operator can perform necessary operations of the equipment with the help of a graphical user interface.

[0169] Fig.17 Another embodiment of the inspection apparatus is schematically illustrated. The system is used to inspect a sample 90 (such as a substrate) on a sample platform 89 and includes a charged particle beam generator 81, a condenser lens module 82, a probe forming objective lens module 83, a charged particle beam deflection module 84, a secondary charged particle detector module 85, and an image forming module 86.

[0170] The charged particle beam generator 81 generates a primary charged particle beam 91. The condenser lens module 82 condenses the generated primary charged particle beam 91. The probe forming objective lens module 83 focuses the converged primary charged particle beam into a charged particle beam probe 92. The charged particle beam deflection module 84 scans the formed charged particle beam probe 92 across the surface of the area of ​​interest on the sample 90 fastened on the sample platform 89. In an embodiment, the charged particle beam generator 81, the condenser lens module 82 and the probe forming objective lens module 83 or their equivalent designs, alternatives or any combination thereof together form a charged particle beam probe generator that generates a scanning charged particle beam probe 92.

[0171] The secondary charged particle detector module 85 detects secondary charged particles 93 emitted from the sample surface after being bombarded by the charged particle beam probe 92 (and possibly together with other reflected or scattered charged particles from the sample surface) to generate a secondary charged particle detection signal 94. An image forming module 86 (e.g., a computing device) is coupled to the secondary charged particle detector module 85 to receive the secondary charged particle detection signal 94 from the secondary charged particle detector module 85 and to form at least one scan image accordingly. In an embodiment, the secondary charged particle detector module 85 and the image forming module 86 or their equivalent designs, alternatives, or any combination thereof together form an image forming device that forms a scan image from the detected secondary charged particles emitted by the sample 90 bombarded by the charged particle beam probe 92.

[0172] As mentioned above, SEM images can be processed to extract contours in the image that describe the edges of objects representing device structures. These contours are then quantified via indicators such as CD. Thus, images of device structures are usually compared and quantified via simplistic indicators such as distance between edges (CD) or simple pixel differences between images. Typical contour models that detect the edges of objects in images in order to measure CD use image gradients. In practice, those models rely on strong image gradients. But in practice, images are usually noisy and have discontinuous boundaries. Techniques such as smoothing, adaptive thresholding, edge detection, erosion and dilation can be used to process the results of image gradient contour models to address noisy and discontinuous images, but will ultimately result in low-resolution quantization of high-resolution images. Thus, in most instances, mathematical processing of images of device structures to reduce noise and automated edge detection will result in a loss of resolution of the image, thereby resulting in a loss of information. Therefore, the result is a low-resolution quantization equivalent to an oversimplified representation of a complex high-resolution structure.

[0173] Thus, it is desirable to have a mathematical representation that can preserve resolution and can also describe the general shape of a structure (e.g., a circuit feature, an alignment mark, or a portion of a metrology target (e.g., a grating feature), etc.) produced or expected to be produced using a patterning process, whether, for example, the structure is in a latent resist image, in a developed resist image, or in a layer transferred to a substrate, such as by etching. In the context of a lithography or other patterning process, the structure can be a device being manufactured or a portion thereof, and the image can be a SEM image of the structure. In some cases, the structure can be a feature of a semiconductor device (e.g., an integrated circuit). In some cases, the structure can be an alignment mark or a portion thereof (e.g., a grating of an alignment mark) used in an alignment measurement process to determine the alignment of an object (e.g., a substrate) with another object (e.g., a patterning device), or a metrology target or a portion thereof (e.g., a grating of a metrology target) used to measure parameters of the patterning process (e.g., overlay, focus, dose, etc.). In an embodiment, the metrology target is a diffraction grating used to measure (e.g.,) overlay.

[0174] In an embodiment, Figure 3 In the method, measured data related to the printed pattern is used to train the model. The trained model can also be used in optimizing the patterning process or adjusting the parameters of the patterning process. For example, the problem solved by OPC is that the final size and positioning of the image of the design layout projected onto the substrate will not be consistent with or not only depend on the size and positioning of the design layout on the pattern forming device. Note that the terms "mask", "reticle", and "patterning device" are interchangeable here. In addition, those skilled in the art will recognize that, especially in the case of lithography simulation / optimization, the terms "mask" / "patterning device" and "design layout" can be interchangeable, because in lithography simulation / optimization, the physical pattern forming device is not necessarily used, but the design layout can be used to represent the physical pattern forming device. For small feature sizes and high feature densities that appear on some design layouts, the location of a particular edge of a given feature will be affected to some extent by the presence or absence of other adjacent features. These proximity effects are caused by the tiny amount of light coupled from one feature to another and / or by non-geometric optical effects (such as diffraction and interference). Similarly, proximity effects may result from diffusion and other chemical effects during the post-exposure bake (PEB), resist development, and etching that typically follows photolithography.

[0175] In order to ensure that the projected image of the design layout is in accordance with the requirements of a given target circuit design, it is necessary to use complex numerical models, corrections or pre-distortions of the design layout to predict and compensate for proximity effects. The paper "Full-Chip Lithography Simulation and Design Analysis-how OPC Is Changing IC Design" (C. Spence, Proc. SPIE, Vol. 5751, pp. 1-14 (2005)) provides an overview of the current "model-based" optical proximity effect correction process. In a typical high-end design, almost every feature of the design layout has some modification in order to achieve high fidelity of the projected image to the target design. These modifications can include offsets or biases of edge positions or line widths and the application of "auxiliary" features that are expected to assist the projection of other features.

[0176] With millions of features typically present in chip designs, applying model-based OPC to target designs involves good process models and considerable computational resources. However, applying OPC is generally not an "exact science" but rather an empirical iterative process that does not always compensate for all possible proximity effects. Therefore, it is necessary to verify the effects of OPC, such as design layout after applying OPC and any other RET, by design checks such as intensive full-chip simulations using calibrated numerical process models, in order to minimize the possibility of building design flaws into the patterning device pattern. This is driven by the huge cost of manufacturing high-end patterning devices, which is in the millions of dollars range; and the impact on turnaround time, which is caused by reworking or repairing the actual patterning devices (once they have been manufactured).

[0177] Both OPC and full-chip RET verification can be based on numerical modeling systems and methods, as described, for example, in U.S. patent application Ser. No. 10 / 815,573, which is incorporated herein by reference in its entirety, and in a paper by Y. Cao et al. entitled “Optimized Hardware and Software For Fast, Full Chip Simulation” (Proc. SPIE, Vol. 5754, No. 405 (2005)).

[0178] One type of RET is related to the adjustment of the global deviation of the design layout. The global deviation is the difference between the pattern in the design layout and the pattern intended to be printed on the substrate. For example, a 25nm diameter circular pattern can be printed on the substrate by a 50nm diameter pattern in the design layout, or printed on the substrate with a large dose by a 20nm diameter pattern in the design layout.

[0179] In addition to the optimization of the design layout or pattern forming device (such as OPC), the illumination source can also be optimized, either together with the pattern forming device optimization or separately, to improve the overall lithography fidelity. In this article, the terms "irradiation source" and "source" can be used interchangeably. Since the 1990s, many off-axis illumination sources (such as annular, quadrupole and dipole) have been introduced, and greater freedom has been provided for OPC design, thereby improving the imaging results. It is known that off-axis illumination is a proven way to distinguish the fine structures (i.e., target features) contained in the pattern forming device. However, when compared with traditional illumination sources, off-axis illumination sources generally provide lower light intensity for aerial images (AI). Therefore, it is necessary to try to optimize the illumination source to obtain an optimized balance between finer resolution and reduced light intensity.

[0180] For example, in the article entitled "Optimum Mask and Source Patterns to Print A Gen Shape" by Rosenbluth et al., Journal of Microlithography, Microfabrication, Microsystems 1(1), pp.13-20, (2002), many illumination source optimization methods can be found. The source is subdivided into multiple regions, each corresponding to a specific region of the pupil spectrum. Afterwards, it is assumed that the source distribution is uniform in each source region, and the brightness of each region is optimized for the process window. However, such an assumption that "the source distribution is uniform in each source region" is not always valid, so the effectiveness of this method is affected. In another example described in the article entitled "Source Optimization for Image Fidelity and Throughput" by Granik, Journal of Microlithography, Microfabrication, Microsystems 3(4), pp.509-522, (2004), several existing source optimization methods are reviewed, and an illuminator pixel-based method is proposed, which converts the source optimization problem into a series of non-negative least squares optimizations. While these methods have demonstrated some success, they typically require multiple complex iterations to converge. Additionally, it may be difficult to determine suitable / optimized values ​​for some of the additional parameters (such as γ in the Granik method) that dictate the tradeoff between optimizing the source for substrate image fidelity and the smoothness requirement of the source.

[0181] For low k1 lithography, optimization of the source and patterning device is very useful to ensure a feasible process window for projection of critical circuit patterns. Some algorithms (e.g., Socha et al., Proc. SPIE, Vol. 5853, 2005, p. 180) discretize the illumination into independent source points and discretize the mask into diffraction orders in the spatial frequency domain, and independently formulate a cost function (which is defined as a function of the selected design variables) based on a process window metric (such as exposure latitude), which can be predicted by the source point intensity and the patterning device diffraction order by an optical imaging model. The term "design variables" used herein includes a set of parameters of a lithographic projection device or a lithographic process, such as parameters that can be adjusted by a user of the lithographic projection device, or image characteristics that can be adjusted by the user by adjusting those parameters. It should be recognized that any characteristics of the lithographic projection device (including those in the source, patterning device, projection optical element and / or resist characteristics) can be among the design variables in the optimization. The cost function is usually a nonlinear function of the design variables. Standard optimization techniques are then used to minimize the cost function.

[0182] Relatedly, the pressure of ever-decreasing design rules has driven semiconductor chipmakers deeper into the era of low k1 lithography with existing 193nm ArF lithography. The move toward lower k1 lithography places high demands on resolution enhancement technology (RET), exposure tools, and the need for lithography-friendly design. Ultra-high numerical aperture (NA) exposure tools of 1.35ArF may be used in the future. To help ensure that the circuit design can be printed onto the substrate with a workable process window, source-patterning device optimization (referred to herein as source-mask optimization or SMO) becomes an important RET required for the 2x nm node.

[0183] Source and pattern forming device (design layout) optimization method and system allow the use of cost functions to simultaneously optimize the source and pattern forming device without constraints and in a practical amount of time. It is described in the commonly assigned International Patent Application No. PCT / US2009 / 065359 filed on November 20, 2009 and published as WO2010 / 059954, entitled "Fast Freeform Source and Mask Co-Optimization Method", which is incorporated herein by reference in its entirety.

[0184] Another source and mask optimization method and system involves optimizing the source by adjusting source pixels, which is described in commonly assigned U.S. patent application No. 12 / 813,456, filed on June 10, 2010, and U.S. patent application publication No. 2010 / 0315614, entitled "Source-Mask Optimization in Lithographic Apparatus," the entire contents of which are incorporated herein by reference.

[0185] In a lithographic projection apparatus, as an example, the cost function is expressed as:

[0186]

[0187] where (z1,z2,...,z N ) are N design variables or their values. p (z1,z2,...,z N ) can be the design variables (z1,z2,...,z N ), such as (z1,z2,...,z N ) is the difference between the actual value and the expected value of a characteristic at the evaluation point of a set of values ​​of the design variable. p is with f p (z1,z2,...,z N ) is associated with a weight constant. A higher w may be assigned to an evaluation point or pattern that is more critical than other evaluation points or patterns. p It is also possible to assign higher w values ​​to patterns and / or evaluation points with larger numbers of occurrences. p Examples of evaluation points may be any physical point or pattern on the substrate, any point on the virtual design layout, or a resist image, or an aerial image, or a combination thereof. p (z1,z2,...,z N ) can also be a function of one or more random effects such as LWR, where the one or more random effects are design variables (z1, z2, ..., z N). The cost function may represent any suitable characteristic of the lithographic projection apparatus or substrate, such as a failure rate of a feature, focal length, CD, image shift, image deformation, image rotation, random effects, throughput, CDU, or a combination thereof. CDU is the local CD variation (e.g., three times the standard deviation of the local CD distribution). CDU may be interchangeably referred to as LCDU. In one embodiment, the cost function represents CDU, throughput, and random effects (i.e., is a function of CDU, throughput, and random effects). In one embodiment, the cost function represents EPE, throughput, and random effects (i.e., is a function of EPE, throughput, and random effects). In one embodiment, the design variables (z1, z2, ..., z N ) includes dose, global bias of the patterning device, shape of the illumination from the source, or a combination thereof. Since the resist image often dictates the circuit pattern on the substrate, the cost function often includes a function representing some characteristic of the resist image. For example, f at such an evaluation point p (z1,z2,...,z N ) can be simply the distance between a point in the resist image and the expected position of that point (i.e., the edge placement error EPE p (z1,z2,...,z N )). The design variables can be any adjustable parameters, such as adjustable parameters of the source, pattern forming device, projection optics, dose, focal length, etc. The projection optics may include components collectively referred to as "wavefront manipulators", which can be used to adjust the shape of the wavefront and intensity distribution and / or phase shift of the illumination beam. The projection optics can preferably adjust the wavefront and intensity distribution at any location along the optical path of the lithographic projection device (such as before the pattern forming device, near the pupil plane, near the image plane, near the focal plane). The projection optics can be used to correct or compensate for certain deformations of the wavefront and intensity distribution caused by (for example) temperature changes in the source, pattern forming device, lithographic projection equipment, thermal expansion of components of the lithographic projection equipment. Adjusting the wavefront and intensity distribution can change the values ​​of the evaluation points and the cost function. These changes can be simulated from a model or actually measured. Of course, CF(z1,z2,...,z N ) is not limited to the form in equation 1. CF(z1,z2,...,z N ) may be in any other suitable form.

[0188] It should be noted that f p (z1,z2,...,z N ) is defined as Therefore, minimizing f p (z1,z2,...,z N) is equivalent to minimizing the cost function defined in Equation 1 Therefore, for simplicity of notation in this article, Equation 1 and f may be used interchangeably. p (z1,z2,...,z N )’s weighted RMS.

[0189] In addition, if the maximum process window (PW) is considered, the same entity part from different PW conditions can be regarded as different evaluation points of the cost function in (Equation 1). For example, if N PW conditions are considered, the evaluation points can be classified according to their PW conditions and the cost function can be written as:

[0190]

[0191] In the uth PW condition u=1,...,U, f pu (z1,z2,...,z N ) is f p (z1,z2,...,z N ) value. When f p (z1,z2,...,z N ) is EPE, then minimizing the above cost function is equivalent to minimizing the edge shift under various PW conditions, and thus, this case leads to maximizing PW. Specifically, if PW also consists of different mask deviations, then minimizing the above cost function also includes minimizing the mask error enhancement factor (MEEF), which is defined as the ratio between the substrate EPE and the induced mask edge deviation.

[0192] The design variables may have constraints which may be expressed as (z1,z2,...,z N)∈Z, where Z is a set of possible values ​​of the design variables. A possible constraint on the design variables can be imposed by the desired throughput of the lithographic projection device. The desired throughput may limit the dose and thus have an impact on the random effect (e.g., impose a lower limit on the random effect). Higher throughput generally leads to lower doses, shorter exposure times, and larger random effects. Consideration of the minimization of substrate throughput and random effects can constrain the possible values ​​of the design variables, because the random effect is a function of the design variables. In the absence of such constraints imposed by the desired throughput, the optimization may obtain an unrealistic set of values ​​for the design variables. For example, if the dose is among the design variables, then in the absence of such constraints, the optimization may obtain a dose value that makes the throughput economically impossible. However, the usefulness of the constraint should not be interpreted as necessity. The throughput may be affected by the adjustment of the parameters of the patterning process based on the failure rate. It is expected to have a lower failure rate of the feature while maintaining a high throughput. The throughput may also be affected by the chemical properties of the resist. Slower resists (e.g., resists that require a higher amount of light to properly expose) lead to lower throughput. Therefore, appropriate parameters for the patterning process may be determined based on an optimization process involving failure rates of features due to resist chemistry or fluctuations, and dose requirements for higher throughput.

[0193] Therefore, the optimization process is performed under the constraints (z1,z2,...,z N )∈Z to find the set of values ​​of the design variables that minimize the cost function, that is, to find:

[0194]

[0195] Fig.18A general method for optimizing the lithographic projection apparatus according to an embodiment is illustrated in FIG. This method includes a step S1202 of defining a multivariate cost function of multiple design variables. The design variables may include any suitable combination selected from the characteristics (1200A) of the illumination source (e.g., the pupil filling ratio, i.e., the percentage of the radiation of the source that passes through the pupil or aperture), the characteristics (1200B) of the projection optics, and the characteristics (1200C) of the design layout. For example, the design variables may include the characteristics (1200A) of the illumination source and the characteristics (1200C) of the design layout (e.g., the global deviation), but not the characteristics (1200B) of the projection optics, which results in SMO. Alternatively, the design variables may include the characteristics (1200A) of the illumination source, the characteristics (1200B) of the projection optics, and the characteristics (1200C) of the design layout, which results in source-mask-lens optimization (SMLO). In step S1204, the design variables are adjusted simultaneously so that the cost function moves toward convergence. In step S1206, it is determined whether a predefined termination condition is met. The predetermined termination condition may include various possibilities, namely, the cost function may be minimized or maximized (as required by the numerical technique used), the value of the cost function has equaled a threshold or has exceeded a threshold, the value of the cost function has reached a preset error limit, or a preset number of iterations has been reached. If any of the conditions in step S1206 is met, the method ends. If none of the conditions in step S1206 is met, steps S1204 and S1206 are iteratively repeated until the desired result is obtained. The optimization does not necessarily result in a single set of values ​​for the design variables, because there may be physical inhibitions caused by factors such as failure rate, pupil fill factor, resist chemistry, throughput, etc. The optimization may provide multiple sets of values ​​for the design variables and associated performance characteristics (e.g., throughput), and allow a user of the lithographic apparatus to select one or more sets.

[0196] In a lithographic projection apparatus, the source, the patterning device, and the projection optics may be optimized alternately (referred to as alternating optimization), or may be optimized simultaneously (referred to as simultaneous optimization). The terms "simultaneous," "simultaneously," "jointly," and "jointly" as used herein mean that the design variables of the characteristics of the source, the patterning device, the projection optics, and / or any other design variables are allowed to be changed simultaneously. The terms "alternating" and "alternatingly" as used herein mean that not all design variables are allowed to be changed simultaneously.

[0197] exist Fig.19 In the process, the optimization of all design variables is performed simultaneously. This process can be called a simultaneous process or a joint optimization process. Alternatively, the optimization of all design variables is performed alternately, such as Fig.19In such a process, in each step, some design variables are fixed while other design variables are optimized to minimize the cost function; then, in the next step, a different set of variables is fixed while other sets of variables are optimized to minimize the cost function. These steps are performed alternately until convergence or some termination conditions are met.

[0198] like Fig.19 As shown in the non-limiting example flow chart of , first, a design layout is obtained (step S1302), then, in step S1304, a source optimization step is performed, in which all design variables of the illumination source (SO) are optimized to minimize the cost function, while all other design variables are fixed. Then in the next step S1306, a mask optimization (MO) is performed, in which all design variables of the pattern forming device are optimized to minimize the cost function, while all other design variables are fixed. Such two steps are performed alternately until certain termination conditions are met in step S1308. Various termination conditions can be used, such as, the value of the cost function becomes equal to a threshold, the value of the cost function crosses beyond a threshold, the value of the cost function reaches a preset error limit, or reaches a preset number of iterations, etc. It should be noted that SO-MO alternating optimization is used as an example of the alternative flow. The alternative flow can take many different forms, such as: SO-LO-MO alternating optimization, in which SO, LO (lens optimization) and MO are performed alternately and iteratively; or the first SMO can be performed once, and then LO and MO are performed alternately and iteratively; etc. Finally, the output of the optimization result is obtained in step S1310, and the process stops.

[0199] As discussed previously, the pattern selection algorithm can be integrated with simultaneous or alternating optimization. For example, when alternating optimization is employed, full chip SO can be performed first, "hot spots" and / or "warm spots" can be identified, and then MO can be performed. In view of the present disclosure, numerous permutations and combinations of sub-optimizations are possible in order to achieve the desired optimization results.

[0200] Fig. 20AAn exemplary optimization method is shown in which a cost function is minimized. In step S502, initial values ​​of the design variables are obtained, including tuning ranges for the design variables (if any). In step S504, a multivariate cost function is set. In step S506, the cost function is expanded in a sufficiently small neighborhood around the starting values ​​of the design variables for the first iteration step (i=0). In step S508, standard multivariate optimization techniques are applied to minimize the cost function. It should be noted that the optimization problem can impose constraints, such as tuning ranges, during the optimization process in S508 or later in the optimization process. Step S520 indicates that each iteration is performed for a given test pattern (also referred to as a "gauge") for the identified evaluation point that has been selected for optimizing the lithography process. In step S510, the lithography response is predicted. In step S512, the results of step S510 are compared with the expected or ideal lithography response value obtained in step S522. If the termination condition is met in step S514, that is, the optimization produces a lithographic response value that is sufficiently close to the desired value, the final value of the design variable is output in step S518. The output step may also include using the final value of the design variable to output other functions, such as outputting a wavefront aberration adjusted map at the pupil plane (or other plane), an optimized source map, and an optimized design layout, etc. If the termination condition is not met, then in step S516, the value of the design variable is updated using the result of the i-th iteration, and the process returns to step S506. The following is a detailed description of the output step. Fig. 20A process.

[0201] In the exemplary optimization process, no assumptions or approximations are made about the design variables (z1, z2, ..., z N ) and f p (z1,z2,...,z N ) except for f p (z1,z2,...,z N ) is sufficiently smooth (e.g., there is a first-order derivative ), which is generally effective in lithography projection equipment. Algorithms such as Gauss-Newton algorithm, Levenberg-Marquardt algorithm, gradient descent algorithm, simulated annealing, genetic algorithm, etc. can be applied to find

[0202] Here, the Gauss-Newton algorithm is used as an example. The Gauss-Newton algorithm is an iterative method suitable for general nonlinear multivariable optimization problems. N ) value (z 1i ,z 2i ,...,z Ni) in the i-th iteration, the Gauss-Newton algorithm 1i ,z 2i ,...,z Ni ) in the neighborhood of p (z1,z2,...,z N ), and then calculate (z 1i ,z 2i ,...,z Ni ) in the neighborhood of N ) value (z 1(i+1) ,z 2(i+1) ,...,z N(i+1) )). Design variables (z1,z2,...,z N ) takes the value ((z 1(i+1) ,z 2(i+1) ,...,z N(i+1) )). This iteration continues until convergence (i.e., CF(z1,z2,...,z N )) no longer decreases) or until the preset number of iterations is reached.

[0203] In particular, in the i-th iteration, in (z 1i ,z 2i ,...,z Ni ) in the neighborhood of

[0204]

[0205] In the case of the approximation of Equation 3, the cost function becomes:

[0206]

[0207] It is the design variable (z1,z2,...,z N ). Except for the design variables (z1, z2, ..., z N ) all other terms are constants.

[0208] If the design variables (z1,z2,...,z N ) is not under any constraints, then (z 1(i+1) ,z 2(i+1) ,...,z N(i+1) )) can be derived by solving N linear equations: Where n=1,2,...N.

[0209] If the design variables (z1,z2,...,z N ) is in the form of J inequalities (for example, (z1,z2,...,z N) is constrained by the tuning range where j = 1, 2, ... J); and under the constraints of K equations (e.g., the interdependencies between design variables) where k = 1, 2, ... K); the optimization process becomes a classic quadratic programming problem, where A nj , B j , C nk , D k is a constant. Additional constraints can be imposed for each iteration. For example, a “damping factor” ΔD can be introduced to limit (z 1(i+1) ,z 2(i+1) ,...,z N(i+1) ) and (z 1i ,z 2i ,...,z Ni ), so that the approximation of Equation 3 holds. Such a constraint can be expressed as z ni -Δ D ≤z n ≤z ni +Δ D (z) can be derived using, for example, the method described in Numerical Optimization (2nd ed.) by Jorge Nocedal and Stephen J. Wright (Berlin-New York: Van den Berg, Cambridge University Press). 1(i+1) ,z 2(i+1) ,...,z N(i+1) )).

[0210] Instead of making f p (z1,z2,...,z N ), the optimization process can minimize the magnitude of the maximum deviation (worst defect) among the evaluation points to their expected values. In such a method, the cost function can be alternatively expressed as:

[0211]

[0212] Among them, CL p It is for f p (z1,z2,...,z N ). This cost function represents the worst defect among the evaluation points. Optimization using this cost function minimizes the magnitude of the worst defect. An iterative greedy algorithm can be used for this optimization.

[0213] The cost function of Equation 5 can be approximated as:

[0214]

[0215] Where q is an even positive integer, such as at least 4, preferably at least 10. Equation 6 mimics the behavior of Equation 5 while allowing the optimization to be performed analytically and accelerated by using methods such as the deepest descent method, conjugate gradient method, and the like.

[0216] Minimizing the worst defect size can also be done with f p (z1,z2,...,z N ). Specifically, as in Equation 3, the approximation f p (z1,z2,...,z N ). Next, the constraint on the worst defect size is written as inequality E Lp ≤f p (z1,z2,...,z N )≤E Up , where E Lp and E Up Is to specify f p (z1,z2,...,z N ) are two constants for the minimum and maximum allowed deviations. Inserting these constraints into equation 3 transforms them into the following equations: (where p = 1, ... P),

[0217]

[0218] and

[0219]

[0220] Because Equation 3 is usually only valid for (z 1i ,z 2i ,...,z Ni ), so the desired constraint E cannot be achieved in this neighborhood. Lp ≤f p (z1,z2,...,z N )≤E Up (which may be determined by any conflicts in the inequalities), then the constant E can be relaxed Lp and E Up Until the constraints can be achieved. This optimization process minimizes (z 1i ,z 2i ,...,z Ni ) neighborhood. Then, each step gradually reduces the worst defect size, and each step is performed iteratively until some termination condition is met. This situation will lead to the best reduction of the worst defect size.

[0221] Another way to minimize the worst defect is to adjust the weight w in each iteration. pFor example, after the i-th iteration, if the r-th evaluation point is the worst defect, then w can be increased in the (i+1)-th iteration. r , so that the reduction of the defect size of the evaluation point is given higher priority.

[0222] In addition, the cost functions in Equations 4 and 5 can be modified by introducing Lagrange multipliers to achieve a trade-off between optimizing the RMS of defect size and optimizing the worst defect size, i.e.,

[0223]

[0224] Wherein λ is a preset constant that specifies the tradeoff between the optimization of the RMS of the defect size and the optimization of the worst defect size. Specifically, if λ=0, this equation becomes Equation 4, and only the RMS of the defect size is minimized; while if λ=1, this equation becomes Equation 5, and only the worst defect size is minimized; if 0<λ<1, the above two cases are considered in the optimization. A variety of methods can be used to solve this optimization. For example, similar to the method described previously, the weighting in each iteration can be adjusted. Alternatively, similar to minimizing the worst defect size from the inequality, the inequalities of Equations 6′ and 6″ can be regarded as constraints on the design variables during the solution of the quadratic programming problem. Then, the bounds on the worst defect size can be incrementally relaxed, or the weights for the worst defect size can be incrementally increased, the cost function values ​​for each achievable worst defect size are calculated, and the design variable values ​​that minimize the overall cost function are selected as the initial point for the next step. By performing this operation iteratively, minimization of this new cost function can be achieved.

[0225] Optimizing the lithography projection equipment can expand the process window. A larger process window provides more flexibility in process design and chip design. The process window can be defined as a set of focal length and dose values ​​that make the resist image within a certain limit of the design target of the resist image. It should be noted that all methods discussed here can also be extended to a generalized process window definition that can be established by different or additional base parameters other than exposure dose and defocus. These base parameters may include (but are not limited to) optical settings such as NA, mean square deviation, aberration, polarization, or optical constants of the resist layer. For example, as described earlier, if the PW is also composed of different mask deviations, the optimization includes minimization of the mask error enhancement factor (MEEF), which is defined as the ratio between the substrate EPE and the induced mask edge deviation. The process window defined for the focal length and dose value is used as an example only in this disclosure. The following describes a method for maximizing the process window according to an example.

[0226] In the first step, starting from the known conditions (f0, ε0) in the process window (where f0 is the nominal focal length and ε0 is the nominal dose), one of the following cost functions in the domain (f0±Δf, ε0±Δε) is minimized:

[0227]

[0228] or

[0229]

[0230] or

[0231]

[0232] If the nominal focal length f0 and the nominal dose ε0 are allowed to shift, they can be related to the design variables (z1,z2,...,z N ) are jointly optimized. In the next step, if (z1,z2,...,z N ,f,ε), then (f0±Δf,ε0±Δε) is accepted as part of the process window so that the cost function is within the preset limits.

[0233] Alternatively, if the focus and dose are not allowed to shift, the design variables (z1, z2, ..., z N In an alternative embodiment, if (z1, z2, ..., z N ), then accept (f0±Δf,ε0±Δε) as part of the process window so that the cost function is within the preset limits.

[0234] The methods described hereinabove can be used to minimize the corresponding cost functions of Eq. 7, 7′, or 7″. If the design variables are characteristics of the projection optics, such as Zernike coefficients, then minimizing the cost functions of Eq. 7, 7′, or 7″ results in maximizing the process window based on projection optics optimization (i.e., LO). If the design variables are characteristics of the source and patterning device in addition to the characteristics of the projection optics, then minimizing the cost functions of Eq. 7, 7′, or 7″ results in maximizing the process window based on SMLO, such as Fig.19 If the design variables are characteristics of the source and the pattern forming device, minimizing the cost function of equation 7, 7' or 7" results in maximizing the SMO-based process window. The cost function of equation 7, 7' or 7" may also include at least one f p (z1,z2,...,z N ), such as f in Equation 7 or Equation 8 p (z1,z2,...,zN ), which is a function of one or more random effects, such as local CD variation or LWR of the 2D features, and throughput.

[0235] Fig.21 A specific example of how a simultaneous SMLO process can use the Gauss-Newton algorithm for optimization is shown. In step S702, starting values ​​for the design variables are identified. The tuning range for each variable can also be identified. In step S704, a cost function is defined using the design variables. In step S706, the cost function is expanded around the starting values ​​for all evaluation points in the design layout. In optional step S710, a full-chip simulation is performed to cover all critical patterns in the full-chip design layout. In step S714, the desired lithography response metrics (such as CD or EPE) are obtained, and in step S712, the desired lithography response metrics are compared with the predicted values ​​of those quantities. In step S716, the process window is determined. Steps S718, S720, and S722 are similar to those described with respect to Fig. 20A The corresponding steps S514, S516 and S518 described. As mentioned before, the final output may be a wavefront aberration map in the pupil plane, which is optimized to produce the desired imaging performance. The final output may also be an optimized source map and / or an optimized design layout.

[0236] Fig. 20B An exemplary method for optimizing a cost function is shown, where the design variables (z1, z2, ..., z N ) includes design variables that can take only discrete values.

[0237] The method begins by defining pixel groups of an illumination source and patterning device pattern blocks of a patterning device (step S802). In general, pixel groups or patterning device pattern blocks may also be referred to as partitions of a lithography process component. In one exemplary method, the illumination source is partitioned into 117 pixel groups, and 94 patterning device pattern blocks are defined for the patterning device (substantially as described above), resulting in a total of 211 partitions.

[0238] In step S804, a lithography model is selected as a basis for lithography simulation. The lithography simulation produces results for calculating lithography indicators or responses. Specific lithography indicators are defined as performance indicators to be optimized (step S806). In step S808, initial (pre-optimization) conditions for the illumination source and the patterning device are set. The initial conditions include the initial states of the patterning device pattern blocks for the pixel groups of the illumination source and the patterning device, so that the initial illumination shape and the initial patterning device pattern can be referenced. The initial conditions may also include mask deviations, NA, and focus slope ranges (or focus gradient ranges). Although steps S802, S804, S806, and S808 are depicted as consecutive steps, it should be understood that in other embodiments of the present invention, these steps may be performed in other orders.

[0239] In step S810, the pixel groups and patterning device blocks are sorted. The pixel groups and patterning device blocks may be interleaved in the sorting. Various sorting methods may be used, including: continuously (e.g., from pixel group 1 to pixel group 117 and from patterning device block 1 to patterning device block 94), randomly, based on the physical location of the pixel groups and patterning device blocks (e.g., sorting pixel groups closer to the center of the illumination source higher), and based on how changes to the pixel groups or patterning device blocks affect performance indicators.

[0240] Once the pixel groups and patterning device pattern blocks are sorted, the illumination source and patterning device are adjusted to improve the performance index (step S812). In step S812, each of the pixel groups and patterning device pattern blocks is analyzed in sorted order to determine whether a change in the pixel group or patterning device pattern block will result in an improved performance index. If it is determined that the performance index will be improved, the pixel group or patterning device pattern block is changed accordingly, and the resulting improved performance index and the modified illumination shape or modified patterning device pattern form a baseline for comparison for subsequent analysis of lower sorted pixel groups and patterning device pattern blocks. In other words, the changes to improve the performance index are maintained. As changes to the states of the pixel groups and patterning device pattern blocks are made and maintained, the initial illumination shape and the initial patterning device pattern are changed accordingly, so that the modified illumination shape and the modified patterning device pattern result from the optimization process in step S812.

[0241] In other methods, patterning device polygon shape adjustment and pairwise polling of pixel groups and / or patterning device pattern blocks are also performed within the optimization process of S812.

[0242] In an alternative embodiment, the staggered simultaneous optimization process may include changing the pixel groups of the illumination source and if an improvement in the performance metric is found, stepping up and down the dose to look for further improvement. In another alternative, the stepping up and down of the dose or intensity may be replaced by a change in the deviation of the patterning device pattern to look for further improvement in the simultaneous optimization process.

[0243] In step S814, it is determined whether the performance metric has converged. For example, if little or no improvement in the performance metric has been witnessed in the last several iterations of steps S810 and S812, the performance metric may be considered to have converged. If the performance metric has not converged, steps S810 and S812 are repeated in the next iteration, with the modified illumination shape and modified patterning device from the current iteration being used as the initial illumination shape and initial patterning device for the next iteration (step S816).

[0244] The optimization method described above may be used to increase the throughput of a lithographic projection apparatus. For example, the cost function may include f as a function of exposure time. p (z1,z2,...,z N ). The optimization of such a cost function is preferably constrained or influenced by the measurement of random effects or other indicators. Specifically, a computer-implemented method for increasing the throughput of a lithography process may include optimizing a cost function as a function of one or more random effects of the lithography process and as a function of the exposure time of the substrate so as to minimize the exposure time.

[0245] In one embodiment, the cost function includes at least one f as a function of one or more random effects p (z1,z2,...,z N ). Random effects can include failure of features, such as Figure 3 The random effects include random variations in features of the resist image, such as SEPE, local CD variations of 2D features, or LWR. In one embodiment, the random effects include random variations in features of the resist image. For example, these random variations may include failure rates of features, line edge roughness (LER), line width roughness (LWR), and critical dimension uniformity (CDU). Including random variation in the cost function allows finding values ​​of the design variables that minimize the random variation, thereby reducing the risk of defects due to random effects.

[0246] Fig. 22A block diagram of a computer system 100 is provided to illustrate a method and process for performing the optimization disclosed herein. The computer system 100 includes a bus 102 or other communication mechanism for communicating information; and a processor 104 (or multiple processors 104 and 105) coupled to the bus 102 for processing information. The computer system 100 also includes a main memory 106 (such as a random access memory (RAM) or other dynamic storage device) coupled to the bus 102 for storing information and instructions executed by the processor 104. The main memory 106 can also be used to store temporary variables or other intermediate information during the execution of instructions executed by the processor 104. The computer system 100 also includes a read-only memory (ROM) 108 or other static storage device coupled to the bus 102 for storing static information and instructions for the processor 104. A storage device 110 (such as a magnetic disk or optical disk) is provided and coupled to the bus 102 for storing information and instructions.

[0247] The computer system 100 may be coupled to a display 112 (such as a cathode ray tube (CRT) or a flat panel or touch panel display) via the bus 102 for displaying information to a computer user. An input device 114 (including alphanumeric keys and other keys) is coupled to the bus 102 for communicating information and command selections to the processor 104. Another type of user input device is a cursor controller 116 (such as a mouse, trackball, or cursor direction keys) for communicating direction information and command selections to the processor 104 and for controlling cursor movement on the display 112. This input device typically has two degrees of freedom in two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a position in a plane. A touch panel (screen) display may also be used as an input device.

[0248] According to one embodiment of the present invention, part of the optimization process can be performed by the computer system 100 in response to the processor 104 for executing one or more sequences of one or more instructions contained in the main storage 106. Such instructions can be read into the main storage 106 from another computer-readable medium (such as storage device 110). The execution of the sequence of instructions contained in the main storage 106 causes the processor 104 to perform the method steps described herein. One or more processors in a multi-processing arrangement can also be used to execute the sequence of instructions contained in the main storage 106. In an alternative embodiment, hard-wired circuits can be used to replace software instructions or in combination with software instructions. Therefore, the description herein is not limited to any specific combination of hardware circuits and software.

[0249] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processor 104 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 110. Volatile media include dynamic memory, such as main memory 106. Transmission media include coaxial cables, copper wires, and optical fibers, including wires that include bus 102. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tapes, any other physical media with hole patterns, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cassettes, carrier waves as described below, or any other media that a computer can read.

[0250] Various forms of computer readable media may be involved in transmitting one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions may initially appear on a disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 may receive data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector connected to bus 102 may receive the data carried in the infrared signal and place the data on bus 102. Bus 102 transmits the data to main memory 106, from which processor 104 retrieves and executes the instructions. The instructions received by main memory 106 may be optionally stored on storage device 110 before or after execution by processor 104.

[0251] The computer system 100 may also preferably include a communication interface 118 coupled to the bus 102. The communication interface 118 provides a two-way data communication coupled to a network link 120, which is connected to a local network 122. For example, the communication interface 118 may be an integrated services digital network (ISDN) card or a modem for providing data communication connected to a corresponding type of telephone line. As another example, the communication interface 118 may be a local area network (LAN) card for providing data communication connected to a compatible LAN. A wireless link may also be implemented. In any such implementation, the communication interface 118 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0252] Typically, network link 120 provides data communication through one or more networks to other data devices. For example, network link 120 may provide a connection to a host computer 124 or data equipment operated by an Internet Service Provider (ISP) 126 through a local network 122. ISP 126 in turn provides data communication services through a global packet data communication network (now commonly referred to as the "Internet") 128. Both local network 122 and Internet 128 use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 120 and through communication interface 118 carry the digital data to and from computer system 100, which is an exemplary form of a carrier wave used to transport the information.

[0253] The computer system 100 can send information and receive data, including program code, through the network, network link 120 and communication interface 118. In the example of the Internet, the server 130 can send a request code for an application through the Internet 128, ISP 126, local area network 122 and communication interface 118. One such downloaded application can provide for, for example, irradiation optimization of an embodiment. When it is received and / or stored in the storage device 110 or other non-volatile storage for later execution, the received code can be executed by the processor 104. In this way, the computer system 100 can obtain the application code in the form of a carrier wave.

[0254] Fig.23 An exemplary lithographic projection apparatus is schematically described, whose illumination source can be optimized using the method described herein. The apparatus comprises:

[0255] - an illumination system IL for conditioning the radiation beam B. In this particular case, the illumination system also comprises a radiation source SO;

[0256] a first stage (eg mask stage) MT having a patterning device holder for holding a patterning device MA (eg reticle) and connected to a first positioner for accurately positioning the patterning device relative to the device PS;

[0257] a second object table (substrate table) WT provided with a substrate holder for holding a substrate W (e.g. a silicon wafer coated with resist) and connected to a second positioner for accurately positioning the substrate relative to the device PS;

[0258] - a projection system ("lens") PS (eg a refractive, reflective or catadioptric optical system) for imaging the illuminated portion of the patterning device MA onto a target portion C of the substrate W (eg comprising one or more dies).

[0259] As depicted in the present invention, the device is of the transmissive type (i.e., with a transmissive mask). However, in general, it may also be of the reflective type, for example (with a reflective mask). Alternatively, the device may use another class of patterning devices as an alternative to using a classical mask; examples include a programmable mirror array or an LCD matrix.

[0260] A source SO (e.g. a mercury lamp or an excimer laser) generates a radiation beam. The beam is fed into an illumination system (illuminator) IL, for example directly or after having traversed an adjustment member such as a beam expander Ex. The illuminator IL may include adjustment members AD for setting the outer radial extent and / or the inner radial extent (commonly referred to as σouter and σinner, respectively) of the intensity distribution in the beam. In addition, the illuminator IL will typically include various other components, such as an integrator IN and a condenser CO. In this way, the beam B impinging on the patterning device MA has a desired uniformity and intensity distribution in its cross-section.

[0261] about Fig.23 It should be noted that the source SO can be within the housing of the lithographic projection device (which is usually the case when the source SO is, for example, a mercury lamp), but it can also be remote from the lithographic projection device, the radiation beam it produces being guided into the device (for example with the aid of appropriate guide mirrors); this latter case is often the case when the source SO is an excimer laser (for example based on KrF, ArF or F2 laser action).

[0262] The beam PB then intercepts the patterning device MA which is held on the patterning device table MT. Having traversed the patterning device MA, the beam B passes through a lens PL which focuses the beam B onto a target portion C of the substrate W. With the aid of the second positioning means (and the interferometry means IF), the substrate table WT can be accurately moved, for example in order to position a different target portion C in the path of the beam PB. Similarly, the first positioning means can be used to accurately position the patterning device MA relative to the path of the beam B, for example after mechanically obtaining the patterning device MA from a patterning device library or during scanning. Typically, the patterning device MA will be accurately positioned with the aid of the interferometry means not in the patterning device library. Fig.23 Movement of the stage MT, WT is achieved by a long-stroke module (coarse positioning) and a short-stroke module (fine positioning) explicitly depicted in the figure. However, in the case of a wafer stepper (as opposed to a step-and-scan tool), the patterning device table MT may be connected to the short-stroke actuator only, or may be fixed.

[0263] The depicted tools can be used in two different modes:

[0264] - in step mode, the patterning device table MT is kept substantially stationary and the entire patterning device image is projected once (i.e. a single "flash") onto a target portion C. The substrate table WT is then shifted in the x-direction and / or y-direction so that a different target portion C can be illuminated by the beam PB;

[0265] - In scan mode, essentially the same situation applies, but a given target portion C is not exposed in a single "flash". Instead, the patterning device table MT is movable in a given direction (the so-called "scanning direction", e.g. the y-direction) at a speed v, so that the projection beam B is scanned across the patterning device image; at the same time, the substrate table WT is simultaneously moved in the same or opposite direction at a speed V=Mv, where M is the magnification of the lens PL (typically M=1 / 4 or =1 / 5). In this way, a relatively large target portion C can be exposed without having to compromise resolution.

[0266] Fig.24 Another exemplary lithographic projection apparatus LA is schematically depicted, the illumination source of which may be optimized using the methods described herein.

[0267] The lithography projection apparatus LA comprises:

[0268] - Source collector module SO;

[0269] - an illumination system (illuminator) IL configured to condition a radiation beam B (e.g. EUV radiation);

[0270] a support structure (eg, mask table) MT configured to support a patterning device (eg, mask or reticle) MA and connected to a first positioning device PM configured to accurately position the patterning device;

[0271] a substrate table (eg wafer table) WT configured to hold a substrate (eg a resist-coated wafer) W and connected to a second positioning device PW configured to accurately position the substrate; and

[0272] - a projection system (eg a reflective projection system) PS configured to project the pattern imparted to the radiation beam B by the patterning device MA onto a target portion C of the substrate W (eg comprising one or more dies).

[0273] As shown here, the apparatus LA is reflective (e.g., employing a reflective mask). It should be noted that since most materials are absorptive in the EUV wavelength range, the mask can have a multilayer reflector, including multiple stacks of, for example, molybdenum and silicon. In one example, the multilayer reflector has 40 layers of paired molybdenum and silicon, each layer being a quarter wavelength thick. Even smaller wavelengths can be produced with X-ray lithography. Since most materials are absorptive in EUV and X-ray wavelengths, a thin sheet of absorbing material patterned on the patterning device topography (e.g., a TaN absorber on top of a multilayer reflector) defines areas where features will be printed (positive resist) or not (negative resist).

[0274] Reference Fig.24 , the illuminator IL receives the extreme ultraviolet radiation beam from the source collector module SO. Methods for generating EUV radiation include, but are not necessarily limited to, converting a material into a plasma state, the material having at least one element having one or more emission lines in the EUV range, such as xenon, lithium, or tin. In one such method, generally referred to as laser produced plasma ("LPP"), the desired plasma can be generated by irradiating a fuel, such as a droplet, a beam, or a cluster of a material having an emission line element, with a laser beam. The source collector module SO may be a laser (in Fig.11 The laser and source collector module may be separate entities, for example when a CO2 laser is used to provide the laser beam for fuel excitation.

[0275] In this case, the laser is not considered to form part of the lithographic apparatus, and the radiation beam is delivered from the laser to the source collector module by means of a beam delivery system comprising, for example, suitable directing mirrors and / or a beam expander. In other cases, the source may be an integral part of the source collector module, for example when the source is a discharge produced plasma EUV generator, commonly referred to as a DPP source.

[0276] The illuminator IL may comprise an adjuster for adjusting the angular intensity distribution of the radiation beam. Typically, at least the outer and / or inner radial extent (generally referred to as σ-outer and σ-inner, respectively) of the intensity distribution in a pupil plane of the illuminator may be adjusted. Furthermore, the illuminator IL may comprise various other components, such as a faceted field mirror arrangement and a faceted pupil mirror arrangement. The illuminator may be used to adjust the radiation beam to have a desired uniformity and intensity distribution in its cross-section.

[0277] The radiation beam B is incident on the patterning device (e.g. mask) MA held on a support structure (e.g. mask table) MT and is patterned by the patterning device. After having been reflected by the patterning device (e.g. mask) MA, the radiation beam B passes through a projection system PS which focuses the radiation beam onto a target portion C of the substrate W. With the help of a second positioning device PW and a position sensor system PS2 (e.g. interferometer device, linear encoder or capacitive sensor), the substrate table WT can be precisely moved, for example in order to position a different target portion C in the path of the radiation beam B. Similarly, the first positioning device PM and a further position sensor system PS1 can be used to precisely position the patterning device (e.g. mask) MA relative to the path of the radiation beam B. The patterning device (e.g. mask) MA and the substrate W can be aligned using pattern shape device alignment marks M1, M2 and substrate alignment marks P1, P2.

[0278] The depicted device LA may be used in at least one of the following modes:

[0279] 1. In step mode, the support structure (e.g. mask table) MT and the substrate table WT are held substantially stationary while an entire pattern imparted to the radiation beam is projected once onto a target portion C (i.e. a single static exposure). The substrate table WT is then moved in the X and / or Y direction so that a different target portion C can be exposed.

[0280] 2. In scan mode, the support structure (e.g. mask table) MT and the substrate table WT are scanned synchronously while a pattern imparted to the radiation beam is projected onto a target portion C (i.e. a single dynamic exposure). The speed and direction of the substrate table WT relative to the support structure (e.g. mask table) MT may be determined by the (de-)magnification and image reversal characteristics of the projection system PS.

[0281] 3. In another mode, the support structure (e.g. mask table) MT holding the programmable patterning device is held substantially stationary and the substrate table WT is moved or scanned while a pattern imparted to the radiation beam is projected onto a target portion C. In this mode, a pulsed radiation source is typically employed and the programmable patterning device is updated as required after each movement of the substrate table WT or between successive radiation pulses during a scan. This mode of operation may be readily applicable to maskless lithography using a programmable patterning device (e.g. a programmable mirror array of the type described above).

[0282] Fig.25The apparatus LA is shown in more detail, including a source collector module SO, an illumination system IL, and a projection system PS. The source collector module SO is constructed and arranged so that a vacuum environment is maintained within an enclosure 220 of the source collector module SO. A plasma 210 for emitting EUV radiation may be formed by a discharge-generated plasma source. EUV radiation may be generated by a gas or vapor, such as xenon, lithium vapor, or tin vapor, wherein an extremely high temperature plasma 210 is formed to emit radiation in the EUV range of the electromagnetic radiation spectrum. The extremely high temperature plasma 210 is formed, for example, by a discharge that causes at least partially ionized plasma. For example, a partial pressure of 10 Pa of Xe, Li, Sn vapor, or any other suitable gas or vapor may be required to effectively generate radiation. In one embodiment, an excited plasma of tin (Sn) is provided to generate EUV radiation.

[0283] The radiation emitted by the high temperature plasma 210 is transferred from the source chamber 211 to the collector chamber 212 via a gas barrier or contaminant trap 230 (referred to in some cases as a contaminant barrier or fin trap) optionally positioned within or behind an opening in the source chamber 211. The contaminant trap 230 may include a channel structure. The contaminant trap 230 may also include a gas barrier or a combination of a gas barrier and a channel structure. The contaminant trap or contaminant barrier 230 further shown herein includes at least a channel structure, as known in the prior art.

[0284] The collector cavity 211 may include a radiation collector CO, which may be a so-called grazing incidence collector. The radiation collector CO has an upstream radiation collector side 251 and a downstream radiation collector side 252. Radiation passing through the collector CO may be reflected off the grating spectral filter 240 to be focused at a virtual source point IF along the optical axis indicated by the dashed line 'O'. The virtual source point IF is often referred to as an intermediate focus, and the source collector module is arranged such that the intermediate focus IF is located at or near the opening of the enclosing structure 220. The virtual source point IF is an image of the plasma 210 for emitting radiation.

[0285] The radiation then passes through an illumination system IL which may comprise a faceted field mirror arrangement 22 and a faceted pupil mirror arrangement 24 arranged to provide a desired angular distribution of the radiation beam 21 at the patterning device MA and a desired uniformity of radiation intensity at the patterning device MA. When the radiation beam 21 is reflected at the patterning device MA held by the support structure MT, a patterned beam 26 is formed and the patterned beam 26 is imaged by the projection system PS via reflective elements 28, 30 onto a substrate W held by a substrate table WT.

[0286] There may typically be more elements present in the illumination optics unit IL and the projection system PS than are shown. The grating spectral filter 240 may be optionally provided, depending on the type of lithographic apparatus. Furthermore, there may be more mirrors than are shown in the figures, for example, there may be more than 100 mirrors in the projection system PS. Fig.25 1-6 additional reflective elements beyond the elements shown.

[0287] Collector optics CO, such as Fig.25 As shown, a nested collector with grazing incidence reflectors 253, 254 and 255 is shown in the figure as just one example of a collector (or collector mirror). The grazing incidence reflectors 253, 254 and 255 are arranged axially symmetrically around the optical axis O. This type of collector optical device CO is preferably used in conjunction with a discharge produced plasma source, usually called a DPP source.

[0288] Alternatively, the source collector module SO may be as follows Fig.26 A portion of the LPP radiation system shown. The laser LA is arranged to inject laser energy into a fuel, such as xenon (Xe), tin (Sn) or lithium (Li), thereby generating a highly ionized plasma 210 with an electron temperature of several tens of eV. The high-energy radiation generated during the deexcitation and recombination of these ions is emitted by the plasma, collected by the near normal incidence collector optics CO and focused onto an opening 221 of the enclosing structure 220.

[0289] The concepts disclosed herein can simulate or mathematically model any general imaging system for imaging sub-wavelength features, and may be particularly useful with the advent of imaging technologies capable of producing ever shorter wavelengths. Existing technologies already in use include EUV (extreme ultraviolet) lithography, which can produce 193 nm wavelengths with ArF lasers, and even 157 nm wavelengths with fluorine lasers. In addition, EUV lithography can produce wavelengths in the 20-5 nm range by using a synchrotron or by bombarding materials (solid or plasma) with high energy electrons to produce photons in this range.

[0290] The embodiments may be further described using the following aspects.

[0291] 1. A method for training a patterning process model, the patterning process model being configured to predict a pattern to be formed on a patterning process, the method comprising:

[0292] obtaining (i) image data associated with a desired pattern, (ii) a measured pattern of a substrate, the measured pattern associated with the desired pattern, (iii) a first model associated with an aspect of the patterning process, the first model comprising a first set of parameters, and (iv) a machine learning model associated with another aspect of the patterning process, the machine learning model comprising a second set of parameters; and

[0293] Iteratively determining values ​​of the first set of parameters and the second set of parameters to train a patterning process model, wherein the iterations include:

[0294] executing the first model and the machine learning model using the image data to collaboratively predict a printed pattern of the substrate; and

[0295] The values ​​of the first set of parameters and the second set of parameters are modified such that a difference between the measured pattern and a predicted pattern of the patterning process model is reduced.

[0296] 2. The method according to aspect 1, wherein the first model and the machine learning model are configured and trained in a deep convolutional neural network framework.

[0297] 3. A method according to aspect 2, wherein the training involves:

[0298] Predicting the printed pattern by forward propagation of the output of the first model and the machine learning model;

[0299] determining a difference between the measured pattern and the predicted pattern of the patterning process model;

[0300] determining a differential of the difference relative to the first set of parameters and the second set of parameters; and

[0301] Based on the differential of the differences, values ​​of the first set of parameters and the second set of parameters are determined by back-propagation of the outputs of the first model and the machine learning model.

[0302] 4. A method according to any one of aspects 1 to 3, wherein the first model is connected to the machine learning model in a series or parallel combination.

[0303] 5. The method according to aspect 4, wherein the series combination of the models comprises:

[0304] The output of the first model is provided as input to the machine learning model.

[0305] 6. The method according to aspect 4, wherein the series combination of the models comprises:

[0306] The output of the machine learning model is provided as input to the first model.

[0307] 7. The method according to aspect 4, wherein the parallel combination of the models comprises:

[0308] providing the same input to the first model and the machine learning model;

[0309] combining outputs of the first model and the machine learning model; and

[0310] A predicted printed pattern is determined based on the combined outputs of the corresponding models.

[0311] 8. The method according to any one of aspects 1 to 7, wherein the first model is a resist model, and / or a space model.

[0312] 9. The method of clause 8, wherein the first set of parameters of the resist model corresponds to at least one of:

[0313] Initial acid distribution;

[0314] Acid diffusion;

[0315] Image contrast;

[0316] long-range pattern loading effects;

[0317] long-range pattern loading effects;

[0318] Acid concentration after neutralization;

[0319] Alkali concentration after neutralization;

[0320] Diffusion due to high acid concentrations;

[0321] Diffusion due to high alkali concentration;

[0322] Resist shrinkage;

[0323] Resist development; or

[0324] Two-dimensional convex curvature effect;

[0325] 10. The method according to any one of aspects 1 to 9, wherein the first model is an empirical model that accurately models the physics of the first aspect of the patterning process.

[0326] 11. A method according to any one of clauses 1 to 10, wherein the first model corresponds to the first aspect related to acid-base diffusion after exposure of the substrate.

[0327] 12. A method according to any one of aspects 1 to 9, wherein the machine learning model is a neural network that models a second aspect of the patterning process having relatively less physics-based understanding.

[0328] 13. A method according to aspect 12, wherein the second set of parameters includes: weights and biases of one or more layers of the neural network.

[0329] 14. The method according to any one of clauses 1 to 13, wherein the patterning process model corresponds to the second aspect of a post-exposure process of the patterning process.

[0330] 15. The method according to any one of aspects 1 to 14, wherein the first aspect and / or the second aspect of the post-exposure process comprises: resist baking, resist development, and / or etching.

[0331] 16. A method for determining an optical proximity effect correction for a patterning process, the method comprising:

[0332] obtaining image data associated with a desired pattern;

[0333] executing a trained patterning process model using the image data to predict a pattern to be printed on a substrate; and

[0334] Optical proximity effect corrections and / or defects are determined using a predicted pattern that will be printed on the substrate subjected to the patterning process.

[0335] 17. A method according to clause 16, wherein the image data is an aerial image and / or a mask image of a desired pattern.

[0336] 18. A method according to aspect 16, wherein the trained patterning process model includes a first model configured to collaboratively predict a first aspect of the patterning process to be printed on the substrate, and a machine learning model of a second aspect of the patterning process.

[0337] 19. A method according to aspect 18, wherein the first model and the machine learning model are combined in series and / or in parallel.

[0338] 20. The method according to any one of clauses 16 to 19, wherein the first model is an empirical model that accurately models the physical properties of the first aspect of the post-exposure process of the patterning process.

[0339] 21. A method according to any one of clauses 16 to 20, wherein the first model corresponds to the first aspect related to acid-base diffusion after exposure of the substrate.

[0340] 22. A method according to any one of aspects 16 to 21, wherein the machine learning model is a neural network that models the second aspect of the patterning process with relatively little physics-based understanding.

[0341] 23. The method of any one of clauses 16 to 22, wherein determining an optical proximity effect correction comprises:

[0342] The desired pattern is adjusted and / or assist features are placed around the desired pattern so that the difference between the predicted pattern and the desired pattern is reduced.

[0343] 24. The method according to any one of aspects 16 to 22, wherein determining the defect comprises:

[0344] A lithography manufacturability check is performed on the predicted pattern.

[0345] 25. A method for training a machine learning model, the machine learning model being configured to determine an etch bias associated with an etch process, the method comprising:

[0346] obtaining (i) resist pattern data associated with a target pattern to be printed on a substrate, (ii) physical effect data characterizing the effect of the etching process on the target pattern, and (iii) measured deviations between the resist pattern and the etched pattern formed on the printed substrate; and

[0347] The machine learning model is trained based on the resist pattern data, the physical effect data, and the measured deviation to reduce a difference between the measured deviation and a predicted etch deviation.

[0348] 26. A method according to aspect 25, wherein the machine learning model is configured to receive the resist pattern data at a first layer of the machine learning model and the physical effect data is received at a last layer of the machine learning model.

[0349] 27. A method according to aspect 26, wherein the output of the last layer is a linear combination of: (i) the etching bias predicted by executing the machine learning model using the resist pattern data as input, and (ii) another etching bias determined based on physical effect data related to the etching process.

[0350] 28. The method of clause 27, wherein the output of the last layer is an etch deviation map from which the etch deviation is extracted, wherein the etch deviation map is generated via:

[0351] executing the machine learning model using the resist pattern data as input to output an etch deviation map, wherein the etch deviation map comprises a deviation resist pattern; and

[0352] The etch bias map is combined with the physical effects data.

[0353] 29. A method according to aspect 28, wherein the machine learning model is configured to receive the resist pattern data and the physical effect data at the first layer of the machine learning model.

[0354] 30. A method according to any one of clauses 25 to 29, wherein training the machine learning model is an iterative process comprising:

[0355] (a) predicting the etch bias by executing the machine learning model using actual resist pattern data and the physical effect data as input;

[0356] (b) determining a difference between the measured deviation and the predicted etch deviation;

[0357] (c) determining the gradient of the difference with respect to a model parameter of the machine learning model;

[0358] (d) using the gradient as a guide to adjust model parameter values ​​so that the difference between the measured deviation and the predicted etch deviation decreases;

[0359] (e) determining whether the difference is minimized or exceeds a training threshold; and

[0360] (f) In response to the difference not being minimized or not violating a training threshold, performing steps (a) to (e).

[0361] 31. The method according to any one of aspects 25 to 30, wherein obtaining resist pattern data comprises:

[0362] One or more process models including a resist model of the patterning process are performed using the target pattern to be printed on the substrate.

[0363] 32. A method according to any one of clauses 25 to 31, wherein the resist pattern data is represented as a resist image, wherein the resist image is a pixelated image.

[0364] 33. The method according to any one of aspects 25 to 32, wherein the physical effect data is data related to an etching term characterizing an etching effect, the etching term comprising at least one of the following:

[0365] a bulk concentration of plasma within trenches of the resist pattern associated with the target pattern;

[0366] a concentration of plasma on top of the resist layer of the substrate;

[0367] a loading effect determined by convolving the resist pattern with a Gaussian kernel having specified model parameters;

[0368] a change in a loading effect on the resist pattern during the etching process;

[0369] a relative position of the resist pattern with respect to an adjacent pattern on the substrate;

[0370] an aspect ratio of the resist pattern; or

[0371] A term related to the combined effect of two or more etch process parameters.

[0372] 34. The method of any one of aspects 25 to 33, wherein obtaining physical effect data comprises:

[0373] A physical effects model is executed, the physical effects model including one or more of the etching terms and a Gaussian kernel specified for corresponding one or more of the etching terms.

[0374] 35. A method according to any one of clauses 25 to 34, wherein the physical effect data is represented as a pixelated image, wherein each pixel intensity indicates a physical effect on the resist pattern associated with the target pattern.

[0375] 36. The method according to any one of aspects 25 to 35, further comprising:

[0376] obtaining a resist profile of the resist pattern; and

[0377] An etch profile is generated by applying the etch bias to the resist profile.

[0378] 37. A system for determining an etch deviation associated with an etch process, the system comprising:

[0379] Semiconductor process equipment; and

[0380] A processor configured to:

[0381] determining physical effect data characterizing an effect of the etching process on the substrate via executing a physical effect model;

[0382] executing a trained machine learning model using the resist pattern and the physical effect data as input to determine the etch bias; and

[0383] A semiconductor device or the etching process is controlled based on the etching deviation.

[0384] 38. A system according to aspect 37, wherein a trained machine learning model is trained using multiple resist patterns, physical effect data associated with each of the resist patterns, and measured deviations associated with each resist pattern so that the difference between the measured deviations and the determined etching deviations is minimized.

[0385] 39. A system according to any one of aspects 37 to 38, wherein the trained machine learning model is a convolutional neural network (CNN) including specific weights and biases, wherein the weights and biases of the CNN are determined through a training process using multiple resist patterns, physical effect data associated with each of the resist patterns, and the measured deviations associated with each resist pattern so that the difference between the measured deviations and the determined etching deviations is minimized.

[0386] 40. The system of any one of clauses 37 to 39, wherein controlling the semiconductor process equipment comprises:

[0387] The value of one or more parameters of the semiconductor device is adjusted such that the yield of the patterning process is improved.

[0388] 41. The system of clause 40, wherein adjusting the value of one or more parameters of the semiconductor process equipment is an iterative process comprising:

[0389] (a) changing a current value of one or more parameters by adjusting a mechanism of the semiconductor process equipment;

[0390] (b) obtaining the resist pattern printed on the substrate via the semiconductor process equipment;

[0391] (c) determining the etch bias by executing a trained machine learning model using the resist pattern, and further determining the etch pattern by applying the etch bias to the resist pattern;

[0392] (d) determining whether a yield of the patterning process is within a desired yield range based on the etched pattern; and

[0393] In response to not being within the yield range, performing steps (a) to (d).

[0394] 42. The system of any one of aspects 37 to 41, wherein controlling the etching process comprises:

[0395] determining the etching pattern by applying the etching bias to the resist pattern;

[0396] determining a yield of the patterning process based on the etch pattern; and

[0397] An etch profile of the etch process is determined based on the etch pattern such that a yield of the patterning process is improved.

[0398] 43. The system of any one of clauses 37 to 42, wherein the yield of the patterning process is the percentage of etched patterns that meet design specifications across the entire substrate.

[0399] 44. A system according to any one of clauses 37 to 43, wherein the semiconductor processing equipment is a lithographic equipment.

[0400] 45. A method for calibrating a process model, the process model being configured to generate a simulation profile, the method comprising:

[0401] obtaining (i) measurement data at a plurality of measurement locations on a pattern, and (ii) contour constraints specified based on the measurement data; and

[0402] The process model is calibrated by adjusting values ​​of model parameters of the process model until the simulation profile satisfies the profile constraints.

[0403] 46. ​​The method of aspect 45, wherein the plurality of measurement locations are edge placement (EP) gauges placed on a printed pattern or on a printed outline of the printed pattern.

[0404] 47. A method according to any one of aspects 45 to 46, wherein the measurement data comprises a plurality of angles, each angle being defined at each measurement location placed on a pattern or on a printed outline of the printed pattern.

[0405] 48. A method according to clause 47, wherein each angle at each measurement location defines a direction in which an edge placement error between the printed contour and the target contour is determined.

[0406] 49. A method according to any of aspects 45 to 48, wherein each contour constraint is a function of a tangent angle between a tangent to the simulated contour at a given measurement location and an angle of the measured data at the given location.

[0407] 50. A method according to any one of aspects 45 to 49, wherein adjusting the value of the model parameter is an iterative process comprising:

[0408] (a) executing the process model using given values ​​of the model parameters to generate the simulation profile, wherein the given values ​​are random values ​​at a first iteration and are adjusted values ​​at subsequent iterations;

[0409] (c) determining a tangent line to the simulated contour at each of the measurement locations;

[0410] (d) determining a tangent angle between an angle of the measurement data at each of the measurement locations and a tangent line;

[0411] (e) determining whether the tangent angle is within a vertical range at one or more of the measurement locations; and

[0412] (f) In response to the tangent angle not being within the vertical range, adjusting the values ​​of the model parameters and performing steps (a) to (e).

[0413] 51. A method according to any one of aspects 45 to 50, wherein the vertical range is a value of an angle between 88° and 92°, preferably 90°.

[0414] 52. A method according to any one of aspects 45 to 51, wherein adjustment is based on a gradient of each tangent angle relative to the model parameter, wherein the gradient indicates how sensitive the tangent angle is to changes in the value of the model parameter.

[0415] 53. A method according to any one of aspects 45 to 52, wherein the process model is a data-driven model including an empirical model and / or a machine learning model.

[0416] 54. A method according to any one of aspects 45 to 53, wherein the machine learning model is a convolutional neural network, and wherein the model parameters are weights and biases associated with multiple layers.

[0417] 55. A method for calibrating a process model, the process model being configured to predict an image of a target pattern, the method comprising:

[0418] obtaining (i) a reference image associated with the target pattern, and (ii) a gradient constraint specified relative to the reference image; and

[0419] The process model is calibrated such that the process model produces a simulated image that (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfies the gradient constraint.

[0420] 56. The method of clause 55, wherein calibrating the process model is an iterative process comprising:

[0421] (a) executing the process model using the target pattern to generate the simulated image;

[0422] (b) determining intensity differences between intensity values ​​of the simulated image and the reference image, and / or transforming the simulated image and the reference image into the frequency domain via a Fourier transform and determining frequency differences between frequencies associated with the simulated image and the reference image;

[0423] (c) determining a simulated gradient of a signal in the simulated image, wherein the signal is a signal along a given line through the simulated image;

[0424] (d) determining whether the following conditions are met: (i) the intensity difference or the frequency difference is minimized, and (ii) the simulated gradient satisfies the gradient constraint associated with the reference image; and

[0425] (e) in response to conditions (i) and (ii) not being satisfied, adjusting values ​​of model parameters of the process model, and performing steps (a) to (d) until conditions (i) and (ii) are satisfied.

[0426] 57. A method according to any one of aspects 55 to 56, wherein the simulated gradient is determined by taking the first derivative of a signal along a given line through the simulated image.

[0427] 58. A method according to any one of clauses 55 to 57, wherein the gradient constraint is obtained by taking the first order derivative of a signal along a given line through the reference image.

[0428] 59. The method according to any one of aspects 55 to 58, further comprising:

[0429] extracting a simulated profile from the simulated image and a reference profile from the reference image, wherein the simulated profile and the reference profile are associated with the target pattern; and

[0430] The process model is calibrated so that the simulated profile satisfies a profile shape constraint, wherein the profile shape constraint ensures that the simulated profile conforms to a shape of the reference profile.

[0431] 60. A method according to any one of aspects 55 to 59, wherein determining whether the contour shape constraint is satisfied comprises:

[0432] The second order derivative of the simulated profile is determined to be within an expected range of the second order derivative of the reference profile.

[0433] 61. A method according to any one of clauses 55 to 60, wherein the reference image is obtained via simulating a physics-based model of a patterning process using the target pattern, the reference image comprising:

[0434] A spatial image of the target pattern;

[0435] a resist image of the target pattern; and / or

[0436] An etched image of the target pattern.

[0437] 62. A method according to any of clauses 55 to 61, wherein the process model is configured to satisfy contour constraints defined relative to a printed contour of a pattern on a printed substrate.

[0438] 63. A method according to any one of aspects 55 to 62, wherein each contour constraint is a function of a tangent angle between a tangent to a simulated contour at a given measurement location and an angle of the measurement data at the given location, wherein the simulated contour is a contour of a simulated pattern determined by executing the process model using the target pattern.

[0439] 64. A system for calibrating a process model, the process model being configured to generate a simulation profile, the system comprising:

[0440] a metrology tool configured to obtain measurement data at a plurality of measurement locations on the pattern; and

[0441] A processor configured to:

[0442] The process model is calibrated by adjusting values ​​of model parameters of the process model until the simulated profile satisfies the profile constraints, the profile constraints being based on the measured data.

[0443] 65. The system of aspect 64, wherein the plurality of measurement locations are edge placement (EP) gauges placed on the printed pattern or on a printed contour of the printed pattern.

[0444] 66. A system according to any of aspects 64 to 65, wherein the measurement data comprises a plurality of angles, each angle being defined at each measurement location placed on a pattern or on a printed contour of the printed pattern.

[0445] 67. The system of aspect 66, wherein each angle at each measurement location defines a direction in which an edge placement error between the printed contour and the target contour is determined.

[0446] 68. A system according to any of aspects 64 to 67, wherein each contour constraint is a function of a tangent angle between a tangent to the simulated contour at a given measurement location and an angle of the measured data at the given location.

[0447] 69. A system according to any one of aspects 64 to 68, wherein adjusting the value of the model parameter is an iterative process comprising:

[0448] (a) executing the process model using given values ​​of model parameters to produce the simulation profile, wherein the given values ​​are random values ​​at a first iteration and are adjusted values ​​at subsequent iterations;

[0449] (c) determining a tangent line to the simulated contour at each of the measurement locations;

[0450] (d) determining a tangent angle between the angle of the measurement data and the tangent at each of the measurement locations;

[0451] (e) determining whether the tangent angle is within a vertical range at one or more of the measurement locations; and

[0452] (f) In response to the tangent angle not being within the vertical range, adjusting the value of the model parameter and performing steps (a) to (e).

[0453] 70. A system according to any one of aspects 64 to 69, wherein the vertical range is a value of an angle between 88° and 92°, preferably 90°.

[0454] 71. A system according to any one of aspects 64 to 70, wherein adjustment is based on a gradient of each tangent angle relative to the model parameter, wherein the gradient indicates how sensitive the tangent angle is to changes in the value of the model parameter.

[0455] 72. A system according to any one of aspects 64 to 71, wherein the process model is a data-driven model including an empirical model and / or a machine learning model.

[0456] 73. A system according to any one of aspects 64 to 72, wherein the machine learning model is a convolutional neural network, and wherein the model parameters are weights and biases associated with multiple layers.

[0457] 74. The system of any one of aspects 64 to 73, wherein the metrology tool is an electron beam device.

[0458] 75. A system according to any one of aspects 64 to 74, wherein the metrology tool is a scanning electron microscope configured to identify and extract contours from a captured image of a pattern on a printed substrate.

[0459] 76. A system for calibrating a process model, the process model being configured to predict an image of a target pattern, the system comprising:

[0460] a metrology tool configured to obtain a reference image associated with the target pattern; and

[0461] A processor configured to:

[0462] The process model is calibrated so that the process model produces a simulated image that (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfies a gradient constraint associated with the reference image.

[0463] 77. The system of aspect 76, wherein calibrating the process model is an iterative process comprising:

[0464] (a) executing the process model using the target pattern to generate the simulated image;

[0465] (b) determining intensity differences between intensity values ​​of the simulated image and the reference image, and / or transforming the simulated image and the reference image into the frequency domain via a Fourier transform and determining frequency differences between frequencies associated with the simulated image and the reference image;

[0466] (c) determining a simulated gradient of a signal in the simulated image, wherein the signal is a signal along a given line through the simulated image;

[0467] (d) determining whether the following conditions are met: (i) the intensity difference or the frequency difference is minimized, and (ii) the simulated gradient satisfies the gradient constraint associated with the reference image; and

[0468] (e) in response to conditions (i) and (ii) not being satisfied, adjusting values ​​of model parameters of the process model, and performing steps (a) to (d) until conditions (i) and (ii) are satisfied.

[0469] 78. A system according to any one of aspects 76 to 77, wherein the simulated gradient is determined by taking the first derivative of a signal along a given line through the simulated image.

[0470] 79. A system according to any one of aspects 76 to 78, wherein the gradient constraint is obtained by taking the first-order derivative of a signal along a given line through the reference image.

[0471] 80. The system of any one of aspects 76 to 79, the processor being further configured to:

[0472] extracting a simulated profile from the simulated image and a reference profile from the reference image, wherein the simulated profile and the reference profile are associated with the target pattern; and

[0473] The process model is calibrated so that the simulated profile satisfies a profile shape constraint, wherein the profile shape constraint ensures that the simulated profile conforms to a shape of the reference profile.

[0474] 81. The system of any one of aspects 76 to 80, wherein determining whether the contour shape constraint is satisfied comprises:

[0475] The second order derivative of the simulated profile is determined to be within an expected range of the second order derivative of the reference profile.

[0476] 82. A system according to any one of clauses 76 to 81, wherein the reference image is obtained via simulating a physics-based model of a patterning process using the target pattern, the reference image comprising:

[0477] A spatial image of the target pattern;

[0478] a resist image of the target pattern; and / or

[0479] An etched image of the target pattern.

[0480] 83. A system according to any of aspects 76 to 82, wherein the process model is configured to satisfy contour constraints defined relative to a printed contour of a pattern on a printed substrate.

[0481] 84. A system according to any one of aspects 76 to 83, wherein each contour constraint is a function of a tangent angle between a tangent to a simulated contour at a given measurement location and an angle of the measurement data at the given location, wherein the simulated contour is a contour of a simulated pattern determined by executing the process model using the target pattern.

[0482] 85. A non-transitory computer readable medium comprising instructions that, when executed by one or more processors, result in operations comprising:

[0483] obtaining (i) resist pattern data associated with a target pattern to be printed on a substrate, (ii) physical effect data characterizing the effect of an etching process on the target pattern, and (iii) a measured deviation between the resist pattern and an etched pattern formed on the printed substrate; and

[0484] The machine learning model is trained based on the resist pattern data, the physical effect data, and the measured deviation to reduce a difference between the measured deviation and a predicted etch deviation.

[0485] 86. A non-transitory computer readable medium comprising instructions that, when executed by one or more processors, result in operations comprising:

[0486] obtaining (i) measurement data at a plurality of measurement locations on a pattern, and (ii) contour constraints specified based on the measurement data; and

[0487] The process model is calibrated by adjusting values ​​of model parameters of the process model until the simulation profile satisfies the profile constraints.

[0488] 87. A non-transitory computer readable medium comprising instructions that, when executed by one or more processors, result in operations comprising:

[0489] obtaining (i) a reference image associated with a target pattern, and (ii) a gradient constraint specified relative to the reference image; and

[0490] A process model is calibrated such that the process model produces a simulated image that (i) minimizes an intensity difference or a frequency difference between the simulated image and the reference image, and (ii) satisfies the gradient constraint.

[0491] Although the concepts disclosed herein may be used for imaging on substrates such as silicon wafers, it should be understood that the disclosed concepts may be used with any type of lithography imaging system, for example, a lithography imaging system for imaging on substrates other than silicon wafers.

[0492] The above description is intended to be illustrative rather than limiting.Thus, it will be appreciated by those skilled in the art that modifications may be made as described without departing from the scope of the claims set forth hereinafter.

Claims

1. A method for training a patterning process model, the patterning process model being configured to predict a pattern to be formed on a patterning process, the method comprising: obtaining (i) image data associated with a desired pattern, (ii) a measured pattern of a substrate, the measured pattern associated with the desired pattern, (iii) a first model associated with a first aspect of the patterning process and being a non-machine learning model that models the first aspect of the patterning process, the first model comprising a first set of parameters, and (iv) a machine learning model comprising a neural network that models a second aspect of the patterning process, the machine learning model comprising a second set of parameters, wherein the first aspect and the second aspect are different from each other, and each of the first aspect and the second aspect is associated with a post-exposure process and comprises: resist baking, resist developing and / or etching; and The values ​​of the first set of parameters and the second set of parameters are iteratively determined through a neural network training process to train a patterning process model, wherein the iteration comprises: executing the first model and the machine learning model using the image data to collaboratively predict a printed pattern of the substrate; and Values ​​of the first set of parameters and the second set of parameters are modified based on the measured pattern and a predicted pattern of the patterning process model.

2. The method of claim 1, wherein the first model and the machine learning model are configured and trained in a deep convolutional neural network framework.

3. The method according to claim 2, wherein the training comprises: Predicting the printed pattern by forward propagation of the output of the first model and the machine learning model; determining a difference between the measured pattern and the predicted pattern of the patterning process model; determining a difference of the difference relative to a first set of parameters and a second set of parameters; and Based on the differential of the differences, values ​​of the first set of parameters and the second set of parameters are determined by back-propagation of the outputs of the first model and the machine learning model.

4. The method according to claim 1, wherein the first model is connected to the machine learning model in a series combination or a parallel combination.

5. The method of claim 4, wherein the serial combination of the first model and the machine learning model comprises: The output of the first model is provided as input to the machine learning model.

6. The method of claim 4, wherein the serial combination of the first model and the machine learning model comprises: The output of the machine learning model is provided as input to the first model.

7. The method of claim 4, wherein the parallel combination of the first model and the machine learning model comprises: providing the same input to the first model and the machine learning model; combining outputs of the first model and the machine learning model; and A predicted printed pattern is determined based on the combined outputs of the corresponding models.

8. The method according to claim 1, wherein the first model is a resist model, and / or an aerial image model.

9. The method of claim 8, wherein the first set of parameters of the resist model corresponds to at least one of: Initial acid distribution; Acid diffusion; Image contrast; long-range pattern loading effects; long-range pattern loading effects; Acid concentration after neutralization; Alkali concentration after neutralization; Diffusion due to high acid concentrations; Diffusion due to high alkali concentration; Resist shrinkage; Resist development; or Two-dimensional convex curvature effect.

10. The method of claim 1, wherein the second parameter set comprises: The weights and biases of one or more layers of the neural network.

Citation Information

Patent Citations

  • System and method for lithography simulation

    US20050076322A1

  • Source-mask optimization in lithographic apparatus

    US20100315614A1

  • Exposure device including an electrically aligned electronic mask for micropatterning

    US5229872A

  • Illumination device

    US5296891A

  • Method and apparatus for patterning and imaging member

    US5523193A