Method of training machine learning model for improving patterning process
By training machine learning models to predict the physical characteristics of the substrate and adjust the patterning process of the lithography device, the problem that lithography devices are difficult to accurately adjust the physical characteristics of the substrate during the patterning process is solved, and the stability and yield of the patterning process are achieved.
Patent Information
- Application Number
- CN202510372495.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-13
- Filing Date
- 2020-07-30
- Publication Date
- 2025-06-13
AI Technical Summary
It is difficult for existing lithography equipment to accurately adjust the physical characteristics of the substrate during the patterning process, resulting in instability and low yields in the patterning process.
Using a machine learning model trained to predict physical characteristics associated with the substrate, the patterning process is adjusted by obtaining reference images and adjusting model parameter values, the model is iteratively trained to reduce the difference in the cost function.
Improves the accuracy and stability of the patterning process, enhances the yield and quality of device manufacturing, and expands design selection and process window.
Smart Images

Figure CN120143543A_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application with an application date of July 30, 2020, an application number of 202080055236.9, and an invention title of "Method for Training a Machine Learning Model for Improving a Patterning Process".
[0002] Cross-reference to related applications
[0003] This application claims the priority of U.S. Application No. 62 / 886,058, filed on August 13, 2019, and the entire content of the U.S. application is incorporated herein by reference. Technical field
[0004] The present disclosure relates to techniques for improving the performance of a device manufacturing process. The techniques can be used in combination with a lithographic apparatus. Background art
[0005] A lithographic apparatus is a machine that applies a desired pattern onto a target portion of a substrate. A lithographic apparatus can be used, for example, in the manufacture of integrated circuits (ICs). In that case, a patterning device (which is alternatively referred to as a mask or a reticle) can be used to generate a circuit pattern corresponding to a single layer of the IC, and this pattern can be imaged onto a target portion (e.g., including part of a die, one or more dies) of a substrate (e.g., a silicon wafer) having a layer of radiation-sensitive material (resist). Generally, a single substrate will contain a network of adjacent target portions that are successively exposed. Known lithographic apparatuses include so-called steppers, in which each target portion is irradiated by exposing the entire pattern onto the target portion at once; and so-called scanners, in which each target portion is irradiated by scanning the pattern via the beam in a given direction ("scan" direction) while synchronously scanning the substrate in a direction parallel or anti-parallel to this direction.
[0006] Before transferring the circuit pattern from the patterning device to the substrate, the substrate can undergo various processes, such as priming, resist coating, and soft baking. After exposure, the substrate can undergo other processes, such as post-exposure bake (PEB), development, hard bake, and measurement / inspection of the transferred circuit pattern. Such an array of processes serves as a basis for manufacturing a single layer of a device (e.g., an IC). The substrate can then undergo various processes such as etching, ion implantation (doping), metallization, oxidation, chemical mechanical polishing, etc., all of which are intended to finish the single layer of the device. If several layers are required in the device, the entire process or a variant thereof is repeated for each layer. Eventually, the devices will be present in each target portion on the substrate. These devices are then separated from each other by techniques such as dicing or sawing, whereby the individual devices can be mounted on a carrier, connected to pins, etc.
[0007] Thus, manufacturing devices such as semiconductor devices typically involves using multiple manufacturing processes to process a substrate (e.g., a semiconductor wafer) to form various features and multiple layers of the device. Such layers and features are typically fabricated and processed using, for example, deposition, lithography, etching, chemical mechanical polishing, and ion implantation. Multiple devices can be fabricated on multiple die on a substrate, and the devices are then separated into individual devices. Such a device manufacturing process can be regarded as a patterning process. The patterning process involves patterning steps such as optical and / or nanoimprint lithography using a patterning device in a lithography apparatus to transfer a pattern on the patterning device to the substrate, and the patterning process typically but optionally involves one or more associated pattern processing steps such as resist development by a developing apparatus, baking the substrate using a baking tool, etching using a pattern with an etching apparatus, etc. Summary of the Invention
[0008] In an embodiment, a method of training a machine learning model is provided, the machine learning model being configured to predict a value of a physical property associated with a substrate for adjusting a patterning process. The method involves: obtaining a reference image associated with a desired pattern to be printed on the substrate; determining a first set of model parameter values of the machine learning model such that a first cost function is reduced from an initial value of the cost function obtained using an initial set of model parameter values, where the first cost function is the difference between the reference image and an image generated via the machine learning model; and training the machine learning model using the first set of model parameter values such that a combination of the first cost function and a second cost function is iteratively reduced. In an embodiment, the second cost function is the difference between a measured value and a predicted value of a physical property associated with the desired pattern, the predicted value being predicted via the machine learning model.
[0009] Furthermore, in an embodiment, a computer program product is provided, including a non-transitory computer-readable medium having instructions recorded thereon, the instructions, when executed by a computer, implementing the foregoing method. Brief Description of the Drawings
[0010] Embodiments will now be described by way of example only with reference to the accompanying drawings, in which:
[0011] Figure 1 A block diagram showing various subsystems of a lithography system according to an embodiment;
[0012] Figure 2 A flow chart depicting an example for modeling and / or simulating at least a portion of a patterning process according to an embodiment;
[0013] Figure 3A flowchart of a method for training a machine learning model configured to predict values of physical properties associated with a substrate for adjusting a patterning process according to an embodiment;
[0014] Figure 4 Illustrates an example of a machine learning model having multiple layers trained according to the method in Figure 3 an embodiment;
[0015] Figure 5A and Figure 5B illustrates an example pattern shift of a grid relative to a grid that results in a grid - dependent error according to an embodiment;
[0016] Figure 6 Schematically depicts an embodiment of a scanning electron microscope (SEM) according to an embodiment;
[0017] Figure 7 Schematically depicts an embodiment of an electron beam inspection device according to an embodiment;
[0018] Figure 8 Is a block diagram of an example computer system according to an embodiment;
[0019] Figure 9 Is a schematic diagram of a lithographic projection apparatus according to an embodiment;
[0020] Figure 10 Is a schematic diagram of an extreme ultraviolet (EUV) lithographic projection apparatus according to an embodiment;
[0021] Figure 11 Is according to an embodiment of Figure 10 a more detailed view of the apparatus in
[0022] Figure 12 Is according to an embodiment of Figure 10 and Figure 11 a more detailed view of the source collector module of the apparatus in DETAILED DESCRIPTION
[0023] Before describing embodiments in detail, it is instructive to present an example environment in which the embodiments can be implemented.
[0024] Figure 1The figure shows an exemplary lithographic projection apparatus 10A. The main components are: a radiation source 12A, which can be a deep ultraviolet excimer laser source or another type of source including an extreme ultraviolet (EUV) source (as discussed above, the lithographic projection apparatus itself does not necessarily have a radiation source); illumination optics, which for example define partial coherence (denoted as sigma or σ) and can include optics 14A, 16Aa and 16Ab for shaping the radiation from source 12A; a patterning device 18A; and projection optics 16Ac, which projects an image of the patterned device pattern onto a substrate plane 22A. An adjustable filter or aperture 20A at the pupil plane of the projection optics can limit the range of beam angles incident on the substrate plane 22A, where the maximum possible angle defines the numerical aperture NA = n sin(Θmax) of the projection optics, where n is the refractive index of the medium between the substrate and the final element of the projection optics, and Θmax is the maximum angle of the beam leaving the projection optics that can still be incident on the substrate plane 22A.
[0025] In a lithographic projection apparatus, the source provides illumination (i.e., radiation) to the patterning device, and the projection optics direct and shape the illumination via the patterning device onto the substrate. The projection optics can include at least some of components 14A, 16Aa, 16Ab and 16Ac. The spatial image (AI) is the radiation intensity distribution at the substrate horizontal plane. The resist layer on the substrate is exposed, and the spatial image is transferred to the resist layer to serve as a potential “resist image” (RI) therein. The resist image can be defined as the spatial distribution of the solubility of the resist in the resist layer. A resist model can be used to calculate the resist image from the spatial image, an example of which can be found in US Patent Application Publication No. US 2009-0157360, the entire content of which is hereby incorporated herein by reference. The resist model is only about the properties of the resist layer (such as the effects of chemical processes occurring during exposure, PEB and development). The optical properties of the lithographic projection apparatus (such as the properties of the source, patterning device and projection optics) govern the spatial image. Since the patterning device used in a lithographic projection apparatus can be changed, it may be desirable to separate the optical properties of the patterning device from the optical properties of the remainder of the lithographic projection apparatus including at least the source and projection optics.
[0026] In an embodiment, assist features (sub-resolution assist features and / or printable resolution assist features) may be placed in a design layout based on how the design layout is optimized according to the methods of the present disclosure. For example, in an embodiment, the method employs a machine learning-based model to determine the pattern of a patterning device. The machine learning model may be a neural network, such as a convolutional neural network, which may be trained in a manner (e.g., as discussed in Figure 3 to obtain accurate predictions at a high rate, thus enabling full-chip simulation of the patterning process.
[0027] A set of training data may be used to train the neural network (i.e., determine the parameters of the neural network). The training data may include or consist of a set of training samples. Each sample may be a pair including an input object (typically a vector, which may be referred to as a feature vector) and a desired output value (also referred to as a supervisory signal), or consisting of the input object and the desired output value. The training algorithm analyzes the training data and adjusts the behavior of the neural network by adjusting the parameters of the neural network (e.g., the weights of one or more layers) based on the training data. The neural network after training may be used to map new samples.
[0028] In the context of determining the pattern of a patterning device, the feature vector may include one or more characteristics of the design layout included or formed by the patterning device (e.g., shape, arrangement, size, etc.), one or more characteristics of the patterning device (e.g., one or more physical properties such as dimensions, refractive index, material composition, etc.), and one or more characteristics of the illumination used in the lithography process (e.g., wavelength). The supervisory signal may include one or more characteristics of the pattern of the patterning device (e.g., CD, profile, etc., of the pattern of the patterning device).
[0029] Given a set of N training samples of the form such that x i is the feature vector of the i-th example and y i is its supervisory signal, the training algorithm searches for a neural network where X is the input space and Y is the output space. The feature vector is an n-dimensional vector representing the numerical features of some object. The vector space associated with these vectors is typically referred to as the feature space. Sometimes it is convenient to operate as follows: use a scoring function to represent g, such that g is defined to return the y value that gives the highest score: . Denote the space of scoring functions by F.
[0030] The neural network may be probabilistic, where g takes the form of a conditional probability model or f takes the form of a joint probability model .
[0031] There are two basic methods for selecting f or g: empirical risk minimization and structural risk minimization. Empirical risk minimization seeks a neural network that best fits the training data. Structural risk minimization includes a penalty function that controls the bias / variance trade-off. For example, in an embodiment, the penalty function can be based on a cost function, which can be mean squared error, number of defects, EPE, etc. The function (or weights within the function) can be modified to reduce or minimize the variance.
[0032] In both cases, it is assumed that the training set includes one or more samples of independent and identically distributed pairs or consists of one or more samples of the pairs. In an embodiment, to measure how well a function fits the training data, a loss function is defined. For a training sample , the loss of the predicted value is .
[0033] The risk R(g) of a function g is defined as the expected loss of g. This can be estimated from the training data as .
[0034] In an embodiment, a machine learning model of a patterning process can be trained to predict, for example, the profile, pattern, CD, and / or the profile, CD, edge placement (such as edge placement error), etc. in a resist and / or an etched image on a wafer of a mask pattern. The goal of training is to achieve accurate prediction of, for example, the profile, spatial image intensity slope, and / or CD of the printed pattern on a wafer. The expected design (such as a wafer target layout to be printed on a wafer) is typically defined as a pre-OPC design layout that can be provided in a standardized digital file format such as GDSII or OASIS or other file formats.
[0035] Figure 2 FIG. shows an exemplary flowchart for modeling and / or simulating multiple parts of a patterning process. As will be appreciated, the model can represent different patterning processes and need not include all of the models described below. The source model 1200 represents the optical characteristics of the illumination of a patterning device (including radiation intensity distribution, bandwidth, and / or phase distribution). The source model 1200 can represent the optical characteristics of the illumination, including but not limited to numerical aperture settings, illumination sigma (σ) settings, and any particular illumination shape (e.g., off-axis radiation shape such as annular, quadrupole, dipole, etc.), where σ (or sigma) is the outer radial extent of the illuminator.
[0036] The projection optical device model 1210 represents the optical characteristics of the projection optical device (including the change in the radiation intensity distribution and / or phase distribution caused by the projection optical device). The projection optical device model 1210 can represent the optical characteristics of the projection optical device, including aberration, distortion, one or more refractive indices, one or more physical sizes, one or more physical dimensions, etc.
[0037] The patterning device / design layout model module 1220 captures how the design features are laid out in the pattern of the patterning device and can include a representation of the detailed physical properties of the patterning device, as described, for example, in U.S. Patent No. 7,587,704, which is incorporated herein by reference in its entirety. In an embodiment, the patterning device / design layout model module 1220 represents the optical characteristics (including the change in the radiation intensity distribution and / or phase distribution caused by a given design layout) of a design layout (e.g., a device design layout corresponding to features of an integrated circuit, a memory, an electronic device, etc.), which is a representation of the arrangement of features on or formed by the patterning device. Since the patterning device used in a lithographic projection apparatus can be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the remainder of the lithographic projection apparatus, which at least includes the illumination and projection optical devices. The purpose of the simulation is typically to accurately predict, for example, edge placement and CD, which can then be compared with the device design. The device design is typically defined as the pre-OPC patterning device layout and will be provided in a standardized digital file format such as GDSII or OASIS.
[0038] The spatial image 1230 can be simulated from the source model 1200, the projection optical device model 1210, and the patterning device / design layout model 1220. The spatial image (AI) is the radiation intensity distribution at the substrate plane. The optical properties of the lithographic projection apparatus (e.g., the properties of the illumination, the patterning device, and the projection optical device) govern the spatial image.
[0039] The resist layer on the substrate is exposed by a spatial image, and the spatial image is transferred to the resist layer to serve as a potential "resist image" (RI) therein. The resist image (RI) can be defined as the spatial distribution of the solubility of the resist in the resist layer. The resist image 1250 can be simulated from the spatial image 1230 using a resist model 1240. The resist model can be used to calculate the resist image from the spatial image, an example of which can be found in the US patent application with publication number US 2009-0157360, the entire disclosure of which is hereby incorporated by reference. The resist model typically describes the effects of chemical processes occurring during resist exposure, post-exposure bake (PEB), and development in order to predict, for example, the profile of the resist features formed on the substrate, and thus it typically only relates to such properties of the resist layer (such as the effects of chemical processes occurring during exposure, post-exposure bake, and development). In an embodiment, the optical properties of the resist layer (such as refractive index, film thickness, propagation, and polarization effects) can be captured as part of the projection optical device model 1210.
[0040] Thus, generally, the connection between the optical model and the resist model is the simulated spatial image intensity within the resist layer, which results from the projection of radiation onto the substrate, refraction at the resist interface, and multiple reflections in the resist film stack. The radiation intensity distribution (spatial image intensity) becomes a potential "resist image" through the absorption of incident energy, which is further modified by diffusion processes and various loading effects. An efficient simulation method fast enough for full-chip applications approximates the true three-dimensional intensity distribution in the resist stack by a two-dimensional spatial (and resist) image.
[0041] In an embodiment, the resist image can be used as an input to the post-pattern transfer process model module 1260. The post-pattern transfer process model 1260 defines the performance of one or more post-resist development processes (such as etching, development, etc.).
[0042] The simulation of the patterning process can, for example, predict the profile, CD, edge placement (such as edge placement error), etc. in the resist and / or the etched image. Thus, the purpose of the simulation is to accurately predict, for example, the edge placement of the printed pattern and / or the slope of the spatial image intensity and / or the CD, etc. These values can be compared with the expected design to, for example, correct the patterning process, identify the locations where defects are predicted to occur, etc. The expected design is typically defined as a pre-OPC design layout that can be provided in a standardized digital file format such as GDSII or OASIS or other file formats.
[0043] Thus, the model formula describes most, if not all, of the known physical and chemical actions of the overall process, and each of the model parameters in the model formula desirably corresponds to a different physical or chemical effect. The model formula thus sets an upper bound on how well the model can be used to simulate the overall manufacturing process.
[0044] In a patterning process (such as lithography, electron beam lithography, directed self-assembly, etc.), an energy-sensitive material (such as a resist) deposited on a substrate typically undergoes a pattern transfer step (such as via exposure). After the pattern transfer step, various post steps such as resist baking and subtractive processes such as resist development and etching are applied. These post-exposure steps or processes impose various effects on the substrate, which result in the patterned layer or the etch having a structure (which has a size different from the target size).
[0045] The computational analysis of the patterning process employs a predictive model that, when properly calibrated, can produce an accurate prediction of the dimensions output from the patterning process. Models of the post-exposure process are typically calibrated based on empirical measurements. The calibration process includes running test wafers with different process parameters, measuring the critical dimensions obtained after the post-exposure process, and calibrating the model to the measurement results. In practice, a well-calibrated model makes quick and accurate predictions of dimensions for improving device performance or yield, enhancing the process window or increasing design choices. In an example, using a deep convolutional neural network (CNN) to model the post-exposure process produces a model accuracy comparable to or better than that produced by traditional techniques, which typically involve modeling using physical term expressions or closed-form equations. Compared to traditional modeling techniques, deep learning convolutional neural networks reduce the requirement for process knowledge in model development and increase the dependence on the personal experience of engineers in model tuning. In short, the deep CNN model for the post-exposure process consists of an input and an output layer and multiple hidden layers such as convolutional layers, normalization layers, and pooling layers. The parameters of the hidden layers are optimized to give the minimum of a loss function. In an embodiment, the CNN model can be trained to model the behavior of any process or a combination of processes related to the patterning process.
[0046] Figure 3FIG. 0 is a flow chart of a method 300 for training a machine learning model 305 (e.g., a CNN), the machine learning model 305 being configured to predict values of physical properties associated with a substrate for adjusting a patterning process. The training method is a more accurate training method compared to existing methods. For example, the training is based on reducing a specific error associated with the model prediction (e.g., via a first cost function, a second cost function, a grid-dependent error, an edge localization error, etc. in one or more training steps) by applying specific weight factors in the CNN, where the weights are related to these errors, thus improving the overall modeling quality.
[0047] After training, the machine learning model 305 can be referred to as the trained machine learning model 305'. Training the machine learning model 305' can also be performed to determine the physical properties. Additionally, patterning process parameters (e.g., dose, focus, OPC, etc.) can be adjusted based on the physical property values to improve the patterning process.
[0048] The method involves training the machine learning model 305 in successive steps to model a process of a patterning process (e.g., a post-exposure process). Successive steps refer to training the machine learning model 305 using a first cost function to determine an initial set of model parameter values, and further training the machine learning model 305 using such initial model parameter values with a second cost function. Such successive step training helps for faster convergence and produces a more accurate model compared to a single-step training process involving a single cost function. The method 300 is discussed in further detail below.
[0049] Process P301 involves obtaining a reference image 301 associated with a desired pattern to be printed on a substrate. In an embodiment, obtaining the reference image 301 involves executing a process model configured to produce the reference image 301 as an output, where the process model models a part of the patterning process. In an embodiment, the process model is a calibrated model of an optical device model, a resist model, and / or an etch model of the patterning process. Thus, in an embodiment, the reference image 301 is a spatial image, a resist image, and / or an etch image of the desired pattern.
[0050] Process P303 involves determining a first set of model parameter values 303 of the machine learning model 305 such that a first cost function is reduced from an initial value of the cost function obtained using an initial set of model parameter values. In an embodiment, the first cost function is the difference between the reference image 301 and the image produced via the machine learning model 305. In an embodiment, the reference image 301 and the produced image are pixelated images. Thus, the first cost function can be the difference in intensity values of the pixelated images. The intensity of a pixel indicates the presence or absence of a feature. For example, a peak intensity signal indicates an edge of a feature (e.g., a contact hole) in the image.
[0051] In an embodiment, determining the first set of model parameter values 303 of the machine learning model 305 is an iterative process. The iteration involves: generating an image by executing the machine learning model 305 using a desired pattern; determining the difference between the generated image and the reference image 301; and adjusting the model parameter values of the machine learning model 305 such that the difference is reduced. In an embodiment, the difference between the generated image and the reference image 301 is minimized.
[0052] Thus, using the first set of initial model values, the machine learning model 305'' (the model 305'' refers to the machine learning model 305 having the model parameter values 303) can accurately predict a spatial image, a resist image, or an etch image associated with a substrate. Additionally, the profile and physical characteristics of the pattern can be obtained from the predicted image for further analysis or improvement of the patterning process.
[0053] In an embodiment, the model parameters are weights and / or biases associated with one or more layers of the machine learning model 305. In an embodiment, the machine learning model 305 is a convolutional neural network including multiple layers, and each layer is associated with weights and / or biases.
[0054] Additionally, the process P305 involves training the machine learning model 305'' using the first set of model parameter values 303 such that the combination of the first cost function and the second cost function is reduced. In an embodiment, the expression is used to calculate the combination of the first cost function (CF1) and the second cost function (CF2), where c1 and c2 are coefficients that can be adjusted to minimize the combination.
[0055] In an embodiment, the second cost function is the difference between the measured value 304 of the physical characteristics associated with the desired pattern and the predicted value, and the predicted value is predicted via the machine learning model 305''. After the training process is completed, a trained machine learning model 305' configured to determine the physical characteristics of the pattern to be imaged in the substrate is obtained.
[0056] In an embodiment, the physical characteristics determined from the predicted image are the critical dimension or the edge placement error associated with the desired pattern. In an embodiment, the physical characteristics are determined using the profile of the pattern in the predicted image of the model. For example, an algorithm can be employed to define the gauge points along the profile and the cutting lines intersecting the profile at the gauge positions. Additionally, to determine the CD, the distance between the gauge points can be measured. Similarly, the EPE can be measured using the gauge points relative to a reference profile (e.g., the reference profile associated with the reference image 301).
[0057] In an embodiment, the measurement value 304 is a CD value obtained, for example, via a metrology tool configured to measure the desired printed pattern on a substrate. In an embodiment, the metrology tool is a scanning electron microscope (SEM) (see, for example, Figures 6 to 7 ) and the measurement value is obtained from the SEM image. In an embodiment, the measurement value 304 is an intensity value of a spatial image associated with the desired pattern. Thus, during the training process, the measurement value 304 (e.g., CD) is compared with the predicted physical property (e.g., the predicted CD). Training is performed such that the predicted value closely matches the measurement value 304.
[0058] In an embodiment, the training of the machine learning model 305 is an iterative process. The iteration involves: initializing the model parameters of the machine learning model 305 with a first set of model parameter values 303; predicting the value of the physical property associated with the substrate by executing the machine learning model 305 with the desired pattern; obtaining the measurement value 304 of the physical property of the desired printed pattern on the substrate via the metrology tool; and adjusting the model parameter values of the machine learning model 305 such that the combination of the first cost function and the second cost function is reduced.
[0059] In an embodiment, the adjustment of the model parameter values is based on gradient descent of the combination of the first cost function and the second cost function. In an embodiment, the sum of the first cost function and the second cost function is minimized. In an embodiment, adjusting the model parameter values of the machine learning model 305 involves determining the gradient map of the sum of the first cost function and the second cost function as a function of the model parameters. Subsequently, based on the gradient map, the model parameter values are determined such that the sum of the cost functions is minimized.
[0060] In an embodiment, adjusting the model parameter values includes adjusting the following values: one or more weights of the layers of the convolutional neural network, one or more biases of the layers of the convolutional neural network, hyperparameters of the CNN, and / or the number of layers of the CNN. In an embodiment, the number of layers is a hyperparameter of the CNN, which can be preselected and may not be changed during the training process. In an embodiment, a series of training processes can be performed with the number of layers being modifiable. In Figure 4 an example of a CNN is illustrated.
[0061] In an embodiment, the training (e.g., Figure 4The CNN) involves: determining the value of a first cost function; and gradually adjusting the weights of one or more layers of the CNN such that the first cost function is reduced (minimized in one embodiment). In an embodiment, the first cost function is the difference between the predicted resist image or the predicted aerial image (e.g., the output vector of the CNN) and the ground truth resist image obtained from the printed substrate (e.g., using an SEM tool). The first cost function or the difference is reduced by modifying the values of the CNN model parameters (e.g., weights, biases, strides, etc.). In an embodiment, the first cost function is calculated as . In such a step, the input to the CNN includes a measured image or a simulated image (e.g., AI / RI) and has an initial value that can be arbitrarily selected. After several iterations of training, the optimized value is obtained and further used as the first set of model parameter values 303 for further training.
[0062] In further training, after reducing (or minimizing) the first cost function, physical properties can be obtained from the predicted image of the machine learning model 305. For example, CD or EPE values can be obtained from the predicted resist image or intensity values can be obtained from the predicted aerial image. These predicted CD, EPE, and / or intensity values are compared with the measured values 304 to further train the machine learning model 305 using a second cost function associated with the physical properties other than the first cost function.
[0063] For example, the second cost function can be the edge placement error (EPE). In such a case, the measured value of EPE and the predicted EPE are used to determine the second cost function. In an embodiment, the second cost function can be expressed as: , where can be EPE, and the function performs contour extraction from the predicted pattern (e.g., via the CNN) and further determines the difference. In an embodiment, the input to such a CNN includes the predicted image (e.g., AI / RI). cnn_parameters can be the weights and biases of the CNN and the values of cnn_parameters are the initial model parameter values obtained based on the first cost function.
[0064] In an embodiment, the gradient corresponding to a cost function (e.g., a first cost function and / or a second cost function) can be dcost / dparameter, where the value of cnn_parameters can be updated based on an equation (e.g., parameter = parameter - learning_rate * gradient). In an embodiment, the parameter can be a weight and / or a bias, and learning_rate can be a hyperparameter used to tune the training process and can be selected by a user or a computer to improve the convergence of the training process (e.g., faster convergence).
[0065] In an embodiment, the trained machine learning model 305' (e.g., Figure 9 the trained CNN) can also be used to correct the simulated pattern or any of its characteristics.
[0066] In an embodiment, method 300 can also involve a process P305 of employing a third cost function for further training the trained machine learning model 305'. Process P305 involves using a first set of model parameter values 303 to train the machine learning model 305' such that the combination of the first cost function, the second cost function, and the third cost function is reduced (minimized in one embodiment). In an embodiment, the third cost function is a grid-dependent function.
[0067] The grid-dependent error is related to the simulation mechanism (e.g., image-based) used during the simulation of the patterning process. In an embodiment, the simulation of one or more process models is image-based, where a grid can be placed on an image (e.g., an image of a substrate pattern) and only the features on the grid are evaluated during the simulation while interpolating off-grid features. Such interpolation may result in inaccurate simulation results (e.g., the substrate pattern). Additionally, the grid size can affect the simulation speed as well as the accuracy of the results. A small grid size gives accurate simulation results but significantly slows down the simulation. Thus, a larger grid can be used for faster simulation, which may adversely affect the accuracy of the simulation results (e.g., the simulated substrate pattern).
[0068] Generally, simulation is an iterative process, so any shift in the pattern placement relative to the grid in each iteration will cause an error in the predicted pattern. Thus, the simulation results including the grid-dependent error can be used to determine the parameters of the patterning process (e.g., dose, focus, mask pattern, etc.), for example, to improve the patterning process. Due to the grid-dependent error, the determined parameters may not result in the desired yield of the patterning process. Therefore, the grid-dependent error should be removed or minimized. According to the present disclosure, such a grid-dependent error is handled via the third cost function.
[0069] Figures 5A to 5BThe figure shows an example pattern shift relative to a grid that results in grid-dependent errors. The accompanying drawings show predicted profiles 501 / 511 (dashed lines) and input profiles 502 / 512 (e.g., design or desired profiles). In Figure 5A , the entire input profile 501 lies on the grid. However, in Figure 5B , a portion of the input profile 511 is off the grid, e.g., at a corner. This results in a difference between the model-predicted profiles 502 and 512. In an embodiment, for example, in an LMC or OPC application, the same pattern can be iteratively presented at different positions on the grid, and it is desirable to have invariant model predictions regardless of the pattern's position. However, no model can achieve perfect shift invariance. Some pathological models may produce large profile differences between pattern shifts.
[0070] In an embodiment, the grid-dependent (GD) error can be measured as follows. To measure the GD error, the pattern and the gauge are shifted together along the profile in sub-pixel steps. For example, for a pixel size = 14 nm, the pattern / gauge can be shifted 1 nm per step in the x and / or y direction. With each shift, the CD of the model prediction is measured along the gauge. Subsequently, the variance in the set of the model-predicted CDs indicates the grid-dependent error.
[0071] In an embodiment, a trained machine learning model can be used for various applications related to a patterning process to improve the yield of the patterning process. For example, method 300 also involves predicting a substrate image of a design layout via the trained machine learning model; determining a mask layout to be used for manufacturing a mask for the patterning process via an OPC simulation using the design layout and the predicted substrate image. In an embodiment, the OPC simulation involves simulating a patterning process model via using the geometry of the design layout and corrections associated with a plurality of segments, determining a simulated pattern to be printed on the substrate; and determining an optical proximity effect correction to the design layout such that the difference between the simulated pattern and the design layout is reduced. In an embodiment, determining the optical proximity effect correction is an iterative process. The iteration involves adjusting the shape and / or size of the geometry of the main features and / or one or more auxiliary features of the design layout such that a performance metric of the patterning process is reduced. In an embodiment, the one or more auxiliary features are obtained from the predicted post-OPC image of the machine learning model.
[0072] In some embodiments, the inspection device can be a scanning electron microscope (SEM) that produces an image of the structures (e.g., some or all of the structures of a device) exposed or transferred onto a substrate. Figure 6An embodiment of an SEM tool is depicted. A primary electron beam EBP emitted from an electron source ESO is converged by a condenser lens CL and then passes through beam deflectors EBD1, an E×B deflector EBD2, and an objective lens OL to irradiate a substrate PSub on a substrate stage ST at a focal point.
[0073] When the substrate PSub is irradiated with the electron beam EBP, secondary electrons are generated from the substrate PSub. The secondary electrons are deflected by the E×B deflector EBD2 and detected by a secondary electron detector SED. A two-dimensional electron beam image can be obtained by detecting electrons generated from a sample synchronously with, for example, a two-dimensional scan of the electron beam by the beam deflector EBD1 or synchronously with a repeated scan of the electron beam EBP in the X or Y direction by the beam deflector EBD1, and by continuously moving the substrate PSub in the other direction in the X or Y direction by the substrate stage ST.
[0074] A signal detected by the secondary electron detector SED is converted into a digital signal by an analog / digital (A / D) converter ADC, and the digital signal is sent to an image processing system IPU. In an embodiment, the image processing system IPU may have a memory MEM to store all or part of the digital image for processing by a processing unit PU. The processing unit PU (such as specially designed hardware or a combination of hardware and software) is configured to convert or process the digital image into a data set representing the digital image. In addition, the image processing system IPU may have a storage medium STOR configured to store the digital image and the corresponding data set in a reference database. A display device DIS may be connected to the image processing system IPU such that an operator can perform necessary operations of the equipment by means of a graphical user interface.
[0075] As mentioned above, an SEM image can be processed to obtain a contour describing the edges of objects representing device structures in the image. These contours are then quantified via metrics such as CD. Thus, images of device structures are typically compared and quantified via simplistic metrics such as edge-to-edge distance (CD) or simple pixel differences between images. Typical contour models for detecting the edges of objects in an image to measure CD use image gradients. In fact, those models rely on strong image gradients. But in practice, images are typically noisy and have discontinuous boundaries. Techniques such as smoothing, adaptive thresholding, edge detection, erosion, and dilation can be used to process the results of an image gradient contour model to address noisy and discontinuous images, but will ultimately result in a low-resolution quantification of a high-resolution image. Thus, in most cases, mathematical operations are performed on an image of a device structure to reduce noise, and automating edge detection results in a loss of resolution of the image, thus resulting in a loss of information. As a result, the outcome is a low-resolution quantification equivalent to a simplistic representation of a complex high-resolution structure.
[0076] Accordingly, it is desirable to have a mathematical representation of a structure (such as circuit features, alignment marks, or metrology target portions (such as grating features), etc.) that is produced or expected to be produced using a patterning process, regardless of whether the structure is located in a latent resist image, in a developed resist image, or in a layer transferred to a substrate by etching, for example, which can maintain resolution and also describe the general shape of the structure. In the context of lithography or other patterning processes, the structure can be a device or a part thereof being manufactured, and the image can be an SEM image of the structure. In some cases, the structure can be a feature of a semiconductor device (such as an integrated circuit). In this case, the structure can be referred to as a pattern or a desired pattern including multiple features of the semiconductor device. In some cases, the structure can be an alignment mark or a portion thereof (such as a grating of the alignment mark) for use in an alignment measurement process to determine the alignment of an object (such as a substrate) with another object (such as a patterning device), or a metrology target or a portion thereof (such as a grating of the metrology target) for measuring parameters of a patterning process (such as overlay, focus, dose, etc.). In an embodiment, the metrology target is a diffraction grating for measuring, for example, overlay.
[0077] Figure 7 Another embodiment of an inspection device is schematically illustrated. The system is configured to inspect a sample 90 (such as a substrate) on a sample platform 88 and includes a charged particle beam generator 81, a condenser lens module 82, a probe forming objective lens module 83, a charged particle beam deflection module 84, a secondary charged particle detector module 85, and an image forming module 86.
[0078] The charged particle beam generator 81 generates a primary charged particle beam 91. The condenser lens module 82 condenses the generated primary charged particle beam 91. The probe forming objective lens module 83 focuses the condensed primary charged particle beam into a charged particle beam probe 92. The charged particle beam deflection module 84 scans the formed charged particle beam probe 92 across the surface of an area of interest on the sample 90 fixed to the sample platform 88. In an embodiment, the charged particle beam generator 81, the condenser lens module 82, and the probe forming objective lens module 83 or their equivalent designs, alternatives, or any combination thereof together form a charged particle beam probe generator that generates a scanned charged particle beam probe 92.
[0079] The secondary charged particle detector module 85 detects secondary charged particles 93 emitted from the sample surface (possibly together with other reflected or scattered charged particles from the sample surface) after being bombarded by the charged particle beam probe 92 to generate a secondary charged particle detection signal 94. The image formation module 86 (such as a computing device) is coupled to the secondary charged particle detector module 85 to receive the secondary charged particle detection signal 94 from the secondary charged particle detector module 85, and thus forms at least one scanned image. In an embodiment, the secondary charged particle detector module 85 and the image formation module 86 or their equivalent designs, alternatives, or any combination thereof together form an image forming device that forms a scanned image from the detected secondary charged particles emitted from the sample 90 bombarded by the charged particle beam probe 92.
[0080] In an embodiment, the monitoring module 87 is coupled to the image formation module 86 of the image forming device to monitor, control, etc. the patterning process using the scanned image of the sample 90 received from the image formation module 86, and / or to derive parameters for patterning process design, control, monitoring, etc. Thus, in an embodiment, the monitoring module 87 is configured or programmed to perform the methods described herein. In an embodiment, the monitoring module 87 includes a computing device. In an embodiment, the monitoring module 87 includes a computer program that is configured to provide the functionality herein and is encoded on a computer-readable medium that forms or is disposed within the monitoring module 87.
[0081] In an embodiment, similar to using a probe to inspect a substrate Figure 6 of an electron beam inspection tool, Figure 7 the electron current in the system of Figure 6 is significantly larger compared to, for example, a CD SEM such as depicted in
[0082] so that the probe spot is large enough for the inspection speed to be fast. However, due to the large probe spot, the resolution may not be as high compared to a CD SEM. In an embodiment, the inspection device discussed above can be a single-beam or multi-beam device without limiting the scope of the present disclosure. Figure 6 and / or Figure 7 The system can process SEM images from, for example,
[0083] Figure 8FIG. is a block diagram of a computer system 100 that can assist in implementing the methods and processes disclosed herein. Computer system 100 includes a bus 102 or other communication mechanism for communicating information, and a processor 104 (or processors 104 and 105) coupled to bus 102 for processing information. Computer system 100 also includes a main memory 106 coupled to bus 102 for storing information and instructions to be executed by processor 104, such as random access memory (RAM) or other dynamic storage device. Main memory 106 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 104. Computer system 100 further includes a read only memory (ROM) 108 or other static storage device coupled to bus 102 for storing static information and instructions for processor 104. A storage device 110, such as a magnetic disk or optical disk, is provided and storage device 110 is coupled to bus 102 for storing information and instructions.
[0084] Computer system 100 may be coupled via bus 102 to a display 112 for displaying information to a computer user, such as a cathode ray tube (CRT), flat panel display, or touch panel display. An input device 114 including alphanumeric keys and other keys is coupled to bus 102 for transmitting information and command selections to processor 104. Another type of user input device is a cursor control 116 for transmitting direction information and command selections to processor 104 and for controlling cursor movement on display 112, such as a mouse, trackball, or cursor direction keys. Such input devices typically have two degrees of freedom in two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a position in a plane. A touch panel (screen) display may also be used as an input device.
[0085] In accordance with one embodiment, portions of a process may be executed by computer system 100 in response to processor 104 executing one or more sequences of one or more instructions contained in main memory 106. Such instructions may be read into main memory 106 from another computer readable medium, such as storage device 110. Execution of the sequences of instructions contained in main memory 106 causes processor 104 to perform the process steps described herein. One or more processors in a multiprocessing arrangement may also be employed to execute the sequences of instructions contained in main memory 106. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Accordingly, the description herein is not limited to any specific combination of hardware circuitry and software.
[0086] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 104 for execution. Such a medium may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 110. Volatile media includes dynamic memory, such as main memory 106. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, DVD, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave as described below, or any other medium from which a computer can read.
[0087] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions may initially be carried on a magnetic disk of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 102 can receive the data carried in the infrared signal and place the data on bus 102. Bus 102 carries the data to main memory 106, from which processor 104 retrieves and executes the instructions. The instructions received by main memory 106 may optionally be stored on storage device 110 either before or after execution by processor 104.
[0088] Computer system 100 also desirably includes a communication interface 118 coupled to bus 102. Communication interface 118 provides two-way data communication coupling to network link 120, which is connected to local area network 122. For example, communication interface 118 may be an integrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 118 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 118 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0089] Network link 120 typically provides data communication to other data equipment via one or more networks. For example, network link 120 may provide a connection from local area network 122 to host computer 124 or to a data device operated by an Internet service provider (ISP) 126. ISP 126 in turn provides data communication services via a global packet data communication network (now commonly referred to as the "Internet" 128). Both local area network 122 and the Internet 128 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals via various networks and signals on network link 120 and via communication interface 118 (which carry digital data to and from computer system 100) are example forms of carriers that convey information.
[0090] Computer system 100 can send messages and receive data including program code via a network, network link 120, and communication interface 118. In the Internet example, server 130 can transmit requested code for an application program via the Internet 128, ISP 126, local area network 122, and communication interface 118. One such downloaded application program can provide, for example, the irradiation optimization of the embodiments. The received code can be executed by processor 104 when it is received, and / or stored in storage device 110 or other non-volatile memory for later execution. In this way, computer system 100 can obtain application code in the form of a carrier wave.
[0091] Figure 9 Schematically depicts an exemplary lithographic projection apparatus that can utilize the techniques described herein. The apparatus includes:
[0092] - An illumination system IL, which is used to condition a radiation beam B. In such a particular case, the illumination system also includes a radiation source SO;
[0093] - A first stage (e.g., a patterning device stage) MT, which is provided with a patterning device holder for holding a patterning device MA (e.g., a mask), and is connected to a first positioner for accurately positioning the patterning device relative to an article PS;
[0094] - A second stage (substrate stage) WT, which is provided with a substrate holder for holding a substrate W (e.g., a silicon wafer coated with resist), and is connected to a second positioner for accurately positioning the substrate relative to an article PS;
[0095] -A projection system (“lens”) PS (e.g., a refractive, reflective, or catadioptric optical system) for imaging an illumination portion of a patterning device MA onto a target portion C (e.g., including one or more dies) of a substrate W.
[0096] As depicted herein, the device is of the transmissive type (i.e., has a transmissive patterning device). However, in general, the device can also be of the reflective type, e.g., (having a reflective patterning device). The device can employ different kinds of patterning devices than classical masks; examples include programmable mirror arrays or LCD matrices.
[0097] A source SO (e.g., a mercury lamp, an excimer laser, a laser-produced plasma (LPP) EUV source) generates a radiation beam. For example, such a beam is fed directly or after having traversed conditioning devices such as an expander Ex into an illumination system (illuminator) IL. The illuminator IL can include conditioning devices AD for setting the outer and / or inner radial extent (commonly referred to as σ - outer and σ - inner respectively) of the intensity distribution in the beam. In addition, the illuminator will generally include various other components, such as an integrator IN and a condenser CO. In this way, the beam B incident on the patterning device MA has the desired uniformity and intensity distribution in its cross-section.
[0098] Regarding Figure 9 It should be noted that the source SO can be within the housing of the lithographic projection apparatus (as is typically the case when the source SO is a mercury lamp), but the source SO can also be remote from the lithographic projection apparatus, and the radiation beam generated by it is directed into the apparatus (e.g., by means of a suitable directing mirror); this is typically the case when the source SO is an excimer laser (e.g., based on KrF, ArF, or F 2 laser).
[0099] The beam PB then intercepts the patterning device MA held on a patterning device table MT. After having traversed the patterning device MA, the beam B passes through a lens PL that focuses the beam B onto the target portion C of the substrate W. By means of a second positioning device (and an interferometric device IF), the substrate table WT can be accurately moved, e.g., so as to position different target portions C in the path of the beam PB. Similarly, e.g., after mechanically retrieving the patterning device MA from a patterning device library or during scanning, a first positioning device can be used to accurately position the patterning device MA relative to the path of the beam B. In general, this will be done by means of Figure 9The long-stroke module (coarse positioning) and short-stroke module (fine positioning) clearly depicted therein are used to move the stages MT and WT. However, in the case of a stepper (relative to a step-and-scan tool), the patterning device stage MT can be connected only to the short-stroke actuator or can be fixed.
[0100] The tool depicted can be used in two different modes:
[0101] - In the step mode, the patterning device stage MT remains substantially stationary, and the entire patterning device image is projected onto the target portion C at once (i.e., a single "flash"). Then, the substrate stage WT is displaced in the x and / or y direction so that different target portions C can be irradiated by the beam PB;
[0102] - In the scan mode, substantially the same situation applies except that a given target portion C is not exposed during a single "flash". Instead, the patterning device stage MT can move at a speed v in a given direction (the so-called "scan direction", e.g., the y direction) so as to scan the projection beam B over the patterning device image; simultaneously, the substrate stage WT moves at a speed V = Mv in the same or opposite direction at the same time, where M is the magnification of the lens PL (typically, M = 1 / 4 or 1 / 5). In this way, a relatively large target portion C can be exposed without compromising the resolution.
[0103] Figure 10 Another exemplary lithographic projection apparatus 1000 is schematically depicted, including:
[0104] - A source collector module SO for providing radiation.
[0105] - An illumination system (illuminator) IL configured to condition a radiation beam B (e.g., EUV radiation) from the source collector module SO.
[0106] - A support structure (e.g., a mask table) MT configured to support a patterning device (e.g., a mask or a reticle) MA and connected to a first positioner PM configured to accurately position the patterning device;
[0107] - A substrate table (e.g., a wafer table) WT configured to hold a substrate (e.g., a wafer coated with a resist) W and connected to a second positioner PW configured to accurately position the substrate; and
[0108] - A projection system (e.g., a reflective projection system) PS configured to project the pattern imparted to the radiation beam B by the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W.
[0109] As depicted herein, the device 1000 is of the reflective type (e.g., employing a reflective mask). It should be noted that since most materials are absorptive in the EUV wavelength range, the patterning device may have a multilayer reflector including, for example, multiple stacked layers of molybdenum and silicon. In one example, the multilayer reflector has 40 pairs of molybdenum and silicon layers, where the thickness of each layer is a quarter wavelength. X-ray lithography can be utilized to generate even smaller wavelengths. Since most materials are absorptive at EUV and x-ray wavelengths, a thin sheet of patterned absorptive material on the topography of the patterning device (e.g., a TaN absorber on top of the multilayer reflector) defines where features will be printed (positive resist) or not printed (negative resist).
[0110] Reference Figure 10 , the illuminator IL receives an extreme ultraviolet radiation beam from the source collector module SO. Methods of generating EUV radiation include but are not necessarily limited to converting a material into a plasma state having at least one element (e.g., xenon, lithium, or tin) using one or more emission spectral lines in the EUV range. In one such method (commonly referred to as laser-produced plasma (“LPP”)), a plasma can be generated by irradiating a fuel (such as droplets, streams, or clusters of a material having a spectral line-emitting element) with a laser beam. The source collector module SO can be part of an EUV radiation system including a laser ( Figure 10 not shown in the figure) for providing the laser beam that excites the fuel. The resulting plasma emits output radiation (e.g., EUV radiation), and the output radiation is collected using a radiation collector disposed in the source collector module. For example, when a CO2 laser is used to provide the laser beam for fuel excitation, the laser and the source collector module can be separate entities.
[0111] In such a case, the laser is not considered part of the lithographic apparatus, and the radiation beam is transmitted from the laser to the source collector module by means of a beam delivery system including, for example, suitable steering mirrors and / or beam expanders. In other cases, for example, when the radiation source is a discharge-produced plasma EUV generator (commonly referred to as a DPP radiation source), the radiation source can be an integral part of the source collector module.
[0112] The illuminator IL may include an adjuster for adjusting the angular intensity distribution of the radiation beam. Generally, at least the outer and / or inner radial ranges of the intensity distribution in the pupil plane of the illuminator can be adjusted (commonly referred to as σ - outer and σ - inner, respectively). In addition, the illuminator IL may include various other components, such as faceted field mirror devices and faceted pupil mirror devices. The illuminator can be used to adjust the radiation beam to have a desired uniformity and intensity distribution in its cross-section.
[0113] A radiation beam B is incident on a patterning device (e.g., a mask) MA held on a support structure (e.g., a mask table) MT and is patterned by the patterning device. After reflection from the patterning device (e.g., mask) MA, the radiation beam B passes through a projection system PS that focuses the beam onto a target portion C of a substrate W. By means of a second locator PW and a position sensor PS2 (e.g., an interferometric device, a linear encoder, or a capacitive sensor), the substrate table WT can be accurately moved, for example, so as to position different target portions C in the path of the radiation beam B. Similarly, a first locator PM and another position sensor PS1 can be used to accurately position the patterning device (e.g., mask) MA relative to the path of the radiation beam B. Patterning device alignment marks M1, M2 and substrate alignment marks P1, P2 can be used to align the patterning device (e.g., mask) MA and the substrate W.
[0114] The depicted apparatus 1000 can be used in at least one of the following modes:
[0115] 1. In a step mode, the support structure (e.g., mask table) MT and the substrate table WT remain substantially stationary while the entire pattern imparted to the radiation beam is projected onto the target portion C in one go (i.e., single static exposure). Subsequently, the substrate table WT is displaced in the X and / or Y direction so that different target portions C can be exposed.
[0116] 2. In a scan mode, the support structure (e.g., mask table) MT and the substrate table WT are scanned synchronously while the pattern imparted to the radiation beam is projected onto the target portion C (i.e., single dynamic exposure). The speed and direction of the substrate table WT relative to the support structure (e.g., mask table) MT can be determined by the (reduction) magnification and image reversal characteristics of the projection system PS.
[0117] 3. In another mode, the support structure (e.g., mask table) MT remains substantially stationary to hold a programmable patterning device, and the substrate table WT is moved or scanned while the pattern imparted to the radiation beam is projected onto the target portion C. In this mode, a pulsed radiation source is typically employed, and the programmable patterning device is updated as required after each movement of the substrate table WT or between successive radiation pulses during the scan. This mode of operation can be readily applied to maskless lithography using a programmable patterning device such as a programmable mirror array of the type mentioned above.
[0118] Figure 11Device 1000 is shown in more detail, the device including a source collector module SO, an illumination system IL, and a projection system PS. The source collector module SO is constructed and arranged such that a vacuum environment can be maintained in the enclosure structure 220 of the source collector module SO. A plasma radiation source can be formed by a discharge to generate a plasma 210 for emitting EUV radiation. EUV radiation can be generated by a gas or vapor (e.g., Xe gas, Li vapor, or Sn vapor), where a very hot plasma 210 is generated to emit radiation in the EUV range of the electromagnetic spectrum. For example, the very hot plasma 210 is generated by a discharge that produces at least partially ionized plasma. To efficiently generate radiation, a partial pressure of, for example, 10 Pa of Xe, Li, Sn vapor, or any other suitable gas or vapor may be required. In an embodiment, an excited tin (Sn) plasma is provided to generate EUV radiation.
[0119] The radiation emitted by the hot plasma 210 is transferred from the source chamber 211 to the collector chamber 212 via an optionally present gas barrier or contaminant trap 230 (also referred to in some cases as a contaminant barrier or foil trap) in or behind an opening located in the source chamber 211. The contaminant trap 230 may include a channel structure. The contaminant trap 230 may also include a gas barrier or a combination of a gas barrier and a channel structure. As is known in the art, the contaminant trap or contaminant barrier 230 as further indicated herein includes at least a channel structure.
[0120] The collector chamber 211 may include a radiation collector CO that may be a so-called grazing incidence collector. The radiation collector CO has an upstream radiation collector side 251 and a downstream radiation collector side 252. Radiation traversing the collector CO may be reflected from the grating spectral filter 240 to be focused at a virtual source point IF along the optical axis indicated by the dashed line "O". The virtual source point IF is generally referred to as an intermediate focus, and the source collector module is arranged such that the intermediate focus IF is located at or near the opening 221 in the enclosure structure 220. The virtual source point IF is an image of the plasma 210 for emitting radiation.
[0121] Subsequently, the radiation traverses the illumination system IL, which may include a faceted field mirror device 22 and a faceted pupil mirror device 24. The faceted field mirror device 22 and the faceted pupil mirror device 24 are arranged to provide a desired angular distribution of the radiation beam 21 at the patterning device MA, and to provide a desired uniformity of the radiation intensity at the patterning device MA. After the radiation beam 21 is reflected at the patterning device MA held by the support structure MT, a patterned beam 26 is formed, and the patterned beam 26 is imaged onto a substrate W held by a substrate stage WT by the projection system PS via reflection elements 28, 30.
[0122] In the illumination optical device unit IL and the projection system PS, there may generally be more elements than those shown. Depending on the type of lithographic apparatus, the grating spectral filter 240 may optionally be present. Additionally, there may be more mirrors than those shown in the figure. For example, in the projection system PS, there may be 1 to 6 additional reflective elements other than those shown in Figure 11 the figure.
[0123] As Figure 11 illustrated in the figure, the collector optical device CO is depicted as a nested collector having grazing-incidence reflectors 253, 254, and 255, merely as an example of a collector (or collector mirror). The grazing-incidence reflectors 253, 254, and 255 are arranged symmetrically about the optical axis O, and this type of collector optical device CO is desirably used in combination with a discharge-produced plasma radiation source.
[0124] Alternatively, the source collector module SO may be part of an LPP radiation system as shown in Figure 12 the figure. The laser LAS is arranged to deposit laser energy into a fuel such as xenon (Xe), tin (Sn), or lithium (Li) to generate a highly ionized plasma 210 having an electron temperature of several tens of eV. The energy radiation generated during the de-excitation and recombination of these ions is emitted from the plasma, collected by the near-normal-incidence collector optical device CO, and focused onto the opening 221 in the enclosure structure 220.
[0125] The concepts disclosed herein can be simulated or mathematically modeled for any conventional imaging system for imaging sub-wavelength features and may be particularly useful for emerging imaging technologies capable of generating wavelengths of increasingly smaller sizes. Emerging technologies already in use include extreme ultraviolet (EUV) lithography capable of using an ArF laser to generate a 193 nm wavelength and even capable of using a fluorine laser to generate a 157 nm wavelength. Additionally, EUV lithography can generate wavelengths in the range of 20 nm to 5 nm by using a synchrotron or by using high-energy electrons to impact a material (solid or plasma) to generate photons in this range.
[0126] Although the concepts disclosed herein can be used for imaging on a substrate such as a silicon wafer, it should be understood that the disclosed concepts can be used with any type of lithographic imaging system, e.g., a lithographic imaging system for imaging on a substrate other than a silicon wafer.
[0127] Although specific reference may be made in this text to the use of embodiments in the manufacture of ICs, it should be understood that the embodiments herein can have many other possible applications. For example, it can be used in the manufacture of integrated optical systems, for guiding and detecting patterns in magnetic domain memories, liquid crystal displays (LCDs), thin film magnetic heads, microelectromechanical systems (MEMs), etc. Those skilled in the art will appreciate that in the context of such alternative applications, any use herein of the terms "reticle", "wafer" or "die" can be considered synonymous with or interchangeable with the more general terms "patterning device", "substrate" or "target portion", respectively. The substrates mentioned herein can be processed in, for example, a track or a coat development system (a tool typically used to apply a resist layer to a substrate and develop the exposed resist) or a metrology or inspection tool, either before or after exposure. Where applicable, the disclosures herein can be applied to such and other substrate processing tools. Additionally, the substrate can be processed more than once, for example to produce, for example, a multi-layer IC, such that the term substrate as used herein can also refer to a substrate that already contains multiple processed layers.
[0128] In this text, as used herein, the terms "radiation" and "beam" encompass all types of electromagnetic radiation, including ultraviolet radiation (e.g., having a wavelength of about 365 nm, about 248 nm, about 193 nm, about 157 nm or about 126 nm) and extreme ultraviolet (EUV) radiation (e.g., having a wavelength in the range of 5 nm to 20 nm), as well as particle beams, such as ion beams or electron beams.
[0129] As used herein, the terms "optimizing" and "optimization" refer to or mean adjusting a patterning device (e.g., a lithographic device), a patterning process, etc., such that the result and / or the process has more desirable characteristics, such as higher accuracy of the projection of a design pattern on a substrate, a larger process window, etc. Thus, as used herein, the terms "optimizing" and "optimization" refer to or mean the process of identifying one or more values of one or more parameters that provide an improvement in at least one relevant metric, such as a local optimum, compared to an initial set of one or more values for those one or more parameters. The terms "optimal" and other related terms should be interpreted accordingly. In an embodiment, the optimization step can be applied iteratively to provide a further improvement in one or more metrics.
[0130] Aspects of the present invention may be implemented in any convenient form. For example, embodiments may be implemented by one or more suitable computer programs which may be carried on a suitable carrier medium which may be a tangible carrier medium (such as a disk) or an intangible carrier medium (such as a communication signal). Embodiments of the present invention may be implemented using suitable apparatus which may specifically take the form of a programmable computer running a computer program arranged to implement the methods as described herein. Thus, embodiments of the present disclosure may be implemented in hardware, firmware, software, or any combination thereof. Embodiments of the present disclosure may also be implemented as instructions stored on a machine-readable medium which may be read and executed by one or more processors. Machine-readable media may include any mechanism for storing or transmitting information in a form readable by a machine (such as a computing device). For example, machine-readable media may include: read-only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustic, or other forms of propagated signals (such as carrier waves, infrared signals, digital signals, etc.), and so on. Additionally, firmware, software, routines, instructions may be described herein as performing certain actions. However, it should be understood that such descriptions are for convenience only and such actions are actually caused by a computing device, processor, controller, or other device executing the firmware, software, routines, instructions, etc.
[0131] In block diagrams, the components illustrated are depicted as discrete functional blocks, but embodiments are not limited to systems that organize the functionality described herein as illustrated. The functionality provided by each of these components may be provided by software or hardware modules that are organized in a different manner than currently depicted, such as doped, combined, replicated, decomposed, distributed (e.g., within a data center or geographically), or otherwise organized differently such software or hardware. The functionality described herein may be provided by one or more processors of one or more computers executing program code stored on a tangible, non-transitory machine-readable medium. In some cases, a third-party content delivery network may host some or all of the information communicated via a network, in which case, to the extent that information (such as content) is purportedly supplied or otherwise provided, the information may be provided by sending instructions to obtain the information from the content delivery network.
[0132] Unless otherwise specifically stated, it should be understood that throughout the present specification, discussions using terms such as "processing", "operating", "computing", "determining", etc. refer to actions or processes of a specific device such as a special-purpose computer or similar special-purpose electronic processing / computing device.
[0133] Embodiments of the present disclosure may be further described in the following aspects.
[0134] 1. A method for training a machine learning model configured to predict a value of a physical property associated with a substrate for adjusting a patterning process, the method comprising:
[0135] Obtaining a reference image associated with a desired pattern to be printed on the substrate;
[0136] Determining a first set of model parameter values of the machine learning model such that a first cost function is reduced from an initial value of the cost function obtained using an initial set of model parameter values, wherein the first cost function is the difference between the reference image and an image generated via the machine learning model; and
[0137] Training the machine learning model using the first set of model parameter values such that a combination of the first cost function and a second cost function is iteratively reduced,
[0138] wherein the second cost function is the difference between a measured value and a predicted value of a physical property associated with the desired pattern, the predicted value being predicted via the machine learning model.
[0139] 2. The method according to aspect 1, wherein obtaining the reference image comprises:
[0140] Executing a process model configured to produce the reference image as an output, wherein the process model models a part of the patterning process.
[0141] 3. The method according to aspect 2, wherein the process model is a calibrated model of an optical device model, a resist model, and / or an etch model of the patterning process.
[0142] 4. The method according to any one of aspects 1 to 3, wherein the reference image is a spatial image, a resist image, and / or an etch image of the desired pattern.
[0143] 5. The method according to any one of aspects 1 to 4, wherein determining the first set of model parameter values of the machine learning model is an iterative process, the iteration comprising:
[0144] Generating the image by executing the machine learning model using the desired pattern;
[0145] Determining the difference between the generated image and the reference image; and
[0146] Adjusting the model parameter values of the machine learning model such that the difference is reduced.
[0147] 6. The method according to any one of aspects 1 to 5, wherein the difference between the generated image and the reference image is minimized.
[0148] 7. The method according to any one of aspects 1 to 6, wherein training the machine learning model is an iterative process, the iteration comprising:
[0149] Initializing the model parameters of the machine learning model with a first set of model parameter values;
[0150] Predicting a value of a physical property associated with the substrate by executing the machine learning model using the desired pattern;
[0151] Obtaining a measured value of the physical property of the desired printed pattern on the substrate via a metrology tool; and
[0152] Adjusting the model parameter values of the machine learning model such that a combination of the first cost function and the second cost function is reduced.
[0153] 8. The method according to aspect 7, wherein adjusting the model parameter values is gradient descent based on a combination of the first cost function and the second cost function.
[0154] 9. The method according to any one of aspects 1 to 8, wherein the sum of the first cost function and the second cost function is minimized.
[0155] 10. The method according to any one of aspects 1 to 9, wherein the model parameters are weights and / or biases associated with one or more layers of the machine learning model.
[0156] 11. The method according to any one of aspects 1 to 10, wherein the machine learning model is a convolutional neural network.
[0157] 12. The method according to any one of aspects 1 to 11, wherein the parameter associated with the substrate is a critical dimension or an edge placement error associated with the desired pattern.
[0158] 13. The method according to any one of aspects 10 to 12, wherein the weights of the convolutional neural network are adjusted to reduce the edge placement error or the model error associated with the model of the patterning process being trained.
[0159] 14. The method according to any one of aspects 1 to 13, wherein the measured value is a CD value obtained via the metrology tool configured to measure the desired printed pattern of the substrate.
[0160] 15. The method according to any one of aspects 7 to 14, wherein the metrology tool is a scanning electron microscope (SEM) and the measured value is obtained from an SEM image.
[0161] 16. The method according to any one of aspects 1 to 15, wherein the measured value is an intensity value of a spatial image associated with the desired pattern.
[0162] 17. The method according to any one of aspects 1 to 11, further comprising:
[0163] training the machine learning model using a first set of the model parameter values such that a combination of the first cost function, the second cost function, and the third cost function is reduced,
[0164] wherein the third cost function is a function dependent on a grid.
[0165] 18. The method according to any one of aspects 1 to 17, further comprising:
[0166] predicting a substrate image for designing a layout via the trained machine learning model;
[0167] determining a mask layout to be used for manufacturing a mask for a patterning process via an OPC simulation using the design layout and the predicted substrate image.
[0168] 19. The method according to aspect 17, wherein the OPC simulation includes:
[0169] simulating a patterning process model via using geometries of the design layout and corrections associated with a plurality of segments to determine a simulated pattern to be printed on a substrate; and
[0170] determining an optical proximity effect correction to the design layout such that a difference between the simulated pattern and the design layout is reduced.
[0171] 20. The method according to aspect 19, wherein the determining the optical proximity effect correction is an iterative process, and the iteration includes:
[0172] adjusting shapes and / or sizes of geometries of main features and / or one or more auxiliary features of the design layout such that a performance metric of the patterning process is reduced.
[0173] 21. The method according to aspect 20, wherein the one or more auxiliary features are obtained from the predicted post-OPC image of the machine learning model.
[0174] 22. The method according to any one of aspects 1 to 21, wherein a combination of the first cost function (CF1) and the second cost function (CF2) is calculated using the expression c1*CF1 + c2*CF2, where c1 and c2 are coefficients that can be adjusted to minimize the combination.
[0175] 23. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions, when executed by a computer, implementing the method of any one of the above aspects.
[0176] It should be understood that the specification and drawings are not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the invention as defined by the appended claims.
[0177] In view of this specification, those skilled in the art will appreciate modifications and alternative embodiments of various aspects of the invention. Accordingly, the specification and drawings should be regarded as merely illustrative and for the purpose of teaching those skilled in the art the general manner of carrying out the invention. It should be understood that the forms of the invention shown and described herein are to be regarded as examples of embodiments. The elements and materials may be substituted for those illustrated and described herein, the parts and processes may be reversed or omitted, certain features may be utilized independently, and the features of the embodiments or the embodiments may be combined, all of which will be apparent to those skilled in the art after having the benefit of this specification. Changes may be made to the elements described herein without departing from the spirit and scope of the invention as described in the appended claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the specification.
[0178] As used throughout this application, the word "may" is used in a permissive sense (i.e., meaning possible) rather than in a mandatory sense (i.e., meaning must). The words "include", "including", "includes", etc. mean including but not limited to. As used throughout this application, the singular forms "a", "an", and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, a reference to "an" element or "a" element includes a combination of two or more elements, although other terms and phrases, such as "one or more", may also be used with respect to one or more elements. Unless otherwise indicated, the term "or" is non-exclusive, i.e., it encompasses both "and" and "or". Terms that describe a conditional relationship, such as "in response to X, Y", "in the case of X, i.e., Y", "if X, then Y", "when X, Y", etc. encompass causal relationships where the antecedent is a necessary causal condition, the antecedent is a sufficient causal condition, or the antecedent is a contributing causal condition to the result, e.g., "in the case where condition Y is obtained, i.e., state X appears" is superordinate to "X appears only in the case of Y" and "X appears in the cases of Y and Z". Such conditional relationships are not limited to results obtained immediately following the antecedent, as some results may be delayed, and in a conditional statement, the antecedent is connected to its result, e.g., the antecedent is related to the likelihood of the result occurring. Unless otherwise indicated, a statement that multiple attributes or functions are mapped to multiple objects (e.g., one or more processors performing steps A, B, C, and D) encompasses both all such attributes or functions mapped to all such objects and subsets of such attributes or functions mapped to subsets of the attributes or functions (e.g., all processors each performing steps A through D, and the case where processor 1 performs step A, processor 2 performs part of steps B and C, and processor 3 performs part of steps C and D). Additionally, unless otherwise indicated, a statement that a value or action is "based on" another condition or value encompasses both the case where the condition or value is the sole factor and the case where the condition or value is one of multiple factors. Unless otherwise indicated, a statement that "each" instance in a collection has some property should not be construed to exclude the case where some otherwise identical or similar members in a larger collection do not have the property (i.e., each does not necessarily mean each and all). A reference to a selection from a range includes the endpoints of that range.
[0179] In the foregoing description, any process, description, or block in a flowchart should be understood to represent a module, segment, or portion of program code that includes one or more executable instructions for implementing a specific logical function or step in the process, and alternative implementations are included within the scope of exemplary embodiments of the present advancement, where functions may be executed not in the order shown or discussed, including substantially concurrently or in the reverse order, as would be understood by those skilled in the art.
[0180] To the extent that certain U.S. patents, U.S. patent applications, or other materials (such as papers) have been incorporated by reference, the text of such U.S. patents, U.S. patent applications, and other materials is incorporated by reference only to the extent that there is no conflict between such materials and the statements and drawings set forth herein. In the event of such a conflict, any such conflicting text in such U.S. patents, U.S. patent applications, and other materials incorporated by reference is specifically not incorporated by reference herein.
[0181] Although certain embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the disclosure. In fact, the novel methods, devices, and systems described herein may be embodied in many other forms; furthermore, various omissions, substitutions, and changes in the form of the methods, devices, and systems described herein may be made without departing from the spirit of the disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.
Claims
1. A method of training a machine learning model, the machine learning model being configured to predict a value of a physical property associated with a substrate for adjusting a patterning process, the method comprises: configuring model parameters of the machine learning model with a first set of model parameter values; predicting a value of a physical property associated with the substrate by executing the machine learning model using a desired pattern; obtaining a measured value of the physical property of the desired printed pattern on the substrate; and adjusting the model parameter values of the machine learning model such that a combination of a first cost function and a second cost function is reduced, wherein the first cost function represents a difference between a reference image and an image generated via the machine learning model, and the second cost function represents a difference between the measured value and the predicted value of the physical property associated with the desired pattern.
2. The method according to claim 1, wherein the reference image is obtained in the following manner: executing a process model to model a part of the patterning process and being configured to generate the reference image as an output, wherein the process model models a part of the patterning process.
3. The method according to claim 2, wherein the process model is a calibrated model of an optical device model, a resist model, and / or an etch model of the patterning process.
4. The method according to claim 1, wherein the reference image is a spatial image, a resist image, and / or an etch image of the desired pattern.
5. The method according to claim 1, wherein the adjusting of the model parameter values is based on gradient descent of the combination of the first cost function and the second cost function.
6. The method according to claim 1, wherein the machine learning model is a convolutional neural network, and wherein the model parameters are weights and / or biases associated with one or more layers of the convolutional neural network.
7. The method according to claim 6, wherein the weights of the convolutional neural network are adjusted to reduce an edge localization error or a model error associated with the model of the patterning process being trained.
8. The method according to claim 1, wherein the parameter associated with the substrate is a critical dimension or an edge localization error associated with the desired pattern, and wherein the measured value is a CD value obtained via a metrology tool.
9. The method according to claim 1, wherein the measured value is an intensity value of a spatial image associated with the desired pattern.
10. The method according to claim 1, further comprises: training the machine learning model by using the first set of model parameter values such that a combination of the first cost function, the second cost function, and a third cost function is reduced, wherein the third cost function is a function dependent on a grid.
11. The method according to claim 1, further comprises: predicting a substrate image for designing a layout via the trained machine learning model; and determining a mask layout to be used for manufacturing a mask for the patterning process via an OPC simulation using the design layout and the predicted substrate image.
12. The method according to claim 11, wherein the OPC simulation comprises: determining a simulated pattern to be printed on a substrate; and determining an optical proximity effect correction to the design layout such that the difference between the simulated pattern and the design layout is reduced.
13. The method according to claim 11, wherein the determining comprises obtaining one or more auxiliary features from the predicted post-OPC image of the machine learning model.
14. The method according to claim 1, wherein the combination of the first cost function (CF1) and the second cost function (CF2) is calculated using the expression where c1 and c2 are coefficients.
15. The method according to claim 1, wherein the first set of model parameter values is predetermined by training the model based on a predetermined cost function.
16. A method of training a machine learning model configured to predict values of physical properties associated with a substrate for adjusting a patterning process, the method comprises: configuring model parameters of the machine learning model with a first set of model parameter values; predicting values of physical properties associated with the substrate by performing the machine learning model with a desired pattern; obtaining measured values of physical properties of the desired printed pattern on the substrate; and adjusting the model parameter values of the machine learning model such that a combination of a first cost function, a second cost function, and a third cost function is reduced, wherein the first cost function represents the difference between a reference image and an image generated via the machine learning model, the second cost function represents the difference between the measured value and the predicted value of the physical property associated with the desired pattern, and wherein the third cost function is a grid-dependent function.
17. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions, when executed by a computer, implement the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Methods and system for lithography process window simulation
US20090157360A1
System and method for mask verification using an individual mask error model
US7587704B2
Cited By
Height compensation value calculation method and device for scanning electron microscope
CN120431166A