Methods for reducing uncertainty in predictions from machine learning models
By using the encoder-decoder architecture to generate multiple posterior distributions and sample to determine variability, quantify and adjust machine learning model parameters, the uncertainty of model prediction during lithography is solved, and the accuracy and consistency of mask layout is improved.
Patent Information
- Application Number
- CN201980078859.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-26
- Filing Date
- 2019-11-19
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2039-11-19
AI Technical Summary
The prediction uncertainty of existing machine learning models is difficult to determine during lithography, resulting in uncertainty in mask layout and affecting the final functionality of the wafer.
Using an encoder-decoder architecture, variability is determined by generating multiple posterior distributions and sampling, model uncertainty is quantified, and model parameters are adjusted to reduce uncertainty.
By quantizing and adjusting model parameters, the certainty of the machine learning model is improved, uncertainty in the lithography process is reduced, and the accuracy and consistency of mask layout are ensured.
Smart Images

Figure CN113168556B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from European application No. 18209496.1 filed on November 30, 2018, and from European application No. 19182658.5 filed on June 26, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to mask manufacturing and patterning processes. More specifically, the present disclosure relates to an apparatus and method for determining and / or reducing uncertainty in parameterized (e.g., machine learning) model predictions. Background Art
[0004] Lithographic projection apparatuses can be used, for example, in the manufacture of integrated circuits (ICs). In such cases, a patterning device (e.g., a mask) can contain or provide a pattern corresponding to an individual layer of the IC (a "design layout"), and this pattern can be transferred to a target portion (e.g., including one or more dies) on a substrate (e.g., a silicon wafer) that has been coated with a layer of radiation-sensitive material ("resist"), such as by irradiating the target portion through the pattern on the patterning device. Typically, a single substrate contains multiple adjacent target portions, to which the pattern is transferred sequentially by the lithographic projection apparatus, one at a time. In one type of lithographic projection apparatus, the pattern on the entire patterning device is transferred to one target portion in a single operation. Such an apparatus is often referred to as a stepper. In an alternative apparatus, often referred to as a stepper-scan apparatus, a projection beam scans the patterning device along a given reference direction (the "scanning" direction) while synchronously moving the substrate parallel or antiparallel to the reference direction. Different portions of the pattern on the patterning device are gradually transferred to one target portion. Since typically the lithographic projection apparatus will have a reduction factor M (e.g. 4), the speed F at which the substrate is moved will be 1 / M times the speed at which the projection beam scans the patterning device. Further information on lithographic apparatus as described herein can be gleaned, for example, from US 6,046,792, which is incorporated herein by reference.
[0005] Before the pattern is transferred from the patterned device to the substrate, the substrate can undergo various processes such as primer coating, resist coating and soft baking. After exposure, the substrate can undergo other processes ("post-exposure processes"), such as post-exposure baking (PEB), development, hard baking and measurement / inspection of the transferred pattern. This series of processes is used as the basis for manufacturing the individual layers of the device (e.g., IC). The substrate can then undergo various processes all intended to complete the individual layers of the device, such as etching, ion implantation (doping), metallization, oxidation, chemical mechanical polishing, etc. If several layers are needed in the device, the entire process or a variation thereof is repeated for each layer. Ultimately, the device will be present in each target portion on the substrate. These devices are then separated from each other by techniques such as cutting or sawing, so that the individual devices can be mounted on a carrier, connected to pins, etc.
[0006] Thus, manufacturing devices such as semiconductor devices typically involves processing a substrate (e.g., a semiconductor wafer) using a number of manufacturing processes to form the various features and multiple layers of the device. Such layers and features are typically manufactured and processed using, for example, deposition, lithography, etching, chemical mechanical polishing, and ion implantation. Multiple devices can be manufactured on multiple dies on a substrate and then separated into individual devices. This device manufacturing process can be considered a patterning process. The patterning process involves a patterning step, such as optical and / or nanoimprint lithography using a patterning device in a lithography apparatus to transfer the pattern on the patterned device to the substrate, and typically but optionally involves one or more associated pattern processing steps, such as resist development by a developing apparatus, baking of the substrate using a baking tool, etching using the pattern using an etching apparatus, and the like. One or more metrology processes are typically involved in the patterning process.
[0007] As noted, photolithography is a core step in the manufacture of devices such as ICs, where patterns formed on a substrate define the functional elements of the device, such as microprocessors, memory chips, etc. Similar photolithographic techniques are also used to form flat panel displays, microelectromechanical systems (MEMS), and other devices.
[0008] As semiconductor manufacturing processes continue to advance, the dimensions of functional elements continue to decrease, while the number of functional elements (such as transistors) per device has steadily increased over the past few decades, following a trend commonly referred to as "Moore's Law." Under current technology, the layers of a device are manufactured using a photolithography projection apparatus that projects a design layout onto a substrate using illumination from a deep ultraviolet radiation source, thereby creating individual functional elements with dimensions well below 100 nm, i.e., less than half the wavelength of the radiation from the radiation source (e.g., a 193 nm radiation source).
[0009] The process in which features having dimensions smaller than the classical resolution limit of a lithographic projection apparatus are printed is generally referred to as low-k1 lithography, according to the resolution formula CD=k1×λ / NA, where λ is the wavelength of the radiation employed (currently 248 nm or 193 nm in most cases), NA is the numerical aperture of the projection optics in the lithographic projection apparatus, CD is the "critical dimension"—typically the minimum feature size to be printed—and k1 is an empirical resolution factor. Generally, the smaller k1 is, the more difficult it is to reproduce on a substrate a pattern similar in shape and size to that planned by the designer to achieve specific electrical functionality and performance. To overcome these difficulties, complex fine-tuning steps are applied to the lithographic projection apparatus, the design layout, or the patterning device. These include, for example, but not limited to, optimization of the NA and optical coherence settings, customized illumination schemes, the use of phase-shifted patterning devices, optical proximity correction (OPC, sometimes also referred to as "optical and process correction") in the design layout, or other methods generally defined as "resolution enhancement techniques" (RET). As used herein, the term "projection optics" should be broadly interpreted to encompass various types of optical systems, including, for example, refractive optics, reflective optics, apertures, and catadioptric optics. The term "projection optics" may also include components that operate collectively or individually according to any of these design types for directing, shaping, or controlling a projection beam of radiation. The term "projection optics" may include any optical component in a lithographic projection apparatus, regardless of where the optical component is located in the optical path of the lithographic projection apparatus. Projection optics may include optical components for shaping, adjusting, and / or projecting radiation from a source before the radiation passes through a patterning device, and / or optical components for shaping, adjusting, and / or projecting radiation after the radiation passes through a patterning device. Projection optics typically do not include a light source and a patterning device. Summary of the Invention
[0010] According to an embodiment, a method for adjusting a lithographic apparatus is provided. The method includes causing a machine learning model to predict multiple posterior distributions for a given input from the machine learning model. The multiple posterior distributions include distributions from a plurality of distributions. The method includes determining the variability of the predicted multiple posterior distributions for a given input by sampling from distributions from a plurality of distributions. The method includes quantifying the uncertainty in the predictions of the machine learning model using the determined variability in the predicted multiple posterior distributions. The method includes adjusting one or more parameters of the machine learning model to reduce the uncertainty in the predictions of the machine learning model. The method includes determining one or more lithographic process parameters based on the predictions from the adjusted machine learning model based on the given input; and adjusting the lithographic apparatus based on the determined one or more lithographic process parameters.
[0011] In one embodiment, the one or more parameters of the machine learning model include one or more weights of the one or more parameters of the machine learning model.
[0012] In one embodiment, the predictions from the tuned machine learning model include one or more of predicted overlay or predicted wafer geometry.
[0013] In one embodiment, the one or more lithography process parameters determined include one or more of mask design, pupil shape, dose, or focus.
[0014] In one embodiment, the determined one or more lithographic process parameters include a mask design, and adjusting the lithographic apparatus based on the mask design includes changing the mask design from a first mask design to a second mask design.
[0015] In an embodiment, the determined one or more lithographic process parameters include a pupil shape, and adjusting the lithographic apparatus based on the pupil shape includes changing the pupil shape from a first pupil shape to a second pupil shape.
[0016] In one embodiment, the determined one or more lithographic process parameters include a dose, and adjusting the lithographic apparatus based on the dose includes changing the dose from a first dose to a second dose.
[0017] In one embodiment, the determined one or more lithographic process parameters include focus, and adjusting the lithographic apparatus based on the focus includes changing the focus from a first focus to a second focus.
[0018] In one embodiment, causing the machine learning model to predict multiple posterior distributions includes causing the machine learning model to generate a distribution among the multiple distributions using parameter dropout.
[0019] In one embodiment, causing the machine learning model to predict multiple posterior distributions for a given input from the machine learning model includes causing the machine learning model to predict a first posterior distribution P Θ The first set of multiple posterior distributions corresponding to (z|x) and the second posterior distribution P φ(y|z) corresponding to a second set of multiple posterior distributions; determining the variability of the predicted multiple posterior distributions for a given input by sampling from the multiple distributions includes: determining the variability of the predicted first set of multiple posterior distributions and the second set of multiple posterior distributions for a given input by sampling from the multiple distributions for the predicted first set of multiple posterior distributions and the predicted second set of multiple posterior distributions; and using the determined variability in the predicted multiple posterior distributions to quantify the uncertainty in the prediction of the machine learning model includes: using the determined variability in the predicted first set of multiple posterior distributions and the predicted second set of multiple posterior distributions to quantify the uncertainty in the prediction of the machine learning model.
[0020] In one embodiment, the given input comprises one or more of: an image, a clip, an encoded image, an encoded clip, or data from a previous layer of a parameterized model.
[0021] In one embodiment, the method further includes: using the determined variability and / or quantified uncertainty in the predicted multiple posterior distributions to adjust the machine learning model to reduce the uncertainty of the machine learning model by making the machine learning model more descriptive or including more diverse training data.
[0022] In one embodiment, sampling includes randomly selecting a distribution from among a plurality of distributions, wherein the sampling is Gaussian or non-Gaussian.
[0023] In one embodiment, determining the variability includes quantifying the variability using one or more statistical operations including one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis, or covariance.
[0024] In one embodiment, the uncertainty of a machine learning model is related to the uncertainty of the weights of one or more parameters of the machine learning model and the size and descriptiveness of a latent space associated with the machine learning model.
[0025] In one embodiment, adjusting the machine learning model to reduce uncertainty of the machine learning model includes increasing the training set size and / or adding dimensions to a latent space associated with the machine learning model.
[0026] In one embodiment, increasing the training set size and / or adding dimensions to the latent space includes: training the machine learning model using more diverse images, more diverse data, and additional clips as input relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the machine learning model.
[0027] In one embodiment, adjusting a machine learning model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the machine learning model includes adding additional dimensions to a latent space associated with the machine learning model.
[0028] In one embodiment, using the determined variability in the predicted multiple posterior distributions to adjust one or more parameters of the machine learning model to reduce the uncertainty of the machine learning model includes: training the machine learning model with additional and more diverse training samples.
[0029] According to another embodiment, a method for quantifying uncertainty in predictions from a parameterized model is provided. The method includes causing the parameterized model to predict multiple posterior distributions for a given input from the parameterized model. The multiple posterior distributions include distributions from a plurality of distributions. The method includes: determining variability in the predicted multiple posterior distributions for the given input by sampling from the distributions from the plurality of distributions; and quantifying uncertainty in the predictions from the parameterized model using the determined variability in the predicted multiple posterior distributions.
[0030] In one embodiment, the parameterized model is a machine learning model.
[0031] In one embodiment, causing the parameterized model to predict the plurality of posterior distributions includes causing the parameterized model to generate a distribution among the plurality of distributions using parameter dropout.
[0032] In one embodiment, causing the parameterized model to predict multiple posterior distributions for a given input from the parameterized model includes: causing the parameterized model prediction to be compared with a first posterior distribution P Θ The first set of multiple posterior distributions corresponding to (z|x) and the second posterior distribution P φ (y|z) corresponding to a second set of multiple posterior distributions; determining the variability of the predicted multiple posterior distributions for a given input by sampling from the multiple distributions includes: determining the variability of the predicted first set of multiple posterior distributions and the second set of multiple posterior distributions for a given input by sampling from the multiple distributions for the predicted first set of multiple posterior distributions and the predicted second set of multiple posterior distributions; and using the determined variability in the predicted multiple posterior distributions to quantify the uncertainty in the parameterized model prediction includes: using the determined variability in the predicted first set of multiple posterior distributions and the predicted second set of multiple posterior distributions to quantify the uncertainty in the parameterized model prediction.
[0033] In one embodiment, the given input comprises one or more of: an image, a clip, an encoded image, an encoded clip, or data from a previous layer of a parameterized model.
[0034] In one embodiment, the method further comprises adjusting the parameterized model using the determined variability and / or the quantified uncertainty in the predicted plurality of posterior distributions to reduce the uncertainty of the parameterized model by making the parameterized model more descriptive or including more diverse training data.
[0035] In one embodiment, the parameterized model comprises an encoder-decoder architecture.
[0036] In one embodiment, the encoder-decoder architecture comprises a variational encoder-decoder architecture, and the method further comprises training the variational encoder-decoder architecture using the probabilistic latent space, the variational encoder-decoder architecture generating realizations in the output space.
[0037] In one embodiment, the latent space comprises a low-dimensional encoding.
[0038] In one embodiment, the method further comprises determining, for a given input, a conditional probability of the latent variable using an encoder portion of an encoder-decoder architecture.
[0039] In one embodiment, the method further comprises determining the conditional probabilities using a decoder portion of an encoder-decoder architecture.
[0040] In one embodiment, the method further includes sampling from the conditional probabilities of the latent variables determined using an encoder portion of the encoder-decoder architecture, and for each sample, predicting an output using a decoder portion of the encoder-decoder architecture.
[0041] In one embodiment, sampling includes randomly selecting a distribution from among a plurality of distributions, wherein the sampling is Gaussian or non-Gaussian.
[0042] In one embodiment, determining the variability includes quantifying the variability using one or more statistical operations including one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis, or covariance.
[0043] In one embodiment, the uncertainty of the parameterized model is related to the uncertainty of the weights of the parameters of the parameterized model and the size and descriptiveness of the latent space.
[0044] In one embodiment, the uncertainty of the parameterized model is related to the uncertainty of the weights of the parameters of the parameterized model and the size and descriptiveness of the latent space, so that the uncertainty of the weights manifests as uncertainty in the output, resulting in increased output variance.
[0045] In one embodiment, adjusting the parameterized model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the parameterized model includes increasing the training set size and / or adding dimensions to the latent space.
[0046] In one embodiment, increasing the training set size and / or adding dimensions to the latent space includes: using more diverse images, more diverse data, and additional clips as input to train the parameterized model relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the parameterized model.
[0047] In one embodiment, adjusting the parameterized model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the parameterized model includes adding additional dimensions to the latent space.
[0048] In one embodiment, adjusting the parameterized model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the parameterized model includes training the parameterized model with additional and more diverse training samples.
[0049] In one embodiment, the additional and more diverse training samples include more diverse images, more diverse data, and additional clips relative to previous training material.
[0050] In one embodiment, the method further includes adjusting a parameterized model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the parameterized model for predicting wafer geometry as part of a semiconductor manufacturing process.
[0051] In one embodiment, using the determined variability in the predicted multiple posterior distributions to adjust a parameterized model to reduce the uncertainty of the parameterized model for predicting wafer geometry as part of a semiconductor manufacturing process includes: using more diverse images, more diverse data, and additional clippings as inputs to train the parameterized model relative to previous training material; and using more dimensions for encoding vectors and more encoding layers in the parameterized model, the more diverse images, more diverse data, additional clippings, more dimensions, and more encoding layers being determined based on the determined variability.
[0052] In one embodiment, the method further includes adjusting a parameterized model using the determined variability in the predicted plurality of posterior distributions to reduce uncertainty in the parameterized model for generating the predicted overlay as part of a semiconductor manufacturing process.
[0053] In one embodiment, using the determined variability in the predicted multiple posterior distributions to adjust a parameterized model to reduce the uncertainty of the parameterized model for generating the predicted overlap as part of a semiconductor manufacturing process includes: using more diverse images, more diverse data, and additional clips as inputs to train the parameterized model; and using more sizes for the encoding vectors and more encoding layers in the parameterized model, the more diverse images, more diverse data, additional clips, more sizes, and more encoding layers being determined based on the determined variability.
[0054] According to another embodiment, a computer program product is provided, the computer program product including a non-transitory computer-readable medium having instructions recorded thereon, the instructions implementing any of the above methods when executed by a computer. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate one or more embodiments and, together with the description, explain these embodiments. Embodiments will now be described, by way of example only, with reference to the accompanying schematic drawings, in which corresponding reference characters indicate corresponding parts, and in which:
[0056] Figure 1 A block diagram of various subsystems of a lithography system according to one embodiment is shown.
[0057] Figure 2 An exemplary flow chart for simulating lithography in a lithographic projection apparatus according to one embodiment is illustrated.
[0058] Figure 3 An overview of the operations of the present method for reducing uncertainty in machine learning model predictions is illustrated according to one embodiment.
[0059] Figure 4 A convolutional encoder-decoder according to one embodiment is illustrated.
[0060] Figure 5 An encoder-decoder architecture within a neural network is illustrated, according to one embodiment.
[0061] Figure 6A This illustrates the case of sampling in the latent space according to one embodiment. Figure 5 A variational encoder-decoder architecture version of .
[0062] Figure 6B An example is shown Figure 4 Another view of the encoder-decoder architecture shown in .
[0063] Figure 6CAn example expected distribution p(z|x) is illustrated, along with the variability of the sampled distributions for p(z|x) from among multiple distributions.
[0064] Figure 7 According to one embodiment, a mask image used as input to a machine learning model, the mean of the predicted output from the machine learning model predicted based on the mask image, an image illustrating the variance in the predicted output, a scanning electron microscope (SEM) image of an actual mask generated using the mask image, and a latent space illustrating the posterior distribution are illustrated.
[0065] Figure 8 According to one embodiment, a second mask image used as input to a machine learning model, a second mean of a predicted output from the machine learning model predicted based on the second mask image, a second image illustrating the variance in the predicted output, a second SEM image of an actual mask generated using the second mask image, and a second latent space illustrating a second posterior distribution are illustrated.
[0066] Figure 9 According to one embodiment, a third mask image is illustrated as being used as input to a machine learning model, a third mean of a predicted output from the machine learning model predicted based on the third mask image, a third image illustrating the variance in the predicted output, a third SEM image of an actual mask produced using the third mask image, and a third latent space illustrating a third posterior distribution.
[0067] Figure 10 is a block diagram of an example computer system according to one embodiment.
[0068] Figure 11 is a schematic diagram of a lithographic projection apparatus according to an embodiment.
[0069] Figure 12 is a schematic diagram of another lithographic projection apparatus according to an embodiment.
[0070] Figure 13 According to one embodiment Figure 12 A more detailed view of the device in .
[0071] Figure 14 According to one embodiment Figure 12 and Figure 13 A more detailed view of the source collector module SO of the device. DETAILED DESCRIPTION
[0072] For previous machine learning models, the certainty of the predictions made by the machine learning model was unclear. That is, given the input, it was unclear whether the previous machine learning model generated accurate and consistent output. Machine learning models that produce accurate and consistent outputs are important in the integrated circuit manufacturing process. As a non-limiting example, when a mask layout is generated from a mask layout design, uncertainty about the predictions of the machine learning model can create uncertainty in the proposed mask layout. For example, these uncertainties may lead to questions about the final functionality of the wafer. Whenever a machine learning model is used to model or make predictions for individual operations in the process, more uncertainty is introduced into the integrated circuit manufacturing process. However, to date, there has been no method to determine the variability (or uncertainty) in the output from the model.
[0073] In order to address these and other shortcomings of previous parameterized (e.g., machine learning) models, (multiple) methods and (multiple) systems include models using encoder-decoder architectures. In the middle of the architecture (e.g., middle layer), the present model plans a low-dimensional encoding (e.g., potential space) that encapsulates information in the input (e.g., image, tensor, and / or other input) to the model. Using variational inference techniques, the encoder determines the posterior probability distribution for the potential vector conditioned on (multiple) inputs. In some embodiments, the model is configured to generate a distribution in multiple distributions for a given input (e.g., using a parameter dropping method). The model samples from the distribution of multiple distributions conditioned on a given input. The model can determine the variation across the sampled distributions. After sampling, the model decodes the sample into the output space. The variability of the output and / or the variation in the sampled distribution defines the uncertainty of the model, which includes the uncertainty of the model parameters (weights) and how simple (small and descriptive) the potential space is.
[0074] Although specific reference may be made herein to the manufacture of ICs, it should be expressly understood that the description herein has many other possible applications. For example, it may be used to manufacture integrated optical systems, guidance and detection patterns for magnetic domain memories, liquid crystal display panels, thin film magnetic heads, and the like. In these alternative applications, those skilled in the art will appreciate that any use of the terms "reticle," "wafer," or "die" herein in the context of such alternative applications should be considered interchangeable with the more general terms "mask," "substrate," and "target portion," respectively. Additionally, it should be noted that the method herein may have many other possible applications in a variety of fields, such as language processing systems, self-driving cars, medical imaging and diagnostics, semantic segmentation, noise reduction, chip design, electronic design automation, and the like. The method may be applied to any field where quantifying uncertainty in predictions from machine learning models is advantageous.
[0075] In this document, the terms "radiation" and "beam" are used to cover all types of electromagnetic radiation, including ultraviolet radiation (e.g., having a wavelength of 365nm, 248nm, 193nm, 157nm or 126nm) and EUV (extreme ultraviolet radiation, e.g., having a wavelength in the range of about 5nm-100nm).
[0076] A patterned device may include or may form one or more design layouts. A CAD (computer-aided design) program may be used to generate the design layout. This process is commonly referred to as EDA (electronic design automation). Most CAD programs follow a set of predetermined design rules in order to create a functional design layout / patterned device. These rules are set based on process and design constraints. For example, design rules define the spatial tolerances between devices (such as gates, capacitors, etc.) or interconnects to ensure that the devices or lines do not interact in an undesirable manner. One or more design rule constraints may be referred to as "critical dimensions" (CDs). The critical dimension of a device may be defined as the minimum width of a line or hole, or the minimum spacing between two lines or two holes. Thus, CDs regulate the overall size and density of the designed device. One of the goals in device manufacturing is to faithfully reproduce the original design intent on a substrate (via patterned devices).
[0077] The terms "mask" or "patterning device" as used herein may be broadly interpreted as referring to a general patterning device that can be used to impart a patterned cross-section to an incoming radiation beam, the patterned cross-section corresponding to the pattern to be created in the target portion of the substrate. In such a context, the term "light valve" may also be used. In addition to classical masks (transmissive or reflective; binary, phase-shifting, hybrid, etc.), other examples of such patterning devices include programmable mirror arrays. An example of such a device is a matrix-addressable surface having a viscoelastic control layer and a reflective surface. The basic principle behind such a device is that (for example) the addressed areas of the reflective surface reflect the incident radiation as diffracted radiation, while the unaddressed areas reflect the incident radiation as undiffracted radiation. Using appropriate filters, the undiffracted radiation can be filtered out of the reflected beam, leaving only the diffracted radiation; in this way, the beam is patterned according to the addressing pattern of the matrix-addressable surface. The required matrix addressing can be performed using suitable electronic components. Other examples of such patterning devices also include programmable LCD arrays. One example of such a structure is given in US Pat. No. 5,229,872, which is incorporated herein by reference.
[0078] As a brief introduction, Figure 1An exemplary lithographic projection apparatus 10A is illustrated. The main components are: a radiation source 12A, which may be a deep ultraviolet (DUV) excimer laser source or other type of source including an extreme ultraviolet (EUV) source (as above, the lithographic projection apparatus itself need not have a radiation source), illumination optics that define partial coherence (denoted as sigma) and may include optics 14A, optics 16Aa, and optics 16Ab that shape the radiation from radiation source 12A; a patterning device 18A; and transmission optics 16Ac that projects an image of the patterning device pattern onto a substrate plane 22A. An adjustable filter or aperture 20A at a pupil plane of the projection optics may limit the range of angles of the beam incident on the substrate plane 22A, with the maximum possible angle defining the numerical aperture NA of the projection optics = nsin(Θ max ), where n is the refractive index of the medium between the substrate and the last element of the projection optics, and Θ max is the maximum angle of a light beam emerging from the projection optics that can still be incident on the substrate plane 22A.
[0079] In a lithographic projection apparatus, a source provides illumination (i.e., radiation) to a patterning device, and projection optics direct the illumination onto a substrate via the patterning device and shape it. The projection optics may include at least some components of optics 14A, optics 16Aa, optics 16Ab, and optics 16Ac. An aerial image (AI) is the radiation intensity distribution at substrate level. A resist model can be used to calculate a resist image from an aerial image, an example of which can be found in U.S. Patent Application Publication No. US2009-0157630, the entire disclosure of which is incorporated herein by reference. The resist model is only related to the properties of the resist layer (e.g., the effects of chemical processes occurring during exposure, post-exposure baking (PEB), and development). The optical properties of the lithographic projection apparatus (e.g., the properties of the illumination, patterning device, and projection optics) dictate the aerial image and can be defined in an optical model. Since the patterning device used in the lithographic projection apparatus can be varied, it is desirable to separate the optical properties of the patterning device from the optical properties of the rest of the lithographic projection apparatus, including at least the source and projection optics. Details of techniques and models used to transform design layouts into various lithographic images (e.g., aerial images, resist images, etc.), apply OPC using these techniques and models, and evaluate performance (e.g., by process window) are described in U.S. Patent Application Publication Nos. US2008-0301620, 2007-0050749, 2007-0031745, 2008-0309897, 2010-0162197, and 2010-0180251, the disclosures of each of which are incorporated herein by reference in their entirety.
[0080] It is often desirable to be able to computationally determine how a patterning process will produce a desired pattern on a substrate. Thus, simulations can be provided to simulate one or more portions of the process. For example, it is desirable to be able to simulate the photolithography process of transferring a patterned device pattern onto a resist layer of a substrate after development of the resist, as well as the resulting pattern in the resist layer.
[0081] Figure 2 , an exemplary flow chart for simulating lithography in a lithographic projection apparatus is illustrated. An illumination model 31 represents the optical properties of the illumination (including the radiation intensity distribution and / or phase distribution). A projection optics model 32 represents the optical properties of the projection optics (including changes in the radiation intensity distribution and / or phase distribution caused by the projection optics). A design layout model 35 represents the optical properties of the design layout (including changes in the radiation intensity distribution and / or phase distribution caused by a given design layout), which is a representation of the arrangement of features on or formed by a patterned device. An aerial image 36 can be simulated using the illumination model 31, the projection optics model 32, and the design layout model 35. A resist image 38 can be simulated from the aerial image 36 using a resist model 37. For example, a simulation of lithography can predict the profile and / or CD in the resist image.
[0082] More specifically, the illumination model 31 can represent the optical properties of the illumination, including but not limited to the NA-sigma (σ) setting and any specific illumination shape (e.g., off-axis illumination such as annular, quadrupole, dipole, etc.). The projection optics model 32 can represent the optical properties of the projection optics model, including, for example, aberrations, distortions, refractive index, physical size or dimensions, etc. The design layout model 35 can also represent one or more physical properties of the physical patterning device, as described, for example, in U.S. Patent No. 7,587,704, the entire contents of which are incorporated herein by reference. The optical properties associated with the lithographic projection apparatus (e.g., properties of the illumination, patterning device, and projection optics) dictate the aerial image. Because the patterning device used in the lithographic projection apparatus can be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the rest of the lithographic projection apparatus, including at least the illumination and projection optics (hence the design layout model 35).
[0083] The resist model 37 can be used to calculate a resist image from an aerial image, an example of which can be found in U.S. Patent No. 8,200,468, the entire contents of which are incorporated herein by reference. The resist model is typically related to the properties of the resist layer (e.g., the effects of chemical processes occurring during exposure, post-exposure baking, and / or development).
[0084] The goal of the simulation is to accurately predict, for example, edge placement, aerial image intensity slope, and / or CD, which can then be compared to the intended design. The intended design is typically defined as a pre-OPC design layout that can be provided in a standardized digital file format such as GDSII, OASIS, or other file formats.
[0085] From the design layout, one or more portions referred to as "clips" can be identified. In one embodiment, a set of clips are extracted that represent complex patterns in the design layout (typically about 50 to 1000 clips, but any number of clips can be used). As will be appreciated by those skilled in the art, these patterns or clips represent small portions of the design (e.g., circuits, cells, etc.), and clips, in particular, represent small portions that require special attention and / or verification. In other words, a clip can be, or can be similar to or have similar behavior to, a portion of the design layout where critical features are identified through experience (including clips provided by the customer), through trial and error, or by running full-chip simulations. A clip typically contains one or more test patterns or gauge patterns. An initial larger set of clips can be provided a priori by the customer based on known critical feature areas in the design layout that require specific image optimization. Alternatively, in another embodiment, an initial larger set of clips can be extracted from the entire design layout using some automatic (such as machine vision) or manual algorithm that identifies critical feature areas.
[0086] For example, simulation and modeling can be used to configure one or more features of the patterned device pattern (e.g., to perform optical proximity correction), one or more features of the illumination (e.g., to change one or more properties of the spatial / angular intensity distribution of the illumination, such as changing the shape), and / or one or more features of the projection optics (e.g., numerical aperture, etc.). Such configurations may be generally referred to as mask optimization, source optimization, and projection optimization, respectively. Such optimizations may be performed individually or combined in different combinations. One such example is source-mask optimization (SMO), which involves the configuration of one or more features of the patterned device pattern and one or more features of the illumination. The optimization technique may focus on one or more clips within the clip. The optimization may use the machine learning models described herein to predict the values of various parameters (including images, etc.).
[0087] In some embodiments, the optimization process of the system can be expressed as a cost function. The optimization process can include finding a set of parameters (design variables, process variables, etc.) of the system that minimizes the cost function. The cost function can have any suitable form, depending on the goal of the optimization. For example, the cost function can be the weighted root mean square (RMS) of the deviations of certain characteristics (evaluation points) of the system relative to the expected values (e.g., ideal values) of these characteristics. The cost function can also be the maximum value (i.e., the worst deviation) among these deviations. The term "evaluation point" should be broadly interpreted to include any characteristic of the system or manufacturing method. Due to the practicality of the implementation of the system and / or method, the design and / or process variables of the system can be limited to a limited range and / or interdependent. In the case of a lithographic projection apparatus, constraints are typically associated with the physical properties and characteristics of the hardware, such as adjustable range and / or patterned device manufacturability design rules. Evaluation points can include, for example, physical points on the resist image on the substrate, as well as non-physical characteristics (such as dose and focus).
[0088] In some embodiments, the illumination model 31, the projection optics model 32, the design layout model 35, the resist model 37, the SMO model, and / or other models associated with and / or included in the integrated circuit manufacturing process may be empirical models for performing the method operations herein. The empirical models may predict outputs based on correlations between various inputs (e.g., one or more characteristics of a mask or wafer image, one or more characteristics of a design layout, one or more characteristics of a patterned device, one or more characteristics of illumination used in a lithography process (such as wavelength), etc.).
[0089] As an example, the empirical model can be a machine learning model and / or any other parameterized model. In some embodiments, the machine learning model (for example) can be and / or include a mathematical equation, an algorithm, a graph, a chart, a network (e.g., a neural network), and / or other tools and machine learning model components. For example, the machine learning model can be and / or include one or more neural networks having an input layer, an output layer, and one or more intermediate or hidden layers. In some embodiments, the one or more neural networks can be and / or include a deep neural network (e.g., a neural network having one or more intermediate or hidden layers between the input layer and the output layer).
[0090] As an example, one or more neural networks can be based on a large collection of neural units (or artificial neurons). One or more neural networks can roughly mimic the way a biological brain works (e.g., large clusters of biological neurons connected via axons). Each neural unit in a neural network can be connected to many other neural units in the neural network. Such connections can strengthen or suppress their influence on the activation state of the connected neural units. In some embodiments, each individual neural unit can have a summation function that combines the values of all its inputs. In some embodiments, each connection (or neural unit itself) can have a threshold function that a signal must exceed before it is allowed to propagate to other neural units. Compared to traditional computer programs, these neural network systems can be self-learning and trained rather than explicitly programmed, and perform much better in certain problem-solving areas. In some embodiments, one or more neural networks can include multiple layers (e.g., in which signal paths traverse from front layers to back layers). In some embodiments, backpropagation techniques can be utilized by neural networks, where forward stimulation is used to reset weights for "front" neural units. In some embodiments, stimulation and inhibition for one or more neural networks can flow more freely, where connections interact in a more chaotic and complex manner. In some embodiments, one or more intermediate layers of the neural network include one or more convolutional layers, one or more recurrent layers, and / or other layers.
[0091] One or more neural networks can be trained (i.e., their parameters determined) using a set of training data. The training data can include a set of training samples. Each sample can be a pair including an input object (usually a vector, which can be called a feature vector) and a desired output value (also called a supervisory signal). The training algorithm analyzes the training data and adjusts the behavior of the neural network by adjusting the parameters of the neural network (e.g., the weights of one or more layers) based on the training data. For example, given a neural network of the form {(x1, y1), (x2, y2), ..., (x N ,y N )} a set of N training samples such that x i is the eigenvector of the ith example and y i The training algorithm finds a neural network g:X→Y, where X is the input space and Y is the output space. A feature vector is an n-dimensional vector representing the numerical characteristics of an object (e.g., a chip design or a clip in the example above). The vector space associated with these vectors is often called the feature space. After training, the neural network can be used to make predictions using new examples.
[0092] As described above, (multiple) present methods and (multiple) systems include parameterized models (e.g., machine learning models such as neural networks) using an encoder-decoder architecture. In the middle (e.g., an intermediate layer) of the model (e.g., a neural network), the model plans a low-dimensional encoding (e.g., a latent space) that encapsulates information in the input to the model (e.g., an image, a tensor, and / or other input). Using variational inference techniques, the encoder determines a posterior probability distribution for the latent vector conditioned on the input (multiple). In some embodiments, the model is configured to generate a distribution from multiple distributions for a given input (e.g., using a parameter dropping method). The model samples from the distribution of multiple distributions of the posterior probability conditioned on the input. In some embodiments, sampling includes randomly selecting a distribution from the distribution of multiple distributions. For example, the sampling can be Gaussian or non-Gaussian. After sampling, the model decodes the sample into the output space. The variability of the output and / or the variability of the sampled distribution defines the uncertainty of the model, which includes the uncertainty of the model parameters (e.g., parameter weights and / or other model parameters) and how simple (small and descriptive) the latent space is. In some embodiments, determining the variability may include quantifying the variability using one or more statistical operations, the one or more statistical operations including one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis, covariance, and / or any other method for quantifying variability. In some embodiments, the uncertainty of the model is related to the uncertainty of the weights of the model's parameters and the size and descriptiveness of the latent space, such that the uncertainty of the weights manifests as uncertainty in the output, resulting in increased output variance.
[0093] This quantification of the variability of the output of a parameterized model (conditional on the input) can be used to determine, among other things, how predictive the model is. This quantification of the variability of the output of a parameterized model can also be used to adjust (e.g., update and improve) the model to make it more descriptive. The adjustments can, for example, include: adding more dimensions to the latent space, adding more diverse training data, and / or other operations. The quantification of the variability of the output of a parameterized model can also be used to guide the type of training data needed to enhance the overall quality of the predictions of the parameterized model. It should be noted that even though machine learning models and / or neural networks are mentioned throughout this specification, machine learning models and / or neural networks are an example of a parameterized model, and the operations described herein can be applied to any parameterized model.
[0094] Figure 3An overview of the operations of the present method for determining, or determining and reducing uncertainty in machine learning model predictions is illustrated. At operation 40, an encoder-decoder architecture of a machine learning model is trained. At operation 42, the machine learning model is caused to predict a plurality of outputs from the machine learning model for a given input (e.g., x and / or z as follows). The given input may include, for example, an image, a clip, an encoded image, an encoded clip, a vector, data from a previous layer of the machine learning model, and / or any other data and / or object that can be encoded.
[0095] In some embodiments, operation 42 includes: the machine learning model uses variational inference techniques to determine a posterior probability distribution for the latent vector and / or model output conditioned on the input(s). In some embodiments, the machine learning model is configured to generate a distribution among multiple distributions for a given input (e.g., using a parameter dropping method). The distribution among the multiple distributions may include, for example, a first posterior distribution among the multiple distributions (e.g., for p described below). θ (z|x)), the second posterior distribution among multiple distributions (e.g., for p described below φ (y|z)), and / or other distributions from multiple distributions. The machine learning model samples from the distributions from multiple distributions, conditioned on a given input. After sampling, the machine learning model can decode the samples into the output space.
[0096] At operation 44, the variability of the predicted multiple output realizations and / or multiple posterior distributions for a given input is determined. At operation 46, the determined variability in the predicted multiple output realizations and / or multiple posterior distributions is used to adjust the machine learning model to reduce the uncertainty of the machine learning model. In some embodiments, operation 46 is optional. In some embodiments, operation 46 includes reporting the determined variability with or without corrective action (e.g., reporting the determined variability in addition to and / or instead of adjusting the machine learning model to reduce the uncertainty of the machine learning model). For example, operation 46 may include outputting an indication of the determined variability. The indication may be an electronic indication (e.g., one or more signals), a visual indication (e.g., one or more graphics for display), a numerical indication (e.g., one or more numbers), and / or other indication.
[0097] Operation 40 includes training an encoder-decoder architecture using samples from a latent space, the samples being decoded into an output space. In some embodiments, the latent space comprises a low-dimensional encoding. As a non-limiting example, Figure 4 A convolutional encoder-decoder architecture 50 is illustrated. The encoder-decoder architecture 50 has an encoding portion 52 (encoder) and a decoding portion 54 (decoder). Figure 4 In the example shown in , the encoder-decoder architecture 50 may output, for example, Figure 4 . The image(s) 56 may have a mean 57 illustrated by a segmented image 58, a variance 59 illustrated by a model uncertainty image 60, and / or other characteristics.
[0098] As another non-limiting example, Figure 5 The encoder-decoder architecture 61 within a neural network 62 is illustrated. The encoder-decoder architecture 61 includes an encoding portion 52 and a decoding portion 54. Figure 5 In , x represents the encoder input (e.g., the input image and / or extracted features of the input image) and x' represents the decoder output (e.g., the predicted output image and / or predicted features of the output image). In some embodiments, x' can represent, for example, the output from an intermediate layer of the neural network (as compared to the final output of the entire model) and / or other outputs. In some embodiments, for example, the variable y can represent the overall output from the neural network. In Figure 5 In [ 15 ], z represents the latent space 64 and / or the low-dimensional encoding (vector). In some embodiments, z is a latent variable or is related to a latent variable. The output x' (and / or in some embodiments, y) is modeled as a (possibly very complex) function of a lower-dimensional random vector z∈Z, whose components are unobserved (latent) variables.
[0099] In some embodiments, the low-dimensional encoding z represents one or more features of an input (e.g., an image). One or more features of the input can be considered key or critical features of the input. For example, features can be considered key or critical features of the input because they are more predictive and / or have other properties than other features of the desired output. The one or more features (dimensions) represented in the low-dimensional encoding can be predetermined (e.g., by a programmer when creating the present machine learning model), determined by previous layers of the neural network, adjusted by a user via a user interface associated with the systems described herein, and / or can be determined by other methods. In some embodiments, the amount of features (dimensions) represented by the low-dimensional encoding can be predetermined (e.g., by a programmer when creating the present machine learning model), determined based on outputs from previous layers of the neural network, adjusted by a user via a user interface associated with the systems described herein, and / or determined by other methods.
[0100] Figure 6A An example is shown Figure 5 The encoder-decoder architecture 61 is constructed such that the latent space 64 is sampled 63 (e.g., Figure 6A Considered as Figure 5 A more detailed version of ). Figure 6A As shown in
[0101] p(z|x)≈q θ (z|x)[1].
[0102] The term p(z|x) is the conditional probability of the latent variable z given the input x. The term q θ (z|x) is or describes the weights of each layer of the encoder. The term p(z|x) is or describes the theoretical probability distribution of z given x. Equation
[0103] z~N(μ,σ 2 I) [2]
[0104] is or describes the prior distribution of the latent variable z, where N represents the normal (e.g., Gaussian) distribution, μ is the mean of the distribution, σ is the covariance, and I is the identity matrix. Figure 6A As shown in 2 are parameters that define the probability. They are just proxies for the true probability that the model is trying to learn given the input. In some embodiments, this proxy can be more descriptive of the task. For example, it can be a standard PDF, or some free-form PDF that can be learned.
[0105] return Figure 3 In some embodiments, operation 42 includes utilizing an encoder-decoder architecture (e.g., Figure 5 61) of the encoder (e.g., Figure 4 52 shown in ) determines or otherwise learns the conditional probability p(z|x) of the latent variable for a given input x. In some embodiments, operation 42 includes utilizing a decoder (e.g., Figure 5 54) determine or otherwise learn the conditional probability p(x'|z) (and / or py|z). In some embodiments, operation 42 includes generating x' in the training set D by maximizing the following equation i The possibility to learn φ (shown in Equation 3 below):
[0106]
[0107] In some embodiments, the conditional probability p(z|x) is determined by the encoder using a variational inference technique. In some embodiments, the variational inference technique includes θ Identify an approximation to p(z|x) in the parametric family of (z|x), where θ is a parameter of the family according to the following equation:
[0108] min KL(p(z|x),q θ (z|x)) [4]
[0109] And substituting into the maximum ELBO(θ), where ELBO represents the evidence of the lower bound, we get
[0110] ELBO(θ)=E qθ(z|x) [log p θ (x|z)]-KL(q θ (z|x), p(z)) [5]
[0111] Where KL is the Kullback-Leibler divergence and is used as a measure of the distance between two probability distributions, θ represents the parameters of the encoding, and φ represents the parameters of the decoding. Through training, the conditional probability q is obtained θ (z|x) (encoder part) and p φ (x'|z) or p φ (y|z)(decoder part).
[0112] In some embodiments, operation 42 comprises sampling from the conditional probability p(z|x) and, for each sample, predicting the output of the predicted plurality of output realizations using a decoder of an encoder-decoder architecture based on the above equations. Additionally: E qθ(z|x) [f(z)] represents the expectation of f(z), where z is sampled from q(z|x).
[0113] In some embodiments, operation 44 includes determining the variability of the predicted multiple output realizations for a given input (e.g., x) based on the predicted output for each sample. Given an input (e.g., x), the machine learning model determines the posterior distribution q θ (z|x) and p φ (x′|q θ (z|x)). Thus, operation 44 comprises determining the posterior distribution q θ (z|x). The distance of the posterior distribution to the origin of the latent space is inversely proportional to the uncertainty of the prediction of the machine learning model (e.g., the closer the distribution is to the origin of the latent space, the less certain the model is). In some embodiments, operation 44 also includes determining another posterior distribution p φ (x′|q θ (z|x)). The variance of the posterior distribution is directly related to the uncertainty of the prediction of the machine learning model (e.g., a larger variance of the second posterior distribution means greater uncertainty). Operation 44 may include determining one or both of these posterior distributions and determining the variability based on one or both of these posterior distributions.
[0114] Figure 6B An example is shown Figure 4 Another view of the encoder-decoder architecture 50 shown in . As above, the machine learning model can learn the posterior distribution p for a given input θ(z|x), and / or p for a given input φ (y|z). In some embodiments, operation 42 includes causing the model to predict multiple posterior distributions p for a given input θ (z|x), multiple posterior distributions p for a given input φ (y|z), and / or other posterior distributions. For example, for p θ (z|x) and / or p φ In some embodiments, for example, the model is configured to generate multiple posterior distributions (e.g., for P θ (z|x) and / or p φ (y|z).
[0115] In some embodiments, operation 44 includes determining the variability of the predicted multiple posterior distributions for a given input by sampling from the multiple distributions, and quantifying uncertainty in the parameterized model prediction using the determined variability in the predicted multiple posterior distributions. For example, causing the machine learning model to predict multiple posterior distributions for a given input from the parameterized model may include aligning the parameterized model prediction with a first posterior distribution p. θ The first set of multiple posterior distributions corresponding to (z|x) and the second posterior distribution p φ Determining the variability of the predicted plurality of posterior distributions for a given input may include sampling from the plurality of distributions for the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions (e.g., by sampling from the plurality of distributions for p θ (z|x) and sample from the distribution for p φ
[0066] In some embodiments, the sampling comprises randomly selecting a distribution from among the plurality of distributions. For example, the sampling can be Gaussian or non-Gaussian.
[0116] In some embodiments, operation 44 includes determining the variability of the sampled distribution. For example, Figure 6C An example expected distribution p(z|x) 600 and variability 602 of sampled distributions for p(z|x) 600 from a distribution in a plurality of distributions are illustrated. For example, variability 602 may be caused by uncertainty in a machine learning model. In some embodiments, quantifying uncertainty in parameterized model predictions using the determined variability in the predicted plurality of posterior distributions includes: using the predicted first plurality of posterior distributions and the second plurality of posterior distributions (e.g., Figure 6C The uncertainty in the machine learning model predictions is quantified by the determined variability in the multiple distributions for p(z|x)600 shown in , and similar distributions for p(y|z) in multiple distributions).
[0117] In some embodiments, determining the variability may include quantifying the variability in the sampled set of distributions using one or more statistical operations, the one or more statistical operations including one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis, covariance, range, and / or any other method for quantifying variability. For example, determining the variability of the sampled set of posterior distributions may include determining the variability of the sampled set of posterior distributions for a given input x o (For example, Figure 6C 600 or a similar distribution for p(y|z) in multiple distributions). As another example, KL distance can be used to quantify how far apart different distributions are.
[0118] In some embodiments, as described above, the uncertainty of the machine learning model prediction is related to the uncertainty of the weights of the machine learning model's parameters and the size and descriptiveness of the latent space. Uncertainty in the weights can manifest as uncertainty in the output, resulting in increased output variance. For example, if the latent space (e.g., as described herein) is low-dimensional, it will not be possible to generalize over a large set of observations. On the other hand, a large latent space will require more data to train the model.
[0119] As a non-limiting example, Figure 7 , a mask image 70 used as an input (e.g., x) to a machine learning model, a mean 72 (image) of the predicted output (image) from the machine learning model predicted based on the mask image 70, an image 74 illustrating the variance in the predicted output, a scanning electron microscope (SEM) image 78 of an actual wafer pattern produced using the mask image, and a latent space 80 illustrating a posterior distribution (e.g., p(y|z) - an example distribution from a plurality of distributions). The latent space 80 illustrates that the latent vector z has seven dimensions 81-87. The dimensions 81-87 are distributed around a center 79 of the latent space 80. The distribution of the dimensions 81-87 in the latent space 80 illustrates a relatively more certain model (smaller variance). This evidence of a relatively more certain model is confirmed by the fact that the mean image 72 and the SEM image 78 appear similar and that there are no dark colors in the variance image 74 or in locations that do not correspond to areas of the structure shown in the SEM image 78.
[0120] In some embodiments (e.g., as described herein), the posterior distribution shown in latent space 80 can be compared (e.g., statistically or otherwise) with other posterior distributions generated using the same input. The method can include determining an indication of the certainty of the model based on the comparison of these posterior distributions. For example, the greater the difference between the compared posterior distributions, the less certain the model is.
[0121] As a comparative, non-limiting example, Figure 8 An example shows Figure 7 This results in greater variation (and greater uncertainty) in the machine learning model output compared to the output shown in . Figure 8 Illustrated is a mask image 88 used as an input (e.g., x) to a machine learning model, a mean 89 of a predicted output from the machine learning model predicted based on the mask image 88, an image 90 illustrating the variance of the predicted output, an SEM image 91 of an actual mask generated using the mask image, and a latent space 92 illustrating a posterior distribution. The latent space 92 illustrates that the latent vector z also has several dimensions 93. The distribution of the dimensions 93 in the latent space 92 now illustrates a relatively more uncertain model. The distribution of the dimensions 93 in the latent space 92 is more concentrated at the origin (narrower), resulting in greater uncertainty in the output (e.g., as described herein, the method includes determining a first posterior distribution p θ (z|x), where the distance of the first posterior distribution to the origin of the latent space is inversely proportional to the uncertainty of the machine learning model. This evidence of a relatively uncertain model is confirmed by the fact that the mean image 89 and the SEM image 91 look very different, and there are large amounts of dark colors in the variance image 90 in locations where corresponding structures are not seen in the SEM image 91.
[0122] Here again, the posterior distribution shown in latent space 92 can be compared (e.g., statistically or otherwise) to other posterior distributions generated using the same input. The method can include determining an indication of the certainty of the model based on the comparison of these posterior distributions.
[0123] As a third non-limiting example, Figure 9 94, a mask image used as input (e.g., x) to a machine learning model, a mean 95 of the predicted output from the machine learning model predicted based on the mask image 94, an image 96 illustrating the variance in the predicted output, an SEM image 97 of the actual mask produced using the mask image 94, and a latent space 98 illustrating several dimensions 99 of the latent vector z. Now, the images 94-97 and the distribution of dimensions 99 in the latent space 98 illustrate a distribution with a ratio of Figure 7 The models shown in the Figure 8For example, mean image 95 looks similar to SEM image 97, but variance image 96 shows stronger colors in region A where no corresponding structure is visible in SEM image 97. In some embodiments, the posterior distribution shown in latent space 98 can be compared to other posterior distributions generated using the same input to determine the uncertainty of the model.
[0124] return Figure 3 In some embodiments, operation 46 is configured such that adjusting the machine learning model using the determined variability and / or the plurality of posterior distributions in the predicted plurality of output realizations includes: determining one or more lithography process parameters based on the predictions from the adjusted machine learning model based on the given inputs; and adjusting the lithography apparatus based on the determined one or more lithography process parameters. In some embodiments, the predictions from the adjusted machine learning model include one or more of: predicted overlay, predicted wafer geometry, and / or other predictions. In some embodiments, the determined one or more lithography process parameters include one or more of: mask design, pupil shape, dose, focus, and / or other process parameters.
[0125] In some embodiments, the one or more determined lithography process parameters include a mask design, and adjusting the lithography apparatus based on the mask design includes changing the mask design from a first mask design to a second mask design. In some embodiments, the one or more determined lithography process parameters include a pupil shape, and adjusting the lithography apparatus based on the pupil shape includes changing the pupil shape from a first pupil shape to a second pupil shape. In some embodiments, the one or more determined lithography process parameters include a dose, and adjusting the lithography apparatus based on the dose includes changing the dose from a first dose to a second dose. In some embodiments, the one or more determined lithography process parameters include a focus, and adjusting the lithography apparatus based on the focus includes changing the focus from a first focus to a second focus.
[0126] In some embodiments, operation 46 is configured such that adjusting the machine learning model using the determined variability and / or multiple posterior distributions in the predicted multiple output realizations to reduce the uncertainty of the machine learning model includes increasing the training set size and / or adding dimensions to the latent space. In some embodiments, increasing the training set size and / or adding dimensions to the latent space includes using more diverse images, more diverse data, and additional clips relative to previous training material as input to train the machine learning model; and using more dimensions for encoding vectors, and using more encoding layers, and / or other training set and / or dimensionality increase operations in the machine learning model. In some embodiments, the additional and more diverse training samples include more diverse images, more diverse data, and additional clips relative to previous training material.
[0127] In some embodiments, operation 46 is configured such that using the determined variability and / or the multiple posterior distributions in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model includes adding additional dimensions to the latent space and / or adding more layers to the machine learning model. In some embodiments, operation 46 is configured such that using the determined variability and / or the multiple posterior distributions in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model includes training the machine learning model with additional and more diverse sampling from the latent space relative to previous sampling from the latent space and / or previous training data used to train the model.
[0128] As a non-limiting example, in some embodiments, operation 46 includes using the determined variability in the predicted plurality of output realizations and / or the plurality of posterior distributions to adjust a machine learning model to reduce uncertainty in the machine learning model for predicting mask geometry in a semiconductor manufacturing process. Figure 7-Figure 9 , if the variability (e.g., as shown in the variability image) of the output from the machine learning model (e.g., the predicted mean image) is high, as shown in Figure 8 As shown, and / or if the distribution-to-distribution variation is relatively high, then as above, the training set size can be increased, and / or the dimensionality of the latent space can be increased. However, if Figure 7 As shown in , if the variability in the output from a machine learning model is low, or if the distribution-to-distribution variation is relatively low, little tuning may be needed.
[0129] In some embodiments, the present method can be used to identify possible defects in the model without adjusting the model, and, for example, to redefine the uncertainty for a particular clip (or image, data, or any other input) using a different (e.g., physical) model. In this example, the uncertainty can be used, for example, to better study the physics of a given process (e.g., resist chemistry, various pattern shapes, the impact of materials, etc.).
[0130] Other examples related to several different aspects of the integrated circuit manufacturing process and / or other processes are contemplated. For example, in some embodiments, operation 46 includes: using the determined variability and / or multiple posterior distributions in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model for predicting wafer geometry as part of the semiconductor manufacturing process. Continuing with this example, using the determined variability to adjust the machine learning model to reduce the uncertainty of the parameterized model for predicting wafer geometry as part of the semiconductor manufacturing process may include: using more diverse images, more diverse data, and additional clippings as inputs to train the machine learning model relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the machine learning model, the more diverse images, more diverse data, additional clippings, more dimensions, and more encoding layers being determined based on the determined variability.
[0131] In some embodiments, operation 46 includes: using the determined variability and / or the plurality of posterior distributions in the predicted plurality of output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model for generating the predicted overlap as part of a semiconductor manufacturing process. Continuing with this example, using the determined variability to adjust the machine learning model to reduce the uncertainty of the machine learning model for generating the predicted overlap as part of a semiconductor manufacturing process includes: training the machine learning model using more diverse images, more diverse data, and additional clips as inputs relative to previous training material; and using more dimensions for the encoding vectors and more encoding layers in the parameterized model, for example, the more diverse images, more diverse data, additional clips, more dimensions, and more encoding layers being determined based on the determined variability.
[0132] Figure 101 is a block diagram illustrating a computer system 100 that can help implement the methods, processes, or apparatus disclosed herein. The computer system 100 includes a bus 102 or other communication mechanism for communicating information and a processor 104 (or multiple processors 104 and 105) coupled to the bus 102 for processing information. The computer system 100 also includes a main memory 106, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 102 for storing instructions and information to be executed by the processor 104. The main memory 106 can also be used to store temporary variables or other intermediate information during the execution of instructions executed by the processor 104. The computer system 100 also includes a read-only memory (ROM) 108 or other static storage device, coupled to the bus 102 for storing static information and instructions for the processor 104. A storage device 110, such as a magnetic disk or optical disk, is provided and coupled to the bus 102 for storing information and instructions.
[0133] The computer system 100 can be coupled to a display 112 via bus 102, such as a cathode ray tube (CRT) or a flat-panel or touch panel display for displaying information to a computer user. An input device 114, including alphanumeric and other keys, is coupled to bus 102 for communicating information and command selections to processor 104. Another type of user input device is a cursor control 116, such as a mouse, trackball, or arrow keys, for communicating directional information and command selections to processor 104 and for controlling cursor movement on display 112. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane. A touch panel (screen) display can also be used as an input device.
[0134] According to one embodiment, portions of one or more methods described herein may be performed by computer system 100 in response to processor 104 executing one or more sequences of one or more instructions contained in main memory 106. Such instructions may be read into main memory 106 from another computer-readable medium, such as storage device 110. Execution of the sequences of instructions contained in main memory 106 causes processor 104 to perform the processing steps described herein. One or more processors in a multi-processing arrangement may also be employed to execute the sequences of instructions contained in main memory 106. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, the description herein is not limited to any specific combination of hardware circuitry and software.
[0135] As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 104 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 110. Volatile media include dynamic memory, such as main memory 106. Transmission media include coaxial cables, copper wire, and optical fibers, including the wires that make up bus 102. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROMs, DVDs, any other optical media, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves described below, or any other medium from which a computer can read.
[0136] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions may initially be carried on a disk on a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 may receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 102 may receive the data carried in the infrared signal and place the data on bus 102. Bus 102 carries the data to main memory 106, from which processor 104 retrieves and executes the instructions. The instructions received by main memory 106 may optionally be stored on storage device 110 before or after execution by processor 104.
[0137] The computer system 100 may also include a communication interface 118 coupled to the bus 102. The communication interface 118 provides bidirectional data communication coupled to a network link 120, which is connected to a local area network 122. For example, the communication interface 118 may be an integrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, the communication interface 118 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, the communication interface 118 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0138] The network link 120 typically provides data communication through one or more networks to other data devices. For example, the network link 120 may provide a connection to a host computer 124 through a local area network 122 or to data equipment operated by an Internet service provider (ISP) 126. The ISP 126, in turn, provides data communication services through a global packet data communication network, now commonly referred to as the "Internet" 128. Both the local area network 122 and the Internet 128 use electrical, electromagnetic, or optical signals to carry digital data streams. The signals through the various networks, as well as the signals on the network link 120 and through the communication interface 118, that carry the digital data to and from the computer system 100, are exemplary forms of carrier waves that transport the information.
[0139] Computer system 100 can send messages and receive data, including program code, through the network(s), network link 120, and communication interface 118. In the Internet example, server 130 can transmit requested code for an application program through Internet 128, ISP 126, local area network 122, and communication interface 118. For example, one such downloaded application can provide all or part of the methods described herein. The received code can be executed by processor 104 as it is received and / or stored in storage device 110 or other non-volatile storage for later execution. In this manner, computer system 100 can obtain application code in the form of a carrier wave.
[0140] Figure 11 An exemplary lithographic projection apparatus that can be used in conjunction with the techniques described herein is schematically depicted. The apparatus comprises:
[0141] - an illumination system IL for conditioning the radiation beam B. In this particular case, the illumination system also comprises a radiation source SO;
[0142] a first object table (e.g., patterned device table) MT provided with a patterned device holder for holding a patterned device MA (e.g., a reticle) and connected to a first positioner for accurately positioning the patterned device relative to the item PS;
[0143] a second object table (substrate table) WT provided with a substrate holder for holding a substrate W (e.g. a silicon wafer coated with resist) and connected to a second positioner for accurately positioning the substrate relative to the item PS; and
[0144] - A projection system ("lens") PS (eg, a refractive, reflective, or catadioptric optical system) for imaging the irradiated portion of the patterned device MA onto a target portion C of the substrate W (eg, comprising one or more dies).
[0145] As described herein, the device is of the transmissive type (i.e., having a transmissive patterning device). However, it can also be of the reflective type (having a reflective patterning device), for example, in general. The device can employ a different kind of patterning device than a classical mask; examples include a programmable mirror array or an LCD matrix.
[0146] A source SO (e.g., a mercury lamp or excimer laser, LPP (laser produced plasma) EUV source) generates a radiation beam. This beam is fed directly or after traversing an adjustment component such as a beam expander Ex into an illumination system (illuminator) IL. The illuminator IL may include adjustment components AD for setting the outer and / or inner radial extent of the intensity distribution in the beam (commonly referred to as σ-outer and σ-inner, respectively). In addition, it will typically include various other components, such as an integrator IN and a condenser CO. In this way, the beam B incident on the patterned device MA has a desired uniformity and intensity distribution in its cross section.
[0147] about Figure 10 It should be noted that the source SO can be within the housing of the lithographic projection apparatus (for example, this is often the case when the source SO is a mercury lamp), but it can also be remote from the lithographic projection apparatus, the radiation beam it generates being introduced into the apparatus (for example, with the aid of suitable guide mirrors); the latter scenario is often the case when the source SO is an excimer laser (for example, based on a KrF, ArF or F2 laser).
[0148] The light beam PB then intercepts the patterned device MA which is held on the patterned device table MT. After traversing the patterned device MA, the light beam B passes through a lens PL which focuses the light beam B onto a target portion C of the substrate W. With the aid of the second positioning component (and the interferometry component IF), the substrate table WT can be accurately moved, for example, so as to position a different target portion C in the path of the light beam PB. Similarly, the first positioning component can be used to accurately position the patterned device MA relative to the path of the light beam B, for example after mechanical retrieval of the patterned device MA from the patterned device library or during scanning. In general, the patterned device MA will be accurately positioned with the aid of the second positioning component (and the interferometry component IF). Figure 11 Movement of the object tables MT, WT is achieved by a long-stroke module (coarse positioning) and a short-stroke module (fine positioning) explicitly depicted in the drawings. However, in the case of a stepper (as opposed to a step and scan tool), the patterned device table MT may be connected only to the short-stroke actuator or may be fixed.
[0149] The depicted tool can be used in two different modes:
[0150] In step mode, the patterned device table MT is held essentially stationary and the entire patterned device image is projected at once (i.e. a single “flash”) onto a target portion C. The substrate table WT is then shifted in the x and / or y direction so that a different target portion C can be irradiated by the beam PB;
[0151] In scan mode, essentially the same scenario applies, except that a given target portion C is not exposed in a single "flash". Instead, the patterning device table MT can be moved in a given direction (the so-called "scanning direction", for example the y-direction) at a speed v, so that the projection beam B is scanned across the patterned device image; simultaneously, the substrate table WT is moved in the same or opposite direction at a speed V=Mv, where M is the magnification of the lens PL (typically M=1 / 4 or 1 / 5). In this way, a relatively large target portion C can be exposed without having to compromise resolution.
[0152] Figure 12 Another exemplary lithographic projection apparatus 1000 that may be utilized in conjunction with the techniques described herein is schematically depicted.
[0153] The lithographic projection apparatus 1000 comprises:
[0154] -Source Collector Module SO
[0155] - an illumination system (illuminator) IL configured to condition the radiation beam B (eg EUV radiation).
[0156] a support structure (e.g., patterned device table) MT configured to support a patterned device (e.g., mask or reticle) MA and connected to a first positioner PM configured to accurately position the patterned device;
[0157] a substrate table (e.g., wafer stage) WT configured to hold a substrate (e.g., a resist-coated wafer) W and connected to a second positioner PW configured to accurately position the substrate; and
[0158] - a projection system (eg, a reflective projection system) PS configured to project the pattern imparted to the radiation beam B by the patterning device MA onto a target portion C of the substrate W (eg, comprising one or more dies).
[0159] like Figure 12As depicted in FIG, the lithographic projection apparatus 1000 is of the reflective type (e.g., employing a reflective patterning device). It should be noted that because most materials are absorptive in the EUV wavelength range, the patterning device can have a multilayer reflector comprising, for example, a multi-stack of molybdenum and silicon. In one example, the multi-stack reflector has 40 pairs of molybdenum and silicon layers, each layer having a thickness of a quarter wavelength. Even smaller wavelengths can be produced using X-ray lithography. Since most materials are absorptive at both EUV and X-ray wavelengths, patterning a thin sheet of absorbing material (e.g., a TaN absorber on top of a multilayer reflector) on the patterning device topography defines where features will be printed (positive resist) or not (negative resist).
[0160] The illuminator IL receives a beam of extreme ultraviolet radiation from a source collector module SO. Methods for generating EUV radiation include, but are not limited to, converting a material into a plasma state having at least one element (e.g., xenon, lithium, or tin) using one or more emission lines in the EUV range. In one such method, often referred to as laser produced plasma ("LPP"), a plasma may be generated by irradiating a fuel (e.g., a droplet, stream, or cluster of material having a line-emitting element) with a laser beam. The source collector module SO may be part of an EUV radiation system that includes Figure 12 A laser (not shown) is used to provide a laser beam for exciting the fuel. The resulting plasma emits output radiation, such as EUV radiation, which is collected using a radiation collector housed in a source collector module. For example, when a CO2 laser is used to provide the laser beam for fuel excitation, the laser and source collector module can be separate entities.
[0161] In such cases, the laser is not considered part of the lithographic apparatus, and the radiation beam is delivered from the laser to the source collector module by means of a beam delivery system including, for example, suitable guide mirrors and / or beam expanders. In other cases, for example, when the source is a discharge produced plasma EUV generator (often referred to as a DPP source), the source may be an integral part of the source collector module. In one embodiment, a DUV laser source may be used.
[0162] The illuminator IL may include an adjuster for adjusting the angular intensity distribution of the radiation beam. Typically, at least the outer and / or inner radial extent of the intensity distribution in a pupil plane of the illuminator (commonly referred to as σ-outer and σ-inner, respectively) may be adjusted. In addition, the illuminator IL may include various other components, such as a faceted field and pupil mirror devices. The illuminator may be used to condition the radiation beam to have a desired uniformity and intensity distribution in its cross-section.
[0163] A radiation beam B is incident on a patterning device (e.g., mask) MA, which is held on a support structure (e.g., patterning device table) MT and is patterned by the patterning device. After reflecting from the patterning device (e.g., mask) MA, the radiation beam B passes through a projection system PS, which focuses the beam onto a target portion C of a substrate W. With the aid of a second positioner PW and a position sensor PS2 (e.g., an interferometric device, a linear encoder, or a capacitive sensor), the substrate table WT can be accurately moved, for example, so that a different target portion C is positioned in the path of the radiation beam B. Similarly, a first positioner PM and another position sensor PS1 can be used to accurately position the patterning device (e.g., mask) MA relative to the path of the radiation beam B. The patterning device (e.g., mask) MA and the substrate W can be aligned using patterning device alignment marks M1, M2 and substrate alignment marks P1, P2.
[0164] The depicted apparatus 1000 may be used in at least one of the following modes:
[0165] In step mode, the support structure (e.g., patterning device table) MT and substrate table WT are held substantially stationary (i.e., a single static exposure) while the entire pattern imparted to the radiation beam is projected at one time onto a target portion C. The substrate table WT is then shifted in the X and / or Y direction so that a different target portion C can be exposed.
[0166] In scan mode, the support structure (e.g., patterning device table) MT and substrate table WT are scanned synchronously (i.e., a single dynamic exposure) as a pattern imparted to the radiation beam is projected onto a target portion C. The speed and direction of the substrate table WT relative to the support structure (e.g., patterning device table) MT may be determined by the (de-)magnification and image reversal characteristics of the projection system PS.
[0167] In another mode, the support structure (e.g., patterning device table) MT is held substantially stationary, thereby holding a programmable patterning device, and the substrate table WT is moved or scanned, while a pattern imparted to the radiation beam is projected onto a target portion. In such a mode, a pulsed radiation source is typically employed, and the programmable patterning device is updated as required after each movement of the substrate table WT or between successive radiation pulses during a scan. Such an operating mode can be readily applied to maskless lithography utilizing a programmable patterning device, such as a programmable mirror array of the type indicated above.
[0168] Figure 13The apparatus 1000 comprising a source collector module SO, an illumination system IL and a projection system PS is shown in more detail. The source collector module SO is constructed and arranged so that a vacuum environment can be maintained in the enclosed structure 220 of the source collector module SO. A plasma 210 emitting EUV radiation can be formed by a plasma source that produces an electrical discharge. EUV radiation can be generated by a gas or vapor (e.g., xenon, lithium vapor, or tin vapor), wherein a very hot plasma 210 is generated to emit radiation in the EUV range of the electromagnetic spectrum. For example, the very hot plasma 210 is generated by an electrical discharge that causes a plasma that is at least partially ionized. In order to effectively generate radiation, a partial pressure of Xe, Li, Sn vapor, or any other suitable gas or vapor of, for example, 10 Pa may be required. In one embodiment, a plasma that excites tin (Sn) is provided to generate EUV radiation.
[0169] The radiation emitted by the hot plasma 210 is transferred from the source chamber 211 to the collector chamber 212 via an optional gas barrier or contaminant trap 230 (also referred to as a contaminant barrier or foil trap in some cases) located in or behind an opening in the source chamber 211. The contaminant trap 230 may include a channel structure. The contaminant trap 230 may also include a gas barrier or a combination of a gas barrier and a channel structure. As is known in the art, the contaminant trap or contaminant barrier 230 further referred to herein includes at least a channel structure.
[0170] The source chamber 211 may include a radiation collector CO, which may be a so-called grazing incidence collector. The radiation collector CO has an upstream radiation collector side 251 and a downstream radiation collector side 252. Radiation that traverses the collector CO may be reflected from the grating spectral filter 240 and then focused along the optical axis indicated by the dotted line "O" into a virtual source point IF. The virtual source point IF is often referred to as an intermediate focus, and the source collector module is arranged such that the intermediate focus IF is located at or near an opening 221 in the enclosure 220. The virtual source point IF is an image of the radiation-emitting plasma 210.
[0171] The radiation then traverses an illumination system IL, which may include a polygonal field mirror device 22 and a polygonal pupil mirror device 24, which are arranged to provide a desired angular distribution of the radiation beam 21 at the patterning device MA, and a desired uniformity of radiation intensity at the patterning device MA. When the radiation beam 21 reflects at the patterning device MA, which is held by the support structure MT, a patterned beam 26 is formed, and the patterned beam 26 is imaged by the projection system PS via reflective elements 28, 30 onto a substrate W held by the substrate table WT.
[0172] There may typically be more elements present in the illumination optics unit IL and the projection system PS than shown. Depending on the type of lithographic apparatus, a grating spectral filter 240 may optionally be present. Furthermore, there may be more mirrors than shown in the figures, for example with Figure 13 Compared to what is shown in FIG, there may be 1-6 additional reflective elements in the projection system PS.
[0173] like Figure 14 As shown in FIG, collector optics CO is depicted as a nested collector having grazing incidence reflectors 253, 254, and 255, merely as an example of a collector (or collector mirror). Grazing incidence reflectors 253, 254, and 255 are arranged axially symmetrically about optical axis O, and this type of collector optics CO can be used in conjunction with a discharge produced plasma source (often referred to as a DPP source).
[0174] Alternatively, the source collector module SO may be as follows Figure 14 . A portion of an LPP radiation system is shown in FIG. A laser LA is arranged to deposit laser energy into a fuel such as xenon (Xe), tin (Sn), or lithium (Li), producing a highly ionized plasma 210 with an electron temperature of tens of electron volts. High-energy radiation generated during the deexcitation and recombination of these ions is emitted from the plasma, collected by a collector optic CO at near normal incidence, and focused onto an opening 221 in an enclosure 220.
[0175] The embodiments may be further described using the following terms:
[0176] 1. A method for quantifying uncertainty in predictions from a machine learning model, the method comprising:
[0177] causing the machine learning model to predict a plurality of output realizations from the machine learning model for a given input;
[0178] determining a variability in the predicted multiple output realizations for a given input; and quantifying uncertainty in the predicted multiple output realizations from a machine learning model using the determined variability in the predicted multiple output realizations.
[0179] 2. A method according to claim 1, wherein causing the machine learning model to predict multiple output realizations includes sampling from conditional probabilities conditioned on a given input.
[0180] 3. A method according to any of clauses 1 to 2, wherein the given input comprises one or more of: an image, a clip, an encoded image, an encoded clip, or data from a previous layer of a machine learning model.
[0181] 4. The method of any one of clauses 1 to 3 further comprises: using the determined variability and / or quantified uncertainty in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model by making the machine learning model more descriptive or including more diverse training data.
[0182] 5. A method according to any one of clauses 1 to 4, wherein the machine learning model comprises an encoder-decoder architecture.
[0183] 6. A method according to clause 5, wherein the encoder-decoder architecture comprises a variational encoder-decoder architecture, the method further comprising: training the variational encoder-decoder architecture using the probabilistic latent space, the variational encoder-decoder architecture generating realizations in the output space.
[0184] 7. The method of clause 6, wherein the latent space comprises a low-dimensional encoding.
[0185] 8. The method of clause 7, further comprising: determining the conditional probability of the latent variable using an encoder portion of an encoder-decoder architecture for a given input.
[0186] 9. The method of clause 8, further comprising: determining the conditional probabilities using a decoder portion of an encoder-decoder architecture.
[0187] 10. The method of clause 9, further comprising: sampling from the conditional probabilities of the latent variables determined using the encoder portion of the encoder-decoder architecture, and for each sample, predicting an output using the decoder portion of the encoder-decoder architecture.
[0188] 11. The method of clause 10, wherein sampling comprises randomly selecting numbers from a given conditional probability distribution, wherein the sampling is Gaussian or non-Gaussian.
[0189] 12. The method of clause 10, further comprising determining a variability of the predicted multiple output realizations for a given input based on the predicted output for each sample in the latent space.
[0190] 13. The method of clause 12, wherein determining the variability comprises quantifying the variability using one or more statistical operations, the one or more statistical operations comprising one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis, or covariance.
[0191] 14. A method according to any one of clauses 8 to 13, wherein the conditional probabilities of the latent variables determined using the encoder part of the encoder-decoder architecture are determined by the encoder part using variational inference techniques.
[0192] 15. The method of clause 14, wherein the variation inference technique comprises: using an encoder portion of an encoder-decoder architecture to identify an approximation of the conditional probability of the latent variable in a parametric family of distributions.
[0193] 16. The method of clause 15, wherein the parametric family of distributions comprises parameterized distributions, wherein the family refers to a type or shape of distributions, or a combination of distributions.
[0194] 17. The method of any one of clauses 1 to 16 further comprises determining a first posterior distribution, wherein a distance of the first posterior distribution to the origin of the latent space is inversely proportional to the uncertainty of the machine learning model.
[0195] 18. The method of any one of clauses 1 to 17, further comprising: determining a second posterior distribution, wherein the variance of the second posterior distribution is directly related to the uncertainty of the machine learning model.
[0196] 19. The method of clause 18, wherein determining the second posterior distribution comprises: directly sampling the latent space.
[0197] 20. The method of clause 18, wherein a second posterior distribution is learned.
[0198] 21. A method according to any one of clauses 1 to 20, wherein the uncertainty of the machine learning model is related to the uncertainty of the weights of the parameters of the machine learning model and the size and descriptiveness of the latent space.
[0199] 22. A method according to claim 21, wherein the uncertainty of the machine learning model is related to the uncertainty in the weights of the parameters of the machine learning model and the size and descriptiveness of the latent space, so that the uncertainty in the weights manifests as uncertainty in the output, resulting in increased output variance.
[0200] 23. A method according to any one of clauses 2 to 22, wherein using the determined variability in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model includes: increasing the training set size and / or adding dimensions to the latent space.
[0201] 24. A method according to claim 23, wherein increasing the training set size and / or adding dimensions to the latent space includes: using more diverse images, more diverse data, and additional clips as input to train the machine learning model relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the machine learning model.
[0202] 25. A method according to any one of clauses 2 to 24, wherein adjusting the machine learning model using the determined variability in the predicted multiple output realizations to reduce the uncertainty of the machine learning model includes adding additional dimensions to the latent space.
[0203] 26. A method according to any one of clauses 2 to 25, wherein using the determined variability in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model includes training the machine learning model using additional and more diverse training samples.
[0204] 27. The method of clause 26, wherein the additional and more diverse training samples include: more diverse images, more diverse data, and additional clips relative to previous training material.
[0205] 28. The method of any one of clauses 2 to 27, further comprising: using the determined variability in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model for predicting wafer geometry as part of a semiconductor manufacturing process.
[0206] 29. A method according to claim 28, wherein using the determined variability in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model for predicting wafer geometry as part of a semiconductor manufacturing process includes: using more diverse images, more diverse data, and additional clippings as inputs to train the machine learning model relative to previous training material; and using more dimensions for encoding vectors and more encoding layers in the machine learning model, the more diverse images, more diverse data, additional clippings, more dimensions, and more encoding layers being determined based on the determined variability.
[0207] 30. The method of any one of clauses 2 to 29, further comprising: using the determined variability in the predicted multiple output realizations to adjust the machine learning model to reduce the uncertainty of the machine learning model for generating the predicted overlap as part of a semiconductor manufacturing process.
[0208] 31. A method according to claim 30, wherein the determined variability in the predicted multiple output realizations is used to adjust the machine learning model to reduce the uncertainty of the machine learning model for generating the predicted overlap as part of a semiconductor manufacturing process, including: using more diverse images, more diverse data, and additional clips as inputs to train the machine learning model; and using more sizes for encoding vectors and more encoding layers in the machine learning model, the more diverse images, more diverse data, additional clips, more sizes, and more encoding layers being determined based on the determined variability.
[0209] 32. A method for quantifying uncertainty in parameterized model predictions, the method comprising:
[0210] causing the parameterized model to predict a plurality of output realizations from the parameterized model for a given input;
[0211] determining the variability of the predicted multiple output realizations for a given input; and
[0212] Uncertainty in the predicted multiple output realizations from the parameterized model is quantified using the determined variability in the predicted multiple output realizations.
[0213] 33. A method according to clause 32, wherein the parameterized model is a machine learning model.
[0214] 34. A computer program product comprising a non-transitory computer-readable medium having recorded thereon instructions, the instructions implementing the method of any one of clauses 1 to 33 when executed by a computer.
[0215] 35. A method for configuring a lithographic apparatus, the method comprising:
[0216] causing the machine learning model to predict a plurality of posterior distributions for a given input from the machine learning model, the plurality of posterior distributions comprising a distribution among the plurality of distributions;
[0217] determining variability in a plurality of predicted posterior distributions for a given input by sampling from the plurality of distributions;
[0218] quantifying uncertainty in machine learning model predictions using the determined variability in the predicted multiple posterior distributions;
[0219] Adjusting one or more parameters of the machine learning model to reduce uncertainty in the machine learning model's predictions; and
[0220] One or more lithography process parameters for adjusting the lithography apparatus are determined based on predictions from the adjusted machine learning model for a given input.
[0221] 36. The method of clause 35, further comprising adjusting the lithographic apparatus based on the determined one or more lithographic process parameters.
[0222] 37. A method according to clause 36, wherein the one or more parameters of the machine learning model include one or more weights of the one or more parameters of the machine learning model.
[0223] 38. A method according to any of clauses 35 to 37, wherein the predictions from the adapted machine learning model include one or more of predicted overlay or predicted wafer geometry.
[0224] 39. The method of any of clauses 35 to 38, wherein the one or more lithography process parameters determined include one or more of mask design, pupil shape, dose, or focus.
[0225] 40. The method of clause 39, wherein the determined one or more lithographic process parameters include a mask design, and adjusting the lithographic apparatus based on the mask design comprises changing the mask design from a first mask design to a second mask design.
[0226] 41. The method of clause 39, wherein the determined one or more lithographic process parameters include a pupil shape, and adjusting the lithographic apparatus based on the pupil shape comprises changing the pupil shape from a first pupil shape to a second pupil shape.
[0227] 42. The method of clause 39, wherein the determined one or more lithographic process parameters include dose, and adjusting the lithographic apparatus based on the dose comprises changing the dose from a first dose to a second dose.
[0228] 43. The method of clause 39, wherein the determined one or more lithographic process parameters include focus, and adjusting the lithographic apparatus based on the focus comprises changing the focus from a first focus to a second focus.
[0229] 44. A method according to any one of clauses 35 to 43, wherein causing the machine learning model to predict multiple posterior distributions includes causing the machine learning model to use parameter dropout to generate a distribution among the multiple distributions.
[0230] 45. A method according to any one of clauses 35 to 44, wherein:
[0231] Causing the machine learning model to predict multiple posterior distributions for a given input from the machine learning model includes: causing the machine learning model to predict a first posterior distribution P Θ The first set of multiple posterior distributions corresponding to (z|x) and the second posterior distribution P φ The second set of multiple posterior distributions corresponding to (y|z);
[0232] Determining the variability of the predicted plurality of posterior distributions for the given input by distribution sampling from the plurality of distributions includes: determining the variability of the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions for the given input by distribution sampling from the plurality of distributions for the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions; and
[0233] Quantifying uncertainty in machine learning model predictions using determined variability in predicted multiple posterior distributions includes quantifying uncertainty in machine learning model predictions using determined variability in predicted first and second multiple posterior distributions.
[0234] 46. A method according to any of clauses 35 to 45, wherein the given input comprises one or more of: an image, a clip, an encoded image, an encoded clip, or data from a previous layer of a parameterized model.
[0235] 47. The method of any one of clauses 35 to 46 further comprises: using the determined variability and / or quantified uncertainty in the predicted multiple posterior distributions to adjust the machine learning model to reduce the uncertainty of the machine learning model by making the machine learning model more descriptive or including more diverse training data.
[0236] 48. A method according to any one of clauses 35 to 47, wherein sampling comprises randomly selecting a distribution from a plurality of distributions, wherein the sampling is Gaussian or non-Gaussian.
[0237] 49. A method according to any one of clauses 35 to 48, wherein determining the variability comprises quantifying the variability using one or more statistical operations, the one or more statistical operations comprising one or more of the following: mean, moment, skewness, standard deviation, variance, kurtosis or covariance.
[0238] 50. A method according to any one of clauses 35 to 49, wherein the uncertainty of the machine learning model is related to the uncertainty of the weights of one or more parameters of the machine learning model and the size and descriptiveness of the latent space associated with the machine learning model.
[0239] 51. A method according to any one of clauses 35 to 50, wherein adjusting the machine learning model to reduce the uncertainty of the machine learning model includes: increasing the training set size and / or adding dimensions to the latent space associated with the machine learning model.
[0240] 52. A method according to claim 51, wherein increasing the training set size and / or adding dimensions to the latent space includes: using more diverse images, more diverse data, and additional clips as input to train the machine learning model relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the machine learning model.
[0241] 53. A method according to any one of clauses 35 to 52, wherein adjusting a machine learning model using the determined variability in the predicted multiple posterior distributions to reduce the uncertainty of the machine learning model includes adding additional dimensions to a latent space associated with the machine learning model.
[0242] 54. A method according to any one of clauses 35 to 53, wherein using the determined variability in the predicted multiple posterior distributions to adjust one or more parameters of the machine learning model to reduce the uncertainty of the machine learning model includes: training the machine learning model using additional and more diverse training samples.
[0243] 55. A method for quantifying uncertainty in parameterized model predictions, the method comprising:
[0244] causing the parameterized model to predict a plurality of posterior distributions for a given input from the parameterized model, the plurality of posterior distributions comprising a distribution among the plurality of distributions;
[0245] determining variability of a plurality of predicted posterior distributions for a given input by sampling from the plurality of distributions; and
[0246] The determined variability in the predicted multiple posterior distributions is used to quantify uncertainty in the parameterized model predictions.
[0247] 56. A method according to clause 55, wherein the parameterized model is a machine learning model.
[0248] 57. The method of any one of clauses 55 to 56, wherein causing the parameterized model to predict a plurality of posterior distributions comprises causing the parameterized model to generate a distribution among the plurality of distributions using parameter dropout.
[0249] 58. A method according to any one of clauses 55 to 57, wherein:
[0250] Causing the parameterized model to predict multiple posterior distributions for a given input from the parameterized model includes: aligning the parameterized model predictions with a first posterior distribution P Θ The first set of multiple posterior distributions corresponding to (z|x) and the second posterior distribution P φ The second set of multiple posterior distributions corresponding to (y|z);
[0251] Determining the variability of the predicted plurality of posterior distributions for the given input by distribution sampling from the plurality of distributions includes: determining the variability of the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions for the given input by distribution sampling from the plurality of distributions for the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions; and
[0252] Quantifying uncertainty in parameterized model predictions using the determined variability in the predicted multiple posterior distributions includes quantifying uncertainty in parameterized model predictions using the determined variability in the predicted first multiple posterior distributions and the predicted second multiple posterior distributions.
[0253] 59. A method according to any of clauses 55 to 58, wherein the given input comprises one or more of: an image, a clip, an encoded image, an encoded clip, or data from a previous layer of a parameterized model.
[0254] 60. The method of any one of clauses 55 to 59 further comprises: using the determined variability and / or quantified uncertainty in the predicted multiple posterior distributions to adjust the parameterized model to reduce the uncertainty of the parameterized model by making the parameterized model more descriptive or including more diverse training data.
[0255] 61. A method according to any one of clauses 55 to 60, wherein the parameterized model comprises an encoder-decoder architecture.
[0256] 62. A method according to clause 61, wherein the encoder-decoder architecture comprises a variational encoder-decoder architecture, the method further comprising: training the variational encoder-decoder architecture using the probabilistic latent space, the variational encoder-decoder architecture generating an implementation in the output space.
[0257] 63. The method of clause 62, wherein the latent space comprises a low-dimensional encoding.
[0258] 64. The method of clause 63, further comprising: determining a conditional probability of the latent variable using an encoder portion of an encoder-decoder architecture for a given input.
[0259] 65. The method of clause 64, further comprising: determining the conditional probabilities using a decoder portion of an encoder-decoder architecture.
[0260] 66. The method of clause 65, further comprising: sampling from the conditional probabilities of the latent variables determined using the encoder portion of the encoder-decoder architecture, and for each sample, predicting an output using the decoder portion of the encoder-decoder architecture.
[0261] 67. The method of clause 55, wherein the sampling comprises: randomly selecting a distribution from a distribution in a plurality of distributions, wherein the sampling is Gaussian or non-Gaussian.
[0262] 68. The method of clause 67, wherein determining the variability comprises quantifying the variability using one or more statistical operations, the one or more statistical operations comprising one or more of: mean, moment, skewness, standard deviation, variance, kurtosis, or covariance.
[0263] 69. A method according to any of clauses 62 to 68, wherein the uncertainty of the parameterized model is related to the uncertainty of the weights of the parameters of the parameterized model and the size and descriptiveness of the latent space.
[0264] 70. A method according to clause 69, wherein the uncertainty of the parameterized model is related to the uncertainty of the weights of the parameters of the parameterized model and the size and descriptiveness of the latent space, so that the uncertainty of the weights manifests as uncertainty in the output, resulting in increased output variance.
[0265] 71. A method according to any one of clauses 62 to 70, wherein adjusting the parameterized model using the determined variability in the predicted multiple posterior distributions to reduce the uncertainty of the parameterized model includes: increasing the training set size and / or adding dimensions to the latent space.
[0266] 72. A method according to claim 71, wherein increasing the training set size and / or adding dimensions to the latent space includes: using more diverse images, more diverse data, and additional clips as input to train the parameterized model relative to previous training material; and using more dimensions for the encoding vectors and using more encoding layers in the parameterized model.
[0267] 73. A method according to any one of clauses 62 to 72, wherein adjusting the parameterized model using the determined variability in the predicted multiple posterior distributions to reduce the uncertainty of the parameterized model comprises adding additional dimensions to the latent space.
[0268] 74. A method according to any one of clauses 60 to 73, wherein adjusting the parameterized model using the determined variability in the predicted multiple posterior distributions to reduce the uncertainty of the parameterized model includes training the parameterized model using additional and more diverse training samples.
[0269] 75. The method of clause 74, wherein the additional and more diverse training samples comprise: more diverse images, more diverse data, and additional clips relative to previous training material.
[0270] 76. The method of any of clauses 60 to 75, further comprising: using the determined variability in the predicted plurality of posterior distributions to adjust a parameterized model to reduce uncertainty in the parameterized model for predicting wafer geometry as part of a semiconductor manufacturing process.
[0271] 77. A method according to claim 76, wherein the determined variability in the predicted multiple posterior distributions is used to adjust a parameterized model to reduce the uncertainty of the parameterized model for predicting wafer geometry as part of a semiconductor manufacturing process, including: using more diverse images, more diverse data, and additional clippings as input to train the parameterized model relative to previous training material; and using more dimensions for encoding vectors and more encoding layers in the parameterized model, the more diverse images, more diverse data, additional clippings, more dimensions, and more encoding layers being determined based on the determined variability.
[0272] 78. The method of any of clauses 60 to 77, further comprising using the determined variability in the predicted plurality of posterior distributions to adjust a parameterized model to reduce uncertainty in the parameterized model for generating the predicted overlay as part of a semiconductor manufacturing process.
[0273] 79. A method according to claim 78, wherein the determined variability in the predicted multiple posterior distributions is used to adjust a parameterized model to reduce the uncertainty of the parameterized model for generating the predicted overlap as part of a semiconductor manufacturing process, including: using more diverse images, more diverse data, and additional clips as input to train the parameterized model relative to previous training material; and using more dimensions for the encoding vectors and more encoding layers in the parameterized model, the more diverse images, more diverse data, additional clips, more dimensions, and more encoding layers being determined based on the determined variability.
[0274] 80. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions implementing the method of any one of clauses 35 to 79 when executed by a computer.
[0275] The concepts disclosed herein can simulate or mathematically model any general imaging system used to image subwavelength features and are particularly useful for emerging imaging technologies that can produce shorter and shorter wavelengths. Emerging technologies already in use include EUV (extreme ultraviolet), DUV lithography, which can produce wavelengths of 193 nm using ArF lasers, and even 157 nm using fluorine lasers. Furthermore, EUV lithography can produce wavelengths in the 20 nm to 5 nm range by using a synchrotron or by bombarding materials (solid or plasma) with high-energy electrons to produce photons in the 20 nm to 5 nm range.
[0276] While the concepts disclosed herein can be used for imaging on substrates such as silicon wafers, it will be understood that the disclosed concepts can be used with any type of lithographic imaging system, for example, a lithographic system for imaging on substrates other than silicon wafers. In addition, combinations and subcombinations of the disclosed elements may comprise separate embodiments. For example, determining the variability of a machine learning model may comprise determining the variability in individual predictions made by the model and / or the variability in a sample set of a posterior distribution generated by the model. These features may comprise separate embodiments, and / or these features may be used together in the same embodiment.
[0277] The above description is intended to be illustrative rather than restrictive. It will therefore be apparent to those skilled in the art that modifications may be made as described without departing from the scope of the claims set forth below.
Claims
1. A method for quantifying uncertainty in parameterized model predictions for lithography simulations, the method comprising: causing a parameterized model related to an integrated circuit manufacturing process to predict a plurality of posterior distributions from the parameterized model for given inputs related to the integrated circuit manufacturing process, the plurality of posterior distributions comprising a distribution from a plurality of distributions; determining variability of the plurality of predicted posterior distributions for the given input by sampling from the distributions of a plurality of distributions; quantifying uncertainty in predictions from the parameterized model associated with the lithography simulation using the determined variability in the predicted plurality of posterior distributions; as well as One or more parameters of the parameterized model are adjusted to reduce the uncertainty of the prediction associated with the lithography simulation.
2. The method of claim 1, wherein the parameterized model is a machine learning model.
3. The method of claim 1 , wherein causing the parameterized model to predict the plurality of posterior distributions comprises: The parameterized model is caused to generate the distribution among a plurality of distributions using parameter dropout.
4. The method according to claim 1, wherein: Causing the parameterized model to predict the plurality of posterior distributions for a given input from the parameterized model includes: causing the parameterized model to predict a first plurality of posterior distributions corresponding to a first posterior distribution, and a second plurality of posterior distributions corresponding to a second posterior distribution; Determining the variability of the predicted plurality of posterior distributions for the given input by sampling from the distributions from a plurality of distributions comprises: determining the variability of the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions for the given input by sampling from the distributions for the predicted first plurality of posterior distributions and the predicted second plurality of posterior distributions from a plurality of distributions; and Quantifying the uncertainty of the prediction using the determined variability in the predicted multiple posterior distributions includes: quantifying the uncertainty of the prediction using the determined variability in the predicted first set of multiple posterior distributions and the predicted second set of multiple posterior distributions.
5. The method of claim 4, wherein the distance of the first posterior distribution to the origin of the latent space is inversely proportional to the uncertainty of the prediction, and the variance of the second posterior distribution is directly related to the uncertainty of the prediction.
6. The method of claim 1, wherein the given input comprises one or more of: an image, a clip, or data from a previous layer of the parameterized model.
7. The method of claim 6, wherein the image comprises an encoded image and the clip comprises an encoded clip.
8. The method according to claim 1, further comprising: The determined variability and / or the quantified uncertainty in the predicted multiple posterior distributions are used to adjust the one or more parameters of the parameterized model to reduce the uncertainty of the prediction by making the parameterized model more descriptive or including more diverse training data.
9. The method of claim 1, wherein the parameterized model comprises an encoder-decoder architecture.
10. The method of claim 9, wherein the encoder-decoder architecture comprises a variational encoder-decoder architecture, the method further comprising: The variational encoder-decoder architecture is trained using a probabilistic latent space, which generates realizations in the output space. The method of claim 10 , wherein the latent space comprises a low-dimensional encoding.
12. The method according to claim 11, further comprising: For the given input, the encoder portion of the encoder-decoder architecture is used to determine the conditional probability of the latent variable.
13. The method according to claim 12, further comprising: The conditional probabilities are determined using the decoder portion of the encoder-decoder architecture.
14. The method of claim 1, wherein sampling comprises: A distribution is randomly selected from the distribution among a plurality of distributions, wherein the sampling is Gaussian or non-Gaussian.
15. The method of claim 10, wherein the uncertainty of the prediction is related to uncertainty in weights of parameters of the parameterized model and the size and descriptiveness of the latent space.
16. The method of claim 10, wherein adjusting the one or more parameters of the parameterized model to reduce the uncertainty of the prediction comprises: Increasing the training set size and / or adding dimensions to the latent space; Adding additional dimensions to the latent space; or • Train the parameterized model with additional and more diverse training samples.
17. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions implementing the method of claim 1 when executed by a computer.
Citation Information
Patent Citations
System and method for creating a focus-exposure model of a lithography process
US20070031745A1
Method for identifying and using process window signature patterns for lithography process control
US20070050749A1
System and method for model-based sub-resolution assist feature generation
US20080301620A1
Multivariable solver for optical proximity correction
US20080309897A1
Method of extracting data and recommending and generating visual displays
US20090157630A1