Method for modeling a system, such as a lithography system, for performing predictive maintenance of the system

By optimizing the prediction model of the lithography system, combining domain knowledge and Bayesian inference, the problems of insufficient data volume and heterogeneity in the predictive maintenance of lithography system are solved, and more accurate fault prediction and optimization are achieved, improving production efficiency.

CN115210651BActive Publication Date: 2025-07-25ASML NETHERLANDS BV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180018809.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-14
Filing Date
2021-02-04
Publication Date
2025-07-25
Estimated Expiration
2041-02-04

AI Technical Summary

Technical Problem

The predictive maintenance of existing lithography systems faces problems with insufficient data volume and heterogeneity, which leads to low prediction accuracy and it is difficult for traditional methods to effectively predict and optimize faults.

Method used

Using a computing framework and method, the joint training and adjustment of the model is achieved by optimizing regularization terms including the first predictive model parameters of a specific configuration and the second predictive model parameters of the other configurations, constructing a predictive maintenance model, and using technologies such as domain knowledge and Bayesian inference.

Benefits of technology

It improves the accuracy of fault prediction and maintenance efficiency of lithography systems, reduces unnecessary downtime and costs, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115210651B_ABST
    Figure CN115210651B_ABST
Patent Text Reader

Abstract

Disclosed is a method for adjusting a prediction model related to at least one specific configuration of a manufacturing device. The method includes: obtaining a function that at least includes a first function of a first prediction model parameter associated with the at least one specific configuration, and a second function of the first prediction model parameter and a second prediction model parameter associated with a configuration other than the at least one specific configuration of the manufacturing device and / or a related device. Obtaining values of the first prediction model parameter based on optimization of the function, and adjusting the prediction model according to these values of the first prediction model parameter to obtain an adjusted prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to European Application No. 20161646.0, filed on March 6, 2020, and European Application No. 20174756.5, filed on March 14, 2020, the entire contents of these European patent applications are incorporated herein by reference. Technical field

[0003] The present invention generally relates to predictive maintenance of systems and to methods for modeling such systems. More specifically, the present invention relates to systems and techniques for processing data from such models. Background art

[0004] A lithographic apparatus is a machine that applies a desired pattern onto a substrate (usually onto a target portion of the substrate). A lithographic apparatus can be used, for example, in the manufacture of integrated circuits (ICs). In such a case, a patterning device (which is alternatively referred to as a mask or reticle) can be used to generate a circuit pattern to be formed on a single layer of the IC. This pattern can be transferred onto a target portion (e.g., a portion including dies, a single die, or several dies) of a substrate (e.g., a silicon wafer). The transfer of the pattern is typically effected via imaging onto a layer of radiation - sensitive material (resist) provided on the substrate. Usually, a single substrate will contain a network of adjacent target portions that are patterned in sequence.

[0005] During the lithographic process, measurements of the structures produced need to be made frequently, for example for process control and verification. Various tools for making these measurements are known, including scanning electron microscopes, which are often used to measure critical dimensions (CDs), and dedicated tools for measuring overlay (the accuracy of alignment of two layers in a device). Recently, various forms of scatterometers have been developed for use in the lithography field. These devices direct a radiation beam onto a target and measure one or more properties of the scattered radiation - such as the intensity at a single reflection angle as a function of wavelength; the intensity at one or more wavelengths as a function of reflection angle; or the polarization as a function of reflection angle - to obtain a diffraction "spectrum" from which the properties of interest of the target can be determined.

[0006] There is a desire to model the operation of a lithographic system or apparatus (or systems in general). This can include monitoring the parameter values of the lithographic system and making predictions of future performance or events based on these parameter values using a model of the system operation. The disclosure herein describes many proposals for solving problems related to such predictive maintenance of lithographic systems, or systems in general. Summary of the invention

[0007] The present invention provides, in a first aspect, a method for adjusting a prediction model related to at least one specific configuration of a manufacturing apparatus, the method comprising: obtaining a function that includes at least a first function of a first prediction model parameter associated with the at least one specific configuration, and a second function of the first prediction model parameter and a second prediction model parameter associated with a configuration other than the at least one specific configuration of the manufacturing apparatus and / or a related apparatus; obtaining a value of the first prediction model parameter based on optimization of the function; and adjusting the prediction model according to the value of the first prediction model parameter to obtain an adjusted prediction model.

[0008] The present invention also provides a computer program product comprising machine-readable instructions for causing a processor to execute the method of the first aspect.

[0009] Additional features and advantages of the present invention, as well as the structure and operation of various embodiments of the present invention, are described in detail below with reference to the accompanying drawings. It should be noted that the present invention is not limited to the specific embodiments described herein. Such embodiments are presented herein for illustrative purposes only. Based on the teachings contained herein, additional embodiments will be apparent to those skilled in the art. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying schematic drawings, in which corresponding reference numerals indicate corresponding parts, and in which;

[0011] Figure 1 a lithographic apparatus is depicted;

[0012] Figure 2 a lithography cell or cluster is depicted, in which an inspection apparatus according to the present invention can be used; and

[0013] Figure 3 is a flowchart depicting a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0014] Before describing embodiments of the present invention in detail, it is instructive to present an example environment in which embodiments of the present invention can be implemented.

[0015] Figure 1Schematically depicts a lithographic apparatus LA. The lithographic apparatus includes: an illumination system (illuminator) IL configured to condition a radiation beam B (e.g., EUV radiation or DUV radiation); a patterning device support or support structure (e.g., a mask table) MT configured to support a patterning device (e.g., a mask) MA and connected to a first positioner PM configured to accurately position the patterning device according to certain parameters; two substrate tables (e.g., wafer tables) WTa and WTb, each substrate table configured to hold a substrate (e.g., a wafer coated with resist) W and each substrate table connected to a second positioner PW configured to accurately position the substrate according to certain parameters; and a projection system (e.g., a refractive projection lens system) PS configured to project a pattern imparted to the radiation beam B by the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W. A reference frame RF connects the various components and serves as a reference or reference object for setting and measuring the positions of the patterning device and the substrate, and the positions of features on the patterning device and the substrate.

[0016] The illumination system may include various types of optical components for guiding, shaping, or controlling the radiation, such as refractive, reflective, magnetic, electromagnetic, electrostatic, or other types of optical components, or any combination thereof.

[0017] The patterning device MT holds the patterning device in a manner that depends on the orientation of the patterning device, the design of the lithographic apparatus, and other conditions such as whether the patterning device is held in a vacuum environment. The patterning device support can take many forms; the patterning device support can ensure that the patterning device (e.g., relative to the projection system) is in a desired position.

[0018] The term "patterning device" as used herein should be construed broadly to mean any device that can be used to impart a pattern to a cross-section of a radiation beam so as to create a pattern in a target portion of the substrate. It should be noted that the pattern imparted to the radiation beam may not exactly correspond to the desired pattern in the target portion of the substrate (e.g., if the pattern includes phase-shifting features or so-called assist features). Generally, the pattern imparted to the radiation beam will correspond to a particular functional layer in a device (such as an integrated circuit) created in the target portion.

[0019] As depicted herein, the apparatus belongs to the transmissive type (e.g., employing a transmissive patterning device). Alternatively, the apparatus may belong to the reflective type (e.g., using a programmable mirror array of the type mentioned above, or using a reflective mask). Examples of patterning devices include masks, programmable mirror arrays, and programmable LCD (liquid crystal display) panels. Any term "reticle" or "mask" used herein may be considered synonymous with the more general term "patterning device". The term "patterning device" may also be construed to mean a device that stores pattern information in digital form for controlling such a programmable patterning device.

[0020] The term "projection system" as used herein should be broadly construed to include any type of projection system, including refractive, reflective, catadioptric, magnetic, electromagnetic, and electrostatic optical systems or any combination thereof, as appropriate for the exposure radiation used or for other factors such as the use of immersion liquid or the use of a vacuum. Any term "projection lens" used herein may be considered synonymous with the more general term "projection system".

[0021] The lithographic apparatus may also be of the type in which at least a portion of the substrate is also covered by a liquid having a relatively high refractive index (e.g., water), so as to fill the space between the projection system and the substrate. The immersion liquid may also be applied to other spaces in the lithographic apparatus, such as the space between the mask and the projection system. It is well known in the art that immersion techniques are used to increase the numerical aperture of the projection system.

[0022] In operation, the illuminator IL receives a beam of radiation from the radiation source SO. The source and the lithographic apparatus may be separate entities, such as when the source is an excimer laser. In such a case, the source is not considered to be part of the lithographic apparatus, and the beam of radiation is transmitted from the source SO to the illuminator IL by means of a beam delivery system BD comprising, for example, suitable directing mirrors and / or beam expanders. In other cases, the source may be an integral part of the lithographic apparatus (e.g., when the source is a mercury lamp). The source SO, the illuminator IL, and, if required, the beam delivery system BD may together be referred to as the radiation system.

[0023] The illuminator IL may include, for example, an adjuster AD for adjusting the angular intensity distribution of the beam of radiation, an integrator IN, and a condenser CO. The illuminator may be used to adjust the beam of radiation so as to have a desired uniformity and intensity distribution in its cross-section.

[0024] The radiation beam B is incident on the patterning device MA held on the patterning device support MT, and is patterned by the patterning device. After having traversed the patterning device (e.g., mask) MA, the radiation beam B passes through the projection system PS, which focuses the beam onto a target portion C of the substrate W. By means of the second positioner PW and the position sensor IF (e.g., an interferometric device, a linear encoder, a 2D encoder or a capacitive sensor), the substrate table WTa or WTb can be moved precisely, for example in order to position different target portions C in the path of the radiation beam B. Similarly, for example after the mechanical retrieval from the mask library or during scanning, the first positioner PM and another position sensor ( Figure 1 not explicitly shown in) can be used to position the patterning device (e.g., reticle / mask) MA accurately relative to the path of the radiation beam B.

[0025] The patterning device (e.g., reticle / mask) MA and the substrate W can be aligned by using mask alignment marks M1, M2 and substrate alignment marks P1, P2. Although the illustrated substrate alignment marks occupy dedicated target portions, they can be located in multiple target portions (the spaces between these are called scribe alignment marks). Similarly, in the case where more than one die is provided on the patterning device (e.g., mask) MA, the mask alignment marks can be located between the dies. Small alignment marks can also be included within the die, between device features, in which case it is desirable for the identification to be as small as possible and not to require any imaging or process conditions different from adjacent features. An alignment system for detecting the alignment marks is described further below.

[0026] The depicted apparatus can be used in various modes. In the scanning mode, the patterning device support (e.g., mask table) MT and the substrate table WT are scanned synchronously while the pattern to be imparted to the radiation beam is projected onto the target portion C (i.e., single dynamic exposure). The speed and direction of the substrate table WT relative to the patterning device support (e.g., mask table) MT can be determined by the magnification (reduction ratio) and image inversion characteristics of the projection system PS. In the scanning mode, the maximum size of the exposure field limits the width (along the non-scanning direction) of the target portion in a single dynamic exposure, while the length of the scanning movement determines the height (along the scanning direction) of the target portion C. As is known in the art, other types of lithographic apparatus and operating modes are possible. For example, the step mode is known. In so-called “maskless” lithography, the programmable patterning device is kept fixed, but has a changing pattern, and the substrate table WT is moved or scanned.

[0027] Combinations and / or variations of the usage patterns described above may also be employed, or completely different usage patterns.

[0028] The lithographic apparatus LA belongs to the so-called dual platform type, which has two substrate tables WTa, WTb, and two stations - an exposure station EXP and a measurement station MEA - between which the substrate tables can be exchanged. When a substrate on one substrate table is being exposed at the exposure station, another substrate can be loaded onto the other substrate table at the measurement station, and various preparatory steps can be performed. This enables a significant increase in the throughput of the apparatus. The preparatory steps may include mapping or profiling the surface height of the substrate using a leveling sensor LS and measuring the position of alignment marks on the substrate using an alignment sensor AS. If a position sensor IF is not able to measure the position of the substrate table while the substrate table is at the measurement station and at the exposure station, a second position sensor may be provided to enable tracking of the position of the substrate table relative to a reference frame RF at both stations. Instead of the dual platform arrangement shown, other arrangements are known and available. For example, other lithographic apparatuses in which a substrate table and a measurement table are provided are known. These substrate tables and measurement tables dock together when performing preparatory measurements and then do not dock when the substrate table undergoes exposure.

[0029] As Figure 2 shown, the lithographic apparatus LA forms part of a lithographic cell LC (sometimes also referred to as a lithocell or cluster), which also includes equipment for performing pre-exposure processes and post-exposure processes on a substrate. Conventionally, this equipment includes a spin coater SC for depositing a resist layer, a developer DE for developing the exposed resist, a chill plate CH, and a bake plate BK. A substrate handling device or robot RO picks up substrates from input / output ports I / O1, I / O2, moves the substrates between different process equipment, and then transfers the substrates to the feed table LB of the lithographic apparatus. These devices, often collectively referred to as a track or a track and coat system, are under the control of a track control unit TCU, which itself is controlled by a management control system SCS that also controls the lithographic apparatus via a lithography control unit LACU. Thus, the different equipment can be operated to maximize throughput and processing efficiency.

[0030] In order to expose the substrate correctly and consistently by the lithographic apparatus, it is necessary to inspect the exposed substrate to measure properties such as overlay errors between subsequent layers, line thickness, critical dimension (CD), etc. Therefore, a manufacturing facility in which a lithography cell LC is located also includes a metrology system MET that receives some or all of the substrates W that have been processed in the lithography cell. The metrology results are provided directly or indirectly to a supervisory control system SCS. If an error is detected, the exposure of subsequent substrates can be adjusted, especially if the inspection can be completed quickly enough such that other substrates of the same batch are still to be exposed. Additionally, an already exposed substrate can be stripped and reworked to improve the yield, or discarded, thereby avoiding further processing of a substrate known to be defective. In the case where only some target portions of the substrate are defective, further exposure can be performed only on those target portions that are good.

[0031] Within the metrology system MET, inspection equipment is used to determine the properties of the substrate, and specifically, how the properties of different substrates or different layers of the same substrate vary between different layers. The inspection equipment can be integrated into the lithography apparatus LA or the lithography cell LC, or can be a separate device. To enable the fastest measurements, it is necessary to have the inspection equipment measure the properties in the exposed resist layer immediately after exposure. However, the latent image in the resist has a very low contrast - there is only a very small refractive index difference between the portion of the resist that has been exposed to radiation and the portion that has not - and not all inspection equipment has sufficient sensitivity to make useful measurements of the latent image. Therefore, the measurement can be made after a post-exposure bake step (PEB), which is typically the first step performed on the exposed substrate and increases the contrast between the exposed and unexposed portions of the resist. At this stage, the image in the resist can be referred to as a semi-latent image. It is also possible to measure the developed resist image - at this time, either the exposed or unexposed portion of the resist has been removed - or to measure the developed resist image after a pattern transfer step such as etching. The latter possibility limits the possibility of reworking a defective substrate, but can still provide useful information.

[0032] Computer modeling techniques can be used to predict, correct, optimize, and / or verify the performance of a system. Such techniques can monitor system data, such as one or more physical quantities and / or machine settings, and predict, correct, optimize, and / or verify system performance based on the values of this system data. Historical data can be used to construct the computer model, and the computer model can be continuously updated, improved, or monitored by comparing the prediction results of the model parameter values with the corresponding system data. In particular, such computer modeling techniques can be used to predict, correct, optimize, and / or verify the system performance of a lithography system or process; and more specifically, can enable predictive maintenance of such a lithography system or process.

[0033] In order for the model to correctly represent a complex system (such as a lithography system), it is required that there be sufficient data available for the model to adequately train, configure, and / or tune the model. However, lithography systems are not ordinary equipment, and as such, the number of operational lithography systems from which data can be obtained is limited. Additionally, these systems are constantly evolving via technological advancements, resulting in system heterogeneity. There are many different types of systems (different platforms or production lines), and even when systems are similar (e.g., the same platform or production line), there may be minor configuration differences (e.g., component differences, software differences) depending on the exact specifications, age, and / or maintenance history. Thus, the datasets associated with each of these subsystems are inherently limited: while the various datasets associated with different subsystems have some commonalities, each dataset has its own specific characteristics that reflect differences or minor variations across systems, as well as drift and evolution over time.

[0034] Sub-component replacement and component maintenance are typically done on a reactive basis, such that only faulty components are investigated and repaired / replaced when a problem is noticed, which inevitably results in undesirable downtime and thus costs in terms of performance and yield. Additionally, undesirable interruptions are prone to cascading, which can disrupt other parties along the manufacturing and supply chain. As an alternative to reactive maintenance, other strategies can include:

[0035] · Periodic maintenance. Components / sub-components are replaced at regular intervals to avoid undesirable machine downtime. This can be wasteful as the machine may be offline and components are replaced at an unnecessary frequency, or if not frequent enough, may not prevent unplanned downtime and repairs.

[0036] · Optimized scheduled maintenance. Components are replaced or overhauled together, which can be even more wasteful in terms of unnecessary component replacement.

[0037] · Transfer entropy assisted root cause analysis. The use of transfer entropy causal methods for root cause analysis of faulty components has been described in US 2018 / 0267523A1, which is incorporated herein by reference. However, the causal graph only informs how the effects propagate within the system; it is not helpful for heterogeneity / isomerism in the data, and thus the amount of available data per configuration is small.

[0038] · Predictive maintenance. This includes building many models that predict the behavior of different elements of the system. In many (e.g., consumer) industrial applications, the systems and components are often mass-produced, providing a large amount of data sources on which the predictive models can be built. However, such a large amount of data is not applicable to lithography systems.

[0039] To provide a specific example, fault prediction (e.g., for a specific sub-component) can be based on online scanner data, provided that the faulty sub-component will produce different patterns in the data compared to normal behavior. For example, the cross-validated metrics on the training set may be good, e.g., where the ROC AUC (Receiver Operating Characteristic - Area Under the Curve) is higher than 0.90; while the same model evaluated based on test or real-time data may show much worse performance (ROC AUC around 0.60). Further investigation may attribute this performance degradation to changes across multiple data sets of key parameters, such as the scanner type or sub-component version. Such configuration changes between and within data sets affect the observed patterns and capture the commonalities between multiple individual instances; however, they are statistically associated with the sub-component faults. However, these parameters cannot be used as predictors, as due to the small sample size, they will not be sufficient to perfectly separate the fault instances from the non-fault instances; the predictive performance of such models will most likely not generalize to new / real-time data.

[0040] Thus, the prediction accuracy for (imminent) faults for a specific configuration is affected by the low data volume. The alternative of simply aggregating all data in all configurations into a cumulative data set will easily lead to overfitting / underfitting problems (failure to achieve model generalization).

[0041] To illustrate this, assume there are data sets D1 to D m , where each of these data sets D i corresponds to a different specific configuration (e.g., different combinations of sub-components and system types). Assume the model corresponding to data set D i includes parameters ω i . In a general scenario, the cost function F(D i , ω i ) can be optimized to be based on data set D ito train the model parameters ω i , for example:

[0042]

[0043] This method solves each optimization problem independently and separately, without taking advantage of any similarities between different datasets. As already mentioned, it is possible to aggregate data from all datasets into a "super" dataset, but the variations across different datasets will be ignored.

[0044] To address this issue, a computational framework and an associated method are provided for predictive maintenance purposes (for any production equipment or tool, such as a lithography tool), wherein the predictive model includes a first component optimized for a population of tools with different configurations, and a second component based on knowledge / behavior of the configurations relative to each other.

[0045] The computational framework and the associated method can use a cost function for the model parameters of a predictive maintenance model, the predictive maintenance model being based on datasets from multiple configurations while taking advantage of the similarities between these datasets. To ensure that configuration-specific behavior is reflected in the solution, the proposed cost function includes a configuration-specific component (the first component) and a regularization term (the second component) that accounts for the coupling of the configuration-specific model parameters. This regularization term can be based on domain knowledge, for example, due to one or more common physical principles that impose similar responses of the scanner to component states.

[0046] Figure 3 is a flowchart depicting a method according to an embodiment. The method includes: at step 300, obtaining a cost function that includes at least a first function of a first predictive model parameter associated with a specific configuration, and a second function (or regularization function) of the first predictive model parameter and a second predictive model parameter associated with other configurations. At step 310, obtaining a value of the first predictive model parameter based on the optimization of the function. At step 320, adjusting the predictive model according to the value of the first predictive model parameter to obtain an adjusted predictive model. Finally, at step 330, using the adjusted predictive model to predict a characteristic (e.g., failure probability or time or other maintenance-related characteristic) of the specific configuration.

[0047] Generally, the cost function to be optimized can (by way of example) take the following form:

[0048] , where the regularization term G(·) involves or links different optimization problems for each dataset, and λ is a weight that can act as a Lagrange multiplier. Thus, the regularization term G(·) can regularize the variation of each dataset based on the similarity of the solutions. In this way, each dataset can still have a separate solution, but now the variation across multiple solutions can be controlled by the regularization term G(·). The weight λ enables simulated annealing during the search for the global minimum or controls the importance of global regularization relative to the contribution of individual models.

[0049] A particular implementation of G(·) may not explicitly depend on all hyperparameters This can be the case where the j-th hyperparameter of the i-th model is not regularized across multiple datasets D1,…,D m or is not regularized at all.

[0050] The definition of the regularization term G(·) can involve domain expertise to encode prior knowledge. In one example, L1 norm regularization or L2 norm regularization is applied. In another example, the regularization takes into account the known (desired) correlation strength between multiple model parameters across multiple configurations. Other examples can use the Bayesian inference framework. In the Bayesian framework, the function G is the logarithm of the prior. The function F is the logarithm of the likelihood function. In the Bayesian framework instead of optimizing the cost function, samples are taken from the posterior corresponding to the prior and the likelihood. In other words, exp(F + G) can be calculated and normalized, which gives the Bayesian joint probability. Other regularization methods can be based on transfer entropy techniques (e.g., any of the transfer entropy causal methods described in US 2018 / 0267523), domain-driven hierarchies, and general complexity control priors. Some of these examples and other examples will now be described in more detail for illustrative purposes.

[0051] In a more straightforward example, the regularization term G(·) can include L2 norm regularization of the solution at some nominal value (e.g., it can encode prior knowledge); for example:

[0052]

[0053] Thus, is the empirical average of the solution; i.e., the aggregated solution at a higher level of the hierarchy.

[0054] In another example, it applies to the case where different datasets have relationships specified by an undirected graph with a Laplacian operator of the graph Then:

[0055]

[0056] Regularized solution with respect to the graph. This function is equal to:

[0057]

[0058] where i~j indicates that nodes i and j are neighbors. W is the adjacency matrix corresponding to the undirected graph. For more details on the general concept, reference can be made to the following publication incorporated by reference: Learning theory and kernel machines by AJ Smola, R Kondor, 2003–Springer.

[0059] The concepts described above can be extended to a hierarchical Bayesian framework for specific choices of F and G. For example, domain knowledge combined with an existing model can provide an initial fault prediction function for new components / configurations (or for one where very little data is available), where this domain knowledge or the confidence inferred therefrom is encoded as a prior. Newly added (e.g., fault) data can be used to update the prior; thus the fault prediction model is gradually improved (updated over time) using the Bayesian inference framework.

[0060] In an embodiment, the effect of the Bayesian framework can be to mix information between different datasets / configurations in order to compensate for configurations that do not have associated data or have very little associated data. For example, if two or more datasets show a common trend, then it can be inferred (with a certain confidence level / degree of confidence based on factors such as how many configurations also show this trend, the strength of the trend, etc.) that the new configuration will also show this trend. Thus, the Bayesian framework can include a common optimization across all datasets such that any common trend will be reflected in the initial fault prediction function for configurations that do not have data or have very little data. Thus, an existing model for an existing component / configuration, together with domain knowledge for a new (e.g., related) product / component, can be used as a prior for a new model for predicting the behavior (e.g., failure time or interruption) of the new component. As a specific example, a model built based on data for a first type of optical element can be used as a prior for a fault prediction model for a second generation type of optical element for which no data is available. When data for this new component becomes available, the model can be updated.

[0061] Thus, a second model (such as a hierarchical Bayesian (HB) model) including a hierarchy / tier having two layers is proposed. The lower layer models multiple individual datasets, and the upper layer models the statistical information of the datasets, which includes the similarity of the data and the variation across multiple datasets. Other model types (whether hierarchical or not) can also be used for the second model to train the first model based on data corresponding to other configurations.

[0062] In other embodiments where both F and G are smooth, the model can be optimized by gradient-based optimization. Alternatively, G(·) can also include non-smooth functions such as the L1 norm that promotes sparsity. In such cases, proximal splitting methods can be used for optimization. Proximal splitting methods are described, for example, in NG Polson, JG Scott, BT Willard's Proximal algorithms in statistics and machine learning (Statistical Science, 2015, projecteuclid.org), which is incorporated herein by reference.

[0063] In cases where different parties are reluctant to share data, the proposed method described herein enables joint learning, where each party (user) only needs to share their model weights (model parameter ω i ), rather than the actual data, for regularization.

[0064] Thus, using the regularization method described herein, domain knowledge can be considered as a prior, co-training of the model can be performed for relevant sub-components, such that multiple individual features of each component are retained while considering generality, and / or co-training of the model can be performed for subgroups, such that multiple individual features of each subgroup (e.g., related to different platforms) are retained while considering generality.

[0065] Additional embodiments are disclosed in the list of numbered aspects below:

[0066] 1. A method of adjusting a prediction model related to at least one specific configuration of a manufacturing apparatus, the method comprising:

[0067] Obtaining a function that includes at least a first function of a first prediction model parameter associated with the at least one specific configuration, and a second function of the first prediction model parameter and a second prediction model parameter associated with a configuration of the manufacturing apparatus and / or a related apparatus other than the at least one specific configuration;

[0068] Obtaining a value of the first prediction model parameter based on optimization of the function; and

[0069] Adjusting the prediction model according to the value of the first prediction model parameter to obtain an adjusted prediction model.

[0070] 2. The method according to aspect 1, comprising: using the adjusted prediction model to probabilistically predict characteristics of any configuration.

[0071] 3. The method according to aspect 2, wherein the characteristic includes the probability of failure of a component associated with the at least one specific configuration.

[0072] 4. The method according to aspect 1, 2 or 3, wherein the step of obtaining the value of the first prediction model parameter includes obtaining the value by Bayesian inference.

[0073] 5. The method according to any one of the preceding aspects, including: obtaining a second model; and using the second model to adjust or train the prediction model based on data corresponding to other configurations showing a common trend or characteristic.

[0074] 6. The method according to aspect 5, wherein the second model includes a hierarchical model.

[0075] 7. The method according to aspect 6, wherein the second model includes a hierarchical Bayesian model.

[0076] 8. The method according to aspect 7, wherein the hierarchical Bayesian model includes a lower layer and an upper layer, the lower layer models individual data sets corresponding to the other configurations, and the upper layer models the statistical information of the data sets.

[0077] 9. The method according to any one of aspect 7 or 8, wherein the hierarchical Bayesian model is used to update the prediction model when data of the at least one specific configuration becomes available.

[0078] 10. The method according to any one of the preceding aspects, wherein the first function includes a model-specific cost function, the model-specific cost function is a cost function specific to the at least one specific configuration, and the second function includes a regularization function.

[0079] 11. The method according to aspect 8, wherein the optimization includes jointly minimizing the first function and the second function according to the first prediction model parameter and the second prediction model parameter to obtain the value of the first prediction model parameter.

[0080] 12. The method according to any one of aspect 8 or 9, wherein the second function includes an L2 norm.

[0081] 13. The method according to any one of aspect 8 or 9, wherein the second function includes an L1 norm and the optimization includes a proximal splitting method.

[0082] 14. The method according to any one of aspect 8 or 9, wherein the corresponding data sets related to each of the configurations have relationships specified by an undirected graph with a graph Laplacian operator, and the second function is based on the Laplacian operator.

[0083] 15. The method according to aspect 12, wherein the second function is approximated by summing over neighboring nodes of an adjacency matrix corresponding to the Laplacian operator.

[0084] 16. The method according to any one of the preceding aspects, wherein the configuration relates to a particular tool, device or a component thereof, or any combination of tools, devices and components thereof.

[0085] 17. The method according to aspect 14, wherein the tool or device comprises a lithography tool or device used in the manufacture of an integrated circuit.

[0086] 18. A computer program comprising program instructions that are operative to perform the method according to any one of the preceding aspects when run on a suitable device.

[0087] 19. A non-transitory computer program carrier comprising the computer program according to aspect 18.

[0088] 20. A processing device that is operative to run the computer program according to aspect 18.

[0089] As used herein, the terms "radiation" and "beam" encompass all types of electromagnetic radiation, including ultraviolet (UV) radiation (e.g., having a wavelength of or about 365 nm, 355 nm, 248 nm, 193 nm, 157 nm or 126 nm) and extreme ultraviolet (EUV) radiation (e.g., having a wavelength in the range of 5 nm to 20 nm), as well as particle beams, such as ion beams or electron beams.

[0090] The term "lens" can, where the context permits, refer to any one or a combination of various types of optical components, including refractive, reflective, magnetic, electromagnetic and electrostatic optical components.

[0091] The foregoing description of the specific embodiments will so fully reveal the general nature of the invention that others can, by applying knowledge within the scope of the art, readily modify and / or adapt such specific embodiments for various applications without undue experimentation, without departing from the general concept of the invention. Therefore, based on the teachings and guidance given herein, these changes and modifications are intended to fall within the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that, for example, the words or terms herein are for the purpose of description and not limitation, such that the terminology or wording of this specification should be interpreted by those skilled in the relevant art in light of the teachings and guidance herein.

[0092] The breadth and scope of the present invention should not be limited by any of the above exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents in terms of aspects.

Claims

1. A method for adjusting a prediction model related to at least one specific configuration of a manufacturing device, the method comprising: Obtaining functions, the functions including at least a first function of a first prediction model parameter associated with the at least one specific configuration, and a second function of the first prediction model parameter and a second prediction model parameter of the manufacturing device and / or related devices associated with configurations other than the at least one specific configuration; Obtaining a value of the first prediction model parameter based on optimization of the functions; And Adjusting the prediction model according to the value of the first prediction model parameter to obtain an adjusted prediction model.

2. The method according to claim 1, comprising: Using the adjusted prediction model to predict the characteristics of any configuration probabilistically.

3. The method according to claim 2, wherein, The characteristics include the probability of failure of components associated with the at least one specific configuration.

4. The method according to claim 1, 2 or 3, wherein The step of obtaining the value of the first prediction model parameter includes obtaining the value by Bayesian inference.

5. The method according to claim 1, comprising: Obtaining a second model; And using the second model to adjust or train the prediction model based on data corresponding to other configurations showing common trends or characteristics.

6. The method according to claim 5, wherein, The second model includes a hierarchical model.

7. The method according to claim 6, wherein, The second model includes a hierarchical Bayesian model.

8. The method according to claim 7, wherein The hierarchical Bayesian model includes a lower layer and an upper layer, the lower layer modeling individual data sets corresponding to the other configurations, and the upper layer modeling the statistical information of the data sets.

9. The method according to any one of claims 7 or 8, wherein The hierarchical Bayesian model is used to update the prediction model when data of the at least one specific configuration becomes available.

10. The method according to claim 1, wherein The first function includes a model-specific cost function, the model-specific cost function being a cost function specific to the at least one specific configuration, and the second function includes a regularization function.

11. The method according to claim 8, wherein The optimization includes jointly minimizing the first function and the second function according to the first prediction model parameter and the second prediction model parameter to obtain the value of the first prediction model parameter.

12. The method according to claim 8 or 9, wherein, The second function includes an L2 norm.

13. The method according to claim 8 or 9, wherein The second function includes an L1 norm and the optimization includes a proximal splitting method.

14. A computer program comprising program instructions that can operate to execute the method according to claim 1 when running on a suitable device.

15. A non-transitory computer program carrier comprising the computer program according to claim 14.

Citation Information

Patent Citations

  • Methods of modelling systems or performing predictive maintenance of lithographic systems

    US20180267523A1

  • Model-based registration and critical dimension metrology

    CN104662543A

  • Integrated lithography method and lithography system

    CN110244523A