Optimization of manufacturing process

Through machine learning algorithms and statistical reasoning technology, modeling and acquisition functions are used to optimize the manufacturing process, the problem of time-consuming and resource-consuming process development in the existing technology is solved, and the process parameter values ​​that meet multiple target specifications are quickly identified.

CN119998740APending Publication Date: 2025-05-13LAM RES CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070296.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-28
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art requires multiple trial and error iterations in process development to identify process parameter values ​​of manufacturing processes, resulting in huge time and resource consumption, especially when it is necessary to comply with multiple target specifications.

Method used

Using machine learning algorithms and statistical inference technology, the next set of process parameter values ​​is selected by modeling the manufacturing process and using the acquisition function to optimize based on the statistical uncertainty of the prediction and the difference in the target specification.

Benefits of technology

It effectively reduces the time and resource consumption of process development, quickly identify the best process parameter values ​​that meet multiple target specifications, and improves the efficiency and accuracy of process optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998740A_ABST
    Figure CN119998740A_ABST
Patent Text Reader

Abstract

Methods, systems, and media for optimizing manufacturing processes are provided. In some implementations, a method of automatically optimizing a manufacturing process includes: (a) providing a first set of process parameter values associated with a first experiment to a model representing the manufacturing process; (b) characterizing the statistical uncertainty of the predictions made by the model; (c) selecting a second set of process parameter values using an acquisition function, wherein the acquisition function identifies the second set of process parameters based on both: (i) predicting a difference between a wafer characteristic and a target specification; and (ii) statistical uncertainty; (d) receiving a result of the manufacturing process performed using the second set of process parameter values; and (e) determining whether the performance of the manufacturing process results in a processed wafer having wafer characteristics that conform to the target specification.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporated by Reference The PCT application form is filed concurrently with this specification as a part of this application. Each application identified in the concurrently filed PCT application form to which this application claims the benefit or priority is incorporated herein by reference in its entirety. Background Art

[0001] Process development (e.g., identifying process parameter values ​​used in a given manufacturing process to achieve target wafer specifications) can be a time-consuming and resource-consuming process. For example, multiple trial and error iterations manually directed by process engineers may be required to identify process parameter values ​​to achieve target specifications.

[0002] The background description provided here is for the purpose of generally presenting the context of the present disclosure. The work of the presently designated inventors is neither explicitly nor implicitly admitted to be prior art against the present disclosure to the extent that it is described in this background section and to aspects of the specification that could not be determined as prior art at the time of filing the application. Summary of the invention

[0003] Systems, apparatus, methods, and media for optimizing manufacturing processes are provided.

[0004] In some implementations, a method for automatically optimizing a manufacturing process includes: (a) providing a first set of process parameter values ​​associated with a first experiment to one or more models representing the manufacturing process to obtain model results, wherein the model results associate a set of candidate process parameter values ​​with corresponding wafer characteristic data, wherein the manufacturing process is an etching process or a deposition process; (b) using the obtained model results to characterize statistical uncertainties of predictions made by the one or more models representing the manufacturing process; and (c) using an acquisition function to select a second set of process parameter values ​​associated with a second experiment, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of candidate process parameter values ​​associated with the first experiment using the model results; and (ii) a set of candidate process parameter values ​​associated with the second experiment using the model results. (i) determining, based on the received results of the manufacturing process, whether the performance of the manufacturing process using the second set of process parameter values ​​produced a processed wafer having one or more wafer characteristics that meet the target specification.

[0005] In some examples, the method further includes: (f) repeating (b)-(e) until a processed wafer meeting the target specification is produced using the manufacturing process.

[0006] In some examples, the target specification is a user-specified target specification.

[0007] In some examples, the acquisition function is a user selected acquisition function. In some examples, the acquisition function is configured to probabilistically quantify the improvement any set of process parameter values ​​will make in approaching the target specification compared to a baseline experiment.

[0008] In some examples, the method also includes receiving a user-selected hyperparameter used by the acquisition function to select the second set of process parameter values, the user-selected hyperparameter specifying a balance between exploration and exploitation.

[0009] In some examples, the target specification includes a plurality of specifications to be achieved in a processed wafer undergoing the manufacturing process.

[0010] In some examples, the one or more models representing the manufacturing process include at least one user-selected model.

[0011] In some examples, the one or more models representing the manufacturing process include physics-based models.

[0012] In some examples, the one or more models representing the manufacturing process include one or more of: a neural network, a Gaussian process model, a decision tree model, a regression model, or any combination thereof.

[0013] In some examples, characterizing the statistical uncertainty of the predictions made by the one or more models includes determining a predicted posterior distribution based on a set of measurement data provided to the one or more models, the predicted posterior distribution indicating a probability distribution of predicted wafer characteristics for a given set of process parameter values. In some examples, for an n-dimensional representation of a process parameter space, a first region of the n-dimensional representation of the process parameter space is associated with a greater statistical uncertainty than a second region of the n-dimensional representation of the process parameter space, and wherein the one or more models have received less experimental data obtained using process parameter values ​​associated with the first region. In some examples, the second region of the n-dimensional representation of the process parameter space is associated with process parameter values ​​that, when utilized by the manufacturing process, produce a processed wafer having wafer characteristics within a predetermined threshold of the target specification. In some examples, the acquisition function is configured to determine whether to select the second set of process parameter values ​​from the first region or the second region. In some examples, the n-dimensional representation of the process parameter space is substantially unbounded for at least one dimension.

[0014] According to some embodiments, a method for automatically optimizing a manufacturing process is provided. The method may include: (a) receiving a plurality of target wafer specifications to be achieved by a manufacturing process; (b) providing a first set of process parameters associated with a first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (c) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; and (d) using an acquisition function to select a second set of process parameter values ​​associated with a second experiment, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points representing differences between a plurality of predicted wafer characteristics associated with the performance of the manufacturing process using the second set of process parameter values ​​and the plurality of target wafer specifications; and (ii) the statistical uncertainty of the predictions made by the one or more models.

[0015] In some examples, the set of points improves on-wafer performance relative to the plurality of target wafer specifications.

[0016] In some examples, the set of points includes a Pareto front. In some examples, the acquisition function determines an expected improvement in a supervolume formed by the Pareto front using the plurality of predicted wafer characteristics related to the performance of the manufacturing process using the second set of process parameter values.

[0017] In some examples, at least two process parameter values ​​in the second set of process parameter values ​​are different from corresponding process parameter values ​​in the first set of process parameter values.

[0018] In some examples, the method further includes determining that the manufacturing process cannot meet at least one target wafer specification of the plurality of target wafer specifications based at least in part on the set of points. In some examples, the method further includes identifying a second plurality of predicted wafer characteristics within a predetermined error threshold of the plurality of target wafer specifications, wherein the second plurality of predicted wafer characteristics are associated with performance of the manufacturing process using a third set of process parameter values.

[0019] In some implementations, a computer program product is provided that includes a non-transitory computer-readable medium on which computer-executable instructions are provided for causing a computer system to perform a method of automatically optimizing a manufacturing process. The method may include: (a) providing a first set of process parameter values ​​associated with a first experiment to one or more models representing a manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data, wherein the manufacturing process is an etching process or a deposition process; (b) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; (c) using an acquisition function to select a second set of process parameter values ​​associated with a second experiment, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a difference between predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and a target specification, wherein the predicted wafer characteristics are generated by the one or more models representing the manufacturing process using the second set of process parameter values; and (ii) the statistical uncertainty of the predictions made by the one or more models; (d) receiving results of the manufacturing process performed using the second set of process parameter values ​​associated with the second experiment; and (e) determining, based on the received results of the manufacturing process, whether the performance of the manufacturing process using the second set of process parameter values ​​produces a processed wafer having one or more wafer characteristics that meet the target specification.

[0020] In some examples, the method further includes: (f) repeating (b)-(e) until a processed wafer meeting the target specification is produced using the manufacturing process.

[0021] In some examples, the target specification is a user-specified target specification.

[0022] In some examples, the acquisition function is a user selected acquisition function. In some examples, the acquisition function is configured to probabilistically quantify the improvement any set of process parameter values ​​will make in approaching the target specification compared to a baseline experiment.

[0023] In some examples, the method also includes receiving a user-selected hyperparameter used by the acquisition function to select the second set of process parameter values, the user-selected hyperparameter specifying a balance between exploration and exploitation.

[0024] In some examples, the target specification includes a plurality of specifications to be achieved in a processed wafer undergoing the manufacturing process.

[0025] In some examples, the one or more models representing the manufacturing process include at least one user-selected model.

[0026] In some examples, the one or more models representing the manufacturing process include physics-based models.

[0027] In some examples, the one or more models representing the manufacturing process include one or more of: a neural network, a Gaussian process model, a decision tree model, a regression model, or any combination thereof.

[0028] In some examples, characterizing the statistical uncertainty of the predictions made by the one or more models includes determining a predicted posterior distribution based on a set of measurement data provided to the one or more models, the predicted posterior distribution indicating a probability distribution of predicted wafer characteristics for a given set of process parameter values. In some examples, for an n-dimensional representation of a process parameter space, a first region of the n-dimensional representation of the process parameter space is associated with a greater statistical uncertainty than a second region of the n-dimensional representation of the process parameter space, and wherein the one or more models have received less experimental data obtained using process parameter values ​​associated with the first region. In some examples, the second region of the n-dimensional representation of the process parameter space is associated with process parameter values ​​that, when utilized by the manufacturing process, produce a processed wafer having wafer characteristics within a predetermined threshold of the target specification. In some examples, the acquisition function is configured to determine whether to select the second set of process parameter values ​​from the first region or the second region. In some examples, the n-dimensional representation of the process parameter space is substantially unbounded for at least one dimension.

[0029] According to some embodiments, a computer program product is provided that includes a non-transitory computer-readable medium on which computer executable instructions are provided for causing a computer system to perform a method for automatically optimizing a manufacturing process. The method may include: (a) receiving a plurality of target wafer specifications to be achieved by a manufacturing process; (b) providing a first set of process parameters associated with a first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (c) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; and (d) using an acquisition function to select a second set of process parameter values ​​associated with a second experiment, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points representing differences between a plurality of predicted wafer characteristics associated with the performance of the manufacturing process using the second set of process parameter values ​​and the plurality of target wafer specifications; and (ii) the statistical uncertainty of the predictions made by the one or more models.

[0030] In some examples, the set of points improves on-wafer performance relative to the plurality of target wafer specifications.

[0031] In some examples, the set of points includes a Pareto front. In some examples, the acquisition function determines an expected improvement in a supervolume formed by the Pareto front using the plurality of predicted wafer characteristics related to the performance of the manufacturing process using the second set of process parameter values.

[0032] In some examples, at least two process parameter values ​​in the second set of process parameter values ​​are different from corresponding process parameter values ​​in the first set of process parameter values.

[0033] In some examples, the method further includes determining that the manufacturing process cannot meet at least one target wafer specification of the plurality of target wafer specifications based at least in part on the set of points. In some examples, the method further includes identifying a second plurality of predicted wafer characteristics within a predetermined error threshold of the plurality of target wafer specifications, wherein the second plurality of predicted wafer characteristics are associated with performance of the manufacturing process using a third set of process parameter values. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Depicted is an exemplary system for optimizing a manufacturing process according to some embodiments.

[0035] Figure 2 is a flow chart of an exemplary process for optimizing a manufacturing process, according to some embodiments.

[0036] Figures 3A-3D is a graph indicating the evolution of selected process parameters during several iterations of an optimization process, according to some embodiments.

[0037] Figure 4 is a flow chart of an exemplary process for optimizing a manufacturing process, for an etching process and / or a deposition process, according to some embodiments.

[0038] Figure 5 is a flow chart of an exemplary process for optimizing a manufacturing process to achieve multiple target specifications, according to some embodiments.

[0039] Figure 6 An exemplary computer system is presented that can be used to implement certain embodiments described herein. DETAILED DESCRIPTION

[0040] In the following description, several specific details are set forth to provide a thorough understanding of the presented embodiments. The disclosed embodiments may be implemented without some or all of these specific details. In other instances, well-known processing operations are not described in detail to avoid unnecessarily obscuring the disclosed embodiments. Although the disclosed embodiments will be described in conjunction with specific embodiments, it should be understood that these specific embodiments are not intended to limit the disclosed embodiments.

[0041] In process development, a process engineer may often need to identify process parameter values ​​for a given manufacturing process to achieve a target (e.g., customer-specified) specification or a set of target specifications. Conventionally, this process may be time and labor intensive. For example, typically, a process engineer may have to perform several experiments to identify the best process parameter values ​​that meet the specifications. Since there may be many process parameters (e.g., temperature parameters, gas flow pressure parameters, gas species parameters, gas flow rate parameters, etc.), the process may require multiple (e.g., tens, hundreds, thousands, etc.) trial and error iterations, especially when multiple specifications need to be met. For example, a change in the value of a process parameter may have an effect (e.g., a favorable effect) on a first target specification and a different effect (e.g., a negative effect) on a second target specification. Therefore, because multiple process parameters can be adjusted, each of which has a different effect on each target specification, manual process development is expensive in terms of time and resources.

[0042] Disclosed herein is a technique for optimizing process development. In particular, the techniques described herein utilize machine learning algorithms and statistical (e.g., Bayesian) reasoning to guide process development. A process model (sometimes referred to herein as a "surrogate model") can be used to model or simulate a given manufacturing process (e.g., an etching process, a deposition process, etc.). Since there may not be enough experimental data (e.g., experimental data obtained via a previous manufacturing process) to limit the process model, there may be statistical model uncertainty in the prediction of the process model. Model uncertainty may be caused by a lack of training data. The techniques disclosed herein use statistical reasoning techniques to optimally select a set of process parameter values ​​to be used in the next experiment. The best selected set of process parameter values ​​is selected based on an acquisition function, which quantitatively considers modeling epistemic uncertainty, and trades off between improving the current best wafer feature set for the target specification, further limiting the statistical uncertainty associated with the underlying process model (e.g., surrogate model). By iteratively selecting the next process parameter value to try in subsequent experiments, an optimal set of process parameter values ​​may be efficiently identified even in the presence of multiple competing target specifications and / or in the presence of many process parameters that may be adjusted.

[0043] Furthermore, as described in more detail below, the acquisition function can allow exploration of a substantially unbounded process parameter space. In this way, process parameter values ​​that differ significantly from previously tried process parameter values ​​can be tried in some experiments (e.g., based on the likelihood that such values ​​will improve the current best processed wafer characteristics close to the target specification). The acquisition function can effectively balance the exploration of previously untested areas of parameter space and the utilization of known areas of parameter space that may improve on-wafer performance.

[0044] Figure 1102 is an example of a system for optimizing a manufacturing process according to some embodiments. As shown, the optimization system 102 may be configured to iteratively determine experiments to be performed by the manufacturing tool 106, wherein each experiment is selected based on a model of the manufacturing process to be performed by the manufacturing tool 106 and used to achieve the target specification 104. In other words, the optimization system 102 may identify an experimental set of process parameters to be used in conjunction with the manufacturing process (e.g., an etching process and / or a deposition process) performed by the manufacturing tool 106, wherein each iterative set of process parameters identified by the optimization system 102 is selected to iteratively approach the target specification 104. In some embodiments, the target specification 104 may indicate various specifications of features of a wafer that is processed by the manufacturing tool 106 using the manufacturing process, the various specifications including, but not limited to, geometric measurements, such as a target etch depth, a target deposition thickness, a target aspect ratio, a target sidewall angle, a target sidewall depth, a target critical dimension, a target spacing, electrical specifications such as resistance and capacitance, and the like. In some implementations, goal specifications 104 may include one or more (eg, two, three, five, ten, etc.) goal specifications, as will be discussed in greater detail below.

[0045] In some implementations, the optimization system 102 may include a surrogate model 110, a reasoning engine 112, and a design of experiment (DoE) engine 114. The surrogate model 110 may be one or more models that model the manufacturing process to be performed by the manufacturing tool 106 to meet the target specification 104. The surrogate model 110 may include any suitable type of model, such as a neural network, a Bayesian neural network, a regression model, a Gaussian process model, a physics-based model, a determination tree, and / or any combination thereof. The surrogate model 110 may be configured to consider historical data 108, which may include metrology results of wafers previously processed (e.g., processed by the manufacturing tool 106) and corresponding process parameters that produced the processed wafers with the metrology results. In some embodiments, the surrogate model 110 may be trained using the historical data 108. In some implementations, the surrogate model 110 may be more than one model (e.g., two models, three models, ten models, etc.), or a collection of multiple models. In some such embodiments, the multiple models may be of the same type (eg, neural network, regression model, physics-based model, etc.) or of different types.

[0046] The output of the surrogate model 110 may be used by the inference engine 112 to constrain the model data. For example, the inference engine 112 may generate information indicating the statistical uncertainty of the output of the surrogate model 110. As a more specific example, as described below in conjunction with Figure 2As described, the inference engine 112 may generate a predicted posterior distribution of data based on model outputs generated by the surrogate model 110 and / or the historical data 108 .

[0047] The DoE engine 114 can utilize the output of the reasoning engine 112 (e.g., the statistical uncertainty information generated by the reasoning engine 112) to generate a next experiment to be performed by the manufacturing tool 106. In some implementations, the next experiment can specify a set of process parameter values ​​to be used in the next experiment. In some embodiments, the DoE engine 114 can select the next process parameter values ​​associated with the next experiment to be performed so that when performed as a manufacturing process by the manufacturing tool 106, the results of the next experiment will produce a processed wafer with the following wafer characteristics: wafer characteristics that are closer to the target specification 104 than the current best process parameters. In other words, the DoE engine 114 can iteratively identify experimental process parameter values ​​that iteratively identify process parameters so that the processed wafer characteristics are close to the target specification 104. Additionally or alternatively, in some implementations, the process parameter values ​​associated with the next experiment can be those used to further constrain the statistical uncertainty estimates generated by the reasoning engine 112. For example, the DoE engine 114 may identify a next experiment in a given iteration that, when executed, will reduce statistical uncertainty, even if the experiment does not produce a processed wafer having wafer characteristics that are closer to the target specification 104 than those produced by the previous experiment. In some implementations, the DoE engine 114 may use a hyperparameter that adjusts explore versus exploit to balance the tradeoff between reducing statistical uncertainty and identifying parameter values ​​that produce wafer characteristics within an acceptable margin of the target specification. For example, in some iterations, the DoE engine 114 may prioritize exploration and may identify a next experiment that may reduce statistical uncertainty. Conversely, in some iterations, the DoE engine 114 may prioritize exploitation and may identify a next experiment that is associated with a parameter value in a low statistical uncertainty region toward the target specification. In some embodiments, the DoE engine 114 may utilize an acquisition function to select a process parameter associated with the next experiment. In some implementations, the hyperparameter of the acquisition function may be used to adjust the balance between exploration and exploitation.

[0048] In some embodiments, the techniques described herein may utilize an alternative model to evaluate a first set of process parameter values. Data generated by the model associated with historical information (e.g., metrology information collected from previously processed parameters) may be used to constrain the data-based model. Then, the statistical uncertainty of the model may be determined, for example in the form of a predicted posterior distribution. The predicted posterior distribution may indicate the likelihood of achieving a specific set of wafer characteristics given a set of process parameter values ​​and underlying data (e.g., underlying data generated by the model and / or historical data). Note that the predicted posterior distribution indicates the certainty associated with a given prediction. The acquisition function may then be used to select the next set of process parameter values ​​to be evaluated (e.g., in the next experiment). As described above, the acquisition function may effectively balance the trade-off between exploration and exploitation, taking into account the statistical uncertainty associated with the model. For example, the acquisition function may select the next set of process parameter values ​​from an area with relatively high uncertainty to further explore the process parameter space. Conversely, in some cases, the acquisition function may select the next set of process parameter values ​​from an area with relatively low uncertainty so that the process parameter values ​​may be driven toward a specification closer to the target wafer. Note that in some implementations, the acquisition function may switch between exploration and exploitation on an iteration-by-iteration basis.

[0049] Figure 2 is a flow chart of an exemplary process 200 for optimizing a manufacturing process according to some embodiments. In some implementations, the blocks of process 200 may be executed on a server device and / or a controller device of a manufacturing tool. In some embodiments, the blocks of process 200 may be executed on a server device and / or a controller device of a manufacturing tool. Figure 2 In some implementations, two or more blocks of process 200 may be performed substantially in parallel. In some embodiments, one or more blocks of process 200 may be omitted.

[0050] Process 200 may begin by receiving historical data at block 202. In some implementations, the historical data may include metrology results associated with previously processed wafers. The historical data may be related to a specific manufacturing process and / or a specific manufacturing tool or class of manufacturing tools. Note that the metrology results may include in-situ and / or ex-situ results. Examples of metrology techniques that may be used to provide historical data include electron microscopy (EM), transmission electron microscopy (TEM), scanning electron microscopy (SEM), critical dimension SEM (CD-SEM), etc. In some embodiments, the historical data may include metrology results paired with process parameter values ​​that produced a given metrology result, thereby allowing the process parameter values ​​to be mapped to the resulting post-processing wafer characteristics.

[0051] At 204, the process 200 may receive an initial set of process parameter values ​​and a target specification. The target specification may indicate a user-specified (e.g., customer-specified) specification to which the processed wafer is to conform. The target specification may include one or more wafer feature specifications, such as a target etch depth, a target deposition thickness, a target sidewall thickness, a target aspect ratio, and the like. The initial set of process parameter values ​​may be selected in any suitable manner. For example, in some embodiments, the initial set of process parameter values ​​may be user-specified. As another example, in some embodiments, the initial set of process parameter values ​​may be randomly selected, or randomly selected from a predetermined range.

[0052] At 206, process 200 may receive user selections of alternative models and inference methods. Exemplary types of alternative models that may be utilized include regression models, neural networks, physically based numerical simulation models, and the like. Exemplary types of inference methods that may be used include variational inference, Markov chain Monte Carlo (MCMC), local Gaussian approximation (LGA), and the like. In some embodiments, alternative models and inference methods may be combined in a model family, such as a Gaussian process (GP) model, a tree Parzen estimator (TPE) model, and the like. User selection of alternative models and / or inference methods may be via a user interface, via values ​​set in a configuration file, and the like. Note that in some implementations, alternative models and inference methods may be fixed or hard-coded. In this case, box 206 may be omitted.

[0053] At 208, the process 200 can use the historical data, the surrogate model, and the inference method to constrain the posterior model by applying the initial set of process parameter values ​​to the surrogate model and using the inference method to constrain the output of the surrogate model. The posterior model can represent the uncertainty in the surrogate model. In some embodiments, the following techniques can be used to constrain the posterior model. Given the initial set of process parameter values ​​and other recipe set points K, the model parameters θ of the surrogate model, the predicted set W of wafer characteristics (which are predicted using the surrogate model with the initial set of process parameter values) can be expressed as Given a set of process parameter values ​​K and surrogate model parameters θ, the probability distribution representing the likelihood of a wafer characteristic W, denoted herein as p(W|K,θ), can be determined based on the distributions of various wafer characteristics. For example, the probability distribution can be determined as follows:

[0054] Continuing with this example, the restricted posterior model, denoted in this paper as p(θ,D) (where D represents the historical data), can be determined as follows:

[0055] In the equation given above, L(D|θ) represents the likelihood of a particular data set D given the surrogate model parameters θ. The restricted posterior model p(θ|D) represents the distribution of the model parameters based on the historical data D.

[0056] At 210, the process 200 may determine a predicted posterior distribution, generally denoted herein as p(W|K,D). The predicted posterior distribution represents the probability of achieving wafer characteristic W using a manufacturing process with process parameter value K given historical data D. The predicted posterior distribution may be determined based on a probability distribution of wafer characteristics given alternative model parameters and process parameter values, denoted as p(W|K,θ), and based on a restricted model, denoted as p(θ|D). For example, in some embodiments, the predicted posterior distribution may be determined as follows: p(W|K,D)=∫p(W|K,θ)p(θ|D)dθ

[0057] It should be understood that the predicted posterior distribution can incorporate the statistical uncertainty in the historical data D as well as the statistical uncertainty associated with the predictions produced by the alternative model.

[0058] At 212, the process 200 may receive a user selection of an acquisition function. The acquisition function may be a function that uses statistical uncertainty in the surrogate model predictions (e.g., using a prediction posterior distribution) to select a next set of process parameter values ​​for a next experiment. As described above, the acquisition function may balance the trade-off between exploration of regions of the process parameter space associated with higher uncertainty and utilization of regions of the process parameter space associated with lower uncertainty. Examples of acquisition functions include a probability of improvement acquisition function (where the next set of process parameter values ​​is selected that has the highest probability of improving wafer characteristics toward a target specification relative to the current best wafer characteristics) and an expected improvement acquisition function (where the magnitude of the improvement of the next set of process parameter values ​​relative to the current best wafer characteristics is considered). In some implementations, the user selection of the acquisition function may be performed via a user interface, via a configuration file, etc. Note that in some implementations, the acquisition function may be hard-coded. In this case, block 212 may be omitted.

[0059] At 214, the process 200 may use the acquisition function to identify a next set of process parameter values, e.g., process parameter values ​​for the next experiment. The definition of the acquisition function may include a utility function, which may quantify the improvement of the on-wafer performance W relative to the current best result W in achieving the desired goal. +For example, when the acquisition function is an expected improvement (EI) acquisition function, the utility function is defined as the increase in the objective function value f(W) relative to the current best observed process result f(W) + ), and can be determined as follows: U EI (W) = max[f(W)-f(W + ),0]

[0060] In some implementations, the acquisition function may use the predicted posterior distribution to determine the next set of process parameter values. For example, using the expected improvement (EI) acquisition function, the next set of process parameter values ​​K' may be determined as follows:

[0061] In the equation given above, K' represents the next set of process parameter values. The process parameter value K'=argmax[EI(K')] that maximizes the acquisition function is selected as a set of process parameter values ​​used in the next experiment. Statistically, it can be expected that the next set of process parameter values ​​K' will produce the greatest improvement in terms of on-wafer performance targets. Following similar ideas, other acquisition functions can be calculated in the system, such as improvement probability, lower confidence bound, etc. As shown in the above equation, in some implementation schemes, the next set of process parameter values ​​can be determined by considering the predicted wafer characteristics that are improved relative to the current best predicted wafer characteristics (for example, relative to the target specification), and by using the predicted posterior distribution to determine the process parameter value that may maximize the improvement of the predicted wafer characteristics, wherein the predicted posterior distribution indicates the statistical uncertainty of the underlying surrogate model.

[0062] Note that while the examples given above are about expected improvement acquisition functions, similar equations can be used for other acquisition functions, such as probability of improvement acquisition functions. Additionally, in some embodiments, various hyperparameters, such as hyperparameters for controlling the tradeoff between exploration and exploitation, can be included in the acquisition function. In some embodiments, such hyperparameters can be adjusted by user input or user selection.

[0063] At 216, the process 200 may perform the manufacturing process represented by the alternative model using the next set of process parameter values. For example, in some implementations, the process 200 may cause instructions to be transmitted to a controller associated with a manufacturing tool, wherein the instructions indicate the next set of process parameter values. The controller may then use the manufacturing tool and perform the manufacturing process on the wafer using the next set of process parameter values. Note that in some implementations, the historical data may be augmented with data obtained using the results of the manufacturing process (e.g., metrology data).

[0064] At 218, the process 200 may use the next set of process parameter values ​​to determine whether the target specification received at block 204 has been met for the wafer processed at block 216. If at 218, the process 200 determines that the target specification has been met ("yes" at 218), the process 200 may end. Conversely, if at 218, the process 200 determines that the target specification has not been met ("no" at 218), the process 200 may loop back to block 208 and may use metrology data associated with the wafer processed at 216 to further constrain the a posteriori model.

[0065] In some embodiments, the process 200 may loop through blocks 208-218 until a target specification is processed. By executing each manufacturing process using a set of process parameter values ​​(which are selected using an acquisition function that is in turn based on the statistical uncertainty of the underlying surrogate model), the target specification may be met with fewer iterations than if each set of process parameter values ​​were manually selected by, for example, a process engineer. In particular, the acquisition function can quickly hone in on a desired region of parameter space, and then identify a set of optimized process parameter values ​​within the desired region by controlling the tradeoff between exploration and exploitation relative to the inherent uncertainty of the surrogate model and the uncertainty of the experimental data. This, in turn, may allow for more efficient use of manufacturing resources by identifying optimal process parameter values ​​using fewer test wafers.

[0066] As above combined Figure 1 and Figure 2 As described, experimental process parameter values ​​may be selected, wherein each set of process parameter values ​​for a given iteration is selected based on predictions of a surrogate model and based on statistical uncertainties associated with the predictions of the surrogate model. As described above, the next set of process parameter values ​​may be selected based on an acquisition function that takes into account the uncertainties associated with the predictions of the surrogate model. Over a series of iterations, the process parameter values ​​may evolve such that wafers processed using the process parameter values ​​achieve (or approach) a target specification.

[0067] Figures 3A-3D The evolution of process parameter values ​​over a series of iterations according to some embodiments is shown. Figure 3AReferring to panel 302a, at iteration 0, the surrogate model utilizes an initial data set (e.g., including data point 310). Referring to panel 304a, the surrogate model generates predictions 312 to fit the data. The target specification is represented by line 301. Referring to panel 306a, the surrogate model's predictions 312 are associated with uncertainty 314. Uncertainty 314 represents the statistical uncertainty associated with prediction 312. In some embodiments, uncertainty 314 can be represented by a prediction posterior distribution, as described above in conjunction with Figure 2 As described above. Note that the data including data point 310 is clustered in a region of parameter space K between -1 and 1. Thus, there is relatively less uncertainty in the region of parameter space K where the data points are clustered, and there is relatively more uncertainty in the region of parameter space K where data has not yet been collected. Referring to panel 308a, an acquisition function 316 is shown, which is a function of parameter space K. Note that the highest value of the acquisition function is about -0.4, which is close to the parameter value (in K space) that yields the best wafer characteristic data to date (e.g., associated with data point 310).

[0068] Steering Figure 3B , a similar graph for Iteration 2 is shown. Note that more data points have been added, such as experimental results acquired from Iteration 0 and Iteration 1, as shown in panels 302b, 306b, and 304b. Note that the acquisition function 318 shown in panel 308b has two major peaks at Iteration 2: the first peak is centered in K-space in the region where data has been collected, and the second peak is centered in K-space in the region where data has not yet been collected, and therefore, there is a relatively high uncertainty in the prediction of the surrogate model. Based on the maximum value of the acquisition function 318, the next set of process parameter values ​​in K-space is selected for Iteration 3.

[0069] Steering Figure 3C , similar screens 302c, 304c, 306c, and 308c are shown. Note that in iteration 3, since the acquisition function after iteration 2 indicates that the next set of process parameter values ​​should be in a region of K space that has not been explored (e.g., K is greater than 0.5), the process parameter values ​​evaluated in iteration 3 are from this unexplored region of K space. As shown in screen 306c, the process parameter values ​​tested in iteration 3 significantly reduce the uncertainty associated with the basic surrogate model. In addition, the process parameter values ​​from iteration 3 produce wafer characteristics that are closer to the target specification 301 relative to the process parameter values ​​used in iterations 0-2.

[0070] Steering Figure 3D, similar screens 302d, 304d, 306d, and 308d are shown for iteration 10. Note that after ten iterations, the uncertainty associated with the surrogate model has been greatly reduced in all regions of K space, as shown in screen 306d. Also, note that a process parameter value K that meets the target specification 301 has been identified.

[0071] It should be noted that Figures 3A-3D As shown, there is essentially no limit to the parameter space from which the acquisition function can select the next process parameter value for the next experiment. In other words, even in the case where all available historical data is clustered in a particular region of K-space, the acquisition function can select process parameter values ​​outside of that region of K-space based on: process parameter values ​​that can reduce statistical uncertainty and / or process parameter values ​​that can result in wafer characteristics that are closer to the target specification. Furthermore, it should be understood that although Figures 3A-3D The target specification is represented as a line (e.g., a single specification to be met), but this is for illustration only. The technique can be applied to a set of target specifications (e.g., multiple target specifications), as described below in conjunction with Figure 5 described.

[0072] Figure 4 is a flow chart of an exemplary process 400 for iteratively identifying process parameter values ​​to achieve a target specification according to some embodiments. In some implementations, the blocks of process 400 may be executed on a server device and / or a controller device associated with a manufacturing tool. In some embodiments, the blocks of process 400 may be executed in accordance with Figure 4 In some embodiments, two or more blocks of process 400 may be performed in a different order than shown. In some embodiments, two or more blocks of process 400 may be performed substantially in parallel. In some embodiments, one or more blocks of process 400 may be omitted.

[0073] Process 400 may begin at 402 by providing a first set of process parameter values ​​associated with a first experiment to one or more models representing a manufacturing process to obtain model results that associate a set of candidate process parameter values ​​with corresponding wafer characteristic data. Figure 1 -3, one or more models may be one or more alternative models representing a manufacturing process. In some embodiments, the manufacturing process may be an etching process, a deposition process, etc. One or more models may be associated with a shared model family, such as a neural network, a regression model, a physics-based model, a determination tree, a Gaussian process model, etc. The first set of process parameter values ​​may be identified based on historical data (e.g., based on a previous manufacturing process), a manual specification by a process engineer, randomly identified from a range of possible values, and the like.

[0074] At 404, process 400 can characterize statistical uncertainty, which is the statistical uncertainty of predictions made by one or more models representing the manufacturing process. Figure 2 The statistical uncertainty of the predictions made by one or more models can be a predicted posterior distribution. The statistical uncertainty can be determined at least in part based on historical data (e.g., metrology data obtained from previous experiments or previous manufacturing process operations). More detailed exemplary techniques for determining statistical uncertainty are described above in conjunction with Figure 2 Display and description.

[0075] At 406, process 400 may select a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a difference between predicted wafer characteristics and target specifications using the second set of process parameter values; and (ii) a statistical uncertainty in the predictions made by one or more models. In some implementations, the acquisition function may be a probability of improvement acquisition function, or an expected improvement acquisition function, as described above in conjunction with Figure 2 In some implementations, the acquisition function may be user specified, for example, via a user interface or configuration file. Note that by selecting the second set of process parameter values ​​based on both the difference between the currently predicted wafer characteristics and the target specification to be met, and the statistical uncertainty of the underlying model, the second set of process parameter values ​​may be selected in a manner that balances improvement toward the target specification and exploration of the process parameter space that has not been experimentally tested. A more detailed example of a manner in which the acquisition function is used to balance the following is provided above in conjunction with Figures 3A-3D Display and describe: exploration of untested process parameter space and exploitation of previously tested process parameter space.

[0076] At 408, process 400 may perform a manufacturing process using a second set of process parameter values ​​associated with a second experiment. For example, in some implementations, a server executing the acquisition function may send instructions to a controller device of a manufacturing tool, where the instructions indicate a second set of process parameter values ​​to be used in the second experiment. The controller may then cause the manufacturing tool to perform the second experiment (e.g., a manufacturing process using the second set of process parameter values) on the wafer.

[0077] At 410, the process 400 can determine whether the target specification is met. In other words, the process 400 can determine whether the processed wafer from the second experiment meets the target specification.

[0078] If at 410, the process 400 determines that the target specification has been met ("yes" at 410), the process 400 can end. Conversely, if at 410, the process 400 determines that the target specification has not been met ("no" at 410), the process 400 can loop back to 404 and can update the statistical uncertainty of the predictions made by the one or more models using, for example, metrology results associated with the wafers processed at block 408. In some implementations, the process 400 can loop through blocks 404-410 until an experiment (e.g., a manufacturing process) performed using the process parameter values ​​meets the target specification. In some embodiments, the process 400 can end after a predetermined number of iterations, regardless of whether the target specification has been met.

[0079] In some embodiments, the techniques described herein can be used to identify process parameter values ​​that are used to meet multiple target wafer specifications rather than a single target specification. In this case, the acquisition function can identify the process parameter values ​​of the next experiment of multiple process parameter values, which may achieve wafer characteristics that are closer to the overall multiple target specifications. It should be understood that in some cases, not all target specifications can be met or realized. However, the techniques described herein can allow the identification of process parameter values ​​that balance the trade-offs between one or more target specifications that need to be met and other target specifications that are not met, thereby achieving the best solution. In other words, the techniques described herein can identify process parameter values ​​that best balance the trade-offs between multiple target specifications. In some cases, when not all target specifications need to be met, the best solution can involve wafer characteristics for each specification, which is within a predetermined threshold of the target of the specification. It should be noted that in some implementations, a set of points that best balance the trade-offs between multiple targets (e.g., multiple specifications to be met) can be called the Pareto Front. The Pareto Front is a hypersurface that separates the known achievable performance space from the unknown achievable performance space. The techniques described herein can iteratively propose a set of consecutive experiments that improve the Pareto front to maximize the process space in various metrology metrics. The process engineer can use the identified and / or characterized Pareto front to prioritize metrology specifications (e.g., target metrology specifications) to balance the trade-offs between multiple competing metrology specification targets for a given process constraint.

[0080] In some embodiments, by identifying a set of non-dominated points, the techniques described herein can identify process parameter values ​​that are used to best balance the trade-offs between multiple target specifications. As used herein, "non-dominated" refers to a point (e.g., a set of process parameter values) that is as good as every other point (e.g., a set of process parameter values) in terms of satisfying multiple target specifications, and a point that is better than at least one other point (e.g., a set of process parameter values) for at least one of the multiple target specifications. In some embodiments, a hypervolume metric H that characterizes the volume formed by a set of non-dominated points can be determined. In some such embodiments, the acquisition function can take into account the improvement in the hypervolume metric H to select the next set of process parameter values. Note that larger hypervolume values ​​correspond to better solution sets. In the current Pareto front, the hypervolume can be expanded by adding experimental points of dominant points. The expansion can be characterized as a hypervolume improvement, as described below in conjunction with Figure 5 Described in more detail.

[0081] Figure 5 is a flow chart of an exemplary process 500 for identifying process parameter values ​​in a manner that balances tradeoffs between multiple target specifications, according to some embodiments. In some implementations, the blocks of process 500 may be performed by a server device and / or controller of a manufacturing tool. In some embodiments, the blocks of process 500 may be performed in accordance with Figure 5 In some implementations, two or more blocks of process 500 may be performed substantially in parallel. In some implementations, one or more blocks of process 500 may be omitted.

[0082] Process 500 may begin at 502 by obtaining a plurality of target wafer specifications to be achieved by a manufacturing process. The plurality of target wafer specifications may include a target etch depth, a target deposition thickness, a target sidewall angle, a target aspect ratio, etc. The target wafer specifications may be user specified.

[0083] At 504, process 500 may provide a first set of process parameter values ​​associated with a first experiment to one or more models to obtain model results that associate a set of candidate process parameter values ​​with corresponding wafer characteristic data. Block 504 may be combined with using the above Figure 4 The technique described in block 402 may be performed using techniques similar to those described in block 402 .

[0084] At 506, process 500 may characterize the statistical uncertainty of predictions made by one or more models (representing the manufacturing process). Block 506 may be used in conjunction with the above Figure 4 The technique described in block 404 may be performed using techniques similar to those described in block 404 .

[0085] At 508, the process 500 may use an acquisition function to select a second set of process parameter values ​​associated with a second experiment. In some implementations, the acquisition function may be selected so that the acquisition function is used to select a next experiment whose results may improve the current best process parameter values ​​in terms of meeting multiple target specifications. In some implementations, the acquisition function may be selected so that the acquisition function is used to select a next experiment whose results will further characterize the process limitations and / or the trade-offs between various target specifications. In such implementations, the acquisition function may be used for target vector estimation. In such implementations, the acquisition function may be a Pareto front-based acquisition function. In some embodiments, the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points that represent differences between predicted wafer characteristics (e.g., using the second set of process parameter values ​​applied to one or more alternative models); and (ii) statistical uncertainties in the predictions made by one or more models. Note that the set of points may be a set of non-dominated points, where each point in the set of points is at least as good as the other points with respect to multiple target wafer specifications and is better than at least one point with respect to at least one target wafer specification. In some embodiments, the set of points may be considered a Pareto front. In some embodiments, the second set of process parameter values ​​can be based on a predicted hypervolumetric improvement of wafer characteristics relative to the current Pareto front. For example, the second set of process parameter values ​​can be selected by maximizing the hypervolumetric improvement.

[0086] In some implementations, the quality of the volume of available process space formed by the set of points (e.g., by the points associated with the Pareto front) can be referred to as a hypervolume metric H. In general, a larger volume (e.g., a larger value of H) can correspond to a better set of solutions, e.g., consisting of process parameter values ​​that produce wafer specifications that are closer to the desired multiple target specifications. Note that as additional experiments are conducted with additional manufacturing processes, observations from the additional experiments can be added to the historical data considered by the alternative model. This can further increase the number of points that are currently dominant in the Pareto front, thereby increasing the hypervolume H.

[0087] In some embodiments, the acquisition function may select a second set of process parameter values ​​by considering the improvement in the hypervolume metric H of the volume formed by the set of points. The improvement in the hypervolume metric may generally be expressed as HVI, which is defined as in represents the measurement of the newly proposed experiment, represents the current Pareto frontier and represents a reference point in the metrology space. The reference point is usually chosen to be an existing process result that is dominated by other better process results. Therefore, represents the excess volume of the available process space currently enclosed by the Pareto front, and represents the improved hypervolume enclosed by the Pareto front of improvements from the new data y. The expected hypervolume improvement (denoted as EHVI) can be determined as follows: Where p(y) is the posterior predictive distribution function. In some implementations, the expected hypervolume improvement EHVI can be efficiently calculated by dividing the non-dominated space into multiple integral slices. In some embodiments, the number of integral slices can be selected to be as small as possible. The integral of the criterion can be calculated within each integral slice. In some embodiments, the value of the integral of the criterion can be the sum of its contributions to each integral slice. For example, the expected hypervolume improvement can be determined by: Where S d Represents the d-dimensional target space The integral slice in the current Pareto frontier The volume of the enclosed usable target space; represents the target space volume for improvement; d express The Lebesgue measure on , i.e. the width of the slice of available process space; Denotes the predicted posterior distribution. Since the acquisition function is defined for multiple targets, the predicted posterior distribution is a multidimensional distribution function over all targets, and the integral is over the entire multidimensional target space of the problem.

[0088] At 510, process 500 may perform a manufacturing process using a second set of process parameters associated with a second experiment. Block 510 may use the same process parameters as above. Figure 4 The technique described in block 408 may be performed using techniques similar to those described in block 408 .

[0089] At 512, the process 500 may determine whether a plurality of target specifications have been met. If at 512, the process 500 determines that a plurality of target specifications have been met ("yes" at 512), the process 500 may end. Conversely, if at 512, the process 500 determines that a plurality of target specifications have not been met ("no" at 512), the process 500 may loop back to 506, and the statistical uncertainty may be updated based on the experimental results (e.g., metrology results) associated with the performance of the second experiment. It should be noted that in some cases, a plurality of target specifications may not all be met. In this case, the process 500 may be configured to loop through blocks 506 to 512 until a different stopping criterion is reached. The different stopping criteria may include a predetermined number of iterations that have been performed, an improvement in wafer characteristics achieved between two consecutive experiments being less than a predetermined improvement threshold, and the like. Background of the Disclosed Computing Implementations

[0090] A system including a manufacturing tool as described herein may include logic for automated control of components.

[0091] The analysis logic can be designed and implemented in any of a variety of ways. For example, the logic can be implemented in hardware and / or software. Examples are given in the controller section of this article. Hardware-implemented control logic can be provided in any of a variety of forms, including hard-coded logic in a digital signal processor, an application-specific integrated circuit, and other devices with algorithms implemented as hardware. The analysis logic can also be implemented as software or firmware instructions configured to execute on a general-purpose processor. The system control software can be provided by "programming" in a computer-readable programming language.

[0092] The computer program code for controlling the processes in the process train can be written in any conventional computer readable programming language: for example, assembly language, C, C++, Pascal, Fortran or other languages. The compiled object code or script is executed by the processor to perform the tasks identified in the program. As also indicated, the program code can be hard-coded.

[0093] Integrated circuits used in the logic may include chips in the form of firmware storing program instructions, digital signal processors (DSPs), chips defined as application specific integrated circuits (ASICs), and / or one or more microprocessors, or microcontrollers that execute program instructions (e.g., software). Program instructions may be instructions delivered in the form of various individual settings (or program files) defining operating parameters for performing a particular analysis or image analysis application.

[0094] Figure 6 6 is a block diagram of an example of a computing device 600 suitable for implementing some embodiments of the present disclosure. For example, the device 600 may be suitable for implementing some or all of the functionality for optimizing the manufacturing processes described herein.

[0095] The computing device 600 may include a bus 602 that directly or indirectly couples the following devices: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, an input / output (I / O) port 612, an input / output component 614, a power supply 616, and one or more presentation components 618 (e.g., a display). In addition to the CPU 606 and the GPU 608, the computing device 600 may also include Figure 6 Additional logic devices not shown include, but are not limited to, an image signal processor (ISP), a digital signal processor (DSP), an ASIC, an FPGA, etc.

[0096] although Figure 6 The various blocks of are shown as being connected by wires via bus 602, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, presentation components 618 such as a display device may be considered I / O components 614 (e.g., if the display is a touch screen). As another example, CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device in addition to the memory of GPU 608, CPU 606, and / or other components). In other words, Figure 6 The computing devices described in the present disclosure are illustrative only. No distinction is made between categories such as "workstations," "servers," "laptops," "desktops," "tablets," "client devices," "mobile devices," "handheld devices," "electronic control units (ECUs)," "virtual reality systems," and / or other device or system types, as all of these are contemplated under Figure 6 within the range of computing devices.

[0097] The bus 602 may represent one or more buses, such as an address bus, a data bus, a control bus, or a combination thereof. The bus 602 may include one or more bus types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or other types of buses.

[0098] Memory 604 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by computing device 600. Computer-readable media may include volatile and non-volatile media, and removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and / or communication media.

[0099] Computer storage media may include volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 604 may store computer readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 600. As used herein, computer storage media itself does not include signals.

[0100] Communication media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery media. The term "modulated data signal" may refer to a signal having one or more of its characteristics that are set or changed in such a way as to encode information in the signal. By way of example and not limitation, communication media may include wired media such as a wired network or a direct wire connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above items should also be included within the scope of computer-readable media.

[0101] The CPU 606 may be configured to execute computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. The CPUs 606 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing multiple software threads simultaneously. The CPU 606 may include any type of processor and may include different types of processors, depending on the type of computing device 600 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor may be an ARM processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or auxiliary coprocessors (e.g., math coprocessors), the computing device 600 may also include one or more CPUs 606.

[0102] GPU 608 can be used by computing device 600 to render graphics (e.g., 3D graphics). GPU 608 may include many (e.g., tens, hundreds, or thousands) cores capable of processing many software threads simultaneously. GPU 608 may generate pixel data of an output image in response to a rendering command (e.g., a rendering command from CPU 606 received via a host interface). GPU 608 may include a graphics memory, such as a display memory, for storing pixel data. Display memory may be included as part of memory 604. GPU 608 may include two or more GPUs operating in parallel (e.g., via a link). When combined, each GPU 608 may generate pixel data of different parts of an output image or different output images (e.g., a first GPU for a first image, a second GPU for a second image). Each GPU may contain its own memory or may share memory with other GPUs.

[0103] In examples where computing device 600 does not include GPU 608 , CPU 606 may be used to render graphics.

[0104] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communications). The communication interface 610 may include components and functionality that enable communication over any of a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating over Ethernet), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0105] I / O ports 612 can enable computing device 600 to be logically coupled to other devices, including coupling to I / O components 614, presentation components 618, and / or other components, some of which can be built into (e.g., integrated into) computing device 600. Illustrative I / O components 614 include microphones, mice, keyboards, joysticks, tracking pads, satellite antennas, scanners, printers, wireless devices, etc. I / O components 614 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by users. In some cases, the input can be transmitted to appropriate network elements for further processing. NUI can implement any combination of speech recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) related to the display of computing device 600. Computing device 600 may include a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, a touch screen technology, and a combination of these, to perform gesture detection and recognition. Additionally, computing device 600 may include an accelerometer or gyroscope that enables detection of motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, computing device 600 may use the output of the accelerometer or gyroscope to render an immersive augmented reality or virtual reality.

[0106] The power supply 616 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 616 may provide power to the computing device 600 to enable the components of the computing device 600 to operate.

[0107] The presentation component 618 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 618 may receive data from other components (e.g., GPU 608, CPU 606, etc.) and output the data (e.g., as images, video, sound, etc.).

[0108] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (e.g., program modules) executed by a computer or other machine (e.g., a personal data assistant or other handheld device). In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The present disclosure may be practiced in a variety of system configurations, including in handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in distributed computing environments, where tasks are performed by remote processing devices that are linked through a communications network. Other Considerations

[0109] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context and the like dictate otherwise. For example, reference to "a unit" includes a combination of two or more such units. Unless otherwise indicated, the "or" conjunction is used in its proper sense as a Boolean logical operator, encompassing both selection of feature alternatives (A or B, where the selection of A is mutually exclusive with B) and selection of feature combinations (A or B, where both A and B are selected).

[0110] It should be understood that the phrases "for each <item> in one or more <items>," "each <item> in one or more <items>," and the like, if used herein, include both single-item groups and multi-item groups, i.e., the phrase "for...each" is used in the sense that it is used in a programming language to refer to each item in any group of items referenced. For example, if the group of items referenced is a single item, then "each" will refer only to that single item (although dictionary definitions of "each" often define the term to mean "each of two or more things"), and does not mean that there must be at least two of those items. Similarly, the terms "set" or "subset" by themselves should not be taken to necessarily cover multiple items - it should be understood that a set or subset may cover only one member or multiple members (unless the context dictates otherwise).

[0111] The use of ordinal indicators such as (a), (b), (c) ... etc., if any, in the present disclosure and claims should be understood not to convey any particular order or sequence, except where such order or sequence is explicitly indicated. For example, if there are three steps labeled (i), (ii), and (iii), it should be understood that unless otherwise indicated, these steps may be performed in any order (or even simultaneously, if not otherwise contraindicated). For example, if step (ii) involves processing of an element created in step (i), then step (ii) may be considered to occur at some point after step (i). Similarly, if step (i) involves processing of an element created in step (ii), the opposite should be understood. It should also be understood that the use of the ordinal indicator "first" herein, such as "the first item," should not be understood to implicitly or inherently imply that there must be a "second" instance, such as "the second item."

[0112] It may be claimed that various computing components including processors, memories, instructions, routines, modules, or other components may be "configured to" perform a task or tasks. In such contexts, the phrase "configured to" is used to refer to a structure by a component including a structure (e.g., stored instructions, circuitry, etc.) that performs a task or tasks during operation. Thus, even when a particular component is not necessarily currently operational (e.g., not in an on state), it may be said that a unit / circuit / component is configured to perform a task.

[0113] The components using the term "configured to" may relate to hardware - for example, circuits, memories storing program instructions executable to perform an operation. In addition, "configured to" may refer to a general structure (such as a general circuit) that is manipulated by software and / or firmware (such as an FPGA or a general processor executing software) so as to be able to operate in a manner to perform the task (multiple tasks). In addition, "configured to" may refer to one or more memories or memory elements storing computer-executable instructions for performing the task (multiple tasks). Such memory elements may include memory on a computer chip with processing logic. In some contexts, "configured to" may also include adapting a manufacturing process (such as a semiconductor manufacturing facility) to manufacture equipment (such as an integrated circuit) that is suitable for implementing or performing one or more tasks.

[0114] Although the above embodiments have been described in detail to a certain extent for the purpose of clear understanding, it is apparent that certain variations and modifications may be implemented within the scope of the appended claims. It should be noted that there are many alternative ways to implement the processes, systems, and equipment of the embodiments of the present invention. Therefore, the embodiments of the present invention should be considered illustrative rather than restrictive, and these embodiments are not limited to the details given here.

Claims

1. A method for automatically optimizing a manufacturing process, the method comprising: (a) providing a first set of process parameter values ​​associated with a first experiment to one or more models representing a manufacturing process to obtain model results, wherein the model results associate a set of candidate process parameter values ​​with corresponding wafer characteristic data, wherein the manufacturing process is an etching process or a deposition process; (b) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; (c) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a difference between predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and a target specification, wherein the predicted wafer characteristics are generated by the one or more models representing the manufacturing process using the second set of process parameter values; and (ii) the statistical uncertainty of the predictions made by the one or more models; (d) receiving results of the manufacturing process performed using the second set of process parameter values ​​associated with the second experiment; as well as (e) determining, based on the received results of the manufacturing process, whether the performance of the manufacturing process using the second set of process parameter values ​​produced a processed wafer having one or more wafer characteristics that meet the target specification.

2. The method according to claim 1, further comprising: (f) Repeating (b)-(e) until a processed wafer meeting the target specification is produced using the manufacturing process.

3. The method according to claim 1, wherein: The target specification is a user-specified target specification.

4. The method according to claim 1, wherein: The acquisition function is an acquisition function selected by a user.

5. The method according to claim 4, wherein: The acquisition function is configured to probabilistically quantify the improvement any set of process parameter values ​​will make in approaching the target specification compared to a baseline experiment. 6 . The method of claim 1 , further comprising receiving a user-selected hyperparameter used by the acquisition function to select the second set of process parameter values, the user-selected hyperparameter specifying a balance between exploration and exploitation.

7. The method according to any one of claims 1 to 6, wherein: The target specifications include a plurality of specifications to be achieved in a processed wafer undergoing the manufacturing process.

8. The method according to any one of claims 1 to 6, wherein: The one or more models representing the manufacturing process include at least one user-selected model.

9. The method according to any one of claims 1 to 6, wherein: The one or more models representing the manufacturing process include physics-based models.

10. The method according to any one of claims 1 to 6, wherein: The one or more models representing the manufacturing process include one or more of: a neural network, a Gaussian process model, a decision tree model, a regression model, or any combination thereof.

11. The method according to any one of claims 1 to 6, wherein: Characterizing the statistical uncertainty of the predictions made by the one or more models includes determining a predicted posterior distribution based on a set of measurement data provided to the one or more models, the predicted posterior distribution indicating a probability distribution of predicted wafer characteristics for a given set of process parameter values.

12. The method according to claim 11, wherein: For an n-dimensional representation of a process parameter space, a first region of the n-dimensional representation of the process parameter space is associated with a larger statistical uncertainty than a second region of the n-dimensional representation of the process parameter space, and wherein the one or more models have received less experimental data obtained using process parameter values ​​associated with the first region.

13. The method according to claim 12, wherein: The second region of the n-dimensional representation of the process parameter space is associated with process parameter values ​​that, when utilized by the manufacturing process, produce processed wafers having wafer characteristics within a predetermined threshold of the target specification.

14. The method according to claim 13, wherein: The acquisition function is configured to determine whether to select the second set of process parameter values ​​from the first region or the second region.

15. The method according to claim 12, wherein: The n-dimensional representation of the process parameter space is substantially unbounded for at least one dimension.

16. A method for automatically optimizing a manufacturing process, the method comprising: (a) receiving a plurality of target wafer specifications to be achieved by a manufacturing process; (b) providing a first set of process parameters associated with the first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (c) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; as well as (d) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points representing differences between a plurality of predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and the plurality of target wafer specifications; and (ii) the statistical uncertainty of the predictions made by the one or more models.

17. The method according to claim 16, wherein: The set of points improves on-wafer performance relative to the plurality of target wafer specifications.

18. The method according to claim 16, wherein: This set of points comprises the Pareto front.

19. The method according to claim 18, wherein: The acquisition function determines an expected improvement in a supervolume formed by the Pareto front using the plurality of predicted wafer characteristics related to the performance of the manufacturing process using the second set of process parameter values.

20. The method according to any one of claims 16 to 19, wherein: At least two process parameter values ​​in the second set of process parameter values ​​are different from corresponding process parameter values ​​in the first set of process parameter values.

21. The method of any of claims 16-19, further comprising determining, based at least in part on the set of points, that the manufacturing process cannot meet at least one target wafer specification of the plurality of target wafer specifications.

22. The method of claim 21 further comprising identifying a second plurality of predicted wafer characteristics within a predetermined error threshold of the plurality of target wafer specifications, wherein the second plurality of predicted wafer characteristics are associated with performance of the manufacturing process using a third set of process parameter values.

23. A computer program product comprising a non-transitory computer readable medium on which are provided computer executable instructions for causing a computer system to perform a method for automatically optimizing a manufacturing process, the method comprising: (a) providing a first set of process parameter values ​​associated with a first experiment to one or more models representing a manufacturing process to obtain model results, wherein the model results associate a set of candidate process parameter values ​​with corresponding wafer characteristic data, wherein the manufacturing process is an etching process or a deposition process; (b) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; (c) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a difference between predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and a target specification, wherein the predicted wafer characteristics are generated by the one or more models representing the manufacturing process using the second set of process parameter values; and (ii) the statistical uncertainty of the predictions made by the one or more models; (d) receiving results of the manufacturing process performed using the second set of process parameter values ​​associated with the second experiment; as well as (e) determining, based on the received results of the manufacturing process, whether the performance of the manufacturing process using the second set of process parameter values ​​produced a processed wafer having one or more wafer characteristics that meet the target specification.

24. The computer program product of claim 23, wherein the method further comprises: (f) Repeating (b)-(e) until a processed wafer meeting the target specification is produced using the manufacturing process.

25. The computer program product of claim 23, wherein: The target specification is a user-specified target specification.

26. The computer program product of claim 23, wherein: The acquisition function is an acquisition function selected by a user.

27. The computer program product of claim 26, wherein: The acquisition function is configured to probabilistically quantify the improvement any set of process parameter values ​​will make in approaching the target specification compared to a baseline experiment.

28. The computer program product of claim 23, the method further comprising receiving a user-selected hyperparameter used by the acquisition function to select the second set of process parameter values, the user-selected hyperparameter specifying a balance between exploration and exploitation.

29. The computer program product according to any one of claims 23 to 28, wherein: The target specifications include a plurality of specifications to be achieved in a processed wafer undergoing the manufacturing process.

30. The computer program product according to any one of claims 23 to 28, wherein: The one or more models representing the manufacturing process include at least one user-selected model.

31. A computer program product according to any one of claims 23 to 28, wherein: The one or more models representing the manufacturing process include physics-based models.

32. A computer program product according to any one of claims 23 to 28, wherein: The one or more models representing the manufacturing process include one or more of: a neural network, a Gaussian process model, a decision tree model, a regression model, or any combination thereof.

33. A computer program product according to any one of claims 23 to 28, wherein: Characterizing the statistical uncertainty of the predictions made by the one or more models includes determining a predicted posterior distribution based on a set of measurement data provided to the one or more models, the predicted posterior distribution indicating a probability distribution of predicted wafer characteristics for a given set of process parameter values.

34. The computer program product of claim 33, wherein: For an n-dimensional representation of a process parameter space, a first region of the n-dimensional representation of the process parameter space is associated with a larger statistical uncertainty than a second region of the n-dimensional representation of the process parameter space, and wherein the one or more models have received less experimental data obtained using process parameter values ​​associated with the first region.

35. The computer program product of claim 34, wherein: The second region of the n-dimensional representation of the process parameter space is associated with process parameter values ​​that, when utilized by the manufacturing process, produce processed wafers having wafer characteristics within a predetermined threshold of the target specification.

36. The computer program product of claim 35, wherein: The acquisition function is configured to determine whether to select the second set of process parameter values ​​from the first region or the second region.

37. The computer program product of claim 34, wherein: The n-dimensional representation of the process parameter space is substantially unbounded for at least one dimension.

38. A computer program product comprising a non-transitory computer readable medium on which are provided computer executable instructions for causing a computer system to perform a method of automatically optimizing a manufacturing process, the method comprising: (a) receiving a plurality of target wafer specifications to be achieved by a manufacturing process; (b) providing a first set of process parameters associated with the first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (c) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; as well as (d) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points representing differences between a plurality of predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and the plurality of target wafer specifications; and (ii) the statistical uncertainty of the predictions made by the one or more models.

39. The computer program product of claim 38, wherein: The set of points improves on-wafer performance relative to the plurality of target wafer specifications.

40. The computer program product of claim 38, wherein: This set of points comprises the Pareto front.

41. The computer program product of claim 40, wherein: The acquisition function determines an expected improvement in a supervolume formed by the Pareto front using the plurality of predicted wafer characteristics related to the performance of the manufacturing process using the second set of process parameter values.

42. A computer program product according to any one of claims 38 to 41, wherein: At least two process parameter values ​​in the second set of process parameter values ​​are different from corresponding process parameter values ​​in the first set of process parameter values.

43. A computer program product according to any one of claims 38 to 41, wherein: The method also includes determining, based at least in part on the set of points, that the manufacturing process cannot meet at least one target wafer specification of the plurality of target wafer specifications.

44. The computer program product of claim 43, wherein: The method also includes identifying a second plurality of predicted wafer characteristics within a predetermined error threshold of the plurality of target wafer specifications, wherein the second plurality of predicted wafer characteristics are associated with performance of the manufacturing process using a third set of process parameter values.

45. A system for automatically optimizing a manufacturing process, the system comprising: at least one fabrication chamber configured to perform a fabrication process, wherein the fabrication process is an etching process or a deposition process; as well as At least one controller configured to: (a) providing a first set of process parameter values ​​associated with a first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (b) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; (c) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a difference between predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and a target specification, wherein the predicted wafer characteristics are generated by the one or more models representing the manufacturing process using the second set of process parameter values; and (ii) the statistical uncertainty of the predictions made by the one or more models; (d) receiving results of the fabrication process performed using the second set of process parameter values ​​associated with the second experiment and using the fabrication chamber; as well as (e) determining, based on the received results of the manufacturing process, whether the performance of the manufacturing process using the second set of process parameter values ​​produced a processed wafer having one or more wafer characteristics that meet the target specification.

46. ​​The system of claim 45, wherein: The at least one controller is further configured to: (f) Repeating (b)-(e) until a processed wafer meeting the target specification is produced using the manufacturing process.

47. The system of claim 45, wherein: The target specification is a user-specified target specification.

48. The system of claim 45, wherein: The acquisition function is an acquisition function selected by a user.

49. The system of claim 48, wherein: The acquisition function is configured to probabilistically quantify the improvement any set of process parameter values ​​will make in approaching the target specification compared to a baseline experiment.

50. The system of claim 45, wherein: The controller is further configured to receive a user-selected hyperparameter used by the acquisition function to select the second set of process parameter values, the user-selected hyperparameter specifying a balance between exploration and exploitation.

51. The system according to any one of claims 45-50, wherein: The target specifications include a plurality of specifications to be achieved in a processed wafer undergoing the manufacturing process.

52. The system according to any one of claims 45-50, wherein: The one or more models representing the manufacturing process include at least one user-selected model.

53. The system according to any one of claims 45-50, wherein: The one or more models representing the manufacturing process include physics-based models.

54. The system according to any one of claims 45-50, wherein: The one or more models representing the manufacturing process include one or more of: a neural network, a Gaussian process model, a decision tree model, a regression model, or any combination thereof.

55. The system according to any one of claims 45-50, wherein: To characterize the statistical uncertainty of the predictions made by the one or more models, the at least one controller is configured to determine a predicted posterior distribution based on a set of measurement data provided to the one or more models, the predicted posterior distribution indicating a probability distribution of predicted wafer characteristics for a given set of process parameter values.

56. The system of claim 55, wherein: For an n-dimensional representation of a process parameter space, a first region of the n-dimensional representation of the process parameter space is associated with a larger statistical uncertainty than a second region of the n-dimensional representation of the process parameter space, and wherein the one or more models have received less experimental data obtained using process parameter values ​​associated with the first region.

57. The system of claim 56, wherein: The second region of the n-dimensional representation of the process parameter space is associated with process parameter values ​​that, when utilized by the manufacturing process, produce processed wafers having wafer characteristics within a predetermined threshold of the target specification.

58. The system of claim 57, wherein: The acquisition function is configured to determine whether to select the second set of process parameter values ​​from the first region or the second region.

59. The system of claim 56, wherein: The n-dimensional representation of the process parameter space is substantially unbounded for at least one dimension.

60. A system for automatically optimizing a manufacturing process, the system comprising: at least one fabrication chamber configured to perform a fabrication process; as well as At least one controller configured to: (a) receiving a plurality of target wafer specifications to be achieved by a manufacturing process; (b) providing a first set of process parameters associated with the first experiment to one or more models representing the manufacturing process to obtain model results, the model results associating a set of candidate process parameter values ​​with corresponding wafer characteristic data; (c) using the obtained model results to characterize the statistical uncertainty of predictions made by the one or more models representing the manufacturing process; as well as (d) selecting a second set of process parameter values ​​associated with a second experiment using an acquisition function, wherein the acquisition function identifies the second set of process parameter values ​​based on: (i) a set of points representing differences between a plurality of predicted wafer characteristics associated with performance of the manufacturing process using the second set of process parameter values ​​and the plurality of target wafer specifications; and (ii) the statistical uncertainty of the predictions made by the one or more models.

61. The system of claim 60, wherein: The set of points improves on-wafer performance relative to the plurality of target wafer specifications.

62. The system of claim 60, wherein: This set of points comprises the Pareto front.

63. The system of claim 62, wherein: The acquisition function determines an expected improvement in a supervolume formed by the Pareto front using the plurality of predicted wafer characteristics related to the performance of the manufacturing process using the second set of process parameter values.

64. The system of any one of claims 60-63, wherein: At least two process parameter values ​​in the second set of process parameter values ​​are different from corresponding process parameter values ​​in the first set of process parameter values.

65. The system of any one of claims 60-63, wherein: The at least one controller is further configured to determine that the manufacturing process cannot meet at least one target wafer specification of the plurality of target wafer specifications based at least in part on the set of points.

66. The system of claim 65, wherein: The at least one controller is further configured to identify a second plurality of predicted wafer characteristics within a predetermined error threshold of the plurality of target wafer specifications, wherein the second plurality of predicted wafer characteristics are associated with performance of the manufacturing process using a third set of process parameter values.