Dynamic Control of Machine Learning-Based Measurement Recipe Optimization

By dynamically controlling convergence trajectories and incorporating domain knowledge, the method enhances measurement performance and reliability in semiconductor manufacturing, addressing challenges of small resolution and complex geometric structures.

JP7710513B2Active Publication Date: 2025-07-18KLA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023520397
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-02
Filing Date
2021-09-23
Publication Date
2025-07-18
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

Current measurement techniques in semiconductor manufacturing face challenges due to increasing demand for small resolution, multiple parameter correlations, and the complexity of geometric structures, particularly with the use of opaque materials, leading to suboptimal measurement performance and inability to dynamically control convergence trajectories.

Method used

A method and system for training a measurement recipe by dynamically controlling the convergence trajectories of multiple performance goals, employing domain knowledge to regularize the optimization process through physics-based measurement performance metrics, and adjusting weighting values in the loss function to achieve a stable and balanced measurement model.

Benefits of technology

The solution improves measurement performance and reliability by reducing computational effort and minimizing errors, ensuring accurate and consistent measurement across various applications, while effectively addressing the challenges of small resolution and complex geometric structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710513000023
    Figure 0007710513000023
  • Figure 0007710513000024
    Figure 0007710513000024
  • Figure 0007710513000025
    Figure 0007710513000025
Patent Text Reader

Abstract

Described herein are methods and systems for training and implementing metrology recipes while dynamically controlling the convergence trajectory of multiple performance objectives. Performance metrics are employed to regularize the optimization process employed during measurement model training, model-based regression, or both. Weighting values ​​associated with each performance objective in the model optimization loss function are dynamically controlled during model training. In this manner, the convergence of each performance objective and the trade-offs between the performance objectives in the loss function are controlled to arrive at a trained measurement model in a stable and balanced manner. The trained measurement model is employed to estimate values ​​of one or more parameters of interest based on measurements of a structure having unknown values ​​for the parameters of interest. In another aspect, weighting values ​​associated with each performance objective in model-based regression on the measurement model are dynamically controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The described embodiments relate to a measurement system and a measurement method, and more particularly, to a method and a system for improved measurement of a semiconductor structure.

Background Art

[0002] Cross - reference to related applications This patent application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 089,550, filed on October 9, 2020, entitled "Method for Machine Learning Recipe Optimization by Dynamic Control of The Training", and the entire subject matter thereof is incorporated herein by reference.

[0003] Semiconductor devices, such as logic devices and memory devices, are typically manufactured by a series of processing steps applied to a sample. Various features and multiple structural layers of a semiconductor device are formed by these processing steps. For example, lithography, among other things, is one semiconductor manufacturing process that includes generating a pattern on a semiconductor wafer. Further examples of semiconductor manufacturing processes include, but are not limited to, chemical mechanical polishing, etching, deposition, and ion implantation. Multiple semiconductor devices may be manufactured on a single semiconductor wafer and then separated into individual semiconductor devices.

[0004] The measurement process is used at various steps during the semiconductor manufacturing process to detect defects on the wafer and improve the yield. Optical and X - ray based measurement techniques offer the potential for high throughput without the risk of damaging the sample. Many measurement - based techniques, including the implementation of light wave scattering measurement, reflectivity measurement, and polarization analysis, as well as related analysis algorithms, are commonly used to characterize the critical dimensions, film thickness, composition, overlay, and other parameters of nanostructures.

[0005] Many measurement techniques are indirect methods of measuring the physical properties of a sample during measurement. In most cases, the raw measurement signal cannot be used to directly determine the physical properties of the sample. Instead, a measurement model is adopted to estimate the values of one or more target parameters based on the raw measurement signal. For example, polarization analysis is an indirect method of measuring the physical properties of a sample during measurement. Generally, a physical-based measurement model or a machine learning-based measurement model is required to determine the physical properties of the sample based on the raw measurement signal (e.g., α meas and β meas ).

[0006] In some examples, a physical-based measurement model is created that attempts to predict the raw measurement signal (e.g., α meas and β meas ) based on assumed values of one or more model parameters. As shown in equations (1) and (2), the measurement model includes parameters related to the measurement tool itself, such as mechanical parameters (P machine ), and parameters related to the sample being measured. When determining the values of the target parameters, some sample parameters are treated as fixed values (P spec-fixed ), and other sample parameters of interest are treated as floating values (P spec-float ), i.e., their values are determined based on the raw measurement signal. α model = f(P machine , P spec-fixed , P spec-float ) (1) β model = g(P machine , P spec-fixed , P spec-float ) (2)

[0007] Mechanical parameters are parameters used to characterize a measurement tool (e.g., polarizer analyzer 101). Exemplary mechanical parameters include angle of incidence (AOI), analyzer angle (A0), polarizer angle (P0), illumination wavelength, numerical aperture (NA), compensator or waveplate (if present), etc. Sample parameters are parameters used to characterize a sample (e.g., material parameters and geometric parameters that characterize the structure during measurement). In the case of a thin film sample, exemplary sample parameters include refractive index, dielectric function tensor, nominal layer thickness of all layers, layer order, etc. In the case of a CD sample, exemplary sample parameters include geometric parameter values related to different layers, refractive indices related to different layers, etc. For measurement purposes, mechanical parameters and many sample parameters are treated as known fixed value parameters. However, one or more values of the sample parameters are treated as unknown floating parameters of interest.

[0008] In some examples, the value of the floating parameter of interest is determined by an iterative process (e.g., regression) that produces the best fit between the theoretical prediction and the experimental data. The value of the unknown floating parameter of interest is changed and the model output values (e.g., α meas and β meas ) are calculated and compared to the raw measurement data, and the process is repeated until a set of sample parameter values is determined that results in a close enough match between the model output values (e.g., α model and β model ) and the experimental measurements (e.g., α and β). In some other examples, the floating parameter is determined by searching a library of pre-calculated solutions to find the closest match.

[0009] In some other examples, a trained machine learning-based measurement model is employed to directly estimate the value of the parameter of interest based on the raw measurement data. In these examples, the machine learning-based measurement model takes the raw measurement signal as the model input and generates the value of the parameter of interest as the model output.

[0010] To generate useful estimates of target parameters for a particular measurement application, both a physics-based measurement model and a machine learning-based measurement model must be trained. In general, model training is based on raw measurement signals collected from samples having known values of the target parameters (design of experiments (DOE) data).

[0011] The machine learning-based measurement model is parameterized by a number of weight parameters. Conventionally, the machine learning-based measurement model is trained by a regression process (e.g., ordinary least squares regression). The values of the weight parameters are iteratively adjusted to minimize the difference between a known reference value of the target parameter and the value of the target parameter estimated by the machine learning-based measurement model based on the measured raw measurement signal.

[0012] As described above, the physics-based measurement model is parameterized by a number of machine parameters and sample parameters. Conventionally, the physics-based measurement model is also trained by a regression process (e.g., ordinary least squares regression). To minimize the difference between the raw measurement data and the modeled measurement data, one or more of the machine parameters and sample parameters are iteratively adjusted. For each iteration, the value of the target sample parameter is maintained at the known DOE value.

[0013] Conventionally, the training of both machine learning-based measurement models and physics-based measurement models (also known as measurement recipe generation) is typically achieved by minimizing the total output error, which is expressed as least squares minimization. The total output error represents the aggregation of all errors resulting from measurements, including the overall measurement uncertainty, i.e., accuracy error, tool-to-tool matching error, parameter tracking error, in-wafer variation, etc. Unfortunately, training the model based on the overall measurement uncertainty without controlling the components of the overall measurement uncertainty leads to suboptimal measurement performance. In many cases, when training is performed based on simulated data, especially due to the mismatch between the simulated data and the actual data, large modeling errors occur.

[0014] Furthermore, the experience, measurement data, and domain knowledge obtained from physics are not directly represented in the objective function that promotes the optimization of the measurement model. As a result, domain knowledge is not fully utilized in the development process of the measurement recipe. In this case as well, it leads to suboptimal measurement performance.

[0015] Moreover, current measurement model training techniques cannot consider the dynamics of the training system. Therefore, even when the measurement model is trained based on multiple goals related to the desired measurement specifications, the current training system cannot dynamically control the convergence trajectories of the multiple performance goals.

Prior Art Documents

Patent Documents

[0016]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0017] Increasing demand for small resolution, multiple parameter correlations, increasing complexity of geometric structures, and increasing use of opaque materials result in measurement challenges in future measurement applications. Therefore, methods and systems for improving measurement recipe generation are desired.

Means for Solving the Problems

[0018] This specification describes methods and systems for training a measurement recipe while dynamically controlling the convergence trajectories of multiple performance goals. Domain knowledge is employed to regularize the optimization process employed during training of the measurement model, during model-based regression, or both. Domain knowledge includes performance metrics employed to quantitatively characterize the measurement performance of a measurement system in a particular measurement application. Thus, the optimization process is physically regularized by one or more representations of the physics-based measurement performance metrics.

[0019] The weighting values associated with each of the performance goals in the loss function of model optimization are dynamically controlled during training of the model. Thus, the convergence of each performance goal and the trade-off between multiple performance goals of the loss function are controlled to reach a trained measurement model in a stable and balanced manner.

[0020] In one aspect, the measurement model is trained by dynamically controlling the weights associated with each regularization term related to the optimization process. The training is based on measurement data related to multiple instances of one or more design of experiments (DOE) measurement targets disposed on one or more wafers, reference values of target parameters related to the DOE measurement targets, actual measurement data collected from multiple instances of one or more regularization structures disposed on one or more wafers, and measurement performance metrics related to the actual measurement data.

[0021] In a further aspect, the trained measurement model is employed to estimate the value of a target parameter based on measurement values of a structure having an unknown value of one or more target parameters. In some embodiments, the measurement system employed to measure the unknown structure is the same measurement system employed to collect DOE measurement data. Generally, the trained measurement model may be employed to estimate the value of a target parameter based on a single measured spectrum or to simultaneously estimate the value of a target parameter based on a plurality of spectra.

[0022] In some embodiments, the regularization structure is the same structure as the DOE measurement target. However, generally, the regularization structure may be different from the DOE measurement target.

[0023] In some embodiments, the actual regularization measurement data is collected by a specific measurement system. In these embodiments, the measurement model is trained for a measurement application that includes measurements performed by the same measurement system.

[0024] In some other embodiments, the actual regularization measurement data is collected by a plurality of instances of a measurement system, i.e., a plurality of measurement systems that are substantially identical. In these embodiments, the measurement model is trained for a measurement application that includes measurements performed by any of the plurality of instances of the measurement system.

[0025] In some examples, measurement data associated with each measurement of a plurality of instances of one or more design of experiments (DOE) measurement targets by a measurement system is simulated. The simulated data is generated from a parameterized model of each measurement of one or more DOE measurement structures by the measurement system.

[0026] In some other examples, the measurement data related to multiple instances of one or more design of experiments (DOE) measurement targets is the actual measurement data collected by a measurement system or multiple instances of a measurement system. In some of these embodiments, the same measurement system or multiple instances of a measurement system are employed to collect the actual regularized measurement data from the regularization structure.

[0027] In some embodiments, the physical measurement performance metric characterizes the actual measurement data collected from each of multiple instances of one or more regularization structures. In some embodiments, the performance metric is based on historical data, domain knowledge about the process involved in creating the structure, physics, or the best guess by the user. In some examples, the measurement performance metric is a single point estimate. In other examples, the measurement performance metric is a distribution of estimated values.

[0028] Generally, the measurement performance metric associated with the measurement data collected from a regularization structure provides information about the value of the physical attributes of the regularization structure. As a non-limiting example, the physical attributes of the regularization structure include any of measurement accuracy, matching between tools, wafer average, in-wafer range, tracking with respect to a reference, matching between wafers, tracking with respect to wafer splitting, etc.

[0029] In a further aspect, the performance of a trained measurement model is verified using test data with error budget analysis. As test data for verification purposes, actual measurement data, simulated measurement data, or both may be employed. Error budget analysis for actual data enables the estimation of the individual contributions to the overall error, such as accuracy, tracking, precision, tool matching error, wafer-to-wafer consistency, wafer signature consistency, etc. In some embodiments, the test data is designed such that the overall model error is divided into each contributing component.

[0030] In another aspect, model-based regression in the measurement model is physically regularized by one or more measurement performance metrics, and the weighting values associated with each regularization term related to each performance metric are dynamically controlled during the regression process.

[0031] The foregoing is a summary and, accordingly, includes simplifications, generalizations, and omissions as necessary, and thus those skilled in the art will understand that the summary is for illustrative purposes only and is in no way limiting. Other aspects, features, and advantages of the devices and / or processes described herein will become apparent in the non-limiting detailed description set forth herein.

Brief Description of the Drawings

[0032]

Fig. 1

Fig. 2

Fig. 3A

Fig. 3B

Fig. 3C

Fig. 4A

Fig. 4B

Fig. 4C

Fig. 5

Fig. 6

Fig. 7

Fig. 8

[0033] Next, examples of the background and several embodiments of the present invention are referred to in detail, and the examples are shown in the accompanying drawings.

[0034] This specification describes methods and systems for training a measurement recipe while dynamically controlling the convergence trajectories of multiple performance goals. Domain knowledge is employed to regularize the optimization process employed during training of the measurement model, during model-based regression, or both. Domain knowledge includes performance metrics employed to quantitatively characterize the measurement performance of a measurement system in a particular measurement application. Thus, the optimization process is physically regularized by one or more representations of the physics-based measurement performance metrics. By way of non-limiting example, probability distributions related to measurement accuracy, matching between tools, tracking, within-wafer variation, etc. are employed to physically regularize the optimization process. Thus, these important metrics are controlled during training of the measurement model, during model-based regression, or both.

[0035] The weighting values associated with each performance goal in the loss function for model optimization are dynamically controlled during the training of the model. In this way, the convergence of each performance goal and the trade-off between multiple performance goals of the loss function are controlled to reach a trained measurement model in a stable and balanced manner. The dynamic control of the weighting values significantly reduces the computational amount required for training the measurement model associated with a specific measurement recipe. Furthermore, the resulting trained measurement model, model-based measurement, or both achieve a significant improvement in measurement performance and reliability.

[0036] By dynamically controlling the weights associated with each regularization term related to the optimization process, the model's consistency is improved and the computational amount related to the training of the model is reduced. The optimization process is less affected by overfitting. When physical regularization and dynamic control of the weighting values are adopted, measurement performance specifications such as accuracy, matching between tools, and parameter tracking are more reliably met across various measurement model architectures and measurement applications. In some embodiments, simulated training data is adopted. In these embodiments, physical regularization and dynamic control of the weighting values significantly reduce the error due to the mismatch between the simulated measurement data and the actual measurement data.

[0037] In one aspect, the measurement model is trained by dynamically controlling the weights associated with each regularization term related to the optimization process. The training is based on measurement data related to multiple instances of one or more design of experiments (DOE) measurement targets arranged on one or more wafers, reference values of target parameters related to the DOE measurement targets, actual measurement data collected from multiple instances of one or more regularization structures arranged on one or more wafers, and measurement performance indicators related to the actual measurement data. Furthermore, one or more measurement performance indicators are adopted to regularize the optimization that facilitates the measurement model training process.

[0038] Figure 1 shows a system 100 for measuring the characteristics of a sample according to an exemplary method presented herein. As shown in Figure 1, system 100 may be used to perform spectroscopic polarization analysis measurements on a structure 101. In this aspect, system 100 may include a spectroscopic polarizer equipped with an illuminator 102 and a spectrometer 104. The illuminator 102 of system 100 is configured to generate illumination in a selected wavelength range (e.g., 100 - 2500 nm) and direct this towards a structure disposed on the surface of the sample over a measurement spot 110. Next, the spectrometer 104 is configured to receive the illumination reflected from the structure 101. It should further be noted that the light generated from the illuminator 102 is polarized using a polarization state generator 107 to generate a polarized illumination beam 106. The radiation reflected by the structure 101 passes through a polarization state analyzer 109 and is sent to the spectrometer 104. The radiation received by the spectrometer 104 in the collection beam 108 is analyzed with respect to the polarization state, thereby enabling spectral analysis of the radiation passing through the analyzer by the spectrometer. These spectra 111 are passed to a computing system 130 for analysis of the structures described herein.

[0039] As shown in Figure 1, system 100 includes a single measurement technique (i.e., SE). However, in general, system 100 may include any number of different measurement techniques. As non-limiting examples, system 100 may be configured as a spectroscopic polarizer (including Mueller matrix polarization analysis), a spectroscopic reflectometer, a spectroscopic scatterometer, an overlay scatterometer, an angle-resolved beam profile reflectometer, a polarization-resolved beam profile reflectometer, a beam profile reflectometer, a beam profile polarizer, any single or multiple wavelength polarizers, or any combination thereof. Further, in general, the measurement data collected by different measurement techniques and analyzed according to the methods described herein may be collected from multiple tools, a single tool integrating multiple techniques, or a combination thereof, including, as non-limiting examples, soft X-ray reflectometry measurements, small-angle X-ray scattering measurements, imaging-based measurement systems, hyperspectral imaging-based systems, scatterometry overlay measurement systems, and the like.

[0040] In a further embodiment, system 100 may include one or more computing systems 130 employed to perform measurements of a structure based on a measurement model developed according to the methods described herein. The one or more computing systems 130 may be communicatively coupled to spectrometer 104. In one aspect, the one or more computing systems 130 are configured to receive measurement data 111 related to the measurement of a structure being measured (e.g., structure 101).

[0041] In one aspect, computing system 130 is configured as a measurement model training engine 150 for training a measurement model based on the measured values of the regularization structure described herein. FIG. 2 is a diagram showing an exemplary measurement model training engine 200 in one embodiment. As shown in FIG. 2, measurement model training engine 200 receives measurement data X DOE 203 related to simulated measured values, actual measured values, or both of a plurality of instances of one or more design of experiments (DOE) measurement targets disposed on one or more wafers. In an example of spectroscopic polarization analysis measurements, the DOE measurement data includes measured spectra, simulated spectra, or both. In one example, DOE measurement data 203 includes measured spectra 111 collected by measurement system 100 from a plurality of instances of one or more DOE measurement targets.

[0042] Furthermore, measurement model training engine 200 receives one or more target parameter reference values Y related to the DOE measurement target DOEReceive 205 from reference source 204. Examples of target parameters include geometric parameters characterizing the measured structure, dispersion parameters characterizing the measured structure, process parameters characterizing the process employed to manufacture the measured structure, electrical properties of the measured structure, and the like. Exemplary geometric parameters include critical dimension (CD), overlay, and the like. Exemplary process parameters include lithography focus, lithography dose, etch time, and the like.

[0043] In some embodiments, reference value 205 is simulated. In these embodiments, reference source 204 is a simulation engine that generates corresponding simulated DOE measurement data 203 for known reference value 205. In some embodiments, reference value 205 is a value measured by a reliable measurement system (e.g., a scanning electron microscope, etc.). In these embodiments, the reference source is a reliable measurement system.

[0044] Measurement model training engine 200 also receives actual regularization measurement data X REG 202 related to its actual regularization measurement data 202 along with measurement performance metric θ REG 201. In one example, the regularization measurement data 202 includes measured spectra 111 collected by measurement system 100 from multiple instances of one or more regularization structures disposed on one or more wafers.

[0045] Measurement model training engine 200 trains the measurement model while dynamically controlling the weights associated with each regularization term of the optimization function based on an optimization function regularized by one or more measurement performance metrics. In some examples, the measurement model is a neural network model. As shown in FIG. 2, neural network module 206 receives dataset X DOE 203 and X REGEvaluate the neural network model h(·) for 202. Further, the neural network module 206 evaluates each regularization term associated with the optimization function related to model training. These results 207 are transmitted to the loss evaluation module 208. The loss evaluation module 208 DOE 205, and X DOE the corresponding values estimated by the neural network model from 203, the regularization terms estimated by the neural network module 206, and the regularization weighting term γ k and the current value of, determine the value of the optimization function. The loss evaluation module 208 updates the neural network weighting value W based on the value of the optimization function. The updated neural network weighting value 212 is transmitted to the neural network module 206. The neural network module 206 updates the neural network model with the updated neural network weighting value for the next iteration of the training process.

[0046] In each iteration of the training process, the control module 210 determines the updated value of each regularization weighting term γ k associated with each measurement target. Each updated value is determined based on the achieved value of the measurement target and the desired value of each measurement target. As shown in Figure 2, the control module 210 receives an indication 209 regarding the value of each achieved measurement target and the desired value 214 of each measurement target. In each iteration, the control module 210 compares the achieved value and the desired value associated with each measurement target and determines the updated value 211 of each regularization weighting term. The updated value 211 of each regularization weighting term is transmitted to the loss evaluation module 208. The loss evaluation module 208 evaluates the optimization function using the updated value 211 in the next iteration.

[0047] By continuously adjusting the weights of each measurement target during the training process, the neural network is trained to achieve the desired specifications of each measurement target with less computational effort.

[0048] The control module 210 employs a controller that optimizes a plurality of measurement targets. As a non-limiting example, the controller can be any of a linear quadratic regulator (LQR) - based controller, a proportional-integral-derivative (PID) controller, an optimal controller, an adaptive controller, a model predictive controller, etc.

[0049] In some embodiments, the parameters of the controller are optimized for robust performance by search algorithms such as genetic algorithms, simulated annealing algorithms, gradient descent algorithms, etc.

[0050] In some examples, each measurement performance metric is represented as an individual distribution. In one example, the distribution of the measurement accuracy related to the regularization structure is an inverse gamma distribution. Equation (1) shows the probability density function p for the measurement accuracy dataset x, where Γ(·) represents the gamma function, the constant a represents the shape parameter, and the constant b represents the scale parameter.

Equation

[0051] In another example, the distribution of the mean value of the measured regularization structure instances on the wafer is described by a normal distribution. Equation (2) shows the probability density function m for the measured wafer mean dataset x, where μ represents a specific mean and σ represents a specific variance related to the distribution.

Equation

[0052] In a further aspect, to regularize the optimization that facilitates the training of the measurement model, in particular, statistical information characterizing the actual measurement data collected from the regularization structure, such as known distributions related to important measurement performance metrics such as measurement accuracy, wafer mean, etc., is adopted. Equation (3) is for the DOE parameter y of interestDOE The joint likelihood and the criteria of the measurement performance index related to the measurement of the regularization structure reg are shown. By maximizing the joint likelihood, the measurement model h(·) evolves during training to maintain the fidelity of the DOE measurement data x DOE while adapting to meet the measurement performance of the measurement data x reg related to the regularization structure. P(y DOE ,criteria reg |h(·),x DOE ,x reg ) (3)

[0053] To maximize the joint likelihood, the DOE measurement data contributes to the mean squared error, and the measurement data related to the regularization structure contributes as a regularization term in the loss function. In summary, maximizing the joint likelihood is equivalent to assuming independence between the DOE measurement data and the regularization data and between different regularization data sets and minimizing the loss function shown in Equation (4), where Reg(h(·)) is the comprehensive regularization with respect to the model parameters weighted by the constant parameter α, and Reg k (x reg,k ,h(·),θ reg,k ) is the k-th regularization term weighted by the regularization weighting parameter γ k , where x reg,k is the k-th regularization data set, and θ reg,k is a vector of parameters that describes the statistical information related to the actual measurement data collected from the regularization structure. J(h(·);x,y,θ)=||h(x DOE )-y DOE || 2 +α·Reg(h(·))+γ1·Reg1(x reg,1 ,h(·),θ reg,1 )+···+γ k ·Reg k (x reg,k ,h(·),θ reg,k ) (4)

[0054] In one example, in the optimization of the measurement model, two different regularization terms, Reg1 and Reg2, are adopted. Reg1 represents the regularization of the measurement accuracy with respect to the measurement accuracy dataset x reg-prec and Reg2 represents the regularization of the wafer average with respect to the in-wafer dataset x WIW . As a non-limiting example, assume that the measurement accuracy is described by an inverse gamma distribution having shape parameters

Number

Number

Number

Number

Number

Number

Number

Number

[0055] In this example, the measurement model optimization function can be described as shown in Equation (7), where h W,b (·) is a neural network model with weighted values W and bias value b, and the model error variance is

Number

Number

Number

[0056] The DOE data set, measurement accuracy data set, and wafer average data set employed in model training using the measurement model optimization function are shown in Equation (8).

Number

[0057] The known parameters of the statistical model that describe the model error, the weight values of the neural network, the measurement accuracy, and the wafer average within the wafer are shown in Equation (9).

Number

[0058] During model training, the optimization function shown by Equation (7) balances between the DOE data estimation error and all other criteria. The first term represents the DOE data estimation error as the mean squared error penalized by the model error variance

Number

Number

[0059] In each iteration, the optimization function promotes changes to the weight values W and bias values b of the neural network model h W,b (·), thereby minimizing the optimization function. When the optimization function reaches a sufficiently low value, the measurement model is considered trained, and the trained measurement model 213 is stored in memory (e.g., memory 132).

[0060] Figures 3A to 3C are plots showing indicators characterizing the measurement tracking performance in each iteration of the measurement model training in an example.

[0061] Figure 3A illustrates a plot 220 showing the achieved R 2 value related to the measurement of the critical dimension of the DOE structure in each iteration of the model training. As shown in Figure 3A, the achieved R 2 value converges rapidly and stably to the desired value R 2 .

[0062] Figure 3B illustrates a plot 221 showing the achieved gradient value related to the measurement of the critical dimension of the DOE structure in each iteration of the model training. As shown in Figure 3B, the achieved gradient value converges rapidly and stably to the desired gradient value.

[0063] Figure 3C illustrates a plot 222 showing the achieved accuracy value related to the measurement of the critical dimension of the DOE structure in each iteration of the model training. As shown in Figure 3C, the achieved accuracy value converges stably to the desired accuracy value.

[0064] Figures 4A to 4C are plots showing the weight values associated with each measurement target shown in Figures 3A to 3C in each iteration of the measurement model training in an example, respectively.

[0065] Figure 4A shows the R of the measured values of the critical dimension of the DOE structure shown in Figure 3A in each iteration of the model training.2 Plot 223 showing the weighted values assigned to the terms of the objective function representing the values is illustrated. As shown in FIG. 4A, R 2 The weighted values associated with the R values converge rapidly and stably to a small number as the desired value of the performance goal is achieved.

[0066] FIG. 4B illustrates plot 224 showing the weighted values assigned to the terms of the objective function representing the gradient values of the measured values of the critical dimensions of the DOE structure shown in FIG. 3B for each iteration of model training. As shown in FIG. 4B, the weighted values associated with the gradient values converge rapidly and stably to a small number as the desired value of the performance goal is achieved.

[0067] FIG. 4C illustrates plot 225 showing the weighted values assigned to the terms of the objective function representing the accuracy values of the measured values of the critical dimensions of the DOE structure shown in FIG. 3C for each iteration of model training. As shown in FIG. 4C, the weighted values associated with the accuracy values converge rapidly and stably to a small number as the desired value of the performance goal is achieved.

[0068] In another further aspect, the trained measurement model is employed to estimate the value of the target parameter based on the measured values of the structure having unknown values of one or more target parameters. In some examples, the trained model provides both an estimated value of the target parameter value and the uncertainty of the measured value. The trained measurement model is employed to estimate the values of one or more target parameters from actual measurement data (e.g., measured spectra) collected by a measurement system (e.g., metrology system 100). In some embodiments, the measurement system is the same measurement system employed to collect the DOE measurement data. In other embodiments, the measurement system is a simulated system for synthetically generating the DOE measurement data. In one example, the actual measurement data includes measured spectra 111 collected by metrology system 100 from one or more metrology targets having unknown values of one or more target parameters.

[0069] Generally, a trained measurement model may be employed to estimate the value of a target parameter based on a single measured spectrum or to simultaneously estimate the values of target parameters based on multiple spectra.

[0070] In some embodiments, the regularization structure is the same as the DOE measurement target. However, generally, the regularization structure may be different from the DOE measurement target.

[0071] In some embodiments, the actual regularization measurement data collected from multiple instances of one or more regularization structures is collected by a specific measurement system. In these embodiments, the measurement model is trained for a measurement application that includes measurements performed by the same measurement system.

[0072] In some other embodiments, the actual regularization measurement data collected from multiple instances of one or more regularization structures is collected by multiple instances of a measurement system, i.e., multiple measurement systems that are substantially identical. In these embodiments, the measurement model is trained for a measurement application that includes measurements performed by any of the multiple instances of the measurement system.

[0073] In some examples, measurement data associated with each measurement of multiple instances of one or more design of experiments (DOE) measurement targets by a measurement system is simulated. The simulated data is generated from a parameterized model of each measurement of one or more DOE measurement structures by the measurement system.

[0074] In some other examples, the measurement data related to multiple instances of one or more design of experiments (DOE) measurement targets is the actual measurement data collected by a measurement system or multiple instances of a measurement system. In some of these embodiments, the same measurement system or multiple instances of a measurement system are employed to collect the actual regularized measurement data from the regularization structure.

[0075] In some embodiments, the physical measurement performance metric characterizes the actual measurement data collected from each of multiple instances of one or more regularization structures. In some embodiments, the performance metric is based on data collected from a reference measurement system, nominal DOE parameter values, historical data, domain knowledge about the process involved in creating the structure, physics, statistical data collected from multiple processes and multiple measurement techniques, or the best guess by the user. In some examples, the measurement performance metric is a single point estimate. In other examples, the measurement performance metric is a distribution of estimated values.

[0076] Generally, the measurement performance metric associated with the measurement data collected from the regularization structure provides information about the value of the physical attributes of the regularization structure. As a non-limiting example, the physical attributes of the regularization structure include any of measurement accuracy, measurement precision, matching between tools, wafer average, in-wafer range, in-wafer variation, wafer signature, tracking with respect to a reference, inter-wafer variation, tracking with respect to wafer splitting, etc.

[0077] In some examples, the measurement performance metric includes a specific value of a parameter of the regularization structure at a specific location on the wafer and the corresponding uncertainty. In one example, the measurement performance metric is the critical dimension (CD) and its uncertainty at a specific location on the wafer, e.g., the CD is 35 nanometers ± 0.5 nanometers.

[0078] In some examples, the measurement performance metric includes the probability distribution of the values of the parameters of the structure across a wafer, within a lot of wafers, or across multiple wafer lots. In one example, the CD exhibits a normal distribution with a mean value and a standard deviation. For example, the mean value of the CD is 55 nanometers and the standard deviation is 2 nanometers.

[0079] In some examples, the measurement performance metric includes the spatial distribution of the values of the target parameter across the entire wafer, such as a wafer map, and the corresponding uncertainty at each location.

[0080] In some examples, the measurement performance metric includes the distribution of the measured values of the target parameter across multiple tools to characterize the matching between tools. The distribution may represent the mean value across each wafer, the value at each site, or both.

[0081] In some examples, the measurement performance metric includes the distribution of the measurement accuracy error.

[0082] In some examples, the measurement performance metric includes a wafer map that matches the estimated value for the entire wafer lot.

[0083] In some examples, the measurement performance metric includes one or more metrics that characterize the tracking of the estimated value of the target parameter using the reference value of the target parameter. In some examples, the metrics that characterize the tracking performance include the R 2 value, the gradient value, and / or the offset value.

[0084] In some examples, the measurement performance metric includes one or more metrics that characterize the tracking of the estimated value of the target parameter with respect to the wafer average for a DOE split experiment. In some examples, the metrics that characterize the tracking performance include the R 2 value, the gradient value, and / or the offset value.

[0085] FIG. 5 shows a plot 180 that characterizes tracking performance. As shown in FIG. 5, the x-position of each data point on plot 180 indicates the predicted value of the target parameter, and the y-position of each data point indicates the known value of the target parameter (e.g., the DOE reference value). Ideal tracking performance is indicated by the dashed line 181. If all predicted values exactly match the corresponding known reliable values, all data points would lie on line 181. However, in reality, the tracking performance is not perfect. Line 182 shows the best-fit line for the data points. As shown in FIG. 5, line 182 is characterized by a slope and a y-intercept value, and the correlation between the known value and the predicted value is R 2 characterized by a value.

[0086] In a further aspect, the performance of the trained measurement model is verified using test data by error budget analysis. As test data for verification purposes, actual measurement data, simulated measurement data, or both may be employed.

[0087] Error budget analysis on real data enables the estimation of the individual contributions to the overall error, such as accuracy, tracking, precision, tool matching error, within-wafer consistency, and within-wafer signature consistency. In some embodiments, the test data is designed such that the total model error is split into each contributing component.

[0088] As a non-limiting example, the real data is the following subset, namely, reference values for the calculation of accuracy and tracking, such as slope, offset, R 2Actual data with reference values including, for example, 3STEYX, mean squared error, 3 sigma error, actual data from measurements of the same site measured multiple times to estimate measurement accuracy, actual data from measurements of the same site measured by different tools to estimate matching between tools, actual data from measurements of sites on multiple wafers to estimate wafer-to-wafer variation of wafer average and wafer variance, and any of the actual data from measurements of multiple wafers to identify typical wafer patterns such as bull's-eye patterns expected to be present on a given wafer, such as wafer signature.

[0089] In some other examples, a parameterized model of the structure is adopted to generate simulated data for error budget analysis. The simulated data is generated such that each parameter of the structure is sampled within its DOE and the other parameters are fixed at their nominal values. In some examples, other parameters of the simulation, such as system model parameters, are included in the error budget analysis. Since the true reference values of the parameters are known from the simulated data, the error due to the change of each parameter of the structure can be separated.

[0090] In some examples, additional simulated data is generated using different noise sampling to calculate accuracy error.

[0091] In some examples, additional simulated data is generated outside the DOE of the parameterized structure to estimate extrapolation error.

[0092] In another aspect, model-based regression in the measurement model is physically regularized by one or more measurement performance metrics, and the weights associated with each regularization term related to the optimization process are dynamically controlled. The estimated values of one or more target parameters are determined based on actual measurement data collected from multiple instances of one or more target structures disposed on one or more wafers, statistical information related to the measurement, and prior estimated values of the target parameters.

[0093] In one aspect, computing system 130 is configured as a measurement model regression engine that performs measurements of the structures described herein. FIG. 6 is a diagram showing an exemplary measurement model regression engine 190 in one embodiment. As shown in FIG. 6, measurement model regression engine 190 includes a regression module 191, a loss evaluation module 193, and a control module 195. As shown in FIG. 6, loss evaluation module 193 receives measurement data X POI 188 from a measurement source 199, such as a spectroscopic polarimeter, etc., related to the measurement of one or more measurement targets. In one example, measurement data 188 includes measured spectra 111 collected by measurement system 100 from one or more measurement targets. Further, loss evaluation module 193 receives a measurement performance index θ REG 187 related to measurement data 188.

[0094] Regression module 191 includes a measurement model that simulates measurement spectrum 192 based on assumed values of one or more target parameters 197. Loss evaluation module 193 iteratively updates the values of one or more target parameters 197 related to the measurement target to be measured based on an optimization function regularized by one or more measurement performance indices. When the iteration stop criterion is met, the current values 198 of the one or more target parameters are stored in a memory (e.g., memory 132).

[0095] The loss function of the regression includes a data reconstruction error term and one or more regularization terms. Equation (10) shows an exemplary loss function of model-based regression for estimating the values of one or more target parameters from actual measurement values.

Equation

[0096] During the iteration of the regression process, the control module 195 determines the updated value of each regularization weighting term γ k associated with each measurement goal. Each updated value is determined based on the achieved value of the measurement goal and the desired value of each measurement goal. As shown in Figure 6, the control module 195 receives an indication 194 regarding the value of the achieved measurement goal and the desired value 189 of each measurement goal. In each iteration, the control module 195 compares the achieved value and the desired value associated with each measurement goal and determines the updated value 196 of each regularization weighting term. The updated value 196 of each regularization weighting term is transmitted to the loss evaluation module 193. The loss evaluation module 193 evaluates the optimization function using the updated value 196 in the next iteration.

[0097] By continuously adjusting the weights of each measurement goal during the regression process, the values of the target parameters are determined so that the desired specifications of each measurement goal are achieved with less computational effort.

[0098] The control module 195 employs a controller that optimizes a plurality of measurement targets. As a non-limiting example, the controller can be any of a linear quadratic regulator (LQR)-based controller, a proportional integral derivative (PID) controller, an optimal controller, an adaptive controller, a model predictive controller, etc.

[0099] In some embodiments, the parameters of the controller are optimized for robust performance by a search algorithm such as a genetic algorithm, a simulated annealing algorithm, a gradient descent algorithm, etc.

[0100] In one example, the regularization term is the measurement accuracy and the wafer average within the wafer as described above. In this example, the regularization term related to the measurement accuracy is shown in Equation (11), where Y reg-prec is the prior estimated value of the target parameter, and σ(Y reg-prec ) represents the standard deviation of Y reg-prec .

Equation

Equation

Equation

Equation

[0101] In some embodiments, the values of the target parameters employed to train the measurement model are derived from measurements of the DOE wafers by a reference measurement system. The reference measurement system is a reliable measurement system that produces sufficiently accurate measurement results. In some examples, the reference measurement system is too slow to be used online to measure wafers as part of the wafer manufacturing process flow, but is suitable for offline use, such as for model training. By way of non-limiting example, the reference measurement system may include a stand-alone optical measurement system, such as a spectroscopic ellipsometer (SE), an SE with multiple illumination angles, an SE that measures Mueller matrix elements, a single-wavelength ellipsometer, a beam profile ellipsometer, a beam profile reflectometer, a broadband reflectance spectrometer, a single-wavelength reflectometer, an angle-resolved reflectometer, an imaging system, a scatterometer, such as a speckle analyzer, an X-ray-based measurement system, such as a small-angle X-ray scattering (SAXS) meter operating in transmission or grazing incidence mode, an X-ray diffraction (XRD) system, an X-ray fluorescence (XRF) system, an X-ray photoelectron spectroscopy (XPS) system, an X-ray reflectometer (XRR) system, a Raman spectroscopy system, an atomic force microscope (AFM) system, a transmission electron microscope system, a scanning electron microscope system, a soft X-ray reflectivity measurement system, an imaging-based measurement system, a hyperspectral imaging-based measurement system, a scatterometry overlay measurement system, or other techniques capable of identifying device geometries.

[0102] In some embodiments, the measurement model trained as described herein is implemented as a neural network model. In other examples, the measurement model may be implemented as a linear model, a non-linear model, a polynomial model, a response surface model, a support vector machine model, a decision tree model, a random forest model, a kernel regression model, a deep network model, a convolutional network model, or other types of models.

[0103] In some examples, the measurement model trained as described herein may be implemented as a combination of models.

[0104] In yet another aspect, the measurement results described herein can be used to provide active feedback to a process tool (e.g., a lithography tool, an etching tool, a deposition tool, etc.). For example, the value of a measurement parameter determined based on the measurement method described herein can be transmitted to an etching tool to adjust the etching time to achieve a desired etching depth. Similarly, etching parameters (e.g., etching time, diffusion rate, etc.) or deposition parameters (e.g., time, concentration, etc.) may be included in the measurement model to provide active feedback to the etching tool or the deposition tool, respectively. In some examples, a correction to a process parameter determined based on the measured device parameter value and the trained measurement model may be transmitted to the process tool. In one embodiment, the computing system 130 determines the value of one or more target parameters during the process based on the measurement signal 111 received from the measurement system. Further, the computing system 130 transmits a control command to a process controller (not shown) based on the determined value of the one or more target parameters. The control command causes the process controller to change the state of the process (e.g., stop the etching process, change the diffusion rate, change the lithography focus, change the lithography dose, etc.).

[0105] In some embodiments, the methods and systems for measuring semiconductor devices described herein are applicable to the measurement of memory structures. These embodiments enable the measurement of optical critical dimensions (CDs), films, and compositions for periodic and planar structures.

[0106] In some examples, the measurement model is implemented as an element of a SpectraShape® optical critical dimension measurement system available from KLA-Tencor Corporation, Milpitas, California, USA. Thus, the model is created and ready to be used immediately after the spectrum is collected by the system.

[0107] In some other examples, the measurement model is implemented offline by a computing system implementing, for example, AcuShape® software available from KLA-Tencor Corporation, Milpitas, California, USA. The resulting trained model may be incorporated as an element of the AcuShape® library accessible by the metrology system performing the measurements.

[0108] In general, dynamic control of the weighting values associated with each term of the optimization loss function may be applied to any machine learning algorithm where the loss function includes multiple objectives and desired specifications are defined for each of these objectives.

[0109] FIG. 7 shows a method 300 for training a measurement model based on one or more metrology performance metrics in at least one novel aspect. Method 300 is suitable for implementation by a metrology system such as the metrology system 100 shown in FIG. 1 of the present invention. In one aspect, it should be understood that the data processing blocks of method 300 may be implemented via a pre-programmed algorithm executed by one or more processors of computing system 130 or any other general-purpose computing system. It should be understood that the specific structural aspects of metrology system 100 are not representative of limitations herein and should be construed as merely illustrative.

[0110] In block 301, a certain amount of design of experiments (DOE) measurement data related to the measurement of one or more DOE measurement targets is received by the computing system.

[0111] In block 302, known reference values of one or more target parameters related to the DOE measurement target are received by the computing system.

[0112] In block 303, a certain amount of regularization measurement data from the measurement of one or more regularization structures arranged on the first wafer by a measurement tool is received by the computing system.

[0113] In block 304, values of one or more measurement performance indicators related to the regularization measurement data are received by the computing system.

[0114] In block 305, a measurement model is iteratively trained based on an optimization function that includes a certain amount of design of experiments (DOE) measurement data, reference values of one or more target parameters, regularization measurement data, and one or more measurement performance indicators. The optimization function includes regularization terms related to each of the one or more measurement performance indicators, and the weighting values associated with each of the regularization terms are dynamically controlled during the iteration of the measurement model.

[0115] FIG. 8 shows a method 400 for estimating values of one or more target parameters based on an optimization function regularized by one or more measurement performance indicators in at least one novel aspect. The method 400 is suitable for implementation by a measurement system such as the measurement system 100 shown in FIG. 1 of the present invention. In one aspect, it should be understood that the data processing blocks of the method 400 can be implemented via pre-programmed algorithms executed by one or more processors of the computing system 130 or any other general-purpose computing system. It should be understood that the specific structural aspects of the measurement system 100 in this specification are not intended to be limiting and should be construed as merely illustrative.

[0116] In block 401, a certain amount of measurement data from the measurement of one or more measurement targets arranged on the wafer by a measurement tool is received by the computing system.

[0117] In block 402, values of one or more measurement performance indicators related to the measurement data are received by the computing system.

[0118] In block 403, based on a regression analysis including an optimization function regularized by one or more measurement performance indicators, values of one or more target parameters characterizing one or more measurement targets are estimated from a certain amount of measurement data. The optimization function includes a regularization term associated with each of the one or more measurement performance indicators, and the weighting value associated with each of the regularization terms is dynamically controlled during the iteration of the measurement model.

[0119] In a further embodiment, system 100 includes one or more computing systems 130 employed to perform measurements of a semiconductor structure based on spectroscopic measurement data collected according to the methods described herein. The one or more computing systems 130 may be communicatively coupled to one or more spectrometers, active optical elements, process controllers, and the like. In one aspect, the one or more computing systems 130 are configured to receive measurement data related to spectral measurements of the structures of wafer 101.

[0120] It should be recognized that one or more of the steps described throughout this disclosure may be performed by a single computer system 130 or, alternatively, by a plurality of computer systems 130. Further, different subsystems of system 100 may include computer systems suitable for performing at least some of the steps described herein. Accordingly, the foregoing description should not be construed as a limitation on the present invention, but merely as an illustration.

[0121] Furthermore, computer system 130 may be communicatively coupled to the spectrometer in any manner known in the art. For example, one or more computing systems 130 may be coupled to a computing system associated with the spectrometer. In another example, the spectrometer may be directly controlled by a single computer system coupled to computer system 130.

[0122] The computer system 130 of system 100 may be configured to receive and / or obtain data or information from a subsystem of the system (such as a spectrometer, etc.) via a transmission medium that may include a wired portion and / or a wireless portion. Thus, the transmission medium may function as a data link between the computer system 130 and other subsystems of system 100.

[0123] The computer system 130 of system 100 may be configured to receive and / or obtain data or information (such as measurement results, modeling inputs, modeling results, reference measurement results, etc.) from other systems via a transmission medium that may include a wired portion and / or a wireless portion. Thus, the transmission medium may function as a data link between the computer system 130 and other systems (such as the on-board memory system 100, an external memory, or other external systems). For example, the computing system 130 may be configured to receive measurement data from a storage medium (i.e., memory 132 or an external memory) via a data link. For example, spectral results obtained using the spectrometer described herein may be stored in a permanent or semi-permanent memory device (such as memory 132 or an external memory). In this regard, the spectral results may be imported from on-board memory or from an external memory system. Further, the computer system 130 may transmit data to other systems via the transmission medium. For example, a measurement model or estimated parameter values determined by the computer system 130 may be transmitted and stored in an external memory. In this regard, the measurement results may be exported to another system.

[0124] Computing system 130 includes, but is not limited to, a personal computer system, a mainframe computer system, a workstation, an image computer, a parallel processor, or any other device known in the art. In general, the term "computing system" can be broadly defined to include any device having one or more processors that execute instructions from a memory medium.

[0125] Program instructions 134 for implementing methods such as the methods described herein may be transmitted via a transmission medium such as a wire, cable, or wireless transmission link. For example, as shown in FIG. 1, program instructions 134 stored in memory 132 are transmitted to processor 131 via bus 133. Program instructions 134 are stored in a computer-readable medium (e.g., memory 132). Exemplary computer-readable media include read-only memory, random access memory, magnetic disk or optical disk, or magnetic tape.

[0126] As used herein, the term "critical dimension" includes any critical dimension of a structure (e.g., a lower critical dimension, an intermediate critical dimension, an upper critical dimension, a sidewall angle, a grating height, etc.), a critical dimension between any two or more structures (e.g., a distance between two structures), and a displacement between two or more structures (e.g., an overlay displacement between overlay grating structures, etc.). Structures may include three-dimensional structures, patterned structures, overlay structures, and the like.

[0127] As used herein, the term "critical dimension application" or "critical dimension measurement application" includes any critical dimension measurement.

[0128] As used herein, the term "measurement system" includes any system that is at least partially employed to characterize a sample in any manner, including measurement applications such as critical dimension measurement, overlay measurement, focus / dose measurement, and composition measurement. However, such technical terms do not limit the scope of the term "measurement system" described herein. Further, system 100 may be configured for measuring patterned wafers and / or unpatterned wafers. The measurement system may be configured as an LED inspection tool, an edge inspection tool, a backside inspection tool, a macro inspection tool, or a multi-mode inspection tool (including data from one or more platforms simultaneously), and any other measurement or inspection tool that benefits from the techniques described herein.

[0129] This specification describes various embodiments of a semiconductor measurement system that can be used to measure a sample within any semiconductor processing tool (e.g., an inspection system or a lithography system). As used herein, the term "sample" is used to refer to a wafer, reticle, or any other sample that can be processed (e.g., printed or inspected for defects) by means known in the art.

[0130] As used herein, the term "wafer" generally refers to a substrate formed of a semiconductor material or a non-semiconductor material. Examples include, but are not limited to, single crystal silicon, gallium arsenide, and indium phosphide. Such substrates are generally found in and / or may be processed in semiconductor manufacturing facilities. In some cases, a wafer may include only the substrate (i.e., a bare wafer). Alternatively, a wafer may include one or more layers of different materials formed on the substrate. One or more layers formed on the wafer may be either "patterned" or "unpatterned". For example, a wafer may include a plurality of dies having repeatable pattern features.

[0131] A "reticle" may be a reticle at any stage of the reticle manufacturing process, or a completed reticle that may or may not have been released for use in semiconductor manufacturing equipment. A reticle, or "mask", is generally defined as a substantially transparent substrate having substantially opaque regions formed thereon and configured in a pattern. The substrate may include, for example, a glass material such as amorphous SiO2. The reticle may be placed on a wafer covered with resist during the exposure step of the lithography process so that the pattern on the reticle is transferred to the resist.

[0132] One or more layers formed on the wafer may or may not be patterned. For example, the wafer may include a plurality of dies each having a repeatable pattern feature. Formation and processing of such material layers can ultimately result in a completed device. Many different types of devices may be formed on the wafer, and the term wafer as used herein is intended to encompass wafers on which any type of device known in the art is being manufactured.

[0133] In one or more exemplary embodiments, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store or hold desired program code means in the form of instructions or data structures and that can be accessed by a general purpose or special purpose computer, or a general purpose or special purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or infrared, radio, and microwave are included in the definition of medium. As used herein, the disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where the disk typically magnetically reproduces data and the disc optically reproduces data using a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0134] While specific particular embodiments have been described above for purposes of illustration, the teachings of this patent document have general applicability and are not limited to the specific embodiments above. Accordingly, various modifications, adaptations, and combinations of the various features of the described embodiments can be made without departing from the scope of the invention as set forth in the claims.

Claims

1. A measurement tool including an illumination source and a detector configured to collect a certain amount of regularization measurement data from measurements of one or more regularization structures disposed on a first wafer; A computing system, Receiving a certain amount of design of experiments (DOE) measurement data related to measurements of one or more design of experiments (DOE) measurement targets; Receiving known reference values of one or more target parameters related to the DOE measurement target; Receiving the regularization measurement data; Receiving values of one or more measurement performance indicators related to the regularization measurement data; Iteratively training a measurement model based on an optimization function including the certain amount of design of experiments (DOE) measurement data, the reference values of the one or more target parameters, the regularization measurement data, and the one or more measurement performance indicators, wherein the optimization function includes regularization terms related to each of the one or more measurement performance indicators, and weighted values associated with each of the regularization terms are dynamically controlled during iterative training of the measurement model, and iteratively training the measurement model; A computing system configured to perform; A system characterized by comprising.

2. The system according to claim 1, wherein at least a part of the certain amount of design of experiments (DOE) measurement data related to measurements of one or more design of experiments (DOE) measurement targets is generated by simulation.

3. The system according to claim 2, wherein the reference values of the one or more target parameters related to the DOE measurement target are known values related to the simulation.

4. The system according to claim 1, wherein the reference values of the one or more target parameters related to the DOE measurement target are measured by a reliable reference measurement system.

5. The system according to claim 1, wherein at least a part of the certain amount of design of experiments (DOE) measurement data is collected from actual measurement values of one or more design of experiments (DOE) measurement targets disposed on a second wafer.

6. The system according to claim 5, wherein the first wafer and the second wafer are the same wafer.

7. The system according to claim 1, wherein the one or more regularization structures and the one or more design of experiments (DOE) measurement targets have the same structure.

8. The system according to claim 1, wherein the measurement tool collects a certain amount of measurement data from measurements of one or more measurement targets disposed on a third wafer, the one or more measurement targets being characterized by one or more target parameters having unknown values, and the computing system estimates values of the target parameters of the one or more measurement targets from the certain amount of measurement data based on the trained measurement model and is further configured as such.

9. The system according to claim 1, wherein the trained measurement model is any one of a neural network model, a linear model, a non-linear model, a polynomial model, a response surface model, a support vector machine model, a decision tree model, a random forest model, a kernel regression model, a deep network model, and a convolutional network model.

10. The system according to claim 1, wherein the measurement tool is a spectroscopic measurement tool.

11. The system according to claim 1, wherein the weighting values associated with each of the regularization terms are dynamically controlled by any one of a linear quadratic regulator (LQR)-based controller, a proportional integral derivative (PID) controller, an optimal controller, an adaptive controller, and a model predictive controller.

12. Receiving a certain amount of design of experiments (DOE) measurement data related to measurements of one or more DOE measurement targets; Receiving known reference values of one or more target parameters related to the DOE measurement targets; Receiving a certain amount of regularization measurement data from measurements of one or more regularization structures disposed on a first semiconductor wafer by a measurement tool; Receiving values of one or more measurement performance indicators related to the regularization measurement data; Iteratively training a measurement model based on an optimization function that includes the fixed amount of design of experiments (DOE) measurement data, the reference values of the one or more target parameters, the regularized measurement data, and the one or more measurement performance indicators, wherein the optimization function includes regularization terms associated with each of the one or more measurement performance indicators, and the weighting values associated with each of the regularization terms are dynamically controlled during iterative training of the measurement model. A method characterized by including the above. **Claim 13** The method according to claim 12, wherein at least a part of the fixed amount of design of experiments (DOE) measurement data related to the measurement of one or more DOE measurement targets is generated by simulation. **Claim 14** The method according to claim 13, wherein the reference values of the one or more target parameters related to the DOE measurement target are known values related to the simulation. **Claim 15** The method according to claim 12, wherein at least a part of the fixed amount of design of experiments (DOE) measurement data is collected from actual measurement values of one or more DOE measurement targets arranged on a second wafer. **Claim 16** The method according to claim 15, wherein the first wafer and the second wafer are the same wafer. **Claim 17** The method according to claim 12, wherein the one or more regularization structures and the one or more DOE measurement targets have the same structure. **Claim 18** The method according to claim 12, Receiving a fixed amount of measurement data from the measurement of one or more measurement targets arranged on a third wafer by the measurement tool, wherein the one or more measurement targets are characterized by one or more target parameters having unknown values. Estimating the values of the target parameters of the one or more measurement targets from the fixed amount of measurement data based on the trained measurement model. A method further characterized by including the above. **Claim 19** A measurement tool including an illumination source and a detector configured to collect a fixed amount of measurement data from the measurement of one or more measurement targets arranged on a wafer. A computing system, Receiving the fixed amount of measurement data; Receiving values of one or more measurement performance indicators related to the measurement data; Estimating values of one or more target parameters characterizing the one or more measurement targets from the fixed amount of measurement data based on a regression analysis including an optimization function regularized by the one or more measurement performance indicators, wherein the optimization function includes regularization terms related to each of the one or more measurement performance indicators, and the weighting values associated with each of the regularization terms are dynamically controlled during iterative training of the measurement model, estimating values of one or more target parameters; A computing system configured to perform; A system characterized by comprising.

20. Receiving a fixed amount of measurement data from measurements of one or more measurement targets disposed on a semiconductor wafer; Receiving values of one or more measurement performance indicators related to the measurement data; Estimating values of one or more target parameters characterizing the one or more measurement targets from the fixed amount of measurement data based on a regression analysis including an optimization function regularized by the one or more measurement performance indicators, wherein the optimization function includes regularization terms related to each of the one or more measurement performance indicators, and the weighting values associated with each of the regularization terms are dynamically controlled during iterative training of the measurement model; A method characterized by including.

Citation Information

Patent Citations

  • Multi-layer / multi-input / multi-output (mlmimo) model, and method of using the same

    JP2009246368A

  • Model-based hotspot monitoring

    JP2018524821A

  • Model-Based Metrology Using Images

    US20190325571A1

  • System, method and computer program product for fast automatic determination of signals for efficient metrology

    WO2017100424A1

  • Semiconductor metrology with information from multiple processing steps

    WO2017176637A1