Creation and Use of Virtual Features for Process Modeling.
By generating standard feature measurements and parameterizing characteristics to create virtual features for process models, the inefficiencies and costs of conventional substrate processing are addressed, resulting in improved accuracy and reduced waste through optimized process modeling.
Patent Information
- Application Number
- JP2025500943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-11
- Filing Date
- 2023-09-12
- Publication Date
- 2025-09-11
AI Technical Summary
Conventional methods for modeling substrate processing operations are time-consuming, costly, and inaccurate due to variations in substrate features, leading to inefficiencies and increased costs in producing substrates that meet target properties.
Generating a standard feature measurement from metrology data, parameterizing characteristics, and creating arrays of virtual features to input into process models for predictive modeling, allowing for corrective actions to optimize the process.
Improves the accuracy and reduces costs by systematically generating input/output mappings, optimizing process recipes, and minimizing waste and equipment wear, thereby enhancing the production of substrates with target properties.
Smart Images

Figure 2025530069000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to methods associated with process models used to evaluate manufactured devices, such as semiconductor devices. More particularly, the present disclosure relates to methods for generating and utilizing virtual features for process modeling. [Background technology]
[0002] A product may be produced by performing one or more manufacturing processes using manufacturing equipment. For example, a semiconductor manufacturing equipment may be used to produce a substrate using a semiconductor manufacturing process. The product is produced to have certain characteristics suitable for a target application. The characteristics of the substrate input to a process operation have an effect on the output of that process operation. A process model may be used to predict the results of a process operation. Summary of the Invention
[0003] The following is a simplified summary of the present disclosure to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure, nor is it intended to limit the scope of particular embodiments or claims of the present disclosure. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] In one aspect of the present disclosure, a method includes receiving profile data for a plurality of features of a substrate. The method further includes generating a representative profile based on the profile data of the plurality of features. The method further includes generating a first array of features. Each of the first array of features is based on the representative profile. The method further includes providing the first array of features to a process model. The method further includes obtaining a first output from the process model based on the first array of features. The method further includes performing a corrective action in consideration of the first output from the process model.
[0005] In another aspect of the present disclosure, a system includes a memory and a processing device coupled to the memory. The processing device is configured to perform an operation. The operation includes receiving profile data for a plurality of features of the substrate. The operation further includes generating a representative profile based on the profile data for the plurality of features. The operation further includes generating a first array of features. Each of the first array of features is based on the representative profile. The operation further includes providing the first array of features to a process model. The operation further includes obtaining a first output from the process model based on the first array of features. The operation further includes performing a corrective action in consideration of the first output from the process model.
[0006] In another aspect of the present disclosure, a non-transitory machine-readable storage medium has stored thereon instructions that, when executed, cause a processing device to perform operations. The operations include receiving profile data for a plurality of features of a substrate. The operations further include generating a representative profile based on the profile data for the plurality of features. The operations further include generating a first array of features. Each of the first array of features is based on the representative profile. The operations further include providing the first array of features to a process model. The operations further include obtaining a first output from the process model based on the first array of features. The operations further include performing a corrective action in consideration of the first output from the process model.
[0007] In the figures of the accompanying drawings, the present disclosure is illustrated by way of example and not by way of limitation. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary system architecture according to some embodiments. [Figure 2A] FIG. 1 is a block diagram of an example dataset generator for generating datasets for one or more supervised models, according to some embodiments. [Figure 2B] FIG. 1 is a block diagram of an example dataset generator for generating datasets for one or more unsupervised models, according to some embodiments. [Figure 3] FIG. 1 is a block diagram illustrating a system for generating output data, according to some embodiments. [Figure 4A] 1 is a flow diagram of a method for generating a dataset for a machine learning model, according to some embodiments. [Figure 4B] 1 is a flow diagram of a method for utilizing measurements of characteristics of a substrate to perform corrective action, according to some embodiments. [Figure 4C]1 is a flow diagram of a method for generating and utilizing a parameterization of features, according to some embodiments. [Figure 4D] 1 is a flow diagram of a method for obtaining predicted outputs from a process model, according to some embodiments. [Figure 5] 1A-1C illustrate exemplary substrates including features, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating a computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Described herein are techniques related to modeling the operation of processing procedures using a range of applicable input features. Manufacturing equipment is used to produce products such as substrates (e.g., wafers, semiconductors). The manufacturing equipment may include fabrication or processing chambers for isolating the substrates from their environment. The properties of the produced substrates should meet target values to facilitate a specific function. Manufacturing parameters are selected to produce substrates that meet the target property values. Many manufacturing parameters (e.g., hardware parameters, process parameters, etc.) contribute to the properties of the processed substrates. The manufacturing system may specify set points for the property values, receive data from sensors located within the manufacturing chambers, and control the parameters by adjusting the manufacturing equipment until the sensor readings match the set points.
[0010] A processing procedure (e.g., a method for manufacturing a substrate) may include many processing operations (e.g., processing steps). For example, a semiconductor wafer may be manufactured by adding material to a substrate in one or more deposition operations, removing material from a substrate in one or more etching operations, changing the properties of the substrate in one or more annealing operations, etc. For example, a deposition operation may deposit material on the surface of the substrate, into holes in the substrate, on the sidewalls of features or structures in the substrate, etc. The output of a processing operation depends on the input product on which the operation is performed. For example, the result of a deposition operation depends on the properties of the substrate on which the material is deposited.
[0011] It may be beneficial to predict the results of a processing operation. The results may be predicted through modeling, for example, using one or more physics-based models to predict the results of the processing operation. The physics-based models may include deposition models, etch models, and / or models for any type of processing performed on a substrate. The models may be provided with data indicating an input substrate, and after performing the modeled processing operation, the models may generate a prediction of an output substrate. The substrate process model may include a simulation model that receives as input simulation parameters such as an etch or deposition rate, or the change in etch or deposition rate over time. The substrate process model may include a model that receives as input process parameters such as gas flow rates, temperatures, and radio frequency (RF) parameters.
[0012] Some conventional systems may model the characteristics of the substrate that is input to a processing operation. Performing the modeling may be an expensive process in terms of time, processing power, subject matter expertise, etc. In some systems, multiple processing operations may be performed sequentially. The results of each operation may affect the next operation. The modeling approach may take into account each of the many processing operations, which further multiplies the time and cost investment for accurate modeling results.
[0013] Some conventional systems measure the characteristics of a substrate input to a processing operation and generate a model from the measurements. The substrate may contain multiple features, such as pillars, gates, trenches, holes, or the like. The multiple features may be nominally or ideally identical. For example, a substrate processing procedure may aim to produce a substrate with a two-dimensional or three-dimensional array of identical features. Differences between features may occur due to inhomogeneities within the process chamber, differences between features of input structures to the processing operation, differences in measurements of features, etc. Applying a process model (e.g., a deposition model) may produce results that are affected by differences between substrate features. The results of the process model may not produce a clear picture of the impact the input feature structure has on the output features. The results of the process model may not produce results with a clear indication of the changes or modifications to be made to the input structure to produce the target output structure. The results of the process model may be limited to the shape, properties, and parameters of the features present on the substrate measured to provide input to the process model.
[0014] In some conventional systems, input features of multiple substrates may be measured and provided to a process model. A range of substrates may provide more data for input / output mapping than modeling based on a smaller number of substrates, a wider range of input features to cover a wider input and output space, etc. Providing a larger number of substrates for measurement may increase costs in terms of materials, energy, time, equipment wear, maintenance, product disposal costs, metrology costs, etc. Generating a larger range of input features may involve modifying one or more process operation recipes, which may reduce the usefulness of the results. Modeling using measurements of many substrates may involve expending time, energy, processing power, etc., to model several essentially identical processes without gaining much new information. Using measurements of many substrates as input for process modeling may have some disadvantages similar to using one or a small number of substrates. For example, certain combinations of feature parameters / shapes may not be included in the sample set, and differences between adjacent features or features on the same substrate may obscure the cause of modeling results.
[0015] The systems and methods of the present disclosure may address one or more of these shortcomings of conventional methods. In some embodiments, one or more measurements of a substrate are provided. The measurements may be generated by one or more metrology tools. The measurements may be obtained from one or more microscopic images. The measurements may be obtained, for example, from one or more scanning electron microscope (SEM) images, one or more cross-sectional scanning electron microscope (XSEM) images, and / or one or more transmission electron microscope (TEM) images.
[0016] From these microscopic images, measurements of one or more features of one or more substrates may be extracted. The one or more substrates may include one or more features. The features may include substrate structures, shapes, profiles, or the like. The features may include gates, holes, trenches, masks, spacers, or any other characteristics that may be present on the substrate of interest. The feature measurements may include data points indicating the boundaries of the feature, a function fit to the shape of the feature, measurements of the size of one or more shapes or portions of the feature, or the like. The measurements of one or more features may be obtained from a partially processed substrate, e.g., a substrate that has undergone some processing operations of a process recipe but not yet undergone other processing operations. The measurements of one or more features may also be obtained from a completed substrate, e.g., a substrate that has undergone all processing operations of a process recipe. The microscopic image from which feature measurements are extracted may be taken from a partially completed substrate or a completed substrate.
[0017] In some embodiments, measurements of one or more features are provided to a model configured to generate a standard feature measurement. The standard feature may be generated by performing a statistical analysis based on measurements of one or more features. The standard feature measurement may be an average of measurements of multiple features. The multiple features may be nominally identical, e.g., a process recipe associated with the production of substrates may aim to produce an array of identical features. The substrate features and / or feature measurements may not be identical due to differences in processing conditions, measurement accuracy, or the like. Generating the standard feature measurement may include generating measurements of features likely to be produced by a substrate processing procedure. Generating the standard feature measurement may include generating measurements of features likely to be produced by a substrate processing procedure. Generating the standard feature measurement may include selecting measurements of a substrate feature to designate as the standard. Generating the standard feature measurement may include combining the standard feature measurements and selecting several measurements from several measured features to designate as standard feature measurements. Generating the standard feature measurement may include generating measurements unrelated to the measured features. Generating measurements of the standard features may include statistical analysis of the features. Generating measurements unrelated to the measured features may include generating an average, median, and / or ideal set of measurements from multiple measurements of the features to serve as the standard features.
[0018] In some embodiments, one or more characteristics of the standard feature are parameterized. The parameterization may include the slope of a portion of the feature, the size of a portion of the feature, the radius of curvature of a portion of the feature, or the like. The parameterization may be performed manually. The parameterization may be performed by a model. The parameterization may be performed by a machine learning model. The parameterization may include parameterizing a characteristic of the standard feature, including the variation among the multiple features used to generate the standard feature. The parameterization may enable the generation of variation of the standard feature. For example, each parameter may have an associated range. The range may be generated from a statistical metric associated with the multiple features used to generate the standard feature, such as a number of standard deviations from the mean, inner quartiles, range, or another metric of the parameter values of the characteristic of the multiple features. Statistical analysis may be performed to determine the range of the parameter. The range may be manually generated and / or adjusted. Combinations of parameter values within the associated range may be systematically or randomly utilized to generate multiple parameterized representations of the feature. The multiple parameterized representations of the feature may substantially span the space of likely feature shapes.
[0019] Each of multiple parameterized representations of a feature (e.g., each feature of a unique shape) may be used to generate an array of identical features. The array may be two-dimensional. The array may be three-dimensional. Each array of identical features may be provided to a process model. The process model may digitally perform one or more process operations on a virtual substrate including the array of features. For example, the process model may perform a deposition operation, an etching operation, or the like, on a substrate including the array of identical features. For example, the process model may model the deposition of material on the surface of the substrate, in holes in the substrate, or the like. One or more outputs of interest may be extracted from the process model, such as the thickness of a deposited layer at one or more locations, the width of an etched hole at one or more depths, or the like. In some embodiments, one or more virtual substrates including arrays of non-identical features may be generated and provided to the process model.
[0020] An input / output mapping may be generated from the results of the process model. The input / output mapping may be generated by a fitting model. The input / output mapping may be generated by a machine learning model. The input / output mapping may include a set or list of inputs to the process model correlated to the output results of the process model. The input / output mapping may include combinations of inputs that are likely to produce a substrate having a target output quality. The inputs to the input / output mapping may include parameters of the substrate features. The inputs to the input / output mapping may include measurements of properties of the substrate features.
[0021] Corrective actions may be taken in light of the input / output mapping. The input / output mapping may inform target input substrate geometries to the process operation to achieve the target output of the process operation. For example, one or more corrective actions may be taken to achieve an input profile or feature shape to the process operation to facilitate the target output after performing the process operation. The corrective action may include updating the process recipe. The corrective action may include performing maintenance on the process chamber and / or one or more chamber parts of the process chamber. The corrective action may include scheduling maintenance on one or more chamber parts. The corrective action may include providing an alert to a user. The corrective action may include performing a cleaning or seasoning operation on the process chamber and / or scheduling a cleaning or seasoning operation on the process chamber.
[0022] Aspects of the present disclosure provide technical advantages over conventional methods. Utilizing metrology data of one or more substrates to generate a virtual substrate to feed into a process model may be an improvement over other methods. Modeling process operations may be expensive in terms of time, processing power, energy, etc. Modeling a substrate for input to a process model may include simulating multiple process operations, which may double the cost of modeling. Utilizing metrology data (e.g., microscope images) as a basis for modeling a substrate to feed into a process model may reduce costs compared to modeling the process operations to generate the substrate.
[0023] Utilizing measurements of multiple features to generate a standard feature may improve the accuracy and / or applicability of the process model results. Substrates may contain an array of slightly different features. Differences may occur due to differences in processing conditions. Variations in the measurements of various features may result in differences in the simulated substrate features. Differences in the virtual features of the simulated substrate (e.g., differences between features represented in the data provided to the process model) may interfere with the interpretation of the model results. This may require expensive modeling and / or measuring additional substrates, performing additional process modeling, and the like, to separate effects due to feature variations from effects due to feature design.
[0024] Parameterizing standard features and generating multiple arrays of features based on varying the parameters of the standard features may improve learning and enable more process improvements than conventional methods. Conventional methods may rely on measured or process-modeled features to generate input / output mappings for process operations. By parameterizing standard features, substrates containing feature variations may be systematically and / or randomly generated and provided to a process model, where the substrates approximately span the applicable space of feature geometry dimensions, for example. Known changes to feature geometry based on the parameters of the features may be utilized to generate input / output mappings. By adjusting the parameters of the standard features, generating substrates containing arrays of the adjusted features, and providing the substrates to a process model, testing different feature geometries, optimizing feature geometries for input to process operations, and the like may be easily performed.
[0025] Taking one or more corrective actions in consideration of the input / output mapping according to the methods of the present disclosure may improve the characteristics of the processed substrate. Optimizing the input substrate to a process operation may increase the likelihood of producing an output from the process operation that meets one or more performance targets. Improving a process recipe, process chamber, or the like in consideration of the input / output mapping has the advantage of reducing the time, energy, chamber wear, chamber maintenance, chamber maintenance time, replacement parts, materials, and / or waste costs associated with producing a product that does not meet target performance standards.
[0026] In one aspect of the present disclosure, a method includes receiving profile data for a plurality of features of a substrate. The method further includes generating a representative profile based on the profile data of the plurality of features. The method further includes generating a first array of features. Each of the first array of features is based on the representative profile. The method further includes providing the first array of features to a process model. The method further includes obtaining a first output from the process model based on the first array of features. The method further includes performing a corrective action in consideration of the first output from the process model.
[0027] In another aspect of the present disclosure, a system includes a memory and a processing device coupled to the memory. The processing device is configured to perform an operation. The operation includes receiving profile data for a plurality of features of the substrate. The operation further includes generating a representative profile based on the profile data for the plurality of features. The operation further includes generating a first array of features. Each of the first array of features is based on the representative profile. The operation further includes providing the first array of features to a process model. The operation further includes obtaining a first output from the process model based on the first array of features. The operation further includes performing a corrective action in consideration of the first output from the process model.
[0028] In another aspect of the present disclosure, a non-transitory machine-readable storage medium has stored thereon instructions that, when executed, cause a processing device to perform operations. The operations include receiving profile data for a plurality of features of a substrate. The operations further include generating a representative profile based on the profile data for the plurality of features. The operations further include generating a first array of features. Each of the first array of features is based on the representative profile. The operations further include providing the first array of features to a process model. The operations further include obtaining a first output from the process model based on the first array of features. The operations further include performing a corrective action in consideration of the first output from the process model.
[0029] 1 is a block diagram illustrating an example system 100 (example system architecture) according to some embodiments. System 100 includes client devices 120, manufacturing equipment 124, sensors 126, measurement equipment 128, a prediction server 112, and a data store 140. Prediction server 112 may be part of a prediction system 110. Prediction system 110 may further include server machines 170 and 180.
[0030] The sensors 126 may provide sensor data 142 associated with the manufacturing equipment 124 (e.g., associated with the manufacturing equipment 124 producing a corresponding product, such as a substrate). The sensor data 142 may be used to ascertain the health of the equipment and / or the health of the product (e.g., product quality). The manufacturing equipment 124 may produce a product according to a recipe or by running for a period of time. In some embodiments, the sensor data 142 may include one or more values of optical sensor data, spectral data, temperature (e.g., heater temperature), spacing (SP), pressure, high frequency radio frequency (HFRF), radio frequency (RF) match voltage, RF match current, RF match capacitor position, electrostatic chuck (ESC) voltage, actuator position, current, flow rate, power, voltage, etc. The sensor data (e.g., portions of the sensor data 142) may relate to a product currently being processed, recently processed products, the number of recently processed products, etc. The sensor data may include data stored related to previously produced products. The sensor data 142 may include attribute data, a label of the status of the manufacturing equipment, etc. Examples of attribute data include manufacturing equipment ID or design label, sensor ID, type and / or location. Examples of manufacturing equipment status labels include current faults, service life, etc.
[0031] The sensor data 142 may be related to, correlated with, or indicative of manufacturing parameters, such as hardware parameters of the manufacturing equipment 124 or process parameters of the manufacturing equipment 124. Examples of hardware parameters include hardware settings or installed components, such as the size and type of installed components. Examples of process parameters include heater settings, gas flow rate settings, pressure settings, etc. Alternatively or additionally, data related to some hardware parameters and / or process parameters may be stored as manufacturing parameters 150. The manufacturing parameters 150 may include historical manufacturing parameters (e.g., associated with historical processing runs) and current manufacturing parameters. The manufacturing parameters 150 may indicate input settings for manufacturing devices (e.g., heater power, gas flow rate, etc.). The sensor data 142 and / or manufacturing parameters 150 may be provided while the manufacturing equipment 124 is performing a manufacturing process (e.g., equipment readings while processing a product). The sensor data 142 may vary from product to product (e.g., from substrate to substrate). The substrate may have properties measured by metrology equipment 128. Examples of properties include film thickness, film strain, critical dimensions, optical properties, electrical properties, etc. The properties may be measured in stand-alone metrology equipment, by integrated or in-line metrology systems, or in other similar manners. Metrology data 160 may be a component of data store 140. Metrology data 160 may include historical metrology data (e.g., metrology data associated with previously processed products).
[0032] In some embodiments, metrology data 160 may be provided without the use of stand-alone metrology equipment. For example, metrology data 160 may be in-situ metrology data (e.g., measurements or surrogates of measurements collected during processing), integrated metrology data (e.g., measurements or surrogates of measurements collected while a product is in a chamber or under vacuum, other than during a processing operation), in-line metrology data (e.g., data collected after a substrate is removed from vacuum), etc. Metrology data 160 may include current metrology data (e.g., metrology data associated with a product currently being processed or a recently processed product).
[0033] Metrology equipment 128 may include microscopes and / or imaging equipment. Metrology equipment 128 may include one or more devices for obtaining images of a substrate, portions of a substrate, features of a substrate, or the like. Metrology equipment 128 may include SEM equipment, XSEM equipment, TEM equipment, and / or other forms of imaging and microscopy equipment. Metrology data 160 may include image data, microscopy data, and the like.
[0034] In some embodiments, the sensor data 142, the metrology data 160, or the manufacturing parameters 150 may be processed (e.g., by the client device 120 and / or the prediction server 112). Processing the sensor data 142 may include generating features. In some embodiments, the features are patterns in the sensor data 142, the metrology data 160, and / or the manufacturing parameters 150. Examples of such features include slope, width, height, peaks, etc. In some embodiments, the features are combinations of values from the sensor data 142, the metrology data, and / or the manufacturing parameters. Examples of such features include power derived from voltage and current, etc. The sensor data 142 may include features, which may be used by the prediction component 114 to perform signal processing and / or to obtain predicted data 168 for taking corrective actions.
[0035] Each instance (e.g., set) of sensor data 142 may correspond to a product (e.g., substrate), a set of manufacturing equipment, a type of substrate produced by the manufacturing equipment, or the like. Similarly, each instance of metrology data 160 and manufacturing parameters 150 may correspond to a product, a set of manufacturing equipment, a type of substrate produced by the manufacturing equipment, or the like. The data store may further store information relating sets of different data types, e.g., information indicating that the set of sensor data, the set of metrology data, and the set of manufacturing parameters all relate to the same product, the same manufacturing equipment, the same type of substrate, etc.
[0036] The data store 140 may further include virtual substrate data 162. The virtual substrate data 162 may include data regarding a simulated substrate, a synthetic substrate, and / or a virtual substrate. The virtual substrate data 162 may include feature measurements, feature parameters, feature images, etc. Various properties and representations of features may be stored as feature data 164. The virtual substrate data 162 may include measurements, parameters, and / or images of a simulated substrate. Properties and representations of the simulated substrate and / or virtual substrate may be stored as substrate data 166. The substrate data 166 may include two-dimensional and / or three-dimensional arrays of features.
[0037] In some embodiments, the prediction system 110 may generate prediction data 168. The prediction data 168 may be generated using one or more models, such as physics-based models, deposition models, machine learning models, etc. The prediction data 168 may include the output of a process model. The prediction data 168 may include predicted results of one or more process operations applied to a virtual substrate. The prediction data 168 may include predicted shortcomings of a process operation, recipe, or equipment. The prediction data 168 may include recommended corrective actions, such as corrective action data. Operation of the prediction system 110 may include the use of one or more supervised models trained using input data labeled with target outcome data. Operation of the prediction system 110 may include the use of one or more unsupervised models trained using input data unlabeled with target output data. Operation of the prediction system 110 may include the use of one or more semi-supervised models that include a mixture of labeled and unlabeled input data during training.
[0038] Client device 120, manufacturing equipment 124, sensors 126, measurement equipment 128, prediction server 112, data store 140, server machine 170, and server machine 180 may be coupled to one another via network 130 to generate predictive data 168 for performing corrective actions. In some embodiments, network 130 may provide access to cloud-based services. Operations performed by client device 120, prediction system 110, data store 140, etc. may be performed by a cloud-based virtual device.
[0039] In some embodiments, network 130 is a public network that provides client device 120 with access to prediction server 112, data store 140, and other public computing devices. In some embodiments, network 130 is a private network that provides client device 120 with access to manufacturing equipment 124, sensors 126, measurement equipment 128, data store 140, and other private computing devices. Network 130 may include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), routers, hubs, switches, server computers, cloud computing networks, and / or combinations thereof.
[0040] Client device 120 may include computing devices such as personal computers (PCs), laptops, mobile phones, smartphones, tablet computers, netbook computers, network-connected televisions (“smart TVs”), network-connected media players (e.g., Blu-ray players), set-top boxes, over-the-top (OTT) streaming devices, operator boxes, etc. Client device 120 may include a corrective action component 122. Corrective action component 122 may receive user input of instructions related to manufacturing equipment 124 (e.g., via a graphical user interface (GUI) displayed via client device 120). In some embodiments, corrective action component 122 transmits the instructions to prediction system 110, receives output (e.g., prediction data 168) from prediction system 110, determines corrective action based on the output, and causes the corrective action to be implemented. In some embodiments, the corrective action component 122 obtains sensor data 142 related to the manufacturing equipment 124 (e.g., from the data store 140) and provides the sensor data 142 related to the manufacturing equipment 124 to the predictive system 110.
[0041] In some embodiments, metrology data 160 may be provided to prediction system 110, prediction server 112, prediction component 114, model 190, or the like. Measurement data 160 may be retrieved from data store 140 by corrective action component 122 and provided to prediction system 110. Prediction system 110 may generate substrate data 166 and / or prediction data 168 as output feature data 164, either of which may be stored in data store 140. Client device 120 may retrieve output from prediction system 110 (e.g., via corrective action component 122) and provide the output to data store 140. In some embodiments, corrective action component 122 stores data in data store 140 for use as input to machine learning models, physics-based models, or other models. In some embodiments, a component of prediction system 110 (e.g., prediction server 112, server machine 170) retrieves input data from data store 140. In some embodiments, the prediction server 112 may store the output of the trained model 190 (e.g., the predicted data 168) in the data store 140, and the client device 120 may retrieve this output from the data store 140.
[0042] In some embodiments, a corrective action component 122 receives corrective action instructions from the predictive system 110 and causes the corrective action to be implemented. Each client device 120 may include an operating system that enables a user to one or more of generating, examining, or editing data. This data may include, for example, instructions related to the manufacturing equipment 124, corrective actions related to the manufacturing equipment 124, etc. The client device 120 may also include a component or system for providing alerts to the user. The alerts may be of potential shortcomings in process operations, process procedures, process recipes, process equipment, etc.
[0043] In some embodiments, metrology data 160 corresponds to historical property data of a product, and prediction data 168 relates to predicted property data. The historical property data of a product may include manufacturing parameters associated with historical sensor data and data of a product processed using the historical manufacturing parameters. The predicted property data may include data of a product that will be produced or that was produced under conditions recorded by current sensor data and / or current manufacturing parameters. In some embodiments, prediction data 168 is or includes predicted metrology data (e.g., virtual metrology data, virtual composite microscope images) of a product that will be produced or that was produced according to conditions recorded as current sensor data, current measurement data, current metrology data, and / or current manufacturing parameters. The prediction data 168 may include results of submitting a simulated substrate to a process model. The prediction data 168 may include predictions of results of applying a process operation to a substrate. The prediction data 168 may include mapping data. The mapping data may include correlating input substrate properties with output substrate properties of a process operation. The mapping data may include predicting properties of an output substrate of a process operation based on properties of an input substrate. Substrate properties may include parameters, features, feature profiles, dimensions, etc. In some embodiments, the predicted data 168 is or includes an indication of any anomalies and, optionally, one or more causes of those anomalies. Anomalies may include an abnormal product, an abnormal component, an abnormal equipment, an abnormal material or energy usage, etc. In some embodiments, the predicted data 168 is an indication of a time change or drift of a component of manufacturing equipment 124, a sensor 126, a metrology device 128, or the like. In some embodiments, the predicted data 168 is an indication of the end of life of a component of manufacturing equipment 124, a sensor 126, a metrology device 128, or the like. In some embodiments, the predicted data 168 is an indication of the progress of a processing operation being performed.In some embodiments, the predictive data 168 may be used for process control.
[0044] Performing a manufacturing process that results in a defective product can be costly in terms of time, energy, product, components, manufacturing equipment 124, costs of identifying defects and discarding the defective product, etc. By inputting metrology data 1602 (e.g., measurements extracted from TEM or XSEM images of a substrate) into the prediction system 110, receiving output of prediction data 168, and performing corrective action based on the prediction data 168, the system 100 may have the technical advantage of avoiding the costs of producing, identifying, and discarding defective product. The system 100 may increase the likelihood of producing substrates having properties within target thresholds. Increasing the likelihood of producing substrates having properties within target thresholds may reduce the production cost per successful substrate. Production costs may be reduced in areas such as production time, materials, energy, equipment component wear, increased process chamber downtime, and increased maintenance costs.
[0045] Running a manufacturing process that results in a component failure of manufacturing equipment 124 can be costly in terms of downtime, damage to product, damage to equipment, rush-ordering replacement components, etc. The systems and / or methods of the present disclosure may mitigate one or more of these deficiencies. By inputting a virtual substrate based on measured characteristic properties into a model, receiving an output, and performing corrective actions, system 100 may have a technical advantage over conventional systems. The virtual substrate may be based on metrology data 160. The virtual substrate may be generated based on one or more microscope images. The corrective actions may include predicted operational maintenance. The corrective actions may include component replacement, treatment, cleaning, etc. System 100 may have the technical advantage of avoiding the cost of unexpected component failure. System 100 may have the advantage of avoiding the cost of unscheduled downtime. System 100 may have the advantage of avoiding the cost of lost productivity due to equipment downtime. System 100 may have the advantage of avoiding the cost of product waste. By utilizing the systems and / or methods of the present disclosure, the system 100 may avoid these costs plus additional costs. Differences between predicted and measured properties of a substrate may include indications of drifting, aging, or faulty equipment. Performance of components, such as manufacturing equipment 124, sensors 126, metrology equipment 128, and the like, may be monitored over time to provide an indication of deteriorating components.
[0046] The manufacturing parameters may be suboptimal for producing a product, and producing the product may have costly consequences, such as increased resource (e.g., energy, coolant, gas, etc.) consumption, increased time to produce the product, increased component failures, and increased amount of defective product. By inputting metrology indications into the prediction system 110 and using the output data to perform corrective actions, the system 100 may have a technical advantage over conventional methods. The metrology indications may include a virtual substrate. The virtual substrate may be based on measured characteristics of the substrate. The output of the prediction system 110 may include predicted data 168. The corrective action may include updating the manufacturing parameters. Updating the manufacturing parameters may include setting the optimal manufacturing parameters to produce the product. The system 100 may have the technical advantage of utilizing more advantageous manufacturing parameters. The manufacturing parameters may include hardware parameters, process parameters, input substrate characteristics, etc. The system 100 may avoid the costly consequences of utilizing suboptimal manufacturing parameters.
[0047] The corrective action may be related to one or more types of process control. Process control may include computational process control (CPC), statistical process control (SPC), advanced process control (APC), model-based process control, etc. SPC may include control of electronic components to determine process progress. SPC may include predicting the useful life of components. SPC may include comparing data with historical data, for example, comparing trace data to historical data to determine whether the trace data is within a three-sigma window of the mean. The corrective action may be related to preventive operational maintenance, design optimization, manufacturing parameter updates, manufacturing recipe updates, feedback control, machine learning corrections, or the like.
[0048] In some embodiments, the corrective action includes issuing an alert to a user. The alert may include an alert to stop or not execute a manufacturing process. An alert may be provided when the forecast data 168 indicates an anomaly. An alert may be provided when the forecast data 168 indicates an abnormal product, component, equipment, etc. In some embodiments, performing the corrective action includes causing updates to one or more manufacturing parameters to be performed. In some embodiments, performing the corrective action may include retraining a machine learning model associated with the manufacturing equipment 124. Performing the corrective action may also include updating other types of models associated with the manufacturing equipment 124, such as adjusting physics-based models, process models, or other similar models. In some embodiments, performing the corrective action may include training a new machine learning model and / or developing a new physics-based model or process model associated with the manufacturing equipment 124.
[0049] The manufacturing parameters 150 may include hardware parameters and / or process parameters. Hardware parameters may include information indicating which components are installed in the manufacturing system, an indication of component age, an indication of software versions or updates, etc. Process parameters may include temperature, pressure, gas flow rate, current, voltage, lift speed, etc. In some embodiments, the corrective action includes performing preventative operational maintenance. Preventative operational maintenance may include replacing, treating, cleaning, etc., components of the manufacturing system. In some embodiments, the corrective action includes performing design optimization. Design optimization may include updating manufacturing parameters, updating the manufacturing process, and / or updating manufacturing equipment to improve the performance of the manufacturing system. In some embodiments, the corrective action includes updating a recipe. Changing the recipe may include changing when manufacturing subsystems enter idle or active modes, changing set points for various characteristic values, or the like.
[0050] Prediction server 112, server machine 170, and server machine 180 may each include one or more computing devices such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc. Operations of prediction server 112, server machine 170, server machine 180, data store 140, etc. may be performed by a cloud computing service, a cloud data storage service, etc.
[0051] The prediction server 112 may include a prediction component 114. In some embodiments, the prediction component 114 may receive metrology data 160 and generate an output based on the current data for performing corrective actions related to the manufacturing equipment 124. The metrology data 160 may be received from, for example, the client device 120, retrieved from a data store 140, or the like. The output of the prediction component 114 may be prediction data 168. In some embodiments, the prediction data 168 may include one or more predicted dimensional measurements of the processed product. In some embodiments, the prediction component 114 may use one or more trained machine learning models 190 to determine an output based on the current data for performing corrective actions.
[0052] In some embodiments, prediction server 110 may receive metrology data 160 (e.g., measurements of one or more substrates) and generate feature data 164 as output. The output feature data may include standard features. Further information regarding the formation of standard features is discussed in connection with FIG. 4B . Prediction system 110 may receive feature data 164 (e.g., measurements of standard features, parameters of standard features, measurements or parameters of non-standard features, etc.) and generate one or more virtual substrates as output. The virtual substrates may be stored as substrate data 166. Prediction system 110 may receive one or more virtual substrates and generate output. This output may be stored as prediction data 168. This output may include predicted properties of the substrates after process operations are performed on the substrates. This output may include one or more effects that the properties of the input substrates have on the output substrates. This output may include an input / output mapping in the form of data points, data trends, multidimensional fits, or the like.
[0053] Manufacturing equipment 124 may be associated with one or more machine learning models, such as model 190. The machine learning models associated with manufacturing equipment 124 may perform many tasks, including process control, classification, performance prediction, etc. Model 190 may be trained using data associated with manufacturing equipment 124 or data associated with products processed by manufacturing equipment 124, such as sensor data 142 (e.g., collected by sensors 126), manufacturing parameters 150 (e.g., associated with process control of manufacturing equipment 124), metrology data 160 (e.g., generated by metrology equipment 128), etc.
[0054] One type of machine learning model that may be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. Artificial neural networks generally include a feature representation component with a classifier or recurrent layer that maps features to a desired output space. For example, a convolutional neural network (CNN) hosts multiple layers of convolutional filters. Pooling may be performed and nonlinearities may be addressed in lower layers, and above the lower layers, a multilayer perceptron is typically added to map upper layer features extracted by the convolutional layers to a decision (e.g., a classification output).
[0055] A recurrent neural network (RNN) is another type of machine learning model. Recurrent neural network models are designed to interpret a series of inputs that are intrinsically related to each other, such as time trace data, sequential data, etc. The output of a perceptron in an RNN is fed back as input to that perceptron to generate the next output.
[0056] Deep learning is a type of machine learning algorithm that uses a cascade of multiple layers of nonlinear processing units to perform feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks may learn in a supervised (e.g., classification) and / or unsupervised manner (e.g., pattern analysis). Deep neural networks include a hierarchy of layers, with different layers learning different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly more abstract and complex representation. For example, in an image recognition application, the raw input may be a matrix of pixels; a first representation layer may extract the pixels and encode edges; a second layer may construct and encode the edge configuration; a third layer may encode higher-level shapes (e.g., teeth, lips, gums, etc.); and a fourth layer may recognize scanning tasks. Notably, the deep learning process can independently learn which features are optimally placed at which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have significant credit assignment path (CAP) depth. A CAP is a chain of transformations from input to output. A CAP describes the potentially causal connections between input and output. For feedforward neural networks, the CAP depth may be the depth of the network or the number of hidden layers + 1. For recurrent neural networks, where signals may propagate through layers more than once, the CAP depth is potentially infinite.
[0057] In some embodiments, the prediction component 114 and / or the model 190 may include a process model. The process model may predict the results of performing one or more process operations. The process model may be a physics-based model, a simulation model, a machine learning model, etc.
[0058] In some embodiments, the prediction component 114 and / or the model 190 may include a model for generating a parameterized representation of a feature. The model may be a machine learning model. The model may receive as input several measurements of a plurality of features, statistical metrics related to the measurements of the plurality of features, measurements of standard features, or the like. The model may determine which characteristics of the feature to parameterize based, for example, on characteristics that vary among the measured features. The characteristics may include the rounding radius of the feature (e.g., top of a gate, corner of a sidewall, bottom of a hole, etc.), slope (e.g., slope of a sidewall), length, critical dimension, or other characteristics.
[0059] In some embodiments, the prediction component 114 and / or the model 190 include a model for generating combinations of parameter values for modeling. Some characteristics of the feature may be parameterized. One or more parameters of the feature may be adjusted to generate an updated feature. For example, parameters may be adjusted to generate a feature having a different shape, different dimensions, etc. from a standard feature. Some parameter values may be impossible, unlikely, or unprofitable for substrate production. Some combinations of parameter values may be impossible or unprofitable to produce, for example, and may not produce a substrate with target properties or performance, etc. The parameterization of the feature may be provided to a model, which may be configured to determine which combinations of parameter values are likely to provide useful information under further analysis. For example, the model may determine which combinations of parameters are impossible or unlikely to correlate to physical structures, which combinations of parameters are unlikely to produce favorable results, which combinations of parameters are cost-prohibitive to produce, etc. The model may be a model based on physical phenomena, a simulation model, a machine learning model, etc.
[0060] In some embodiments, the prediction component 114 and / or the model 190 include a model for input / output mapping. The prediction system 110 may be configured to generate multiple virtual substrates, each including an array of features, and perform operations to provide the substrates to a process model. Differences between the substrates provided to the process model may be correlated with differences in the output results of the process model. For example, there may be an input / output mapping associated with the procedure of the prediction system 110. The model may be used to extract the input / output mapping from input and output data of the process model. The model may be used to provide several input parameters that have a significant impact on the output results. The model may be used to generate input designs (e.g., feature parameters, feature shapes, etc.) that are likely to produce target outputs. For example, the model may be used to optimize input parameters to produce target output parameters. Models such as these models may be models based on physical phenomena, deformation models, e.g., principal component analysis models, machine learning models, etc.
[0061] In some embodiments, the prediction component 114 receives the metrology data 160, performs signal processing to decompose the data into data sets, provides the current data sets as inputs to the trained model 190, and obtains output from the trained model 190 that indicates predicted data 168. In some embodiments, the prediction component 114 receives metrology data for the substrate (e.g., predicted metrology data based on sensor data) and provides the metrology data to the trained model 190. The model 190 may be configured to accept data indicative of substrate metrology and generate predicted input / output mapping data as an output. In some embodiments, the predicted data indicates metrology data (e.g., a prediction of substrate quality). In some embodiments, the predicted data indicates component health.
[0062] In some embodiments, the various models discussed in connection with model 190 may be combined into one model or may be separate models. For example, supervised machine learning models, unsupervised machine learning models, and / or models based on physical phenomena may be combined into one or more ensemble models.
[0063] Data may be passed back and forth between several separate models included in models 190 and prediction component 114. In some embodiments, some or all of these operations may instead be performed by different devices, such as client device 120, server machine 170, server machine 180, etc. Those skilled in the art will understand that variations in data flow, which components perform which processes, which data is provided to which models, and the like, are within the scope of this disclosure.
[0064] The data store 140 may be a memory (e.g., random access memory), a drive (e.g., a hard drive, a flash drive), a database system, a cloud-accessible memory system, or another type of component or device capable of storing data. The data store 140 may include multiple storage components (e.g., multiple drives or multiple databases) that may reside across multiple computing devices (e.g., multiple server computers). The data store 140 may store sensor data 142, manufacturing parameters 150, metrology data 160, virtual substrate data 162, and prediction data 168.
[0065] The sensor data 142 may include historical sensor data and / or current sensor data. The sensor data may include time tracking of sensor data throughout the duration of the manufacturing process, association of data with physical sensors, preprocessed data such as averages and composite data, and data indicating sensor performance over time (i.e., across many manufacturing processes). The manufacturing parameters 150 and metrology data 160 may include similar features. For example, the metrology data 160 may include historical metrology data and / or current metrology data. The historical sensor data, historical metrology data, and historical manufacturing parameters may be historical data. At least a portion of the historical data may be used to train the model 190. The current sensor data, current metrology data may be current data for which predicted data 168 is generated. The current data may be provided as input to one or more trained models. The predicted data 168 may be used to perform one or more corrective actions. The virtual substrate data 162 may include data to be provided to a process model to generate process model outputs, associated with generating virtual, synthetic, and / or digital substrates. The virtual substrate data 162 may include data indicative of features, feature properties, feature parameter representations, substrates, substrates containing arrays of features, and the like.
[0066] In some embodiments, prediction system 110 further includes server machine 170 and server machine 180. Server machine 170 includes dataset generator 172 that can generate datasets for training, validating, and / or testing model 190. Dataset generator 172 may generate datasets for models, including one or more machine learning models. A dataset may include a set of data inputs. A dataset may include a set of target outputs. Some operations of dataset generator 172 are described in detail below with reference to FIGS. 2A-B and 4A. In some embodiments, dataset generator 172 may divide historical data into a training set, a validation set, and a test set. For example, the training set may include 60 percent of the historical data used to generate the model. The validation set may include 20 percent of the historical data used to generate the model. The test set may include 20 percent of the historical data used to generate the model.
[0067] In some embodiments, the prediction system 110 generates multiple sets of attributes (e.g., via the prediction component 114). The attributes may be related to the partitioning or preprocessing of the data input to the machine learning model. For example, a first set of attributes may correspond to a first set of sensor data types corresponding to each of the datasets, and a second set of attributes may correspond to a second set of sensor data types corresponding to each of the datasets. The set of attributes may include data such as sensor data from a set of sensors, a set of metrology measurements, etc. The set of attributes may include combinations of values from a set of measurements. The set of attributes may include patterns of values from the first set of measurements. Each of the training, validation, and / or test datasets may use the same set of attributes to train, validate, and / or test the model.
[0068] In some embodiments, historical data is provided to the machine learning model 190 as training data. In some embodiments, virtual substrate data, synthetic substrate data, or other similar data is provided to the machine learning model 190 as training data. In some embodiments, this historical and / or synthetic data may be or include microscopic image data. The type of data provided depends on the application of the machine learning model. For example, the machine learning model may be trained by providing it with a set of feature parameters as training inputs and providing an indication of a non-physical combination of the parameters as a target output. In some embodiments, a large amount of data may be used to train the model 190. For example, sensor and metrology data for hundreds of substrates may be used.
[0069] Server machine 180 includes a training engine 182, a verification engine 184, a selection engine 185, and / or a test engine 186. Engines (e.g., training engine 182, verification engine 184, selection engine 185, and test engine 186) may refer to hardware (e.g., circuitry, dedicated logic circuitry, programmable logic circuitry, microcode, processing device, etc.), software (e.g., instructions executing on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. Training engine 182 may be capable of training a model 190 using one or more sets of features associated with a training set from dataset generator 172. Training engine 182 may generate multiple trained models 190, where each trained model 190 corresponds to a different set of attributes of the training set (e.g., sensor data from a different set of sensors, a subset of metrology measurements, etc.). For example, a first trained model may be trained using all attributes (e.g., X1-X5), a second trained model may be trained using a first subset of attributes (e.g., X1, X2, X4), and a third trained model may be trained using a second subset of attributes (e.g., X1, X3, X4, and X5), where the second subset of attributes may overlap with the first subset of features. Dataset generator 172 may receive the output of the trained models and assemble the data into training, validation, and test datasets and use those datasets to train a second model.
[0070] The validation engine 184 may be capable of validating the trained models 190 using a corresponding set of attributes in the validation set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of attributes in the training set may be validated using the first set of attributes in the validation set. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of attributes in the validation set. The validation engine 184 may discard trained models 190 with accuracies that do not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 with accuracies that meet the threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting the trained model 190 with the highest accuracy among the trained models 190.
[0071] The testing engine 186 may be capable of testing the trained models 190 using a corresponding set of attributes of a test set from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of attributes of a training set may be tested using a first set of attributes of a test set. The testing engine 186 may determine the trained model 190 that has the highest accuracy of all of the trained models based on the test set.
[0072] In the case of a machine learning model, model 190 may refer to a model artifact generated by training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for each corresponding training input). Patterns in the dataset that map data inputs to target outputs (correct answers) can be found, and machine learning model 190 is provided with a mapping that captures these patterns. Machine learning model 190 may use one or more of support vector machines (SVMs), radial basis functions (RBFs), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-nearest neighbor algorithms (k-NNs), linear regression, random forests, neural networks (e.g., artificial neural networks, recurrent neural networks), and the like.
[0073] In some embodiments, one or more machine learning models 190 may be trained using historical data (e.g., historical metrology data). In some embodiments, the models 190 may have been trained using virtual substrate data 162, feature data 164, substrate data 166, etc.
[0074] Generating and utilizing virtual substrate data 162 has significant technical advantages over other methods. Developing an understanding of the relationships between process operation inputs and process operation results may improve process design, product design, operation design, process operation results, etc. Improving process operation results may reduce the cost of the process in terms of the percentage of defective products produced, the percentage of materials, time, energy, etc. spent producing defective products, product performance, etc. By comparing predicted results of process operations with measured results, defects in models, processing equipment components, process recipes, or the like may be discovered, diagnosed, and corrected. Accurate correction of defects may, for example, improve the performance of the manufacturing system, improve the predictive power of one or more models, reduce unplanned maintenance events, etc.
[0075] In some systems, measurements of substrate features may vary. For example, microscopic images of substrate features may vary in unpredictable or adverse ways. Different images may have different characteristics, such as different contrast, brightness, clarity, etc. This may be due to operator error, microscopy procedures, etc. Different features of a substrate designed to be identical may not be identical, for example, due to differences in processing conditions near the feature's location. Even identical features of a substrate may be measured or imaged differently, for example, due to equipment limitations. Feature data 164 may be generated in response to receiving data for several features. Feature data 164 may record likely feature characteristics, average feature characteristics, target feature characteristics, or the like. Applying a process model to an array of identical features may improve the reliability of input / output mappings based on process model results. For example, applying a process model to an array of identical features may eliminate the possibility that observed results depend on differences between substrate features, which may not be included in the target substrate design. By systematically varying a feature, generating several arrays of identical features, and providing the arrays of features to a process model, inferences may be drawn about relationships between various characteristics of the feature inputs to a process operation and characteristics of the output of the process operation. Parameterizing the features (e.g., feature shapes, feature properties, etc.) may enable robust exploration of feature property space that provides a more complete input / output mapping than may be obtained through random chance using metrology of generated physical substrates. Parameterizing the features may enable exploration of input feature properties with respect to output results of a process operation at a lower cost than developing and implementing adjustments to a process recipe to generate physically different features and providing substrates with physically different features to the process operation.
[0076] The prediction component 114 may provide current data to the model 190 and may execute the model 190 on the inputs to obtain one or more outputs. For example, the prediction component 114 may provide current measurement data to the model 190 and may execute the model 190 on the inputs to obtain one or more outputs. The prediction component 114 may be able to determine (e.g., derive) predicted data 168 from the output of the model 190. From the output, the prediction component 114 may determine (e.g., derive) confidence data that indicates the degree of confidence that the predicted data 168 is an accurate predictor of the process associated with the input data for products produced or to be produced using the manufacturing equipment 124. The prediction component 114 or the corrective action component 122 may use this confidence data to determine whether to cause a corrective action associated with the manufacturing equipment 124 to be performed based on the predicted data 168.
[0077] The confidence data may include or indicate a degree of confidence that the prediction data 168 is an accurate prediction for a product or component associated with at least a portion of the input data. In one example, the confidence is a real number between 0 and 1, inclusive, where 0 indicates no confidence that the prediction data 168 is an accurate prediction for a product processed according to the input data or an accurate prediction for the component health of a component of the manufacturing equipment 124, and 1 indicates absolute confidence that the prediction data 168 accurately predicts a characteristic of a product processed according to the input data or a component health of a component of the manufacturing equipment 124. In response to the confidence data indicating a confidence below a threshold level for a predetermined number of instances (e.g., a percentage of instances, a frequency of instances, a total number of instances, etc.), the prediction component 114 may retrain the trained model 190 (e.g., based on current sensor data 146, current manufacturing parameters, etc.). In some embodiments, retraining may include generating one or more datasets (e.g., via dataset generator 172) utilizing historical and / or synthetic data.
[0078] For purposes of illustration and not limitation, aspects of the present disclosure describe training one or more machine learning models 190 using historical data and inputting current data into one or more trained machine learning models to determine predicted data 168. The historical data used for training may include historical metrology data, historical virtual substrate data, etc. The current data may include current metrology data, current virtual substrate data, etc. In other embodiments, heuristic, physics-based, or rule-based models are used to determine predicted data 168 (e.g., without using trained machine learning models). In some embodiments, such models may be trained using historical and / or synthetic data. In some embodiments, these models may be retrained using a combination of true historical data and synthetic data. The prediction component 114 may monitor historical sensor data 144, historical manufacturing parameters, and metrology data 160. Any of the information described with respect to data inputs 210A-B in FIGS. 2A-B may be monitored or otherwise used by the heuristic, physics-based, or rule-based models.
[0079] In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, and server machine 180 may be provided by fewer machines. For example, in some embodiments, server machines 170 and 180 may be combined into a single machine, and in other embodiments, server machine 170, server machine 180, and prediction server 112 may be combined into a single machine. In some embodiments, client device 120 and prediction server 112 may be combined into a single machine. In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, server machine 180, and data store 140 may be performed by a cloud-based service.
[0080] In general, functions described as being performed by client device 120, prediction server 112, server machine 170, and server machine 180 in one embodiment may, in other embodiments, be performed on prediction server 112, where appropriate. Furthermore, functions attributed to particular components may be performed by different or multiple components operating together. For example, in some embodiments, prediction server 112 may determine corrective actions based on prediction data 168. In another example, client device 120 may determine prediction data 168 based on output from a trained machine learning model.
[0081] Additionally, the functionality of a particular component may be performed by different or multiple components operating together. One or more of prediction server 112, server machine 170, or server machine 180 may be accessed as a service offered to other systems or devices through an appropriate application programming interface (API).
[0082] In embodiments, a "user" may be described as a single individual. However, other embodiments of the present disclosure encompass a "user" being an entity controlled by multiple users and / or automated sources. For example, a collection of individual users united as a group of administrators may be considered a "user."
[0083] Embodiments of the present disclosure may be applied to data quality assessment, feature enhancement, model validation, virtual metrology (VM), predictive maintenance (PdM), marginal optimization, process control or the like.
[0084] 2A-2B illustrate block diagrams of example dataset generators 272A-B (e.g., dataset generators 172 of FIG. 1 ) that generate datasets for training, testing, validation, etc. of a model (e.g., model 190 of FIG. 1 ), according to some embodiments. Each dataset generator 272 may be part of server machine 170 of FIG. 1 . In some embodiments, several machine learning models associated with manufacturing equipment 124 may be trained, used, and maintained (e.g., within a manufacturing facility). Each machine learning model may be associated with one of the dataset generators 272, or multiple machine learning models may share one dataset generator 272, etc.
[0085] 2A illustrates a system 200A that includes a dataset generator 272A for generating a dataset for one or more supervised models (e.g., model 190 of FIG. 1). The dataset generator 272A may use historical data and / or labeled historical data to generate the dataset (e.g., data input 210A, target output 220A). In some embodiments, a dataset generator similar to dataset generator 272A may be utilized to train an unsupervised machine learning model; for example, the target output 220A may not be generated by dataset generator 272A.
[0086] The dataset generator 272A may generate datasets for training, testing, and / or validating a model. In some embodiments, the dataset generator 272A may generate a dataset for a machine learning model. As an example, the dataset generator 272A is described in the context of a machine learning model configured to parameterize one or more characteristics of a feature. Similar dataset generation may be performed for supervised machine learning models performing other functions, with appropriate substitutions of input data, target output data, etc. The machine learning model may be provided with a set of feature data 264A as data input 210A. The machine learning model may be configured to accept one or more substrate feature measurements as inputs and generate a parameterized representation of one or more characteristics of the feature as output. The parameterized representation may include the parameterized characteristic, a standard, average, or expected set of parameter values, upper and lower parameter values, etc.
[0087] In some embodiments, a dataset generator 272A generates a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210A (e.g., training inputs, validation inputs, test inputs). The data inputs 210A may be provided to the training engine 182, the validation engine 184, or the test engine 186. This dataset may be used to train, validate, or test a model (e.g., model 190 of FIG. 1).
[0088] In some embodiments, data input 210A may include one or more sets of data. As an example, system 200A may generate sets of feature data that may include one or more of feature data related to one or more properties of the feature, combinations of feature data for one or more feature properties, patterns from feature data from one or more measurements of the properties of the feature, feature properties from different sets of substrates, etc.
[0089] In some embodiments, dataset generator 272A may generate a first data input corresponding to first set of feature data 264A for training, validating, or testing a first machine learning model. Dataset generator 272A may generate a second data input corresponding to second set of feature data 264B (not shown) for training, validating, or testing a second machine learning model. Additional sets (e.g., including any number of sets of feature data up to a final set, set of feature data 264Z) may be generated by dataset generator 272A for training, validating, or testing additional machine learning models. Any number of sets of feature data may be utilized as data input 210A, for example, according to the target performance of the associated model.
[0090] In some embodiments, dataset generator 272A generates a dataset (e.g., a training set, a validation set, a test set), which includes one or more data inputs 210A (e.g., training inputs, validation inputs, test inputs) and may include one or more target outputs 220A corresponding to the data inputs 210A. The dataset may further include mapping data that maps the data inputs 210A to the target outputs 220A. In some embodiments, dataset generator 272A may generate data for training a machine learning model configured to output a feature parameter representation. Depending on the context, data inputs 210A may also be referred to as “features,” “attributes,” “information,” or “vectors.” In some embodiments, dataset generator 272A may provide the dataset to training engine 182, validation engine 184, or testing engine 186, where the dataset is used to train, validate, or test a machine learning model (e.g., one of the machine learning models included in model 190, ensemble model 190, etc.).
[0091] System 200B, including dataset generator 272B (e.g., dataset generator 172 in FIG. 1 ), generates datasets for one or more unsupervised machine learning models (e.g., model 190 in FIG. 1 ). Dataset generator 272B may use historical data to generate datasets (e.g., data input 210B). Dataset generator 272B is configured to generate datasets for machine learning models configured to receive as input data a set of substrates provided to a process model and a corresponding set of substrates output by the process model, as described above, and to generate as output an indication of an appropriate or valid input / output mapping. Dataset generator 272B may be associated with a machine learning model that provides a list of input feature characteristics having the strongest impact on output characteristics, a set of characteristic parameters associated with generating a target output substrate, or the like. For any unsupervised machine learning model, a dataset generator similar to dataset generator 272B may be utilized with corresponding permutations of data inputs. Dataset generator 272B may share one or more functions with dataset generator 272A.
[0092] The dataset generator 272B may generate a dataset for training, testing, and validating a machine learning model. The machine learning model is provided with set process model data 262A (e.g., inputs and outputs of a process model based on a substrate including an array of features) as data input 210B. The machine learning model may include two or more separate models (e.g., the machine learning model may be an ensemble model). The machine learning model may be configured to generate output data indicating influential input substrate feature parameters, combinations of input substrate feature parameters that are likely to enable a target output, etc. In some embodiments, training may not include providing a target output to the machine learning model. The dataset generator 272B may generate a dataset for training an unsupervised machine learning model.
[0093] In some embodiments, dataset generator 272B generates a dataset (e.g., a training set, a validation set, a test set), which includes one or more data inputs 210B (e.g., training inputs, validation inputs, test inputs). Data inputs 210B may also be referred to as "features," "attributes," or "information." In some embodiments, dataset generator 272B may provide this dataset to training engine 182, validation engine 184, or test engine 186, where the dataset is used to train, validate, or test a machine learning model (e.g., model 190 of FIG. 1). Some operations for generating a training set are further described with respect to FIG. 4A.
[0094] In some embodiments, the dataset generator 272B may generate a first data input corresponding to the first set of process model data 262A to train, validate, or test a first machine learning model, and the dataset generator 272B may generate a second data input corresponding to the second set of process model data 262B to train, validate, or test a second machine learning model. Additional sets of data (e.g., any target number of datasets up to the final set of process model data 262Z) may be generated by the dataset generator 272B to train, validate, or test additional machine learning models. Any number of sets of process model data may be utilized as data input 210A according to the target performance of the associated model.
[0095] The data input 210B for training, validating, or testing a machine learning model may include information for a particular manufacturing chamber (e.g., of a particular piece of substrate manufacturing equipment). In some embodiments, the data input 210B may include information for a particular type of manufacturing equipment, e.g., manufacturing equipment sharing particular characteristics. The data input 210B may include data related to a certain type of device, e.g., intended function, device design, devices produced using a particular recipe, etc. Training a machine learning model based on one type of equipment, device, recipe, etc. may enable the trained model to generate plausible predictive data in several settings (e.g., for several different facilities, products, etc.).
[0096] In some embodiments, following generating a dataset and using the dataset to train, validate, or test a machine learning model, the model may be further trained, validated, or tested, or adjusted (e.g., weights or parameters, such as connection weights of a neural network, associated with the model's input data).
[0097] FIG. 3 is a block diagram illustrating a system 300 for generating output data (e.g., predicted data 168 of FIG. 1 ) according to some embodiments. In some embodiments, system 300 may be used with a machine learning model. The machine learning model may perform several functions related to generating standard features, parameterizing features, generating an array of features, utilizing output from a process model, and / or performing corrective actions. In some embodiments, system 300 may be used with a machine learning model to determine corrective actions associated with manufacturing equipment. In some embodiments, system 300 may be used with a machine learning model to determine defects in manufacturing equipment. In some embodiments, system 300 may be used with a machine learning model to cluster or classify process operation results. System 300 may be used with a machine learning model associated with a manufacturing system that has different functionality than those listed above. System 300 is described as being used with a model configured to parameterize features of a substrate. Other models having different functionality may be used with system 300 or appropriate analogs.
[0098] At block 310, the system 300 (e.g., a component of the prediction system 110 of FIG. 1 ) performs data partitioning (e.g., via the dataset generator 172 of the server machine 170 of FIG. 1 ) of data for use in training, validating, and / or testing the machine learning model. In some embodiments, the feature data 364 includes historical data, such as historical metrology data, measurements extracted from microscope images of historical substrates, etc. The feature data may further include associated parameter representations, e.g., parameter representations of historical features performed by subject matter experts. The feature data 364 may undergo data partitioning at block 310 to generate a training set 302, a validation set 304, and a test set 306. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data.
[0099] The generation of the training set 302, validation set 304, and test set 306 may be tailored to a particular application. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data. The system 300 may generate multiple sets of attributes for each of the training, validation, and test sets. For example, if the feature data 364 includes 20 measures of characteristics of one or more features, the feature data may be divided into a first set of attributes including measures 1-10 and a second set of attributes including measures 11-20. The target output data (e.g., parameter representations) may also be divided into multiple sets. The training inputs may be divided into multiple sets, the target outputs may be divided into multiple sets, both the training inputs and target outputs may be divided into multiple sets, or neither the training inputs nor the target outputs may be divided into multiple sets. Multiple models may be trained on different sets of data.
[0100] At block 312, the system 300 performs model training (e.g., via the training engine 182 of FIG. 1 ) using the training set 302. Training of machine learning models and / or training of models based on physical phenomena (e.g., digital twins) may be achieved with supervised learning methods, which involve feeding a training dataset consisting of labeled inputs through the model, observing its outputs, defining an error (by measuring the difference between the output and the label values), and tuning the model's weights using techniques such as deep gradient descent and backpropagation to minimize the error. In many applications, repeating this process across many labeled inputs in the training dataset yields a model that can generate correct outputs when presented with inputs different from those present in the training dataset. In some embodiments, training of machine learning models may be achieved in an unsupervised manner, e.g., no labels or classifications may be provided during training. Unsupervised models may be configured to perform anomaly detection, outcome clustering, and the like.
[0101] For each training data item in the training dataset, the training data item may be input to a model (e.g., a machine learning model). The model may then process the input training data item (e.g., several measured dimensions of a manufactured device, a cartoon picture of a manufactured device, etc.) to generate an output. The output may include, for example, a parameterized representation of a feature of the substrate. This output may be compared to the label of the training data item (e.g., a valid parameterized representation of the feature generated by another method).
[0102] Processing logic may then compare the generated output (e.g., parametric representation) with the labels included in the training data items (e.g., the provided target parametric representation). Processing logic determines an error (i.e., classification error) based on the difference between the output and the labels. Processing logic adjusts one or more weights and / or values of the model based on this error.
[0103] When training a neural network, an error term or delta may be determined for each node of the artificial neural network. Based on this error, the artificial neural network adjusts one or more of its parameters (weights for one or more inputs of a node) for one or more of its nodes. The parameters may be updated in a backpropagation manner, with the nodes in the top layer updated first, followed by the nodes in the next layer, and so on. An artificial neural network includes multiple layers of "neurons," each of which receives as input values from neurons in the previous layer. The parameters for each neuron include weights associated with the values received from each of the neurons in the previous layer. Adjusting the parameters may therefore include adjusting the weights assigned to each of the inputs to one or more neurons in one or more layers of the artificial neural network.
[0104] The system 300 may train multiple models using multiple sets of attributes from the training set 302 (e.g., a first set of attributes from the training set 302, a second set of attributes from the training set 302, etc.). For example, the system 300 may train models to generate a first trained model using a first set of attributes in the training set (e.g., feature measurements 1-10, measurements 1-10 from a substrate, measurements from one or more locations on multiple substrates, etc.) and to generate a second trained model using a second set of attributes in the training set (e.g., feature measurements 11-20, etc.). In some embodiments, the first and second trained models may be combined to generate a third trained model (e.g., which may, alone, be a better predictor or synthetic data generator than either the first or second trained models). In some embodiments, the sets of attributes used in comparing models may overlap (e.g., the first set of attributes may be from feature measurements 1-15 and the second set of attributes may be from feature measurements 5-20). In some embodiments, hundreds of models may be generated, including models with various permutations of attributes and combinations of models.
[0105] At block 314, the system 300 performs model validation (e.g., via the validation engine 184 of FIG. 1 ) using the validation set 304. The system 300 may validate each of the trained models using a corresponding set of features in the validation set 304. For example, the system 300 may validate a first trained model using a first set of attributes (e.g., metrology measurements 1-10) in the validation set and a second trained model using a second set of attributes (e.g., metrology measurements 11-20) in the validation set. In some embodiments, the system 300 may validate hundreds of models (e.g., models with various permutations of features, combinations of models, etc.) generated at block 312. At block 314, the system 300 may determine the accuracy of each of the one or more trained models (e.g., via model validation) and may determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, flow returns to block 312, where the system 300 performs model training using a different set of attributes from the training set. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, flow proceeds to block 316. The system 300 may discard trained models that have an accuracy lower than the threshold accuracy (e.g., based on a validation set).
[0106] At block 316, the system 300 performs model selection (e.g., via selection engine 185 of FIG. 1 ) to determine which model of the one or more trained models that meet the threshold accuracy has the highest accuracy (e.g., selected model 308 based on the check of block 314). In response to determining that two or more models of the trained models that meet the threshold accuracy have the same accuracy, flow may return to block 312, where the system 300 performs model training to determine the trained model with the highest accuracy using a further refined training set corresponding to the further refined set of attributes.
[0107] At block 318, the system 300 performs model testing (e.g., via the test engine 186 of FIG. 1 ) using the test set 306 to test the selected model 308. The system 300 may test the first trained model using a first set of attributes in the test set (e.g., sensor data from sensors 1-10) and determine that the first trained model meets a threshold accuracy (e.g., based on the first set of attributes in the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 is too well-fitted to the training set 302 and / or validation set 304 and cannot be applied to other data sets, such as the test set 306), flow proceeds to block 312, where the system 300 performs model training (e.g., retraining) using a different training set (e.g., different feature measurements) corresponding to a different set of attributes. In response to determining that the selected model 308 has an accuracy that meets the threshold accuracy based on the test set 306, flow proceeds to block 320. At least in block 312, the model may learn patterns in the training data to make predictions or generate feature parameterizations, and in block 318, the system 300 may apply the model to the remaining data (e.g., test set 306) to test the predictions or the generation of parameterizations.
[0108] In block 320, the system 300 uses the trained model (e.g., the selected model 308) to receive current data 322 (e.g., current metrology data, such as measurements from a substrate whose recipe is being optimized) and determines (e.g., extracts) a feature parameter representation 324 from the output of the trained model. Corrective actions associated with the manufacturing equipment 124 of FIG. 1 may be performed in light of the feature parameter representation 324. For example, based on the feature parameter representation 324, several features may be generated that differ from a center or standard feature in one or more parameter values. A multidimensional grid of features may be generated, where each dimension of the grid corresponds to a parameter and each location on the dimension corresponds to a value of the corresponding parameter within a range (e.g., the range may be included in the feature parameter representation 324). In some embodiments, the current data 322 may correspond to the same type of attribute of the historical data used to train the machine learning model. In some embodiments, the current data 322 may correspond to a subset of the types of attributes of the historical data used to train the selected model 308 (e.g., a machine learning model may be trained using some metrology measurements and may be configured to generate output based on a subset of the metrology measurements).
[0109] In some embodiments, the performance of a machine learning model trained, validated, and tested by system 300 may deteriorate. For example, the manufacturing system associated with the trained machine learning model may undergo gradual or sudden changes. The design of the substrates provided to the process operation in question may change. The details of the process operation may change, or the corresponding process model may change. Such changes in the manufacturing system may result in a deterioration in the performance of the trained machine learning model. A new model may be generated to use in place of the degraded machine learning model. This new model may be generated by modifying the old model through retraining, generating a new model, etc. Retraining the model may be performed by providing additional training data including training input data and target output data. Retraining the model may be performed by providing updated feature data 346 as the additional training data. The updated feature data 346 may include data related to an updated processing system, such as an updated substrate design, an updated process recipe, an updated process equipment, an updated process model, or the like.
[0110] In some embodiments, one or more of operations 310-320 may be performed in various orders and / or with other operations not shown and described herein. In some embodiments, one or more of operations 310-320 may not be performed. For example, in some embodiments, one or more of data partitioning of block 310, model validation of block 314, model selection of block 316, or model testing of block 318 may not be performed.
[0111] 3 illustrates a system configured to train, validate, test, and use one or more machine learning models. The machine learning models are configured to accept data (e.g., set points provided to manufacturing equipment, sensor data, measurement data, etc.) as input and provide data (e.g., prediction data, corrective action data, classification data, etc.) as output. The partition, training, validation, selection, testing, and use blocks of system 300 may be similarly performed to train a second model using a different type of data. Additionally, retraining may be performed using current data 322 and / or updated feature data 346.
[0112] 4A-D are flowcharts of methods 400A-D related to generating input / output mappings of a process operation using feature measurements to perform corrective actions, according to some embodiments. Methods 400A-D may be performed by processing logic, which may include hardware (e.g., circuits, dedicated logic circuitry, programmable logic circuitry, microcode, a processing device, etc.), software (e.g., instructions executing on a processing device, a general-purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 400A-D may be performed in part by prediction system 110. Method 400A may be performed in part by prediction system 110 (e.g., server machine 170 and dataset generator 172 in FIG. 1 , dataset generators 272A-B in FIGS. 2A-2B ). Prediction system 110 may use method 400A to generate datasets for at least one of training, validating, or testing a machine learning model according to embodiments of the present disclosure. Methods 400B-D may be executed by prediction server 112 (e.g., prediction component 114) and / or server machine 180 (e.g., training, validation, and testing operations may be performed by server machine 180). In some embodiments, a non-transitory machine-readable storage medium stores instructions that, when executed by a processing device (e.g., a processing device of prediction system 110, a processing device of server machine 180, a processing device of prediction server 112, etc.), cause the processing device to perform one or more of methods 400A-D.
[0113] For ease of explanation, methods 400A-D are illustrated and described as a series of operations. However, operations in accordance with the present disclosure may be performed in various orders and / or simultaneously, and with other operations not shown and described herein. Moreover, not all illustrated operations may be performed to implement methods 400A-D in accordance with the disclosed subject matter. Furthermore, those skilled in the art will understand and appreciate that methods 400A-D may also be represented as a series of interrelated states via a state diagram or events.
[0114] 4A is a flow diagram of a method 400A for generating a dataset for a machine learning model, according to some embodiments. Referring to FIG. 4A, in some embodiments, at block 401, processing logic performing method 400A initializes a training set T to be an empty set.
[0115] At block 402, processing logic generates a first data input (e.g., a first training input, a first validation input), which may include one or more of sensor, manufacturing parameter, metrology data, etc. In some embodiments, this first data input may include a first set of attributes for a type of data, and the second data input may include a second set of attributes for the type of data (e.g., as described with respect to FIG. 3 ). In some embodiments, the input data may include historical data and / or synthetic data. The input data may include feature data, feature parameter data, process model input and output data, etc.
[0116] In some embodiments, at block 403, processing logic optionally generates a first target output for one or more of these data inputs (e.g., a first data input). In some embodiments, this input includes one or more metrology measurements and the target output is a parameter representation of the substrate characteristics. In some embodiments, this first target output is predicted data. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model that can group or find correlations in input data without providing a target output).
[0117] At block 404, processing logic optionally generates mapping data indicating an input / output mapping. This input / output mapping (or mapping data) may relate to a data input (e.g., one or more of the data inputs described herein), a target output for the data input, and an association between the data input and the target output. In some embodiments, such as those associated with machine learning models in which no target output is provided, block 404 may not be performed.
[0118] In some embodiments, at block 405, processing logic adds the mapping data generated at block 404 to the dataset T.
[0119] At block 406, processing logic branches based on whether dataset T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as model 190 in FIG. 1. If so, execution proceeds to block 407; if not, execution returns to block 402. It should be noted that in some embodiments, whether dataset T is sufficient may be determined simply based on the number of inputs in the dataset, and in some embodiments based on the number of inputs in the dataset that are mapped to outputs; in other embodiments, whether dataset T is sufficient may be determined based on one or more other criteria (e.g., measures of diversity of data examples, accuracy, etc.) in addition to or instead of the number of inputs.
[0120] At block 407, processing logic provides dataset T (e.g., to server machine 180) for training, validating, and / or testing a machine learning model, such as machine learning model 190. In some embodiments, dataset T is a training set, and dataset T is provided to training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set, and dataset T is provided to validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set, and dataset T is provided to test engine 186 of server machine 180 to perform testing. For example, in the case of a neural network, input values (e.g., numerical values associated with data input 210A) of a given input / output mapping are input to the neural network, and output values (e.g., numerical values associated with target output 220A) of the input / output mapping are stored at output nodes of the neural network. The connection weights of the neural network are then adjusted according to a learning algorithm (e.g., backpropagation, etc.), and the procedure is repeated for the remaining input / output mappings of dataset T. After block 407, the model (e.g., model 190) may be at least one of trained using training engine 182 of server machine 180, validated using validation engine 184 of server machine 180, or tested using test engine 186 of server machine 180. The trained model may be implemented by prediction component 114 (of prediction server 112) to generate predicted data 168 for performing signal processing, to generate synthetic data 162, or to perform corrective actions related to manufacturing equipment 124.
[0121] FIG. 4B is a flow diagram of a method 400B for performing corrective action using measurements of features of a substrate, according to some embodiments. At block 410, processing logic receives profile data of multiple features of a substrate. The feature profile may be a shape, one or more characteristics, data points along an edge, a function describing a boundary, or the like. The multiple features may be features of a substrate. Further, the multiple features may include features of multiple substrates. The multiple features may be nominally identical, e.g., designed to have similar geometries, properties, performance, etc. Furthermore, the multiple features may be features of multiple nominally identical substrates, e.g., multiple substrates that may have been produced using the same process recipe, multiple substrates that may have been produced using the same process equipment, multiple substrates that may have been produced using the same type of equipment, multiple substrates that may have been designed to perform the same function, etc.
[0122] Profile data for multiple features may be extracted from one or more microscopic images. The microscopic images may be microscopic images of one or more substrates. The microscopic images may be, for example, microscopic images of one or more features, of portions of features, or may include profiles of features. The microscopic images may be TEM images, SEM images, XSEM images, or images generated by other imaging techniques. From these images, data describing one or more profiles of features may be extracted using a model, such as a machine learning model.
[0123] At block 412, processing logic generates a typical profile based on the profile data of the plurality of features. The typical profile may be a profile of typical features. The typical profile and / or typical features may be generated by obtaining the mean, median, mode, or some other metric of one or more characteristics of the feature and / or profile (e.g., the plurality of features) under consideration. The typical profile may be a profile of measured features, e.g., a profile of features having measurements closest to the median or mean of the measured features.
[0124] Generating the representative profile may include generating a parameter representation of the feature and / or feature profile. The parameter representation may describe the characteristics of the feature using several adjustable parameters. For example, parameters describing the feature, feature profile, etc. may include slope, distance, radius of curvature, etc. Generating the parameter representation may be performed manually, by a model, or by a machine learning model. Generating the parameter representation may include considering statistics of the provided profile data, e.g., a range of radii of curvature for the characteristics of the plurality of features. Generating the parameters may include generating a representative value and / or a range of values for the parameter. A representative profile or a representative feature may include a combination of parameter values within the generated range. Generating and using a parameter representation of a feature is discussed in more detail with respect to FIG. 4C .
[0125] At block 414, processing logic generates a first array of features, each based on the exemplary profile. The first array of features may include a virtual or synthetic substrate. The virtual substrate may include data indicating properties of the substrate. The virtual substrate may include an array of identical features, e.g., an array of features having the exemplary profile. The first array of features may be a two-dimensional array. The first array of features may be a three-dimensional array. The virtual substrate may include a two-dimensional or three-dimensional array of features. The virtual substrate may include an array of features arranged in a line (e.g., the features may include properties in two dimensions parallel and perpendicular to the arrangement of the array). The virtual substrate may include an array of features arranged in a grid (e.g., the features may include properties in three dimensions parallel and perpendicular to the two-dimensional grid or array of features).
[0126] At block 416, processing logic provides the first array of features to a process model. The process model predicts the results of applying one or more process operations to an input substrate. The input substrate may include the first array of features. The output may be a prediction of a property of an output substrate of a physical substrate processing procedure given the input substrate. The process model may be a model based on physical phenomena. The process model may be a deposition model. The process model may be an etch model. The process model may be configured to predict the results of any process operation or combination of process operations.
[0127] At block 418, the processing logic obtains a first output from the process model based on the first array of features. The output may include data indicating predicted characteristics of the substrate after additional processing, after one or more additional process operations, etc. The output of the process model may indicate one or more effects that an input feature profile to a process operation has on the output of the process operation. In some embodiments, multiple arrays of features may be provided to the process model. Each of the multiple arrays may include somewhat different features, e.g., features with different profiles, features with different parameter values, or other similar features. The output received by the processing logic may include input / output mapping data, e.g., a collection of data indicating the effect of inputs to the process model on outputs from the process model. Modeling a set of arrays of features is discussed in more detail with respect to FIG. 4D.
[0128] At block 419, processing logic causes corrective action to be performed in consideration of the first output from the process model. The corrective action may include an update. The update may be an update to the design of the input product to the process operation, an update to the process recipe, etc. The corrective action may include maintenance, such as corrective or preventive maintenance. The corrective action may include providing an alert to a user. The alert may include notifying a user of a recommended update to the process operation. The alert may include notifying a user of a recommended update to the product design. The alert may include notifying a user of an impact that a characteristic of the input substrate has on one or more characteristics of an output substrate of the process operation.
[0129] FIG. 4C is a flow diagram of a method 400C for generating and utilizing parameterized representations of features, according to some embodiments. At block 420, processing logic receives profile data for a plurality of features. The features may be features of one or more substrates. The profile data may include data indicating one or more shapes, boundaries, or regions occupied by the features. The profile data may be extracted from a microscope image of the features. The profile data may be extracted from a microscope image of the features as input to a target process operation (e.g., measurements of a substrate may be obtained before the substrate is subjected to the target process operation). The profile data may be extracted from measurements obtained after the target process operation is performed (e.g., properties of the input substrate may be extrapolated from, for example, XSEM measurements of an output substrate of the process operation).
[0130] At block 422, processing logic generates a parametric representation of the exemplary profile based on the profile data. In some embodiments, generation of the parametric representation may instead be performed manually. Generation of the parametric representation may be performed by a machine learning model. The parametric representation may include abstract parameters, such as coefficients of fit of the feature profile. The parametric representation may include physical parameters, such as the rounding radius, slope, distance, etc., of the parameterized feature. The properties to parameterize may be selected based on the range of inputs to the parameterization process. For example, no rounding radius may be selected for the parameterization from a set of features in response to little variation in rounding radius for the set of features (or set of feature profiles).
[0131] At block 424, parameter values of the exemplary profile are altered to generate a second profile, for example, the radius of rounding or the slope may be altered compared to the exemplary profile to generate a new profile, new features, etc.
[0132] At block 426, first and second arrays of features are generated. The first and second arrays of features may comprise first and second substrates. Each of the first arrays of features may be identical to a representative feature, e.g., each of the first arrays of features may share a representative profile. The first substrate may comprise a first array of features, each including the representative profile. Each of the second arrays of features may be identical to a second feature, e.g., each of the 21 arrays of features may include the second feature. The second substrate may comprise a second array of features, each including the second profile.
[0133] At block 428, the first array of features (e.g., the first substrate) and the second array of features (e.g., the second substrate) are provided to a process model. The process model may predict results of performing process operations. For example, the process model may predict results of performing deposition or etch operations on substrates corresponding to the first and second substrates.
[0134] In some embodiments, a first virtual substrate provided to a process model may include an array of identical features. A second virtual substrate provided to a process model may include a second array of identical features, where the second array is different from the first array. Each feature in the second array may be different from each feature in the first array. The arrangement of features in the second array may be different from the arrangement of features in the first array. More virtual substrates may be generated and provided to the process model. Beyond the first and second substrates, many virtual substrates different from the first and second substrates may be generated. Each virtual substrate may include an array of features. Each array of features may be different from the remaining arrays of features. Each array of features may include features that differ in shape, profile, properties, arrangement, or the like from the remaining arrays of features. All features of a single substrate may be identical. In some embodiments, the effect of differently shaped features on a single substrate may be important, and the virtual substrate may include features with different shapes, profiles, properties, etc.
[0135] At block 429, processing logic receives output from a process model based on the first and second arrays of features. The output of the process model may be a predicted result of a process operation performed on a physical substrate. The output of the process model may include an input / output mapping, e.g., the output of the process model may include an indication of the effect that changing parameter values in the exemplary profile has on the output. The output of the process model may be used in performing corrective action, e.g., in updating a substrate design, a process operation recipe, or the like. In some embodiments, more arrays of features may be generated and provided to the process model. A multidimensional grid of profiles may be generated, each of which has values for one or more parameters that differ from the exemplary profile. For example, exploration of the input substrate property space to the process operation may be explored in this manner, and input / output mappings spanning portions of the input property space and the output property space may be analyzed.
[0136] 4D is a flow diagram of a method 400D for obtaining predicted outputs from a process model, according to some embodiments. At block 430, processing logic generates a first set of profiles, each of which differs from a parametric representation of an exemplary profile in the value of at least one parameter. The operations of block 430 may optionally include additional operations shown in blocks 432-436.
[0137] At block 432, processing logic optionally generates a second set of profiles. Each of the profiles in the second set is generated by adjusting one or more parameter values from the values of the parametric representation of the exemplary profile. Each of the profiles in the second set is unique, e.g., each profile in the second set is different from any other profile in the second set. The second set of profiles may include a complete exploration of the parametric representation of the profile. For example, the second set of profiles may include parameter combinations ranging from a lower parameter value to an upper parameter value for each parameter. In some embodiments, a range of values within a range may be generated for each parameter. The second set of profiles may include profiles having each combination of values in the range for each parameter. The second set of profiles may include fewer profiles, e.g., some combinations of parameter values may be excluded.
[0138] At block 434, processing logic provides the second set of profiles to the trained machine learning model. At block 436, processing logic obtains output from the trained machine learning model. The output includes the first set of profiles, e.g., the first set of profiles is a subset of the second set of profiles. The trained machine learning model is configured to determine one or more of the second set of profiles not to use in generating the array of features. The trained machine learning model may be configured to exclude parameter value combinations that are unphysical, unlikely to occur, extremely expensive to generate (e.g., costing above a threshold in terms of cost, time, energy, materials, reliability, etc.), or the like. The trained machine learning model may be a supervised model, e.g., the trained machine learning model may be trained using labeled training data. The trained machine learning model may be provided with manually labeled training data. The trained machine learning model may be provided with training data divided into categories, such as likely physically occurring structures and unlikely physically occurring structures. The trained machine learning model may be an unsupervised model, e.g., the trained machine learning model may be provided with unlabeled training data and may be configured to exclude combinations of parameter values that are not within the parameter space spanned by the training data.
[0139] At block 438, processing logic generates a set of feature arrays. Each array is associated with one of the first set of profiles. Each array may be an array of identical features. Each feature in the array of features may include a corresponding profile. Each of the arrays of features may be or may include a virtual substrate. At block 440, processing logic provides each of the set of feature arrays to a process model.
[0140] At block 442, processing logic obtains a set of outputs from the process model, each of which is associated with one of the set of feature arrays. The outputs may be predicted results of subjecting a substrate including the feature arrays to a process operation.
[0141] At block 444, the processing logic optionally provides the first set of profiles and the set of outputs to a trained machine learning model. The trained machine learning model may be configured to generate one or more indications of the impact of inputs to the process model (e.g., details of the feature profiles) on outputs from the model (e.g., execution of process operations associated with the process model). The trained machine learning model may be configured to extract several influential input parameters. The trained machine learning model may be configured to enumerate input parameters having the most significant impact on output characteristics. The trained machine learning model may be configured to generate an input / output mapping, such as a multidimensional fit.
[0142] At block 446, processing logic obtains one or more instructions of mapping between profile parameters and process model outputs from the trained machine learning model. One or more corrective actions may be performed in light of the outputs from the trained machine learning model.
[0143] 5 illustrates an exemplary substrate 500 including features, according to some embodiments. The substrate 500 may be a physical substrate. The substrate 500 may be a virtual substrate. The substrate 500 may resemble a microscopic image of a device, such as an XSEM or TEM image. Aspects of the present disclosure include providing data indicative of properties of the substrate to a process model corresponding to one or more process operations. The substrate 500 may be a substrate that has not yet undergone the corresponding process operation. The substrate 500 may be a substrate that has undergone the corresponding process operation.
[0144] Substrate 500 includes several features. Substrate 500 includes nominally identical features 580 and 582. A device feature may, for example, include multiple components or be defined by multiple characteristics. A portion of feature 580 stands on pedestal 570. The device may include a feature having a gate 572. The gate may be surrounded by spacers 574 and may have a mask 576 thereon. A deposition material 578 may be disposed on mask 576. Other devices, designs, etc. are within the scope of this disclosure.
[0145] The process model may be an etch model, a deposition model, or another model configured to predict the results of one or more process operations. For example, the process model may predict the results of a process operation that results in the deposition of deposition material 578. Measurements of features 580 and 582 may be performed before the deposition of deposition material 578 or after the deposition of deposition material 578. Some measurement techniques, such as XSEM, may be able to measure the properties and / or profile of features that existed before a process operation was performed. For example, an XSEM metrology system may extract the shape of feature 580 before deposition therefrom and provide data that may be provided to the process model.
[0146] The characteristics of the features may include radius of curvature, slope, distance, thickness, and other characteristics. For example, the radius of curvature of the bow of pedestal 570, the slope of various edges of components such as spacer 574 and / or gate 572, etc. may be characteristics of feature 580. The characteristics of feature 580 may be parameterized based on, for example, variations between the characteristics of feature 580 and feature 582, variations between the characteristics of feature 580 and other features of substrate 500, variations between the characteristics of feature 580 and other features of other substrates, etc.
[0147] FIG. 6 is a block diagram illustrating a computer system 600 according to some embodiments. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network such as a local area network (LAN), an intranet, an extranet, or the Internet). Computer system 600 may operate in the capacity of a server or a client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a server, a network router, a switch, or a bridge, or any other device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the device. Furthermore, the term “computer” is intended to include a collection of computers that individually or jointly execute a set (or sets) of instructions to perform any one or more of the methods described herein.
[0148] In additional aspects, computer system 600 may include a processing device 602, a volatile memory 604 (e.g., random access memory (RAM)), a non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.
[0149] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing another type of instruction set, or a microprocessor implementing a combination of instruction set types), or a specialized processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[0150] Computer system 600 may further include a network interface device 622 (e.g., coupled to a network 674). Computer system 600 may further include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generating device 620.
[0151] In some embodiments, data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) having stored thereon instructions 626 encoding one or more of the methods or functions described herein, including instructions encoding the components of FIG. 1 (e.g., prediction component 114, corrective action component 122, model 190, etc.) and instructions for implementing the methods described herein.
[0152] The instructions 626 may further reside, completely or partially, within the volatile memory 604 and / or within the processing device 602 during execution of the instructions 626 by the computer system 600; thus, the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.
[0153] While the computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" is intended to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" is also intended to include any tangible medium that can store or encode a set of instructions for execution by a computer, causing the computer to perform one or more of the methods described herein. The term "computer-readable storage medium" is intended to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0154] The methods, components, and features described herein may be implemented by discrete hardware components or may be integrated into the functionality of other hardware components, such as an ASIC, FPGA, DSP, or similar device. Further, the methods, components, and features described herein may be implemented by firmware modules or by functional circuitry within a hardware device. Furthermore, the methods, components, and features described herein may be implemented in any combination of hardware devices and computer program components, or in a computer program.
[0155] Unless otherwise specified, terms such as "receive," "perform," "provide," "acquire," "perform," "access," "determine," "add," "use," "train," "reduce," "generate," "modify," or other similar terms refer to operations and processes performed or implemented by a computer system that manipulate and convert data in computer system registers and memory, represented as physical (electronic) quantities, into other data in the computer system memory or registers, or other such information storage, transmission, or display device, also represented as physical quantities. Furthermore, as used herein, terms such as "first," "second," "third," "fourth," etc., are meant as labels to distinguish between different elements and may not have an ordinal meaning based on the numerical designation of those terms.
[0156] The examples described herein further relate to apparatus for performing the methods described herein. The apparatus may be specially constructed to perform the methods described herein, or may include a general-purpose computer system that is selectively programmed by a computer program stored on the computer system. Such a computer program may be stored on a computer-readable tangible storage medium.
[0157] The methods and illustrative examples described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized apparatus to perform the methods described herein and / or each of the individual functions, routines, subroutines, or operations of those methods. Example structures for these various systems are provided in the description above.
[0158] The above description is intended to be illustrative, and not limiting. While the present disclosure has been described with reference to particular illustrative examples and embodiments, it will be recognized that the present disclosure is not limited to the described examples and embodiments. The scope of the present disclosure should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. receiving profile data of a plurality of features of the substrate; generating an exemplary profile based on the profile data of the plurality of features; generating a first array of features, each of the first array of features being based on the exemplary profile; providing the first array of features to a process model; obtaining a first output from the process model based on the first array of features; and performing corrective action in consideration of the first output from the process model; A method comprising:
2. receiving a microscopic image of the substrate including the plurality of features; and extracting the profile data of the plurality of features from the microscopic image. The method of claim 1 further comprising:
3. generating the representative profile, expressing the profile data of the plurality of features as a plurality of sets of characteristic parameters; and performing a statistical analysis to generate a set of characteristic parameters comprising said representative profile; The method of claim 1 , comprising:
4. The method of claim 1 , wherein each feature in the first array of features comprises the representative profile.
5. generating a parametric representation of said exemplary profile; modifying parameters of the exemplary profile to generate a second profile; generating a second array of features, each of the second array of features based on the second profile; providing the second array of features to the process model; and obtaining a second output from the process model based on the second array of features, wherein performing the corrective action also takes into account the second output from the process model. The method of claim 1 further comprising:
6. 6. The method of claim 5, wherein generating the parametric representation comprises providing the profile data of the plurality of features to a trained machine learning model, and obtaining parameters of the parametric representation as output from the trained machine learning model.
7. generating a first set of profiles, each of the first set of profiles differing from the parametric representation of the exemplary profile in a value of at least one parameter; generating a set of feature arrays, each array associated with one of the first set of profiles; providing each of the set of arrays of features to the process model; and obtaining a set of outputs from the process model, each of the set of outputs associated with one of the set of arrays of features; The method of claim 5 further comprising:
8. generating the first set of profiles; generating a second set of profiles, each of the second set of profiles generated by adjusting one or more parameter values associated with the parametric representation of the exemplary profile; providing the second set of profiles to a trained machine learning model; and obtaining the first set of profiles from the trained machine learning model, wherein the trained machine learning model is configured to determine one or more profiles from the second set of profiles that will not be used to generate an array of features. The method of claim 7, comprising:
9. providing the first set of profiles to a trained machine learning model; providing the set of outputs to the trained machine learning model; and obtaining, from the trained machine learning model, one or more indications of a mapping between profile parameters and process model outputs; The method of claim 7 further comprising:
10. The corrective action is: scheduling maintenance of substrate processing systems; Updating the substrate processing recipe, or Providing alerts to users The method of claim 1 , comprising one or more of:
11. The method of claim 1 , wherein the process model comprises a deposition model based on physical phenomena.
12. 1. A system including a memory and a processing device coupled to the memory, the processing device comprising: receiving profile data for a plurality of features, each of the features being a feature of the substrate; generating an exemplary profile based on the profile data of the plurality of features; generating a first array of features, each of the first array of features being based on the exemplary profile; providing the first array of features to a process model; obtaining a first output from the process model based on the first array of features; and performing corrective action in consideration of the first output from the process model; The system that is supposed to run
13. The system of claim 12 , wherein each feature in the first array of features comprises the representative profile.
14. the processing device further comprising: generating a parametric representation of said exemplary profile; Varying a first parameter of the exemplary profile to generate a second profile; generating a second array of features, each of the second array of features based on the second profile; providing the second array of features to the process model; and obtaining a second output from the process model based on the second array of features, wherein performing the corrective action also takes into account the second output from the process model.
13. The system of claim 12, wherein the system is adapted to execute:
15. the processing device further comprising: generating a first set of profiles, each of the first set of profiles differing from the parametric representation of the exemplary profile in a value of at least one parameter; generating a set of feature arrays, each array in the set of arrays associated with one of the first set of profiles; providing each of the set of arrays of features to the process model; and obtaining a set of outputs from the process model, each of the set of outputs associated with one of the set of arrays of features; 15. The system of claim 14, wherein the system is adapted to execute:
16. the processing device further comprising: receiving one or more microscopic images including the plurality of features; and extracting the profile data of the plurality of features from the one or more microscopic images.
13. The system of claim 12, wherein the system is adapted to execute:
17. A non-transitory machine-readable storage medium having instructions stored thereon that, when executed, receiving profile data of a plurality of features of the substrate; generating an exemplary profile based on the profile data of the plurality of features; generating a first array of features, each of the first array of features being based on the exemplary profile; providing the first array of features to a process model; obtaining a first output from the process model based on the first array of features; and performing corrective action in consideration of the first output from the process model; A non-transitory machine-readable storage medium that causes a processing device to perform operations including
18. The operation further comprises: generating a parametric representation of said exemplary profile; modifying parameters of the exemplary profile to generate a second profile; generating a second array of features, each of the second array of features based on the second profile; providing the second array of features to the process model; and obtaining a second output from the process model based on the second array of features, wherein performing the corrective action also takes into account the second output from the process model.
20. The non-transitory machine-readable storage medium of claim 17, comprising:
19. 20. The non-transitory machine-readable storage medium of claim 18, wherein generating the parametric representation comprises providing the profile data of the plurality of features to a trained machine learning model, and obtaining parameters of the parametric representation as output from the trained machine learning model.
20. The operation further comprises: generating a first set of profiles, each of the first set of profiles differing from the parametric representation of the exemplary profile in a value of at least one parameter; generating a set of feature arrays, each array associated with one of the first set of profiles; providing each of the set of arrays of features to the process model; and obtaining a set of outputs from the process model, each of the set of outputs associated with one of the set of arrays of features; 20. The non-transitory machine-readable storage medium of claim 18, comprising:
Citation Information
Patent Citations
Polishing equipment using neural networks for monitoring
JP2020518131A
Correcting component failures in ion implantation semiconductor manufacturing tools
JP2022523101A
Correcting component failures in ion implant semiconductor manufacturing tool
WO2020159730A1