Piecewise function fitting of substrate profiles for process learning
The piecewise functional fit of substrate profiles addresses the challenge of accurately representing manufacturing profiles by dividing data into regions and enforcing smoothness, enabling efficient and cost-effective manufacturing process optimization.
Patent Information
- Application Number
- JP2025500944
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-09
- Filing Date
- 2023-08-07
- Publication Date
- 2025-08-07
AI Technical Summary
Conventional methods for describing substrate profiles in manufacturing processes face challenges in accurately representing the profile, particularly due to the difficulty in isolating physical effects from input changes and correlating target profile generation with nonlinear relationships.
A piecewise functional fit is generated by dividing substrate data into regions and fitting each region with a specific function from a library, enforcing continuity and smoothness at boundaries, allowing for a concise and accurate description of the substrate profile using fewer parameters with physical meaning.
This approach provides a complete and accurate description of the substrate profile with reduced parameters, facilitating easier analysis of input changes and reducing experimental costs by correlating input conditions with profile parameters, thus optimizing manufacturing processes.
Smart Images

Figure 2025525725000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to methods related to machine learning models used in evaluating manufactured devices, such as semiconductor devices, and more particularly, to methods for generating and utilizing piecewise functional fits of substrate profiles for process characterization and process learning. [Background technology]
[0002] Manufacturing equipment may be used to perform one or more manufacturing processes to produce a product. For example, semiconductor manufacturing equipment may be used to produce substrates through a semiconductor manufacturing process. The product is produced to have certain characteristics suitable for a target application. Machine learning models are used in various process control and predictive functions associated with manufacturing equipment. Machine learning models are trained using data associated with manufacturing equipment. Measurements of the product (e.g., a manufactured device) may be taken, and images of these may, for example, enhance understanding of device function, failures, performance, or may be used for metrology or inspection. Summary of the Invention
[0003] The following is a simplified summary of the present disclosure to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure, nor is it intended to limit the scope of particular embodiments or claims of the present disclosure. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] In one aspect of the present disclosure, a method includes receiving, by a processing device, data indicating a plurality of measurements of a profile of a substrate. The method further includes dividing the data into a plurality of sets of data, a first set of the plurality of sets associated with a first region of the profile and a second set of the plurality of sets associated with a second region of the profile. The method further includes fitting the first set of data to a first function to generate a first fitted function. The first function is selected from a library of functions. The method further includes fitting the second set of data to a second function to generate a second fitted function. The method further includes generating a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitted function and a second fitted function.
[0005] In another aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include receiving data indicating a plurality of measurements of a profile of the substrate. The operations further include dividing the data indicating the plurality of measurements into a plurality of sets of data. A first set of the plurality of sets is associated with a first region of the profile. A second set of the plurality of sets is associated with a second region of the profile. The operations further include fitting the first set of data to a first function to generate a first fitting function. The first function is selected from a library of functions. The operations further include fitting the second set of data to a second function to generate a second fitting function. The second function is selected from the library of functions. The second function is different from the first function. The operations further include generating a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitting function and a second fitting function.
[0006] In another aspect of the present disclosure, a system includes a memory and a processing device coupled to the memory. The processing device is configured to perform an operation. The operation includes receiving data indicating a plurality of measurements of a profile of the substrate. The operation further includes dividing the data indicating the plurality of measurements into a plurality of sets of data. A first set of the plurality of sets is associated with a first region of the profile. A second set of the plurality of sets is associated with a second region of the profile. The operation further includes fitting the first set of data to a first function to generate a first fitted function. The first function is selected from a library of functions. The operation further includes fitting the second set of data to a second function to generate a second fitted function. The second function is selected from the library of functions. The second function is different from the first function. The operation further includes generating a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitted function and a second fitted function.
[0007] In the figures of the accompanying drawings, the present disclosure is illustrated by way of example and not by way of limitation. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary system architecture according to some embodiments. [Figure 2A] FIG. 1 is a block diagram of a system including an example dataset generator for generating datasets for one or more supervised models, according to some embodiments. [Figure 2B] FIG. 1 is a block diagram of an example dataset generator for generating datasets for one or more unsupervised models, according to some embodiments. [Figure 3] FIG. 1 is a block diagram illustrating a system for generating output data, according to some embodiments. [Figure 4A]1 is a flow diagram of a method for generating a dataset for a machine learning model, according to some embodiments. [Figure 4B] 1 is a flow diagram of a method for generating a profile piecewise function fit, according to some embodiments. [Figure 5A] FIG. 1 is a block diagram of a substrate measurement generation system according to some embodiments. [Figure 5B] 1A-1C illustrate an exemplary substrate and an exemplary function fit of a profile of the substrate, according to some embodiments. [Figure 5C] 1 is a flow diagram of system components of a system for generating and utilizing a piecewise function fit of a substrate profile, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating a computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Described herein are techniques related to generating a functional description of a profile of a substrate. In some embodiments, the techniques described herein relate to generating a piecewise functional description of a manufactured or simulated substrate. The profile of a substrate may relate to the shape of one or more features of the substrate. For example, the substrate may include one or more critical dimensions, e.g., one or more critical dimensions related to the width of a hole, groove, or trench in the substrate. The profile of a substrate may represent the shape of the substrate, e.g., represent the critical dimension as a function of depth. The techniques described herein may enable the generation of a piecewise function that concisely and accurately describes the shape of a feature, the profile of a substrate, etc.
[0010] Manufacturing equipment is used to produce products such as substrates (e.g., wafers, semiconductors). Manufacturing equipment may include fabrication or processing chambers for isolating (e.g., separating) substrates from the surrounding environment for processing. The properties of the produced substrates should meet target values to facilitate a specific function. Manufacturing parameters are selected to produce substrates that meet the target property values. Many manufacturing parameters (e.g., hardware parameters, process parameters, etc.) contribute to the properties of the processed substrates. A manufacturing system may control the parameters by specifying set points for property values, receiving data from sensors located in the fabrication chambers, and adjusting manufacturing equipment until the sensor readings match the set points. A manufacturing system may generate, produce, process, or manufacture substrates. A substrate may be analyzed (e.g., measured, tested) to predict substrate performance (e.g., quality), evaluate the quality of a manufacturing system / process, etc. One or more profiles (e.g., cross-sections of structures / features of a substrate) may be measured. One or more critical dimensions may be measured for holes in the substrate (e.g., holes etched in the substrate). As used herein, critical dimension refers to the width of a hole (e.g., the width measured perpendicular to the centerline of the hole), and often refers to the width of a hole (e.g., the width measured perpendicular to the centerline of the hole) as a function of depth (e.g., the depth at which the width measurement line intersects the centerline). Many measurements of the critical dimension may be taken at various depths of the hole. The measurements as a function of depth may describe the profile of the substrate.
[0011] A physics-based model may be utilized to generate a simulated substrate (e.g., a set of data that predicts the properties, geometry, etc., of a substrate fabricated according to parameters provided to the physics-based model). Inputs to the physics-based model may differ from inputs to the manufacturing system. For example, the manufacturing system may include setpoints such as power supplied to various components, frequencies of one or more radio frequency components, gas flow settings, etc. Inputs to the physics-based model may include etch rates, deposition rates, gas compositions, energy transfer, etc. The physics-based model may be configured to receive one or more simulation inputs as inputs and generate as output a simulated substrate (e.g., predicted data indicative of the geometry and / or properties of a substrate processed according to the simulation inputs). The simulated substrate data may include data indicative of one or more profiles of the simulated substrate. The simulated substrate data may include one or more measured values of critical dimensions.
[0012] A machine learning model may be utilized to generate the simulated substrate. Inputs to the machine learning model may be different from or the same as (or include one or more of) the inputs to the physics-based model and / or the manufacturing system. The machine learning model may receive as inputs manufacturing parameters (e.g., setpoints, inputs of the manufacturing system), conditions (e.g., inputs of the physics-based simulation), sensor data (e.g., data received by sensors associated with the manufacturing system), combinations thereof, or the like. The machine learning model may be configured to generate a simulated substrate, e.g., predicted data indicative of characteristics of the substrate. The simulated substrate data may include data indicative of one or more profiles of the simulated substrate. The simulated substrate data may include one or more measured values of critical dimensions.
[0013] A profile of a substrate (e.g., a fabricated substrate, a simulated substrate, etc.) may be extracted. The profile may be represented as a series of data points, a series of measurements, etc. For example, a critical dimension may be represented as several points, each corresponding to a measurement at an associated depth. This may be used to generate a plot of the critical dimension against depth. Other dimensions, geometries, profiles, etc. of the substrate may also be represented.
[0014] In conventional systems, several indicators may be extracted from the profile measurements to describe the profile. For example, the profile may be related to a critical dimension of the substrate. From the profile, several measures may be extracted, for example, as an approximation of the profile. The extracted values may include, for example, a maximum value, a minimum value, the location of the maximum or minimum value (e.g., the depth at which the maximum critical dimension occurs), a value at the lowest or highest value of the domain, a slope between two points on the profile, or the like.
[0015] A point-by-point description of a substrate profile (e.g., critical dimensions versus depth) can have several disadvantages. Inputs to the substrate generation system (e.g., process knobs for a substrate manufacturing system, simulation knobs for physics-based models, inputs for machine learning models, etc.) can affect many points in the profile. In a point-by-point description, it can be difficult to isolate the physical (e.g., geometric) effects contributed by changes in the inputs. Furthermore, generating a target profile can be difficult, e.g., many points can be assigned target values, and those target values can be correlated in nonlinear, nontrivial ways.
[0016] Indicator representations of substrate profiles may have several disadvantages: they may not fully represent the profile, they may not accurately represent the profile, they may not represent all portions of the profile, etc. The indicator representation may not adequately represent one or more portions of the profile; for example, multiple profiles that differ in some region of the profile, multiple profiles that differ in some geometric way, multiple profiles that differ by some value, etc. may be represented equally within the indicator representation.
[0017] Aspects of the present disclosure may address one or more of these shortcomings of the prior art. Aspects of the present disclosure may enable the generation of a functional description of key features of one or more profiles of a substrate (e.g., a manufactured substrate, a simulated substrate, etc.). Measurements of the profile of the substrate may be provided to a fitting tool (e.g., fitting software on a general-purpose computer, dedicated hardware, etc.). For example, a series of data points, each corresponding to a critical dimension measurement and an associated depth, may be provided to the fitting tool. The fitting tool may generate a piecewise function describing the profile, e.g., a piecewise function describing the critical dimension as a function of depth.
[0018] In some embodiments, a processing device (e.g., configured to perform the operations of a profile fitting tool) may receive measurements of a profile of a substrate. The data may be divided into portions (e.g., each portion may correspond to a physical region of the profile of the substrate). Each portion may be described by a function (e.g., fit to a function). The entire profile may be described as a piecewise collection of functions that describe these portions.
[0019] In some embodiments, one or more constraints may be applied to regional fit functions, piecewise fit functions, etc. For example, boundaries between portions of the profile data corresponding to boundaries between regions of the substrate may have enforced conditions. Enforced boundary conditions may include continuity (e.g., forcing functions describing adjacent portions of the profile data to have the same value at the boundary, within a threshold error). Enforced boundary conditions may include smoothness (e.g., forcing first derivatives of functions describing adjacent portions of the profile data to have the same value at the boundary, within a threshold error). Enforced boundary conditions may include higher-order conditions (e.g., forcing higher-order derivatives of functions describing adjacent portions of the profile data to have the same value at the boundary, within a threshold error), etc.
[0020] The functions used to fit each portion of the profile may be selected from a library. This selection may be made by a processing device (e.g., by a fitting tool). This selection may be made by a user. In some embodiments, the substrate may include a semiconductor device. In some embodiments, the substrate may include a semiconductor memory device.
[0021] Aspects of the present disclosure may provide technical advantages over conventional techniques. In some embodiments, a complete and accurate (e.g., error, such as sum-of-squares error, is within a target value / threshold) description of a substrate's profile may be generated with a small number of parameters (e.g., fewer than the number of data points describing the profile). In some embodiments, the parameters of the fit may have physical meaning, e.g., the concavity of a portion of the profile (e.g., the coefficients of a second-order polynomial term) may have physical meaning when describing the effective radius of curvature of a portion of the profile. Adjusting various inputs to a substrate generation system (e.g., a manufacturing system, model, etc.) may result in changes in the profile that can be easily analyzed, changes in the profile that can be easily related from parameters to feature dimensions, etc. Fitting the profile to a piecewise fitting function may smooth and / or denoise the profile measurement data.
[0022] In some embodiments, parameters (e.g., fit coefficients) of multiple profiles may be provided to a model (e.g., a statistical model, a clustering model, a machine learning model) to generate additional information about the substrate profile space. In some embodiments, the model may generate data indicating correlations between parameters, which may be readily related to correlations between physical changes in the substrate profile.
[0023] In some embodiments, the profile parameters may be correlated with inputs to the substrate generation system (e.g., by providing the inputs and parameters to train a machine learning model), hi some embodiments, a model may be developed that correlates input parameters to the substrate generation system with profile parameters of the substrate.
[0024] In some embodiments, a profile (e.g., a particular shape of the profile) may be targeted. A substrate generation system may be operated to obtain the target profile. Utilizing techniques of the present disclosure may simplify this process, for example, by correlating fit parameters with geometric characteristics of the substrate profile, correlating input conditions with profile parameters, providing validation of experimental designs, etc.
[0025] The techniques of the present disclosure provide advantages over conventional methods for operating substrate production systems. Designing a processing procedure to produce a target profile can be an expensive process in terms of experimental time, energy, material costs, waste of defective products, costs of developing expertise in experimental design, etc. Designing a procedure to target a profile described by parameters (e.g., parameters that have physical meaning) can reduce these costs.
[0026] Performing clustering on the parameters of the profile fit may enable a more complete understanding of the available output space of the board generation system. For example, the target profile may be outside the accessible output space according to one or more constraints of the board generation system. Easy access to information indicating such constraints may reduce the time, materials, energy, etc. spent on experimental design, testing, or the like.
[0027] In one aspect of the disclosure, a method includes receiving, by a processing device, data indicating a set of measurements of a profile of a substrate. The method further includes dividing, by the processing device, the data indicating the set of measurements into successive sets of data. A first set of the successive sets is associated with a first region of the profile. A second set of the successive sets is associated with a second region of the profile. The method further includes fitting the first set of data to a first function to generate a first fitted function. The first fitted function is selected from a library of functions. The method further includes fitting the second set of data to a second function to generate a second fitted function. The second function is selected from the library of functions. The second function is different from the first function. The method further includes generating a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitted function and a second fitted function.
[0028] In another aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform an operation. The operation includes receiving, by the processing device, data indicating a set of measurements of a profile of a substrate. The operation further includes dividing, by the processing device, the data indicating the set of measurements into series of sets of data. A first set of the series of sets is associated with a first region of the profile. A second set of the series of sets is associated with a second region of the profile. The operation further includes fitting the first set of data to a first function to generate a first fitting function. The first fitting function is selected from a library of functions. The operation further includes fitting the second set of data to a second function to generate a second fitting function. The second function is selected from the library of functions. The second function is different from the first function. The operation further includes generating a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitting function and a second fitting function.
[0029] In another aspect of the present disclosure, a system includes a memory and a processing device coupled to the memory. The processing device is configured to receive data indicating a set of measurements of a profile of a substrate. The processing device is further configured to divide the data indicating the set of measurements into series of sets of data. A first set of the series of sets is associated with a first region of the profile. A second set of the series of sets is associated with a second region of the profile. The processing device is further configured to fit the first set of data to a first function to generate a first fitting function. The first fitting function is selected from a library of functions. The processing device is further configured to fit the second set of data to a second function to generate a second fitting function. The second function is selected from the library of functions. The second function is different from the first function. The processing device is further configured to generate a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fitting function and a second fitting function.
[0030] 1 is a block diagram illustrating an example system 100 (example system architecture) according to some embodiments. System 100 includes client devices 120, manufacturing equipment 124, sensors 126, measurement equipment 128, a prediction server 112, and a data store 140. Prediction server 112 may be part of a prediction system 110. Prediction system 110 may further include server machines 170 and 180.
[0031] The sensors 126 may provide sensor data 142 related to the manufacturing equipment 124 (e.g., related to the manufacturing equipment 124 producing a corresponding product, such as a substrate). The sensor data 142 may be used to ascertain the health of the equipment and / or the health of the product (e.g., product quality). The manufacturing equipment 124 may produce a product according to a recipe or by running over a period of time. In some embodiments, the sensor data 142 may include one or more values of optical sensor data, spectral data, temperature (e.g., heating device temperature), spacing (SP), pressure, high frequency radio frequency (HFRF), radio frequency (RF) match voltage, RF match current, RF match capacitor position, electrostatic chuck (ESC) voltage, actuator position, current, flow rate, power, voltage, etc. The sensor data 142 may include historical sensor data 144 and current sensor data 146. The current sensor data 146 may relate to a product currently being processed, recently processed products, number of recently processed products, etc. The current sensor data 146 may be used as input to a model, such as a trained machine learning model, for example, to generate predicted data 168. The historical sensor data 144 may include data stored related to previously produced products. The historical sensor data 144 may be used to train a model, such as a machine learning model, for example, model 190. The current sensor data 146 may be provided to the model, which may generate as output one or more predictions of properties of a substrate processed under the conditions described by the current sensor data 146. The property predictions may include predictions of critical dimensions (CDs), including a profile of the substrate. The historical sensor data 144 and / or current sensor data 146 may include attribute data, such as manufacturing equipment ID or design label, sensor ID, type and / or location, manufacturing equipment status, e.g., defects present, service life, etc.
[0032] The sensor data 142 may be related to or indicative of manufacturing parameters, such as hardware parameters of the manufacturing equipment 124 (e.g., hardware settings or installed components, e.g., size, type, etc.) or process parameters of the manufacturing equipment 124 (e.g., heater settings, gas flow, etc.). Alternatively or additionally, data related to some hardware and / or process parameters may be stored as manufacturing parameters 150, which may include historical manufacturing parameters 152 (e.g., associated with historical processing runs) and current manufacturing parameters 154. The manufacturing parameters 150 may indicate input settings for the manufacturing devices (e.g., heater power, gas flow rate, etc.). The manufacturing parameters 150 may be provided as model inputs to a model, such as a physics-based model or a machine learning model. The model output may include a simulated substrate, e.g., data representing one or more predicted characteristics of a substrate that would result from processing via the input parameters. The historical parameters 152 may be provided to train a model (e.g., a physics-based model, a machine learning model, etc.). Current parameters 154 may be provided to the model to obtain predicted properties of a substrate manufactured according to the provided current parameters. The predicted properties may include the profile of the simulated substrate, for example, CD as a function of hole depth in the substrate.
[0033] The sensor data 142 and / or manufacturing parameters 150 may be provided while the manufacturing equipment 124 is performing a manufacturing process (e.g., equipment readings while processing a product). The sensor data 142 may vary from product to product (e.g., from substrate to substrate). A substrate (e.g., a substrate produced according to the current parameters 154, a substrate processed under conditions associated with the current sensor data 146, etc.) may have characteristic values (e.g., film thickness, film strain, critical dimensions, etc.) measured by the metrology equipment 128, for example, at a stand-alone metrology facility. Metrology data 160 may be a component of the data store 140. The metrology data 160 may include historical metrology data 164 (e.g., metrology data associated with previously processed products). The metrology data 160 may include current metrology data 166 (e.g., associated with one or more current products). The metrology data may include one or more measurements of the CD of the substrate. The metrology data may include a point-by-point representation of the profile of the substrate (e.g., a CD profile).
[0034] In some embodiments, metrology data 160 may be provided without the use of stand-alone metrology equipment, such as in-situ metrology data (e.g., measurements or surrogates of measurements collected during processing), integrated metrology data (e.g., measurements or surrogates of measurements collected while the product is in the chamber or under vacuum, other than during a processing operation), in-line metrology data (e.g., data collected after the substrate is removed from vacuum), etc. Metrology data 160 may include current metrology data 166 (e.g., metrology data associated with a product currently being processed or a recently processed product).
[0035] In some embodiments, the sensor data 142, the metrology data 160, or the manufacturing parameters 150 may be processed (e.g., by the client device 120 and / or the prediction server 112). Processing the sensor data 142 may include generating features. In some embodiments, the features are patterns in the sensor data 142, the metrology data 160, and / or the manufacturing parameters 150 (e.g., slope, width, height, peaks, board profile, etc.), or combinations of values from the sensor data 142, the metrology data, and / or the manufacturing parameters (e.g., power derived from voltage and current, etc.). The sensor data 142 may include features, which may be used by the prediction component 114 to perform signal processing and / or to obtain predicted data 168 for taking corrective action.
[0036] Each instance (e.g., set) of sensor data 142 may correspond to a product (e.g., a substrate), a set of manufacturing equipment, a type and / or design of substrates produced by the manufacturing equipment, or the like. Similarly, each instance of metrology data 160 and manufacturing parameters 150 may correspond to a product, a set of manufacturing equipment, a type of substrate produced by the manufacturing equipment, or the like. The data store may further store information relating sets of different data types, e.g., information indicating that a set of sensor data, a set of metrology data, and a set of manufacturing parameters all relate to the same product, the same manufacturing equipment, the same type of substrate, etc.
[0037] In some embodiments, a substrate is generated by one or more components of system 100. A physical substrate may be generated by manufacturing equipment 124. For example, properties of the physical substrate may be measured, quantified, and stored as metrology data 160 by metrology equipment 128. A simulated substrate may be generated by prediction system 110. The simulated substrate may be generated using sensor data 142, e.g., current sensor data 146 may be provided to a model, and outputs obtained from the model may indicate one or more properties of a predicted substrate to be generated based on the input conditions. The simulated substrate may be generated using manufacturing parameters 150, e.g., current parameters 154 may be provided to a model, and outputs obtained from the model may indicate one or more properties of a predicted substrate to be generated based on the input parameters. The simulated substrate may be generated based on simulation inputs, such as inputs describing process parameters to which the substrate will be subjected. The simulation inputs may include inputs, such as etch rates, deposition rates, and rate gradients, that are not directly measured or controlled by the physical system. The simulation inputs may include one or more parameters, for example, some inputs may not have explicit physical meaning. Properties of the simulated substrate may be stored, for example, as metrology data 160. The simulated substrate may be generated by one or more physical phenomenon-based models, one or more machine learning models, etc.
[0038] In some embodiments, the prediction system 110 may generate the predicted data 168 using supervised machine learning (e.g., the predicted data 168 includes output from a machine learning model trained using labeled data, such as sensor data labeled with measurement data (which may include, for example, synthetic microscope images generated according to embodiments described herein)). In some embodiments, the prediction system 110 may generate the predicted data 168 using unsupervised machine learning (e.g., the predicted data 168 includes output from a machine learning model trained using unlabeled data, where the output may include clustering results, principal component analysis, anomaly detection, etc.). In some embodiments, the prediction system 110 may generate the predicted data 168 using semi-supervised learning (e.g., the training data may include a mixture of labeled and unlabeled data, etc.). In some embodiments, the prediction system 110 may generate the predicted data 168 using a model based on physical phenomena.
[0039] The prediction data 168 may include data related to a simulated substrate, such as predicted characteristics of a substrate processed at process conditions associated with the simulation inputs. The prediction data 168 may include data related to a profile of the substrate. The prediction data 168 may include functional parameters describing the substrate profile. The prediction data 168 may include output of a model that predicts the functional parameters describing the substrate profile. The prediction data 168 may include output of a model that receives as input one or more parameters of a functional description of the substrate profile and outputs one or more conditions (e.g., process conditions, process recipe operations, fabrication equipment sensor measurements, etc.) predicted to be associated with producing a substrate using the input profile.
[0040] Client device 120, manufacturing equipment 124, sensors 126, measurement equipment 128, prediction server 112, data store 140, server machine 170, and server machine 180 may be coupled to one another via network 130 to generate predictive data 168 for performing corrective actions. In some embodiments, network 130 may provide access to cloud-based services. Operations performed by client device 120, prediction system 110, data store 140, etc. may be performed by a cloud-based virtual device.
[0041] In some embodiments, network 130 is a public network that provides client device 120 with access to prediction server 112, data store 140, and other public computing devices. In some embodiments, network 130 is a private network that provides client device 120 with access to manufacturing equipment 124, sensors 126, measurement equipment 128, data store 140, and other private computing devices. Network 130 may include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), routers, hubs, switches, server computers, cloud computing networks, and / or combinations thereof.
[0042] Client device 120 may include computing devices such as personal computers (PCs), laptops, mobile phones, smartphones, tablet computers, netbook computers, network-connected televisions (“smart TVs”), network-connected media players (e.g., Blu-ray players), set-top boxes, over-the-top (OTT) streaming devices, operator boxes, etc. Client device 120 may include a corrective action component 122. Corrective action component 122 may receive user input of instructions related to manufacturing equipment 124 (e.g., via a graphical user interface (GUI) displayed via client device 120). In some embodiments, corrective action component 122 transmits the instructions to prediction system 110, receives output (e.g., prediction data 168) from prediction system 110, determines corrective action based on the output, and causes the corrective action to be implemented. In some embodiments, the corrective action component 122 obtains sensor data 142 (e.g., current sensor data 146) related to the manufacturing equipment 124 (e.g., from a data store 140, etc.) and provides the sensor data 142 (e.g., current sensor data 146) related to the manufacturing equipment 124 to the predictive system 110.
[0043] In some embodiments, the prediction component 114 may facilitate generation of the predicted data 168 (e.g., by providing input to one or more models 190). To generate the predicted data 168, the corrective action component 122 may retrieve data from the data store 140 and provide the data to the prediction system 110. The sensor data 142 may be provided to the prediction system 110 to generate one or more simulated substrates as output. The manufacturing parameters 150 may be provided to the prediction system 110 to generate one or more simulated substrates as output. The substrate profile data 162 may be provided to the prediction system 110 to generate a functional description of the profile (e.g., a piecewise functional fit of the profile). The profile fit parameters may be provided to the prediction system 110 to generate a predicted procedure for producing a substrate having the input profile. The profile fit parameters may be provided to the prediction system 110 for analysis, such as clustering analysis, parameter space analysis, etc.
[0044] In some embodiments, corrective action component 122 receives output from models 190, prediction component 114, prediction system 110, etc. Corrective action component 122 may store the output data in data store 140, e.g., as prediction data 168, profile data 162, etc. Data output by prediction system 110 may be used as input to another component of prediction system 110. Client device 120 may store data to use as input to one or more models 190. Client device 120 may store output data from one or more models 190. Components of prediction system 110 (e.g., prediction server 112, server machine 170, etc.) may retrieve data (e.g., from client device 120, data store 140, etc.). Prediction server 112 may store output of one or more models 190, e.g., in data store 140, and client device 120 may retrieve the output data. The corrective action component 122 may perform one or more corrective actions based on data, for example, data retrieved from the data store 140, data received from the predictive system 110, and the like.
[0045] In some embodiments, a corrective action component 122 receives corrective action instructions from the predictive system 110 and causes the corrective action to be implemented. Each client device 120 may include an operating system that enables a user to perform one or more of generating, reviewing, or editing data (e.g., instructions related to the manufacturing equipment 124, corrective actions related to the manufacturing equipment 124, etc.).
[0046] In some embodiments, metrology data 160 (e.g., historical metrology data 164) corresponds to historical property data of a product (e.g., a product processed using manufacturing parameters associated with historical sensor data 144 and historical manufacturing parameters in manufacturing parameters 150), and predicted data 168 relates to predicted property data (e.g., predicted property data of a product that will be produced or a product that was produced under conditions recorded by current sensor data 146 and / or current manufacturing parameters). In some embodiments, predicted data 168 is or includes predicted metrology data (e.g., virtual metrology data, simulated substrate data) of a product that will be produced or a product that was produced according to conditions recorded as current sensor data 146, current measurement data, current metrology data 166, and / or current parameters 154. In some embodiments, the predictive data 168 is or includes an indication of any anomalies (e.g., an abnormal product, an abnormal component, an abnormal manufacturing equipment 124, an abnormal energy usage, etc.), and optionally an indication of one or more causes of those anomalies. In some embodiments, the predictive data 168 is an indication of time change or drift of certain components of the manufacturing equipment 124, sensors 126, measurement devices 128, and the like. In some embodiments, the predictive data 168 is an indication of the end of life of components of the manufacturing equipment 124, sensors 126, measurement devices 128, and the like. In some embodiments, the predictive data 168 is an indication of the progress of a processing operation being performed, for example, an indication of the progress of a processing operation being performed, for use in process control.
[0047] Running a manufacturing process that results in a defective product can be costly in terms of time, energy, product, components, manufacturing equipment 124, costs of identifying the defects and scrapping the defective product, etc. By inputting sensor data 142 (e.g., manufacturing parameters that are or will be used to manufacture the product) into the predictive system 110, receiving output of predicted data 168, and performing corrective actions based on the predicted data 168, the system 100 may have the technical advantage of avoiding the costs of producing, identifying, and scrapping defective product. Products that are not predicted to meet performance thresholds can be identified, production can be stopped, corrective actions can be performed, users can be alerted, recipes can be updated, etc.
[0048] Running a manufacturing process that results in a failure of a component of manufacturing equipment 124 can be costly in terms of downtime, damage to the product, damage to the equipment, rush-ordering replacement components, etc. By inputting sensor data 142 (e.g., manufacturing parameters that are or will be used to manufacture the product), metrology data, measurement data, etc. into predictive system 110, receiving output of predictive data 168, and performing corrective action (e.g., predicted operational maintenance, e.g., replacing, treating, cleaning, etc.) based on the predictive data 168, system 100 may have the technical advantage of avoiding the costs of one or more of unexpected component failures, unscheduled downtime, lost productivity, unexpected equipment failures, product waste, or the like. Component performance, e.g., of manufacturing equipment 124, sensors 126, measurement devices 128, and the like, may be monitored over time to provide an indication of deteriorating components.
[0049] The manufacturing parameters may be less than optimal for producing a product, and producing the product may have costly consequences such as increased resource (e.g., energy, coolant, gas, etc.) consumption, increased time to produce the product, increased component failures, increased amount of defective product, etc. By inputting measurements into the predictive system 110, receiving the output of the predictive data 168, and taking corrective action (e.g., based on the predictive data 168) to update the manufacturing parameters (e.g., set optimal manufacturing parameters), the system 100 may have the technical advantage of using optimal manufacturing parameters (e.g., hardware parameters, process parameters, optimal design) to avoid the costly consequences of suboptimal manufacturing parameters.
[0050] Designing and / or updating a recipe can be a costly process, involving experimental design, manufacturing, and metrology operations, each of which are repeated multiple times until a target result is achieved. The target output may be a substrate including a target profile. Metrology data 160 may be input into the prediction system 110 to generate a piecewise functional fit of the substrate's profile. The parameters of the fit may have physical meaning; for example, coefficients of terms in a square polynomial may indicate the sharpness of curvature of a portion of the substrate profile. The fit parameters may be adjusted to generate an updated target profile. By providing the updated fit parameters to the prediction system 110 and receiving as output prediction data 168 corresponding to predicted processing conditions, predicted manufacturing parameters, predicted simulation parameters, or the like, which may result in the input profile, the system 100 may reduce costs associated with recipe design and updates, costs associated with experimental procedures, including material costs, energy costs, equipment usage, etc., costs associated with waste of experimental product, or the like.
[0051] The corrective action may be related to one or more of Computational Process Control (CPC), Statistical Process Control (SPC) (e.g., SPC on electronic components to determine processes under control, SPC to predict the useful life of components, SPC for comparison to 3 sigma graphs, etc.), Advanced Process Control (APC), model-based process control, preventative operational maintenance, design optimization, manufacturing parameter updates, manufacturing recipe updates, feedback control, machine learning corrections, or the like.
[0052] In some embodiments, the corrective action includes issuing an alert (e.g., an alert to stop or not run a manufacturing process if the predictive data 168 indicates a predicted anomaly, such as an anomaly in a product, component, or manufacturing equipment 124). In some embodiments, a machine learning model is trained to monitor the progress of a processing run (e.g., monitor in-situ sensor data to predict whether the manufacturing process has reached completion). In some embodiments, the machine learning model may send instructions to terminate a processing run when the model determines that the process is complete. In some embodiments, the corrective action includes providing feedback control (e.g., feedback control that modifies a manufacturing parameter in response to the predictive data 168 indicating a predicted anomaly). In some embodiments, performing the corrective action includes causing an update to one or more manufacturing parameters. In some embodiments, performing the corrective action may include retraining a machine learning model associated with the manufacturing equipment 124. In some embodiments, performing the corrective action may include training a new model (e.g., a machine learning model) associated with the manufacturing equipment 124.
[0053] The manufacturing parameters 150 may include hardware parameters (e.g., information indicating which components are installed on the manufacturing equipment 124, information indicating component replacement, information indicating the age of a component, information indicating a software version or update, etc.) and / or process parameters (e.g., temperature, pressure, flow rate, rate, current, voltage, gas flow rate, lift speed, etc.). In some embodiments, the corrective action includes performing preventive operational maintenance (e.g., replacing, treating, cleaning, etc., components of the manufacturing equipment 124). In some embodiments, the corrective action includes performing design optimization (e.g., updating manufacturing parameters, manufacturing processes, manufacturing equipment 124, etc. to optimize the product). In some embodiments, the corrective action includes updating a recipe (e.g., changing when a manufacturing subsystem enters idle or active mode, changing set points for various characteristic values, etc.).
[0054] Prediction server 112, server machine 170, and server machine 180 may each include one or more computing devices such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc. Operations of prediction server 112, server machine 170, server machine 180, data store 140, etc. may be performed by a cloud computing service, a cloud data storage service, etc.
[0055] The prediction server 112 may include a prediction component 114. In some embodiments, the prediction component 114 may receive current sensor data 146 and / or current manufacturing parameters (e.g., received from the client device 120 and retrieved from the data store 140) and generate output (e.g., prediction data 168) based on the current data for performing corrective actions associated with the manufacturing equipment 124. The prediction component 114 may receive current data (e.g., current metrology data 166) and generate as output a piecewise functional fit of a profile of the substrate. The prediction component 114 may receive a piecewise functional fit of a target profile of the substrate and generate as output an indication of conditions for producing a substrate having the target profile. In some embodiments, the prediction data 168 may include one or more predicted dimensional measurements of a processed product. In some embodiments, the prediction component 114 may use one or more trained models 190 to determine the output based on the current data. The models 190 may include machine learning models, models based on physical phenomena, statistical models, etc.
[0056] Manufacturing equipment 124 may be associated with one or more machine learning models, such as model 190. The machine learning models associated with manufacturing equipment 124 may perform many tasks, including process control, classification, performance prediction, etc. Model 190 may be trained using data associated with manufacturing equipment 124 or data associated with products processed by manufacturing equipment 124, such as sensor data 142 (e.g., collected by sensors 126), manufacturing parameters 150 (e.g., associated with process control of manufacturing equipment 124), metrology data 160 (e.g., generated by metrology equipment 128), etc.
[0057] One type of machine learning model that may be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. Artificial neural networks generally include a feature representation component with a classifier or recurrent layer that maps features to a desired output space. For example, a convolutional neural network (CNN) hosts multiple layers of convolutional filters. Pooling may be performed and nonlinearities may be addressed in lower layers, and above the lower layers, a multilayer perceptron is typically added to map upper layer features extracted by the convolutional layers to a decision (e.g., a classification output).
[0058] A recurrent neural network (RNN) is another type of machine learning model. Recurrent neural network models are designed to interpret a series of inputs that are intrinsically related to each other, such as time trace data, sequential data, etc. The output of a perceptron in an RNN is fed back as input to that perceptron to generate the next output.
[0059] Deep learning is a type of machine learning algorithm that uses a cascade of multiple layers of nonlinear processing units to perform feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks may learn in a supervised (e.g., classification) and / or unsupervised manner (e.g., pattern analysis). Deep neural networks include a hierarchy of layers, with different layers learning different levels of representation corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into a slightly more abstract and complex representation. For example, in an image recognition application, the raw input may be a matrix of pixels; a first representation layer may extract the pixels and encode edges; a second layer may construct and encode the edge configuration; a third layer may encode higher-level shapes (e.g., teeth, lips, gums, etc.); and a fourth layer may recognize scanning tasks. Notably, the deep learning process can independently learn which features are optimally placed at which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have significant credit assignment path (CAP) depth. A CAP is a chain of transformations from input to output. A CAP describes the potentially causal connections between input and output. For feedforward neural networks, the CAP depth may be the depth of the network or the number of hidden layers + 1. For recurrent neural networks, where signals may propagate through layers more than once, the CAP depth is potentially infinite.
[0060] In some embodiments, the prediction component 114 receives the current sensor data 146, the current metrology data 166, and / or the current manufacturing parameters 154, performs signal processing to decompose the current data into sets of current data, provides the sets of current data as inputs to the trained model 190, and obtains output from the trained model 190 that indicates predicted data 168. In some embodiments, the prediction component 114 receives metrology data for the substrate (e.g., predicted metrology data based on sensor data) and provides the metrology data to the trained model 190. For example, the current sensor data 146 may include sensor data that indicates metrology values (e.g., geometry, profile, etc.) of the substrate. In some embodiments, the prediction data indicates metrology data (e.g., a prediction of substrate quality). In some embodiments, the prediction data indicates component health. In some embodiments, the prediction data indicates process progress (e.g., utilized to terminate a processing operation). In some embodiments, the prediction data indicates a predicted substrate production procedure that will produce a substrate having target characteristics. The prediction data may dictate a procedure for generating a physical or simulated substrate having the target profile.
[0061] In some embodiments, the various models discussed in connection with model 190 (e.g., supervised machine learning models, unsupervised machine learning models, models based on physical phenomena, etc.) may be combined into one model (e.g., an ensemble model) or may be separate models.
[0062] Data may be passed bidirectionally between several separate model and prediction components 114 included in model 190. In some embodiments, some or all of these operations may instead be performed by different devices, such as client device 120, server machine 170, server machine 180, etc. Those skilled in the art will understand that variations in data flow, which components perform which processes, which data is provided to which models, and the like, are within the scope of this disclosure.
[0063] Data store 140 may be memory (e.g., random access memory), a drive (e.g., a hard drive, a flash drive), a database system, a cloud-accessible memory system, or another type of component or device capable of storing data. Data store 140 may include multiple storage components (e.g., multiple drives or multiple databases) that may reside across multiple computing devices (e.g., multiple server computers). Data store 140 may store sensor data 142, manufacturing parameters 150, metrology data 160, synthetic data 162, and prediction data 168.
[0064] The sensor data 142 may include historical sensor data 144 and current sensor data 146. The sensor data may include time tracking of sensor data throughout the duration of the manufacturing process, association of data with physical sensors, preprocessed data such as averages and composite data, and data indicating sensor performance over time (i.e., across many manufacturing processes). The manufacturing parameters 150 and metrology data 160 may also include similar features, such as historical metrology data 164 and current metrology data 166. The historical sensor data 144, historical metrology data 166, and historical manufacturing parameters may be historical data (e.g., at least portions of these data may be used to train one or more models 190). The current sensor data 146, current metrology data 166 may also be current data (e.g., at least portions that follow the historical data and are input to the learning model 190) for which predictive data 168 is generated (e.g., to perform corrective action). The profile data 162 may include measurements such as a physical substrate profile, physical substrate fit parameters, simulated substrate profile data, simulated substrate fit parameters, and target profiles.
[0065] In some embodiments, prediction system 110 further includes server machine 170 and server machine 180. Server machine 170 includes dataset generator 172 that can generate datasets (e.g., a set of data inputs and a set of target outputs) for training, validating, and / or testing model 190, including one or more machine learning models. Some operations of dataset generator 172 are described in detail below with respect to FIGS. 2A-B and 4A. In some embodiments, dataset generator 172 may divide historical data (e.g., historical sensor data 144, historical manufacturing parameters, historical metrology data 164) into a training set (e.g., 60 percent of the historical data), a validation set (e.g., 20 percent of the historical data), and a test set (e.g., 20 percent of the historical data).
[0066] In some embodiments, prediction system 110 generates multiple sets of features (e.g., via prediction component 114). For example, a first set of features may correspond to a first set of types of sensor data (e.g., from a first set of sensors, a first combination of values from the first set of sensors, a first pattern of values from the first set of sensors) corresponding to each of the datasets (e.g., a training set, a validation set, and a test set), and a second set of features may correspond to a second set of types of sensor data (e.g., from a second set of sensors different from the first set of sensors, a second combination of values different from the first combination, a second pattern different from the first pattern) corresponding to each of the datasets.
[0067] In some embodiments, historical data is provided to the machine learning model 190 as training data. In some embodiments, this historical sensor data may be or include microscopic image data. The type of data provided will depend on the application of the machine learning model. For example, the machine learning model may be trained by providing the historical sensor data 144 as training inputs and the corresponding metrology data 160 as target outputs. In some embodiments, a large amount of data may be used to train the model 190. For example, sensor and metrology data for hundreds of substrates may be used.
[0068] Server machine 180 includes a training engine 182, a verification engine 184, a selection engine 185, and / or a test engine 186. Engines (e.g., training engine 182, verification engine 184, selection engine 185, and test engine 186) may refer to hardware (e.g., circuitry, dedicated logic circuitry, programmable logic circuitry, microcode, processing device, etc.), software (e.g., instructions executing on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. Training engine 182 may be capable of training model 190 and / or synthetic data generator 174 using one or more sets of features associated with the training set from dataset generator 172. Training engine 182 may generate multiple trained models 190, where each trained model 190 corresponds to a different set of features of the training set (e.g., sensor data from a different set of sensors). For example, a first trained model may be trained using all features (e.g., X1-X5), a second trained model may be trained using a first subset of features (e.g., X1, X2, X4), and a third trained model may be trained using a second subset of features (e.g., X1, X3, X4, and X5), where the second subset of features may overlap with the first subset of features. Dataset generator 172 may receive the output of the trained models (e.g., prediction data 168, profile data 162, etc.), assemble the data into training, validation, and testing datasets, and use those datasets to train a second model (e.g., a machine learning model configured to output prediction data, corrective actions, etc.).
[0069] The validation engine 184 may be capable of validating the trained models 190 using a corresponding set of validation set features from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of training set features may be validated using a first set of validation set features. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of validation set features. The validation engine 184 may discard trained models 190 with accuracies that do not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 with accuracies that meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting the trained model 190 with the highest accuracy among the trained models 190.
[0070] The testing engine 186 may be capable of testing the trained models 190 using a corresponding set of test set features from the dataset generator 172. For example, a first trained machine learning model 190 trained using a first set of training set features may be tested using a first set of test set features. The testing engine 186 may determine the trained model 190 with the highest accuracy among all of the trained models based on the test set.
[0071] In the case of a machine learning model, model 190 may refer to a model artifact generated by training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for each corresponding training input). Patterns in the dataset that map data inputs to target outputs (correct answers) can be found, and machine learning model 190 is provided with a mapping that captures these patterns. Machine learning model 190 may use one or more of support vector machines (SVMs), radial basis functions (RBFs), clustering, supervised machine learning, semi-supervised machine learning, unsupervised machine learning, k-nearest neighbor algorithms (k-NNs), linear regression, random forests, neural networks (e.g., artificial neural networks, recurrent neural networks), and the like.
[0072] The prediction component 114 may provide current data to the model 190 and may execute the model 190 on the inputs to obtain one or more outputs. For example, the prediction component 114 may provide current sensor data 146 to the model 190 and may execute the model 190 on the inputs to obtain one or more outputs. The prediction component 114 may be able to determine (e.g., extract) predicted data 168 from the output of the model 190. From the output, the prediction component 114 may determine (e.g., extract) confidence data indicating the degree of confidence that the predicted data 168 is an accurate predictor of a process associated with the current sensor data 146 and / or input data for products produced or to be produced under current manufacturing parameters using the manufacturing equipment 124. The prediction component 114 or the corrective action component 122 may use this confidence data to determine whether to cause a corrective action associated with the manufacturing equipment 124 to be performed based on the predicted data 168.
[0073] The confidence data may include or indicate a degree of confidence that the prediction data 168 is an accurate prediction for a product or component associated with at least a portion of the input data. In one example, the confidence is a real number between 0 and 1, inclusive, where 0 indicates no confidence that the prediction data 168 is an accurate prediction for a product processed according to the input data or an accurate prediction for the component health of a component of the manufacturing equipment 124, and 1 indicates absolute confidence that the prediction data 168 accurately predicts a characteristic of a product processed according to the input data or a component health of a component of the manufacturing equipment 124. In response to the confidence data indicating a confidence below a threshold level for a predetermined number of instances (e.g., a percentage of instances, a frequency of instances, a total number of instances, etc.), the prediction component 114 may retrain the trained model 190 (e.g., based on current sensor data 146, current manufacturing parameters, etc.). In some embodiments, retraining may include generating one or more datasets (e.g., via dataset generator 172) utilizing historical and / or simulated data.
[0074] For purposes of illustration and not limitation, aspects of the present disclosure describe training one or more machine learning models 190 using historical data (e.g., historical sensor data 144, historical manufacturing parameters) and synthetic data 162, and inputting current data (e.g., current sensor data 146, current manufacturing parameters 154, and current metrology data 164) into the one or more trained machine learning models to determine predicted data 168. In other embodiments, heuristic, physics-based, or rule-based models are used (e.g., without using trained machine learning models) to determine predicted data 168. In some embodiments, such models may be trained using historical and / or simulated data. In some embodiments, these models may be retrained using a combination of true historical data and simulated data. The prediction component 114 may monitor historical sensor data 144, historical manufacturing parameters, and metrology data 160. Any of the information described with respect to data inputs 210A-B in FIGS. 2A-B may be monitored or otherwise used by the heuristic, physics-based, or rule-based models.
[0075] In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, and server machine 180 may be provided by fewer machines. For example, in some embodiments, server machines 170 and 180 may be combined into a single machine, and in other embodiments, server machine 170, server machine 180, and prediction server 112 may be combined into a single machine. In some embodiments, client device 120 and prediction server 112 may be combined into a single machine. In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, server machine 180, and data store 140 may be performed by a cloud-based service.
[0076] In general, functions described as being performed by client device 120, prediction server 112, server machine 170, and server machine 180 in one embodiment may, in other embodiments, be performed on prediction server 112, where appropriate. Furthermore, functions attributed to particular components may be performed by different or multiple components operating together. For example, in some embodiments, prediction server 112 may determine corrective actions based on prediction data 168. In another example, client device 120 may determine prediction data 168 based on output from a trained machine learning model.
[0077] Additionally, the functionality of a particular component may be performed by different or multiple components operating together. One or more of prediction server 112, server machine 170, or server machine 180 may be accessed as a service offered to other systems or devices through an appropriate application programming interface (API).
[0078] In embodiments, a "user" may be described as a single individual. However, other embodiments of the present disclosure encompass a "user" being an entity controlled by multiple users and / or automated sources. For example, a collection of individual users united as a group of administrators may be considered a "user."
[0079] Embodiments of the present disclosure may be applied to data quality assessment, feature enhancement, model validation, virtual metrology (VM), predictive maintenance (PdM), marginal optimization, process control or the like.
[0080] 2A-B show block diagrams of exemplary dataset generators 272A-B (e.g., dataset generator 172 of FIG. 1 ) that generate datasets for training, testing, validation, etc. of a model (e.g., model 190 of FIG. 1 ), according to some embodiments. Each dataset generator 272 may be part of server machine 170 of FIG. 1 . In some embodiments, several models associated with manufacturing equipment 124 may be trained, used, and maintained (e.g., within a manufacturing facility). Each model may be associated with one dataset generator 272, multiple models may share one dataset generator 272, etc. Dataset generators 272A-B may be used to generate datasets for machine learning models, statistical models, models based on physical phenomena, etc.
[0081] 2A illustrates a system 200A that includes a dataset generator 272A for generating a dataset for one or more models (e.g., model 190 of FIG. 1). The dataset generator 272A may use historical data to generate the dataset (e.g., data input 210A, target output 220A). The historical data may include simulated data, for example, metrology data for a simulated substrate. In some embodiments, a dataset generator similar to dataset generator 272A may be utilized to train an unsupervised machine learning model; for example, the target output 220A may not be generated by dataset generator 272A.
[0082] The dataset generator 272A may generate datasets for training, testing, and validating a model. In some embodiments, the dataset generator 272A may generate datasets for a machine learning model. In some embodiments, the dataset generator 272A may generate datasets for training, testing, and / or validating a model configured to generate a synthetic (e.g., digital or virtual) substrate. The dataset generator 272A may generate the sets of historical sensor data 244A-244Z as data inputs 210A to the machine learning model. The sets of historical sensor data 244A-244Z may be provided as data inputs 210A to the machine learning model. The machine learning model may be configured to receive the sensor data as data inputs and generate profile parameters indicative of properties of the simulated substrate as model outputs.
[0083] In some embodiments, the dataset generator 272A may generate a set of target outputs 220A for the machine learning model. The target outputs 220A may include output substrate profile parameter data 268. The set of target outputs 220A may be related to a set of data inputs 210A. For example, the set of target outputs 220A may describe a profile of a substrate processed under conditions described by the corresponding set of data inputs 210A. The target outputs 220A may be provided to the machine learning model to train, validate, test, etc. the machine learning model.
[0084] A dataset generator similar to dataset generator 272A may be utilized to generate a dataset for a model including a range of functions. The dataset generator may generate a dataset for a model configured to generate a simulated substrate (e.g., configured to generate data indicative of properties of the simulated substrate). The dataset generator may generate a set of data including sensor data, manufacturing parameters, simulation parameters, etc. as data inputs 210A, and may generate a set of data describing substrate properties as target outputs 220A for such a model.
[0085] A dataset generator may generate a dataset for a model configured to generate a profile function fit from substrate profile data. The dataset generator may generate as data input 210A a set of data including a substrate profile, e.g., a set of data including a collection of data points / measurements describing the substrate profile. The dataset generator may generate as target output 220A a set of data including a classification of functions (e.g., selected from a library of functions) to use to fit one or more portions of the profile, fit parameters, boundaries between regions described by different fit parameters, boundary conditions between the regions, etc. The target output 220A may include human-labeled data, machine-labeled data (e.g., the best fit found by a processing device by searching the available data space of fit functions, parameters, boundary locations, boundary conditions, etc.), or the like.
[0086] A dataset generator may generate a dataset for a model configured to generate parameters for a profile function fit from substrate generation data. The dataset generator may generate a set of data as data input 210A including data related to substrate generation. The substrate generation data may include data related to the generation of physical and / or simulated substrates. The substrate generation data may include sensor data related to substrate manufacturing, substrate manufacturing parameters, simulation inputs, etc. The dataset generator may generate a set of data as target output 220A including substrate profile function fit parameters. The target output 220A may include a classification of a function (e.g., selected from a library of functions) to use in the fit, fit parameters, boundary locations between fit regions, boundary constraints / conditions, etc.
[0087] A dataset generator may generate a dataset for a model configured to generate substrate production inputs from a substrate profile. The substrate profile may include data points, measurements, profile function fit parameters, etc. The dataset generator may generate a set of data as data input 210A including a set of substrate profile data. The dataset generator may generate a set of data as target output 220A including substrate production data. The substrate production data may include data related to the generation of a simulated or physical substrate. The substrate production data may include sensor data, manufacturing parameters, simulation inputs, etc.
[0088] In some embodiments, a dataset generator 272A generates a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210A (e.g., training inputs, validation inputs, test inputs). The data inputs 210A may be provided to the training engine 182, the validation engine 184, or the test engine 186. This dataset may be used to train, validate, or test a model (e.g., model 190 of FIG. 1).
[0089] In some embodiments, data input 210A may include one or more sets of data. As an example, system 200A may generate a set of sensor data that may include one or more of sensor data from one or more types of sensors, combinations of sensor data from one or more types of sensors, patterns from sensor data from one or more types of sensors, and / or composite versions thereof. Similar subsets of data may be generated for machine learning models configured to receive different data as input.
[0090] In some embodiments, the dataset generator 272A may generate a first data input corresponding to the first set of historical sensor data 244A for training, validating, or testing a first machine learning model. The dataset generator 272A may generate a second data input corresponding to the second set of historical sensor data 244B for training, validating, or testing a second machine learning model.
[0091] In some embodiments, dataset generator 272A generates a dataset (e.g., a training set, a validation set, a test set), which includes one or more data inputs 210A (e.g., training inputs, validation inputs, test inputs) and may include one or more target outputs 220A corresponding to the data inputs 210A. The dataset may also include mapping data that maps the data inputs 210A to the target outputs 220A. The data inputs 210A may also be referred to as “features,” “attributes,” or “information.” In some embodiments, dataset generator 272A may provide the dataset to training engine 182, validation engine 184, or testing engine 186, where the dataset is used to train, validate, or test a machine learning model (e.g., one or more of the machine learning models included in model 190, ensemble model 190, etc.).
[0092] In some embodiments, a dataset generator, such as dataset generator 272A, may be utilized to generate a dataset for one or more models that are not machine learning models. For example, the dataset generator may be utilized to generate a dataset for a physics-based model. The dataset generator may generate a dataset for a physics-based model configured to generate a simulated substrate. The dataset generator may generate data inputs 210A and / or target outputs 220A for the non-machine learning model. The physics-based model may utilize the dataset generated by the dataset generator to assign and / or adjust values for one or more parameters that define a relationship between an input to and an output from the physics-based model.
[0093] FIG. 2B shows a block diagram of an exemplary dataset generator 272B for generating a dataset for an unsupervised model configured to analyze clustering of input data, according to some embodiments. A system 200B including the dataset generator 272B (e.g., dataset generator 172 of FIG. 1 ) generates a dataset for one or more machine learning models (e.g., model 190 of FIG. 1 ). The dataset generator 272B may use historical data to generate a dataset (e.g., data input 210B). The exemplary dataset generator 272B is configured to generate a dataset for a machine learning model configured to receive function profile fit parameters as input and generate data indicative of fit parameter clustering as output. Similar dataset generators (or similar operations of dataset generator 272B) may be utilized for machine learning models configured to perform different functions, such as a machine learning model configured to receive sensor data and predicted metrology data as input, a machine learning model configured to receive target metrology data (e.g., target microscope images) as input, and generate estimated conditions or processing operation recipes as output that are likely to produce a device consistent with the input target data, etc. Data set generator 272B may share one or more features and / or functionality with data set generator 272A.
[0094] The dataset generator 272B may generate datasets for training, testing, and validating a model. The model may be a machine learning model, a model based on physical phenomena, a statistical model, etc. The model may be provided with the set of profile fit data 262A-262Z (e.g., output from a model trained using the dataset from the dataset generator 272A) as data input 210B. The machine learning model may include two or more separate models (e.g., the machine learning model may be an ensemble model). The machine learning model may be configured to generate output data indicative of a pattern of profile fit parameters, clustering, correlation, outlier data, anomalous data detection, etc.
[0095] In some embodiments, dataset generator 272B generates a dataset (e.g., a training set, a validation set, a test set), which includes one or more data inputs 210B (e.g., training inputs, validation inputs, test inputs). Data inputs 210B may also be referred to as "features," "attributes," or "information." In some embodiments, dataset generator 272B may provide this dataset to training engine 182, validation engine 184, or test engine 186, where the dataset is used to train, validate, or test a machine learning model (e.g., model 190 of FIG. 1). Some embodiments of generating a training set are further described with respect to FIG. 4A.
[0096] In some embodiments, the dataset generator 272B may generate a first data input corresponding to the first set of profile match data 262A for training, validating, or testing a first machine learning model, and the dataset generator 272B may generate a second data input corresponding to the second set of profile match data 262B for training, validating, or testing a second machine learning model.
[0097] The data input 210B for training, validating, or testing a machine learning model may include information for a particular manufacturing chamber (e.g., of a particular piece of substrate manufacturing equipment). In some embodiments, the data input 210B may include information for a particular type of manufacturing equipment, e.g., manufacturing equipment sharing particular characteristics. The data input 210B may include data related to a certain type of device, e.g., intended function, device design, devices produced using a particular recipe, etc. Training a machine learning model based on one type of equipment, device, recipe, etc. may enable the trained model to generate clustered data useful for several substrates (e.g., for several different facilities, products, etc.).
[0098] In some embodiments, following generating a dataset and using the dataset to train, validate, or test a machine learning model, the model may be further trained, validated, or tested, or adjusted (e.g., weights or parameters, such as connection weights of a neural network, associated with the model's input data).
[0099] FIG. 3 is a block diagram illustrating a system 300 for generating output data (e.g., predicted data 168 of FIG. 1 ) according to some embodiments. In some embodiments, system 300 may relate to the generation and use of models. System 300 may relate to the generation and use of machine learning models. System 300 may relate to the generation and use of models based on physical phenomena, statistical models, etc. While the description of FIG. 3 is directed to machine learning models, similar techniques may be applicable to other types of models. System 300 may be used with one or more additional models. In some embodiments, system 300 may be used with machine learning models to generate simulated substrates. System 300 may be used to determine piecewise fit functions for substrate profiles. System 300 may be used for analysis of fit parameters, e.g., clustering, outlier analysis, etc. System 300 may be used to predict substrate production operations that may result in a target substrate profile. System 300 may be used with machine learning models associated with manufacturing systems that have different functionality than those listed above.
[0100] At block 310, the system 300 (e.g., a component of the prediction system 110 of FIG. 1 ) performs data partitioning (e.g., via the dataset generator 172 of the server machine 170 of FIG. 1 ) of data to use in training, validating, and / or testing the machine learning model. In some embodiments, the training data 364 includes historical data, such as historical metrology data, historical design rule data, historical classification data (e.g., classification of whether a product meets a performance threshold), historical substrate profile data, etc. The training data 364 may also include historical sensor data. In some embodiments, the training data 364 may also include synthetic data, such as data associated with a simulated substrate. The training data 364 may undergo data partitioning at block 310 to generate the training set 302, the validation set 304, and the test set 306. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data.
[0101] The generation of training set 302, validation set 304, and test set 306 may be tailored to a particular application. For example, the training set may be 60% of the training data, the validation set may be 20% of the training data, and the test set may be 20% of the training data. System 300 may generate multiple sets of features for each of the training, validation, and test sets. For example, if training data 364 includes sensor data and manufacturing parameters, and the sensor data and manufacturing parameters include sensor data from 20 sensors (e.g., sensor 126 in FIG. 1 ) and features derived from 10 manufacturing parameters (e.g., manufacturing parameters corresponding to the same processing run as the sensor data from these 20 sensors), the sensor data may be divided into a first set of features including sensors 1-10 and a second set of features including sensors 11-20. Furthermore, the manufacturing parameters may be divided into multiple sets, such as a first set of manufacturing parameters including parameters 1-5 and a second set of manufacturing parameters including parameters 6-10. The target inputs may be divided into multiple sets, the target outputs may be divided into multiple sets, both the target inputs and the target outputs may be divided into multiple sets, or neither the target inputs nor the target outputs may be divided into multiple sets. Multiple models may be trained on different sets of data.
[0102] At block 312, the system 300 performs model training (e.g., via the training engine 182 of FIG. 1 ) using the training set 302. Training of machine learning models and / or training of models based on physical phenomena (e.g., digital twins) may be achieved with supervised learning methods, which involve feeding a training dataset consisting of labeled inputs through the model, observing its outputs, defining an error (by measuring the difference between the output and the label values), and tuning the model's weights using techniques such as deep gradient descent and backpropagation to minimize the error. In many applications, repeating this process across many labeled inputs in the training dataset yields a model that can generate correct outputs when presented with inputs different from those present in the training dataset. In some embodiments, training of machine learning models may be achieved in an unsupervised manner, e.g., no labels or classifications may be provided during training. Unsupervised models may be configured to perform anomaly detection, outcome clustering, outlier analysis, etc.
[0103] For each training data item in the training dataset, the training data item may be input to a model (e.g., a machine learning model). The model may then process the input training data item (e.g., sensor data related to a substrate processing procedure, etc.) to generate an output. The output may include, for example, parameters related to parameters of a substrate profile fit. This output may be compared to the label of the training data item (e.g., a substrate profile fit not generated by the model, a fit generated by a subject matter expert, etc.).
[0104] Processing logic may then compare the generated output (e.g., profile fit parameters) with the labels included in the training data items (e.g., human-generated fit parameters). Processing logic determines an error (i.e., classification error) based on the difference between the output and the labels. Processing logic adjusts one or more weights, biases, and / or other values of the model based on this error.
[0105] When training a neural network, an error term or delta may be determined for each node of the artificial neural network. Based on this error, the artificial neural network adjusts one or more of its parameters (weights for one or more inputs of a node) for one or more of its nodes. The parameters may be updated in a backpropagation manner, with the nodes in the top layer updated first, followed by the nodes in the next layer, and so on. An artificial neural network includes multiple layers of "neurons," each of which receives as input values from neurons in the previous layer. The parameters for each neuron include weights associated with the values received from each of the neurons in the previous layer. Adjusting the parameters may therefore include adjusting the weights assigned to each of the inputs to one or more neurons in one or more layers of the artificial neural network.
[0106] System 300 may train multiple models using multiple sets of features from training set 302 (e.g., a first set of features from training set 302, a second set of features from training set 302, etc.). For example, system 300 may train a model to generate a first trained model using a first set of features in the training set (e.g., sensor data from sensors 1-10, metrology measurements 1-10, etc.) and to generate a second trained model using a second set of features in the training set (e.g., sensor data from sensors 11-20, metrology measurements 11-20, etc.). In some embodiments, the first trained model and the second trained model may be combined to generate a third trained model (e.g., which may, alone, be a better predictor or synthetic data generator than either the first or second trained model). In some embodiments, the feature sets used in comparing models may overlap (e.g., the first set of features may be sensor data from sensors 1-15 and the second set of features may be sensor data from sensors 5-20). In some embodiments, hundreds of models may be generated, including models with various permutations of features and combinations of models.
[0107] At block 314, the system 300 performs model validation (e.g., via validation engine 184 of FIG. 1 ) using the validation set 304. The system 300 may validate each of the trained models using a corresponding set of features in the validation set 304. For example, the system 300 may validate a first trained model using a first set of features in the validation set (e.g., sensor data from sensors 1-10 or metrology measurements 1-10) and a second trained model using a second set of features in the validation set (e.g., sensor data from sensors 11-20 or metrology measurements 11-20). In some embodiments, the system 300 may validate hundreds of models (e.g., models with various permutations of features, combinations of models, etc.) generated at block 312. At block 314, the system 300 may determine the accuracy of each of the one or more trained models (e.g., via model validation) and may determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, flow returns to block 312, where system 300 performs model training using a different set of features from the training set. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, flow proceeds to block 316. System 300 may discard trained models that have an accuracy lower than the threshold accuracy (e.g., based on a validation set).
[0108] At block 316, the system 300 performs model selection (e.g., via selection engine 185 of FIG. 1 ) to determine which model of the one or more trained models that meet the threshold accuracy has the highest accuracy (e.g., selected model 308 based on the check of block 314). In response to determining that two or more models of the trained models that meet the threshold accuracy have the same accuracy, flow may return to block 312, where the system 300 performs model training to determine the trained model with the highest accuracy using a further refined training set corresponding to the further refined set of features.
[0109] At block 318, the system 300 performs model testing (e.g., via the test engine 186 of FIG. 1 ) using the test set 306 to test the selected model 308. The system 300 may test the first trained model using a first set of features in the test set (e.g., sensor data from sensors 1-10) and determine that the first trained model meets a threshold accuracy (e.g., based on the first set of features of the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 is too well-fitted to the training set 302 and / or the validation set 304 and cannot be applied to other data sets, such as the test set 306), flow proceeds to block 312, where the system 300 performs model training (e.g., retraining) using a different training set (e.g., sensor data from a different sensor) corresponding to a different set of features. In response to determining, based on the test set 306, that the selected model 308 has an accuracy that meets the threshold accuracy, flow proceeds to block 320. At least in block 312, the model may learn patterns in the training data to make predictions or generate synthetic data, and in block 318, the system 300 may apply the model to the remaining data (e.g., the training set 306) to test the predictions or generate synthetic data.
[0110] At block 320, the system 300 uses the trained model (e.g., the selected model 308) to receive current data 322 (e.g., current sensor data 146 of FIG. 1 ) and determine (e.g., extract) output profile data 324 (e.g., profile data 162 of FIG. 1 ) from the output of the trained model. Corrective action related to the manufacturing equipment 124 of FIG. 1 , such as updating a process recipe or scheduling or performing maintenance on the manufacturing equipment and / or sensors, may be taken into account. In some embodiments, the current data 322 may correspond to the same types of features of the historical data used to train the machine learning model. In some embodiments, the current data 322 may correspond to a subset of the types of features of the historical data used to train the selected model 308 (e.g., the machine learning model may be trained using some sensor measurements and configured to generate output based on a subset of the sensor measurements).
[0111] In some embodiments, data other than current data 322 may be provided as model input, and data other than output profile data 324 may be received as model output. Models that perform other functions, such as the models described in connection with Figures 2A-B, may also be used with systems similar to system 300.
[0112] In some embodiments, the performance of a machine learning model trained, validated, and tested by system 300 may deteriorate. For example, a manufacturing system associated with the trained machine learning model may undergo gradual or sudden changes. The changes in the manufacturing system may result in a degradation of the performance of the trained machine learning model. A new model may be generated to replace the degraded machine learning model. This new model may be generated by modifying the old model through retraining, generating a new model, etc. Retraining may be performed by introducing additional training data 346, including retraining data, and performing additional model training, for example, at block 312.
[0113] In some embodiments, one or more of operations 310-320 may be performed in various orders and / or with other operations not shown and described herein. In some embodiments, one or more of operations 310-320 may not be performed. For example, in some embodiments, one or more of data partitioning of block 310, model validation of block 314, model selection of block 316, or model testing of block 318 may not be performed.
[0114] 3 illustrates a system configured to train, validate, test, and use one or more machine learning models. The machine learning models are configured to accept data (e.g., set points provided to manufacturing equipment, sensor data, measurement data, etc.) as input and provide data (e.g., prediction data, corrective action data, classification data, etc.) as output. The partition, training, validation, selection, testing, and use blocks of system 300 may be similarly performed to train a second model using a different type of data. Additionally, retraining may be performed using current data 322 and / or additional training data 346.
[0115] 4A-B are flow diagrams of methods 400A-B related to training and utilizing a machine learning model, according to certain embodiments. Methods 400A-B may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic circuitry, programmable logic circuitry, microcode, a processing device, etc.), software (e.g., instructions executing on a processing device, a general-purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 400A-B may be performed in part by prediction system 110. Method 400A may be performed in part by prediction system 110 (e.g., server machine 170 and dataset generator 172 of FIG. 1 , dataset generators 272A-B of FIGS. 2A-B). Prediction system 110 may use method 400A to generate datasets for at least one of training, validating, or testing a machine learning model according to embodiments of the present disclosure. Method 400B may be executed by prediction server 112 (e.g., prediction component 114) and / or server machine 180 (e.g., training, validation, and testing operations may be performed by server machine 180). In some embodiments, a non-transitory machine-readable storage medium stores instructions that, when executed by a processing device (e.g., a processing device of prediction system 110, a processing device of server machine 180, a processing device of prediction server 112, etc.), cause the processing device to perform one or more of methods 400A-B.
[0116] For ease of explanation, methods 400A-B are illustrated and described as a series of operations. However, operations in accordance with the present disclosure may be performed in various orders and / or simultaneously, and with other operations not shown and described herein. Moreover, not all illustrated operations may be performed to implement methods 400A-B in accordance with the disclosed subject matter. Furthermore, those skilled in the art will understand and appreciate that methods 400A-B may also be represented as a series of interrelated states via a state diagram or events.
[0117] 4A is a flow diagram of a method 400A for generating a dataset for a model, according to some embodiments. Method 400A may be used to generate a dataset for a machine learning model, a model based on physical phenomena, etc. Referring to FIG. 4A, in some embodiments, at block 401, processing logic performing method 400A initializes dataset T (e.g., a dataset for training, validating, or testing a model) to be an empty set.
[0118] At block 402, processing logic generates a first data input (e.g., a first training input, a first validation input) that may include one or more of sensor, manufacturing parameter, metrology data, substrate profile data, etc. In some embodiments, this first data input may include a first set of features for a type of data, and the second data input may include a second set of features for a type of data (e.g., as described with respect to FIG. 3 ). In some embodiments, the input data may include historical data and / or synthetic data.
[0119] In some embodiments, at block 403, processing logic optionally generates a first target output for one or more of these data inputs (e.g., a first data input). In some embodiments, the input includes one or more instructions for substrate processing (e.g., sensor data, manufacturing parameter data), and the target output includes simulated substrate characteristics. In some embodiments, the input includes substrate profile data, and the output includes a fit of the profile, such as a piecewise function fit. In some embodiments, the input includes profile fit parameters, and the output includes substrate production conditions for producing a substrate that matches the provided profile. In some embodiments, no target output is generated (e.g., an unsupervised machine learning model that can group, cluster, or find correlations in input data without the need to provide a target output).
[0120] At block 404, processing logic optionally generates mapping data indicating an input / output mapping. This input / output mapping (or mapping data) may relate to a data input (e.g., one or more of the data inputs described herein), a target output for the data input, and an association between the data input and the target output. In some embodiments, such as those associated with machine learning models in which no target output is provided, block 404 may not be performed.
[0121] In some embodiments, at block 405, processing logic adds the mapping data generated at block 404 to the dataset T.
[0122] At block 406, processing logic branches based on whether dataset T is sufficient for at least one of training, validating, and / or testing a machine learning model, such as model 190 in FIG. 1. If so, execution proceeds to block 407; if not, execution returns to block 402. It should be noted that in some embodiments, whether dataset T is sufficient may be determined based on the number of inputs in the dataset, in some embodiments based on the number of inputs in the dataset that are mapped to outputs, and in some other embodiments, whether dataset T is sufficient may be determined based on one or more other criteria (e.g., measures of diversity of data examples, accuracy, etc.) in addition to or instead of the number of inputs.
[0123] At block 407, processing logic provides dataset T (e.g., to server machine 180) for training, validating, and / or testing machine learning model 190. In some embodiments, dataset T is a training set, and dataset T is provided to training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set, and dataset T is provided to validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set, and dataset T is provided to test engine 186 of server machine 180 to perform testing. For example, in the case of a neural network, input values (e.g., numerical values associated with data input 210A) of a given input / output mapping are input to the neural network, and output values (e.g., numerical values associated with target output 220A) of the input / output mapping are stored in output nodes of the neural network. The connection weights of the neural network are then adjusted according to a learning algorithm (e.g., backpropagation, etc.), and the procedure is repeated for the remaining input / output mappings of dataset T. After block 407, the model (e.g., model 190) may be at least one of trained using training engine 182 of server machine 180, validated using validation engine 184 of server machine 180, or tested using test engine 186 of server machine 180. The trained model may be implemented by prediction component 114 (of prediction server 112) to generate prediction data 168 for performing signal processing, to generate profile data 162, or to perform corrective actions related to manufacturing equipment 124.
[0124] 4B is a flow diagram of a method 400B for generating a profile piecewise function fit according to some embodiments. At block 410 of method 400B, processing logic receives data indicating a plurality of measurements of a profile of a substrate. The substrate may be a physical substrate, e.g., a physical substrate manufactured by manufacturing equipment 124 of FIG. 1 , and the profile may be measured by metrology equipment 128, etc. The substrate may be a simulated substrate, e.g., the measurements may be generated by a machine learning model, a model based on physical phenomena, etc. The substrate may be a semiconductor device. The substrate may be a memory device, e.g., a semiconductor memory device.
[0125] Generating the simulated substrate may include providing input to a model. Generating the simulated substrate may include providing one or more machine learning inputs to a machine learning model. Generating the simulated substrate may include providing one or more simulation inputs to a physically-based model. Generating the simulated substrate may include obtaining one or more indications of properties of the simulated substrate as output from the model. The indications of the properties of the simulated substrate may include values related to a profile of the substrate, e.g., measurements of geometry of the simulated substrate. Processing logic may perform data processing to determine the measurements of the profile of the substrate from the indications of the properties of the substrate, e.g., in a format related to measuring the profile of a physical substrate, a format accepted by the model as input, etc.
[0126] At block 412, processing logic divides the data indicating the plurality of measurements into a plurality of sets of data, where a first set of the plurality of sets is associated with a first region of the profile, and a second set of the plurality of sets is associated with a second region of the profile, and in some embodiments, the boundaries between the sets may be boundaries between regions of the profile described by different fitting functions.
[0127] At block 414, processing logic fits the first set of data to a first function to generate a first fitting function. The first function may be selected from a library of functions. Generating the first fitting function may include determining one or more parameters of the fitting function. A procedure for generating a fitting function may be used, for example, a procedure for generating a fitting function to minimize an error function between the fit and measurements of a region of the profile. The library of functions may include polynomial functions (e.g., zeroth-order polynomials or constants, first-order linear polynomials, second-order square polynomials, higher-order polynomials, etc.), exponential functions, logarithmic functions, and / or other types of functions that may be applicable to the substrate profile. The library of functions may include combinations of functions, for example, additive combinations of functions, multiplicative combinations of functions, etc. The fitting procedure may select values for coefficients and / or other parameters, for example, to minimize an error function between the fit and data points associated with the substrate profile. In some embodiments, a user selects the first function from a library of functions. In some embodiments, processing logic selects a first function from a library of functions.
[0128] At block 416, processing logic fits the second set of data to a second function to generate a second fitting function. The second function may be selected from a library of functions, e.g., the same library from which the first function was selected. In some embodiments, a first region of the profile is adjacent to a second region of the profile. Generating the first and second fitting functions may include considering one or more boundary conditions, e.g., constraints enforced at the boundary between the two regions. For example, a continuity constraint may be enforced (the value of the first fitting function as it approaches the boundary is equal to the value of the second fitting function as it approaches the boundary from the second region). A smoothness constraint may be enforced (the value of the first derivative of the first fitting function is equal to the value of the first derivative of the second fitting function as it approaches the boundary). Higher-order constraints may also be enforced, e.g., concavity constraints associated with the second derivatives of the first and second fitting functions. Enforcing constraints may produce a more realistic (e.g., physically plausible) fitting model. Enforcing constraints may produce a better statistical fit by removing free variables (e.g., degrees of freedom) from the fitting procedure.
[0129] At block 418, processing logic generates a piecewise function fit of the profile of the substrate. The piecewise function fit includes a first fit function and a second fit function.
[0130] In some embodiments, additional operations may be performed. For example, processing logic may obtain a plurality of piecewise function fits (e.g., parameters of the fits). Processing logic may provide the plurality of piecewise function fits to a machine learning model. Processing logic may receive data as output from the model that directs an analysis of the parameters. For example, the model may be configured to perform clustering analysis, outlier analysis, etc. on the provided fit parameters.
[0131] In some embodiments, the piecewise profile fit may be utilized for further learning. For example, parameters of the fit may be utilized to generate an understanding of the impact of substrate production inputs (e.g., process conditions, simulation conditions, etc.) on the profile shape. Parameters of the piecewise profile fit function may have physical meaning. The relationship between the input parameters and physical results may improve system learning, recipe generation, anomaly detection, corrective action recommendations, etc. For example, processing logic may provide one or more input conditions associated with producing the substrate to the model. Processing logic may provide the piecewise function fit to the model. Processing logic may obtain from the model an indication of the impact of a first input condition of the one or more input conditions on a first parameter of the piecewise function fit.
[0132] 5A illustrates a substrate measurement generation system 500A according to some embodiments. The substrate generation system 500A includes generation of a physical substrate and a simulated substrate. For example, the physical substrate and the simulated substrate may be used together to increase the accuracy of the substrate generation system over using the simulated substrate alone, to reduce the cost of the substrate generation system over using the physical substrate generation alone, etc.
[0133] When creating a physical substrate, process parameters 520 are generated. The process parameters 520 may be provided to a processing device (e.g., a controller of a substrate manufacturing system). The process parameters 520 may be input by a user or generated by a model, for example. The process parameters 520 may be related to manufacturing parameters, for example, manufacturing parameters stored in data store 140 of FIG. 1.
[0134] The process parameters are provided to a processing tool 522. The processing tool 522 may be or include the fabrication equipment 124 of FIG. 1. The processing tool 522 may include one or more processing chambers. The processing tool 522 may be configured to perform processing operations on one or more substrates. The processing operations may include etching operations, deposition operations, annealing operations, etc. The processing tool 522 may produce a substrate 524, which is a physical substrate.
[0135] When generating the simulated substrate, model inputs 528 are obtained by the processing device. The model inputs 528 may include inputs provided to a model configured to generate the simulated substrate as an output. The model inputs 528 may include sensor data. The model inputs 528 may include manufacturing parameters. The model inputs 528 may include other simulation inputs.
[0136] The model inputs are provided to a model 530. The model 530 may be a physics-based model, a machine learning model, a combination thereof, etc. The model 530 may take the model inputs into account to generate a simulated substrate, substrate 524. The simulated substrate may include data indicative of properties of the substrate, for example, data indicative of the geometry of the substrate.
[0137] The substrate 524 is provided to a system for generating substrate measurements 526. The substrate 524 may be provided to a system for generating measurements of the profile of the substrate, for example, measurements of the critical dimension (CD) profile of the substrate. The physical substrate 524 may be provided to a metrology tool to perform the substrate measurements. The simulated substrate may be provided to a digital tool to extract corresponding measurements (e.g., the CD profile of the substrate) from the data provided by the model 530.
[0138] FIG. 5B illustrates a piecewise function fit 560 of a substrate 540 and a substrate profile, according to some embodiments. The substrate 540 includes a recess 542, e.g., a hole, a trough, etc. The profile of the substrate 540 may describe the shape of the recess, e.g., a shape generated while performing one or more etching operations on the substrate. The profile of the substrate 540 may also describe the CD of the substrate 540. The profile of the substrate 540 may be a measure of how wide the recess 542 is as a function of depth; for example, the profile may be an indication of the distance from sidewall to sidewall perpendicular to the centerline 544. The profile may include several measurements of the width of the recess 542 measured at various depths of the recess 542, several measurements of the width of the recess 542 measured perpendicular to various positions of the centerline 544, etc.
[0139] The profile of substrate 540 may be considered to consist of several regions, e.g., regions 546-552. Some regions of the profile may be curved (e.g., regions 546 and 550), while some regions may be substantially straight (e.g., regions 548 and 552). The generation of a fitting function that describes the substrate profile may be improved by generating a piecewise function, e.g., by generating a piecewise function that includes different fitting functions that target the fitting portions, where the fitting portions have shapes such that the fitting portions of the profile may be accurately described by the fitting functions.
[0140] Function fit 560 shows a fit of the profile of substrate 540. Function fit 560 shows a piecewise profile function fit (e.g., CD as a function of depth). This fit is shown as profile fit 570. The plot of function fit 560 includes boundaries 562, 564, and 568. Boundaries 562-568 correlate to the boundaries between regions 546-552 of substrate 540.
[0141] In some embodiments, a user may select functions corresponding to regions of the function fit. For example, a user may be presented with a profile (e.g., data points representing a profile of substrate 540), and the user may select a series of functions that may be used to fit the profile. For example, the user may select a square function for the region before boundary 562, a linear function for the region before boundary 564, a square function for the region before boundary 568, and a linear function for the remainder of the profile. A processing device may determine the locations of boundaries 562-568 (e.g., to minimize an error function), parameters of the fit functions, etc.
[0142] In some embodiments, a user may select constraints (e.g., boundary conditions) to which the piecewise function fit should adhere. In some embodiments, a user may select one set of constraints to use for each boundary (e.g., boundaries 562-568). In some embodiments, a user selects a set of constraints for each boundary individually (e.g., the boundary conditions associated with boundary 562 may be different from the boundary conditions associated with boundary 564).
[0143] In some embodiments, a processing device may select functions to use to fit regions of the profile. The processing device may utilize a machine learning model (e.g., configured to divide the profile into multiple fitting regions) to select functions to use to fit the profile. The processing device may utilize a model based on physical phenomena to select functions to use to fit the profile. The processing device may utilize a fitting model (e.g., which may select functions for fitting that produce fits having minimized error functions, minimized merit functions, etc.). In some embodiments, a hybrid system may be utilized, e.g., a user may select a subset of a library of functions for consideration, and the processing device may determine which subset to use for fitting.
[0144] In some embodiments, a processing device selects boundary constraints to which the piecewise function fit should obey. The processing device may select boundary conditions to minimize an error function, to achieve a target number of free variables, to minimize another merit function, etc. Some embodiments may utilize a hybrid system, e.g., a user may select a subset of the list of boundary conditions that may be enforced, and the processing device may further refine the selection of boundary conditions to apply at each boundary.
[0145] FIG. 5C is a flow diagram of system components of a system 500C for generating and utilizing a piecewise function fit of a substrate profile, according to some embodiments.
[0146] The system 500C includes a function library 502. The function library 502 may include a set of functions that may be fitted to various portions of the substrate profile data corresponding to physical regions of the substrate profile. The function library 502 may include, for example, polynomial functions (e.g., constant or zeroth order polynomials, linear or first order polynomials, square or second order polynomials, cubic or third order polynomials, higher order polynomials, etc.), exponential functions (e.g., y=A Bx Function library 502 may include functions of the form (functions of the form (x, y, y, y), logarithmic functions, distribution functions (e.g., Gaussian, Lorentzian, Voigt, etc.), trigonometric functions (e.g., sine, cosine, hyperbolic sine, etc.), logistic functions, etc. Function library 502 may allow for and / or include combinations of functions, such as additive combinations (e.g., polynomial functions plus exponential functions), multiplicative combinations (e.g., exponential functions multiplied by logarithmic functions), etc. In some embodiments, a subset of library 502 may be utilized; for example, user selection may limit the number and / or types of functions available to fit a profile, the number and / or types of functions available to fit one or more portions of a profile, etc. In some embodiments, a user may select one or more functions from function library 502 to use in fitting a profile.
[0147] System 500C further includes a constraint set 504. Constraint set 504 may include one or more constraints, e.g., one or more constraints for use in fitting the piecewise function fit (e.g., for use by fitting tool 506). The constraints may include conditions enforced at boundaries of the piecewise function fit (e.g., boundaries between regions fitted by different functions, boundaries between regions having different shapes, etc.). The constraints may include that the continuity of the function across the boundary and / or the continuity of the derivative of the function across the boundary is within a threshold. For example, a constraint may enforce continuity of the piecewise function fit across the boundary, a second constraint may enforce smoothness of the fit across the boundary (e.g., continuity of the first derivative), a third constraint may enforce smooth curvature of the fit across the boundary (e.g., continuity of the second derivative), etc. The constraints may be different for different boundaries of the piecewise function fit. The constraints may be selected by a user or may be selected (e.g., optimized) by a processing device. Constraints may be used to reduce the fit space (e.g., reduce the number of floating parameters, the number of free variables, etc.).
[0148] In some embodiments, the elements of group 512 may be selected by a user. For example, a user may select functions from a function library to represent different parts of the profile and select boundary conditions to enforce at each boundary. This user selection may be passed to fitting tool 506.
[0149] The fitting tool 506 may generate a piecewise function fit of the substrate profile. The fitting tool 506 may be configured to reduce an error function, such as a least-squares error, an error function that penalizes non-zero coefficients and reduces some terms, etc. In some embodiments, elements of group 514 may be performed by a processing device, for example, the processing device may select the boundary points between regions of the profile, the function to use to fit the regions, and the constraints to employ (potentially subject to one or more user selections). The processing device may select the functions, boundaries, parameter values, etc. to minimize the error of the fit.
[0150] The piecewise function fits may be provided to a profile representation tool 508. The profile representation tool may be utilized to visualize, analyze, etc., the piecewise function fits of the substrate profile. The piecewise function fits may be further provided to a synthesis and analysis tool 510. The synthesis and analysis tool may assist a user in designing experiments to match the profile, designing experiments to vary one or more features of the profile, mining for clustering of parameter values, selected functions, or other aspects of one or more piecewise function fits, correlating substrate generation inputs with function fit outputs, correlating parameters of the piecewise function fits, etc. The synthesis and analysis tool 510 may be utilized to extract various relationships between data associated with the piecewise function fits. For example, the synthesis and analysis tool 510 may be utilized to generate plots relating input parameters to output parameters (e.g., scatter plots of one parameter against model inputs).
[0151] FIG. 6 is a block diagram illustrating a computer system 600 according to some embodiments. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network such as a local area network (LAN), an intranet, an extranet, or the Internet). Computer system 600 may operate in the capacity of a server or a client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a server, a network router, a switch, or a bridge, or any other device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the device. Furthermore, the term “computer” is intended to include a collection of computers that individually or jointly execute a set (or sets) of instructions to perform any one or more of the methods described herein.
[0152] In additional aspects, computer system 600 may include a processing device 602, a volatile memory 604 (e.g., random access memory (RAM)), a non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.
[0153] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing another type of instruction set, or a microprocessor implementing a combination of instruction set types), or a specialized processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[0154] Computer system 600 may further include a network interface device 622 (e.g., coupled to a network 674). Computer system 600 may further include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generating device 620.
[0155] In some embodiments, data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable medium) having stored thereon instructions 626 encoding one or more of the methods or functions described herein, including instructions encoding the components of FIG. 1 (e.g., prediction component 114, corrective action component 122, model 190, etc.) and instructions for implementing the methods described herein.
[0156] The instructions 626 may further reside, completely or partially, within the volatile memory 604 and / or within the processing device 602 during execution of the instructions 626 by the computer system 600; thus, the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.
[0157] While the computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" is intended to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" is also intended to include any tangible medium that can store or encode a set of instructions for execution by a computer, causing the computer to perform one or more of the methods described herein. The term "computer-readable storage medium" is intended to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0158] The methods, components, and features described herein may be implemented by discrete hardware components or may be integrated into the functionality of other hardware components, such as an ASIC, FPGA, DSP, or similar device. Further, the methods, components, and features described herein may be implemented by firmware modules or by functional circuitry within a hardware device. Furthermore, the methods, components, and features described herein may be implemented in any combination of hardware devices and computer program components, or in a computer program.
[0159] Unless otherwise specified, terms such as "receive," "perform," "provide," "acquire," "perform," "access," "determine," "add," "use," "train," "reduce," "generate," "correct," or other similar terms refer to operations and processes performed or implemented by a computer system that manipulate and convert data in computer system registers and memory, represented as physical (electronic) quantities, into other data in the computer system memory or registers, or other such information storage, transmission, or display device, also represented as physical quantities. Furthermore, as used herein, terms such as "first," "second," "third," "fourth," etc., are meant as labels to distinguish between different elements and may not have an ordinal meaning based on the numerical designation of those terms.
[0160] The examples described herein further relate to apparatus for performing the methods described herein. The apparatus may be specially constructed to perform the methods described herein, or may include a general-purpose computer system that is selectively programmed by a computer program stored on the computer system. Such a computer program may be stored on a computer-readable tangible storage medium.
[0161] The methods and illustrative examples described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized apparatus to perform the methods described herein and / or each of the individual functions, routines, subroutines, or operations of those methods. Example structures for these various systems are provided in the description above.
[0162] The above description is intended to be illustrative, and not limiting. While the present disclosure has been described with reference to particular illustrative examples and embodiments, it will be recognized that the present disclosure is not limited to the described examples and embodiments. The scope of the present disclosure should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. receiving, by a processing device, data indicative of a plurality of measurements of a profile of the substrate; dividing, by the processing device, the data indicative of the plurality of measurements into a plurality of sets of data, a first set of the plurality of sets associated with a first region of the profile and a second set of the plurality of sets associated with a second region of the profile; fitting the first set of data to a first function to generate a first fitted function, the first function being selected from a library of functions; fitting the second set of data to a second function to generate a second fitted function, the second function being selected from the library of functions, the second function being different from the first function; and generating a piecewise function fit of the profile of the substrate, the piecewise function fit including the first fit function and the second fit function. A method comprising:
2. generating the piecewise functional fit of the profile, applying one or more constraints to data points associated with a boundary between the first region and the second region; The method of claim 1 , comprising:
3. The constraint is: the continuity of the piecewise function fit across the boundary; the continuity of the first derivative of the piecewise function fit across the boundary; and Continuity of the second derivative of the piecewise function fit across the boundary 3. The method of claim 2, selected from the group comprising:
4. said library of functions: zero degree polynomial, first degree polynomial, quadratic polynomial, exponential function, or Logarithmic functions The method of claim 1 , comprising at least one of:
5. the plurality of measurements of the profile of the substrate are associated with a simulated substrate, and generating the simulated substrate; providing one or more simulation inputs to a physics-based model; receiving data indicative of the simulated substrate from the physics-based model; and extracting the plurality of measurements of the profile of the substrate from the data indicative of the simulated substrate. The method of claim 1 , comprising:
6. the plurality of measurements of the profile of the substrate are associated with a simulated substrate, and generating the simulated substrate; providing one or more machine learning inputs to the machine learning model; receiving data from the machine learning model indicative of a geometry of the simulated substrate; and extracting the plurality of measurements of the profile of the substrate from the data indicative of the simulated substrate geometry; The method of claim 1 , comprising:
7. The method of claim 1 , wherein the substrate comprises a semiconductor memory device.
8. receiving a plurality of piecewise function fits, the plurality of piecewise function fits relating to a plurality of profiles of a plurality of substrates; providing the plurality of piecewise function fits and the piecewise function fits to a machine learning model; and receiving data from the machine learning model instructing the plurality of piecewise function fits and clustering of fit parameters of the piecewise function fits; The method of claim 1 further comprising:
9. receiving a user selection of the first function; and receiving a user selection of the second function; The method of claim 1 further comprising:
10. selecting, by the processing device, the first function from the library of functions; and selecting, by the processing device, the second function from the library of functions. The method of claim 1 further comprising:
11. providing a model with one or more input conditions associated with generating said substrate; providing the piecewise function fit to the model; receiving from the model an indication of an effect of a first input condition of the one or more input conditions on a first parameter of the piecewise function fit; The method of claim 1 further comprising:
12. A non-transitory machine-readable storage medium having instructions stored thereon that, when executed, receiving data indicating a plurality of measurements of a profile of the substrate; dividing the data indicative of the plurality of measurements into a plurality of sets of data, a first set of the plurality of sets being associated with a first region of the profile and a second set of the plurality of sets being associated with a second region of the profile; fitting the first set of data to a first function to generate a first fitted function, the first function being selected from a library of functions; fitting the second set of data to a second function to generate a second fitted function, the second function being selected from the library of functions, the second function being different from the first function; and generating a piecewise function fit of the profile of the substrate, the piecewise function fit including the first fit function and the second fit function. A non-transitory machine-readable storage medium that causes a processing device to perform operations including
13. 13. The non-transitory machine-readable storage medium of claim 12, wherein generating the piecewise function fit of the profile comprises applying one or more constraints to data points associated with a boundary between the first region and the second region.
14. said library of functions: zero degree polynomial, first degree polynomial, quadratic polynomial, exponential function, or Logarithmic functions 13. The non-transitory machine-readable storage medium of claim 12, comprising at least one of:
15. The non-transitory machine-readable storage medium of claim 12 , wherein the substrate comprises a semiconductor memory device.
16. The operation further comprises: providing a model with one or more input conditions associated with generating said substrate; providing the piecewise function fit to the model; receiving from the model an indication of an effect of a first input condition of the one or more input conditions on a first parameter of the piecewise function fit; 13. The non-transitory machine-readable storage medium of claim 12, comprising:
17. 1. A system including a memory and a processing device coupled to the memory, the processing device comprising: receiving data indicating a plurality of measurements of a profile of the substrate; dividing the data indicative of the plurality of measurements into a plurality of sets of data, a first set of the plurality of sets being associated with a first region of the profile and a second set of the plurality of sets being associated with a second region of the profile; fitting the first set of data to a first function to generate a first fitted function, the first function being selected from a library of functions; fitting the second set of data to a second function to generate a second fitted function, the second function being selected from the library of functions, the second function being different from the first function; and generating a piecewise function fit of the profile of the substrate, the piecewise function fit including the first fit function and the second fit function. A system configured to run
18. 20. The system of claim 17, wherein generating the piecewise function fit of the profile includes applying one or more constraints to data points associated with a boundary between the first region and the second region.
19. The constraint is: the continuity of the piecewise function fit across the boundary; the continuity of the first derivative of the piecewise function fit across the boundary; and Continuity of the second derivative of the piecewise function fit across the boundary 20. The system of claim 18, selected from the group comprising:
20. the processing device further comprising: selecting the first function from the library of functions; and selecting the second function from the library of functions; 20. The system of claim 17 configured to perform:
Citation Information
Patent Citations
Estimating a parameter of a substrate
US20210271171A1