Generating Indications of Learning of Models for Semiconductor Processing
A GUI is used to visualize high-dimensional input and output spaces, addressing the challenge of understanding model learning in semiconductor processing, enabling optimized substrate production by intuitively representing input-output relationships and nonlinearities.
Patent Information
- Application Number
- JP2025500875
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-26
- Filing Date
- 2023-07-24
- Publication Date
- 2025-09-02
AI Technical Summary
Conventional systems struggle to provide an intuitive and efficient way to understand the correlations and nonlinearity between inputs and outputs of machine learning models used in semiconductor processing, making it difficult to optimize substrate production.
A graphical user interface (GUI) is employed to visualize high-dimensional input and output spaces, allowing users to see how inputs affect outputs and facilitating the understanding of model learning, including the display of nonlinearities and the impact of multiple inputs on multiple outputs.
Enables users to gain intuition for model learning and make adjustments to processing parameters, optimizing substrate production by visually representing the relationships between inputs and outputs, thereby improving the predictability and consistency of semiconductor wafer properties.
Smart Images

Figure 2025528672000001_ABST
Abstract
Description
[Technical Field]
[0001] This specification relates to trained models, and more particularly, to generating indications indicative of model learning for models associated with semiconductor processing. [Background technology]
[0002] Chambers are used in many types of processing systems. Chambers include etch chambers, deposition chambers, anneal chambers, etc. Typically, a substrate, such as a semiconductor wafer, is placed on a substrate support within the chamber, and conditions within the chamber are established and maintained to process the substrate. Models are often utilized to improve processing procedures. Models can include trained machine learning models and physics-based models. Summary of the Invention
[0003] The following is a simplified summary of the disclosure to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure, nor to delineate the scope or claims of particular implementations of the disclosure. Its sole purpose is to present some concepts of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] In one aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include receiving a first value associated with a first input parameter of a model. The first input parameter is associated with a first processing condition of a semiconductor wafer processing procedure. The operations further include receiving a first plurality of values. The first plurality of values range from a lowest value of the first plurality of values to a highest value of the first plurality of values. Each of the first plurality of values is associated with a second input parameter of the model. The second input parameter is associated with a second processing condition of the semiconductor wafer processing procedure. The operations further include providing the first value and the first plurality of values to the model. The operations further include receiving a first plurality of outputs from the model. Each of the first plurality of outputs is associated with the first value and one of the first plurality of values. Each of the first plurality of outputs is associated with a first feature of one of the first plurality of simulated substrates. The operations further include preparing the first plurality of outputs for presentation via a presentation element of a graphic user interface (GUI). The presentation element includes two axes. A first of the two axes corresponds to a first characteristic of the first feature. A second of the two axes corresponds to a second characteristic of the first feature. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element indicating a value of the first characteristic of the first feature and a value of the second characteristic of the first feature associated with each of the first plurality of outputs.
[0005] In another aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include receiving a first plurality of outputs from a first model. Each of the first plurality of outputs is associated with a first input value and one of the first plurality of input values. Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates. The operations further include receiving a second plurality of outputs from a second model. Each of the second plurality of outputs is associated with the first value and one of the first plurality of values. Each of the second plurality of outputs is associated with a second characteristic of one of the first plurality of simulated substrates. The operations further include preparing the first plurality of outputs and the second plurality of outputs for presentation via a presentation element of a GUI. The presentation element includes two axes. The first axis corresponds to a first characteristic of the first characteristic. The second axis corresponds to a second characteristic of the second characteristic. Preparing the first plurality of outputs and the second plurality of outputs for presentation includes facilitating generation of a graphic for display in a presentation element indicating, for each simulated substrate of the first plurality of simulated substrates, a value of the first property of the first feature and a value of the second property of the second feature.
[0006] In another aspect of the present disclosure, a method includes receiving a first value associated with a first input parameter of a first model. The first input parameter is associated with a process recipe for processing a substrate. The method further includes receiving, by one or more processors, a first plurality of values. The first plurality of values range from a lowest value to a highest value. Each of the first plurality of values is associated with a second input parameter of the first model. The second input parameter is associated with the process recipe. The method further includes providing the first value and the first plurality of values to the first model. The method further includes receiving a first plurality of outputs from the first model. Each of the first plurality of outputs is associated with the first value and one of the first plurality of values. Each of the first plurality of outputs is associated with a first characteristic of the simulated substrate. The method further includes preparing the first plurality of outputs for presentation via a presentation element of a GUI. The presentation element includes two independent axes. A first of the two independent axes corresponds to a first characteristic of the first characteristic. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic in a presentation element that visually displays a relationship between an output of the first plurality of outputs and the first characteristic of the first feature.
[0007] The present disclosure is illustrated by way of example and not limitation in the figures of the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example system (example system architecture) according to some embodiments. [Figure 2] FIG. 1 is a block diagram of an example dataset generator used to create a dataset for a model, according to some embodiments. [Figure 3] FIG. 1 is a block diagram illustrating a system for generating output data for analysis of model training, according to some embodiments. [Figure 4A]FIG. 1 is a flow diagram of a method for generating a dataset for a machine learning model, according to some embodiments. [Figure 4B] FIG. 1 is a flow diagram of a method for generating a visual representation of the training of a model, according to some embodiments. [Figure 4C] FIG. 1 is a flow diagram of a method for generating a visual representation of the training of multiple models, according to some embodiments. [Figure 4D] FIG. 1 is a flow diagram of a method for generating graphics that demonstrate model learning, according to some embodiments. [Figure 5A] FIG. 10 illustrates an example presentation element that displays an indication of model training for one or more models, according to some embodiments. [Figure 5B] FIG. 10 illustrates an example presentation element illustrating the training of one or more models, according to some embodiments. [Figure 5C] FIG. 1 illustrates an example GUI, according to some embodiments. [Figure 5D] FIG. 10 illustrates an example presentation element illustrating model learning associated with multiple varying input parameters, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating a computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Described herein are techniques, methods, systems, and devices related to systematically tracking and / or reporting on learning performed by machine learning models and / or other types of models, such as statistical models. In examples, machine learning models associated with substrate manufacturing and / or processing are discussed in some embodiments. Visual indications of model learning are presented in some embodiments. In some embodiments, visual indications are provided for high-dimensional input and output spaces, such as those used for machine learning models and statistical models. In embodiments, a graphic user interface (GUI) is provided that allows a user to easily visualize high-dimensional input and output spaces and the relationships between them. The GUI allows a user to see curves showing how an input follows through the output space for any input condition. Thus, the GUI allows a user to see how changing an input affects the output of a machine learning or statistical model. For example, in embodiments, the GUI can provide clues as to what the machine learning model has learned and can provide information about the process being modeled.
[0010] Manufacturing equipment is used to create substrates, such as semiconductor wafers. The properties of these substrates are determined by the conditions under which the substrates are processed. Accurate knowledge of the property values within an operating manufacturing chamber, especially in the immediate vicinity of the substrate, can be used to predict the properties of the finished product, consistently produce substrates with the same properties, and adjust process parameters to optimize substrate production.
[0011] Machine learning models and statistical models may be utilized to understand and / or predict substrate processing procedures. For example, the machine learning model and / or statistical model may be trained to receive as input values indicative of processing conditions for a substrate. In some cases, process parameters may be provided as inputs (e.g., processing recipe setpoints such as heater power, plasma generation power, duration of processing, etc.). In some cases, processing conditions may be provided to the model as inputs (e.g., sensor data collected during substrate processing, such as temperature, pressure, component operation, etc.). In some cases, a combination of data types may be provided to the model as inputs. The model may be configured to generate an indication of an output of the processing procedure (e.g., one or more predicted measurements of the substrate resulting from the processing conditions indicated by the inputs to the model).
[0012] In some cases, generating an understanding of model learning can be inconvenient. For example, in some cases, a user may consider utilizing a trained machine learning model to develop an understanding (e.g., an intuitive understanding) of how an output (e.g., metrology of a fabricated substrate) relates to one or more inputs. In conventional systems, a user may provide a set of input conditions to a model (e.g., a machine learning model) and receive one or more predictions of output substrates from the model. The user may then provide a different set of input conditions to the model and receive one or more predictions of output substrates associated with the new input conditions from the model. Using conventional systems, generating an understanding of model learning, such as concisely displaying correlations between inputs and outputs, can be cumbersome, time-consuming, etc. For example, it can be difficult to recognize the impact of one input on multiple possible outputs (e.g., multiple predicted measurements of a simulated substrate, multiple predicted properties of a simulated substrate, simulated metrology measurements, etc.). It can be difficult to recognize the impact of changing multiple inputs (e.g., changing multiple processing recipe parameters) on one or more outputs. It can be difficult to ascertain the strength of the correlation between inputs and outputs. It can be difficult to ascertain any nonlinearity in the correlation between inputs and outputs. It can be difficult to ascertain the size of a model's available input or output space and the relationship of particular inputs and / or outputs to that size of space.
[0013] In some conventional systems, the influence of an input on an output may be displayed, for example, by a bar graph. Multiple inputs may be placed on one axis (e.g., the x-axis) of the bar graph, and the output may be placed on an orthogonal axis (e.g., the y-axis). The variation of one input at a time according to some "center" or "reference" set of inputs may be represented by the bars on the graph (e.g., the slope of the output space in the direction of the input space associated with one input may be represented by the size of the bar). In such systems, the variation of the output in response to changes in the input across the entire range of the input space may not be represented (e.g., the output may respond differently to changes in the same input from different sets of reference input conditions). Additionally, the influence on a second feature of interest (e.g., a second output) that may be correlated with the first output is not captured in such a representation. In some cases, it may be difficult to ascertain the magnitude of the model's input and / or output space from such a representation. It may be difficult to ascertain the relationship between the output associated with a set of inputs and the magnitude of the output space.
[0014] The techniques, methods, and systems of the present disclosure may address one or more of these shortcomings of conventional systems. In some embodiments, one or more models (e.g., machine learning models, physics-based models, statistical models, etc.) may receive a set of inputs indicating processing parameters of a substrate. The models may generate as outputs one or more indications of predicted characteristics of a substrate (e.g., a semiconductor wafer) associated with the set of inputs. The characteristics may include measurable quantities on any substrate, such as thickness, resistance, refractive index, attenuation coefficient, sheet resistance, geometric measurements (critical dimensions, line widths, etch depths, sidewall heights, etc.), reflectivity, surface properties, etc. In some embodiments, the models may be configured to generate data indicative of characteristics of the substrate; for example, the models may generate as outputs an indication of one or more predicted thickness measurements of the substrate. In some embodiments, characteristics of the characteristics may be of interest. For example, the model may predict the thickness of the substrate at several spatial locations on the substrate. Statistical characteristics of the predicted characteristics, such as the mean, median, standard deviation, uniformity of the thickness, etc., may be calculated.
[0015] In some embodiments, a graphic user interface (GUI) may be utilized to update settings of a tool (e.g., a software tool) utilized in model training and / or presenting relationships between model inputs and outputs. The GUI may further display results of the model training and / or relationships between model inputs and outputs. The GUI may include a presentation element for displaying one or more plots presenting information related to the model training and / or relationships between model inputs and outputs. The GUI may receive one or more instructions (e.g., via user input). The GUI may receive, for example, a user indication of features of interest and / or characteristics of the features of interest (e.g., a user may select one or more outputs of one or more models to be plotted via the presentation element), one or more configuration settings (e.g., to customize the appearance of a plot illustrating model training, customize calculated or displayed data ranges), etc. In some embodiments, various data may be displayed on the same plot via the presentation element, e.g., outputs associated with multiple sets of input conditions may be displayed on the same plot.
[0016] In some embodiments, the output of one or more machine learning models and / or other models (e.g., statistical models) may be presented on a plot, e.g., a scatter plot. In some embodiments, the output from the models may be represented on two axes, e.g., one axis may be associated with one characteristic of the output features (e.g., a statistical characteristic) and the second axis may be associated with a second characteristic of the output features. In some embodiments, one axis may be associated with a characteristic of a first output feature and the second axis may be associated with a characteristic of a second output feature (e.g., output from a second model).
[0017] In some embodiments, the magnitude of the output space of one or more models may be represented. For example, many combinations of inputs (e.g., spanning the input space for which one or more models were trained or configured) may be provided to one or more models. Each of the outputs associated with the input combinations may be presented on the same plot as one or more outputs associated with a particular set of inputs of interest. In some embodiments, data points representing the magnitude of the model output space may be presented differently from a particular model output (e.g., presented against the background of model outputs associated with inputs specified by a user, inputs associated with process optimization, inputs associated with the substrate processing recipe being optimized, etc.).
[0018] In some embodiments, a median or reference set of inputs may be provided to one or more models. Outputs associated with the reference set of inputs may be presented on the same plot as outputs associated with the magnitude of the output space. The outputs associated with the magnitude of the output space may provide a visual indication of a constraint on the magnitude of the output of one or more models (e.g., the training output space of one or more models) compared to a particular output of interest (e.g., associated with a particular set of inputs, associated with model training, etc.).
[0019] In some embodiments, one or more sets of inputs of interest may be provided to one or more models and associated with outputs displayed on a plot. For example, multiple values of one input parameter may be provided to one or more models while holding other input parameters at baseline values. A series of values associated with one input that is not fixed may be generated. For example, the series of values may range from a minimum input value to a maximum input value (e.g., from the minimum value of the input within the trained input space of one or more models to the maximum value of the input within the trained input space of one or more models). In some embodiments, the range of outputs may be displayed via a presentation element in a GUI. In some embodiments, a visual indicator may be displayed to distinguish output values associated with the minimum value of a varying input from outputs associated with the maximum value of the varying input. The series of outputs may generate a curve indicative of model learning or configuration, and the shape, length, and density of data points along the curve may indicate the effect in output space of varying the input from the minimum input value to the maximum input value.
[0020] In some embodiments, multiple inputs may be varied. For example, for a first input, a range of values may be generated. A set of outputs may be received from one or more models, each associated with a different value of the first input and a reference value of all other inputs. Then, a second input may be varied. A range of values may be generated for the second input. A set of outputs may be received from one or more models, each associated with a different value of the second input and a reference value of all other inputs (e.g., including the first input). Then, all outputs (e.g., outputs associated with a described size of the output space of one or more models, outputs associated with varying the first input, outputs associated with varying the second output, etc.) may be displayed via a presentation element. In some embodiments, the presentation of outputs associated with varying the first input may be visually distinguished from the presentation of outputs associated with varying the second input, e.g., by color, shape, pattern, etc. In some embodiments, the presentation element may display the results of varying several input parameters. A plot may be generated that includes several curves (e.g., a collection of points associated with various values ranging from minimum to maximum of the input parameter), each of which meets / intersects at an output associated with a reference input condition.
[0021] In some embodiments, the system may receive a user indication of a reference condition (e.g., via a GUI). In some embodiments, the system may receive a user indication of a range of one or more inputs (e.g., minimum and maximum values) to vary (e.g., via a GUI). In some embodiments, the system may receive an indication of an interval of the varying inputs (e.g., how many values of the varying inputs are provided to one or more models, the interval between input values, etc.). In some embodiments, the system may provide one or more sets of inputs to one or more models in response to a user selection (e.g., via a GUI) of one or more settings.
[0022] Aspects of the present disclosure provide technical advantages over conventional methods. In some embodiments, the impact on one or more output metrics (e.g., as presented by placing a set of data points on multiple axes) may be displayed as an input parameter varies through a range of values. Nonlinearities (e.g., the difference in change in output value for a given change in input value depending on the initial input value) may be visually captured. Value ranges associated with multiple input parameters may be displayed simultaneously or together. A display (e.g., a plot displayed via a presentation element in a GUI) may include a visual record of learning associated with training one or more machine learning models (or, in some embodiments, physics-based models, etc.), e.g., a visual record of the effect of varying one or more input values on output values as understood by the one or more models. In some embodiments, a user may easily change settings / parameters of a presentation element, plot, etc. For example, user selection of one or more features (e.g., one or more predicted characteristics of simulated wafers, simulated metrology measurements, etc.), user selection of one or more characteristics of the features (e.g., statistical metrics, geometric constraints, etc.), user selection of a criteria set of input values, etc. may be easily adjusted. The plot can be quickly updated to display a record of model learning associated with updated user selections. In this way, the user can manipulate the model's output space, gain intuition for model learning, gain intuition for processing inputs to output property mappings, or make adjustments to process parameters or processing recipes for target properties.
[0023] In one aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include receiving a first value associated with a first input parameter of a model. The first input parameter is associated with a first processing condition of a substrate processing procedure. The operations further include receiving a first plurality of values. The first plurality of values range from a lowest value of the first plurality of values to a highest value of the first plurality of values. Each of the first plurality of values is associated with a second input parameter of the model. The second input parameter is associated with a second processing condition of the substrate processing procedure. The operations further include providing the first value and the first plurality of values to the model. The operations further include receiving a first plurality of outputs from the model. Each of the first plurality of outputs is associated with the first value and one of the first plurality of values. Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates. The operations further include preparing the first plurality of outputs for presentation via a presentation element of a graphic user interface (GUI). The presentation element includes two axes. A first of the two axes corresponds to a first characteristic of the first feature. A second of the two axes corresponds to a second characteristic of the first feature. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element indicating a value of the first characteristic of the first feature and a value of the second characteristic of the first feature associated with each of the first plurality of outputs.
[0024] In another aspect of the present disclosure, a non-transitory machine-readable storage medium stores instructions that, when executed, cause a processing device to perform operations. The operations include receiving a first plurality of outputs from a first model. Each of the first plurality of outputs is associated with a first input value and one of the first plurality of input values. Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates. The operations further include receiving a second plurality of outputs from a second model. Each of the second plurality of outputs is associated with the first value and one of the first plurality of values. Each of the second plurality of outputs is associated with a second characteristic of one of the first plurality of simulated substrates. The operations further include preparing the first plurality of outputs and the second plurality of outputs for presentation via a presentation element of a GUI. The presentation element includes two axes. The first axis corresponds to a first characteristic of the first characteristic. The second axis corresponds to a second characteristic of the second characteristic. Preparing the first plurality of outputs and the second plurality of outputs for presentation includes facilitating generation of a graphic for display in a presentation element indicating, for each simulated substrate of the first plurality of simulated substrates, a value of the first property of the first feature and a value of the second property of the second feature.
[0025] In another aspect of the present disclosure, a method includes receiving a first value associated with a first input parameter of a first model. The first input parameter is associated with a process recipe for processing a substrate. The method further includes receiving, by one or more processors, a first plurality of values. The first plurality of values range from a lowest value to a highest value. Each of the first plurality of values is associated with a second input parameter of the first model. The second input parameter is associated with the process recipe. The method further includes providing the first value and the first plurality of values to the first model. The method further includes receiving a first plurality of outputs from the first model. Each of the first plurality of outputs is associated with the first value and one of the first plurality of values. Each of the first plurality of outputs is associated with a first characteristic of the simulated substrate. The method further includes preparing the first plurality of outputs for presentation via a presentation element of a GUI. The presentation element includes two independent axes. A first of the two independent axes corresponds to a first characteristic of the first characteristic. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic in a presentation element that visually displays a relationship between an output of the first plurality of outputs and the first characteristic of the first feature.
[0026] FIG. 1 is a block diagram illustrating an example system 100 (example system architecture) according to some embodiments. System 100 includes client device 120, manufacturing equipment 124, sensors 126, measurement equipment 128, prediction server 112, and data store 140. Prediction server 112 may be part of prediction system 110. Prediction system 110 may further include server machines 170 and 180. Client device 120 may include a presentation component 115 that may perform one or more of the methods described in connection with FIGS. 4A-4D . In some embodiments, presentation component 115 may be included in whole or in part in a different component of system 100, such as prediction server 112, server machine 170, or server machine 180.
[0027] In some embodiments, the manufacturing equipment 124 (e.g., a cluster tool) is part of a substrate processing system (e.g., an integrated processing system). The manufacturing equipment 124 includes one or more of a controller, an enclosure system (e.g., a substrate carrier, a front-opening unified pod (FOUP), a FOUP, a process kit enclosure system, a substrate enclosure system, a cassette, etc.), a side storage pod (SSP), an aligner device (e.g., an aligner chamber), a factory interface (e.g., an equipment front-end module (EFEM)), a load lock, a transfer chamber, one or more processing chambers, and / or a robot arm (e.g., disposed in the transfer chamber, disposed in the front interface, etc.), etc. The enclosure system, the SSP, and the load lock are attached to the factory interface, and the robot arm disposed in the factory interface transfers contents (e.g., substrates, process kit rings, carriers, validation wafers, etc.) between the enclosure system, the SSP, the load lock, and the factory interface. An aligner device is disposed in the factory interface to align the contents. The load locks and processing chambers are mounted in a transfer chamber, and a robotic arm disposed in the transfer chamber transfers the contents (e.g., substrates, process kit rings, carriers, validation wafers, etc.) between the load locks, processing chambers, and transfer chambers. In some embodiments, the manufacturing equipment 124 includes components of a substrate processing system. In some embodiments, the manufacturing equipment 124 is used to fabricate one or more products (e.g., substrates, semiconductors, wafers, etc.). In some embodiments, the manufacturing equipment 124 is used to fabricate one or more components used in a substrate processing system. In some embodiments, the manufacturing equipment 124 is used to fabricate and / or includes bonded metal plate structures (e.g., showerheads used in processing chambers of a substrate processing system).
[0028] The sensors 126 may provide sensor data 142 associated with the manufacturing equipment 124 (e.g., associated with the manufacturing of a corresponding product, such as a wafer, by the manufacturing equipment 124). The sensor data 142 may be used, for example, for equipment health and / or product health (e.g., product quality). The manufacturing equipment 124 may perform a run according to a recipe or over a period of time to manufacture a product. In some embodiments, the sensor data 142 may include one or more values of temperature (e.g., heater temperature), spacing (SP), pressure, high frequency radio frequency (HFRF), electrostatic chuck (ESC) voltage, current, flow (e.g., of one or more gases), power, voltage, etc. The sensor data 142 may include historical sensor data 144 and current sensor data 146. The historical sensor data 144 may relate to a historical process, e.g., a manufacturing or processing run associated with a previously manufactured product (e.g., a substrate, a semiconductor wafer, etc.). The historical sensor data 144 may be utilized as training data for training one or more models, such as model 190. The model 190 may be a machine learning model, a physics-based model, a statistical model, etc. The current sensor data 146 may be associated with non-historical operations, such as a substrate currently undergoing processing, a recently processed substrate, a target substrate of interest, etc.
[0029] The manufacturing equipment 124 may be configured according to manufacturing parameters 150. The manufacturing parameters 150 may be associated with or indicative of parameters, such as hardware parameters (e.g., settings or components (e.g., size, type, etc.) of the manufacturing equipment 124) and / or process parameters of the manufacturing equipment. The manufacturing parameters 150 may include historical manufacturing data (historical parameters 152) and / or current manufacturing data (current parameters 154). The manufacturing parameters 150 may indicate input settings for the manufacturing devices (e.g., heater power, gas flow, etc.). The sensor data 142 and / or manufacturing parameters 150 may be provided while the manufacturing equipment 124 is performing a manufacturing process (e.g., equipment readings may be taken while processing products / substrates). The sensor data 142 may be different for each product (e.g., each wafer). The manufacturing parameters 150 may be the same or substantially the same (e.g., product design, processing recipe, etc.) for a family of products (e.g., except for metadata, etc.). The historical parameters 152 may relate to historical processes, e.g., manufacturing or processing executions associated with previously fabricated products (e.g., substrates, semiconductor wafers, etc.). The historical parameters 152 may be utilized as training data for training one or more models, e.g., model 190. The model 190 may, in embodiments, be a machine learning model, a physics-based model, or a statistical model. The current parameters 154 may relate to non-historical operations, e.g., a substrate currently undergoing processing, a recently processed substrate, a target substrate of interest, etc.
[0030] Metrology data 160 may include measurements of characteristics of products fabricated by manufacturing equipment 124. Historical sensor data 144, historical parameters 152, and metrology data 160 may be associated with fabricated substrates. Metrology data 160 may include data indicating associations between historical data and / or sets of metrology data, e.g., sets of data corresponding to the same fabricated product. Metrology data 160 may include measured and / or predicted metrology (e.g., virtual metrology) associated with any substrate characteristic of interest. Metrology data 160 may include data corresponding to product thickness, resistivity, sheet resistance (e.g., the electrical resistance of a thin film in a direction parallel to the plane of the thin film), critical dimensions (CD, e.g., feature width), linewidth, feature depth, or sidewall height, etc. Metrology data 160 may include multi-point metrology data, e.g., a feature (such as thickness) may be measured at multiple points on the substrate, e.g., at various locations distributed throughout a spatial dimension of the substrate.
[0031] In some embodiments, the sensor data 142, the metrology data 160, and / or the manufacturing parameters 150 may be processed (e.g., by the client device 120 and / or by the prediction server 112). Processing the sensor data 142 may include generating attributes (e.g., data features, vectors, feature vectors, etc.). In some embodiments, the attributes are patterns (e.g., slope, width, height, peaks, etc.) of the sensor data 142, the metrology data 160, and / or the manufacturing parameters 150, or combinations of values (e.g., power derived from voltage and current, etc.) from the sensor data 142, the metrology data 160, and / or the manufacturing parameters 150. The sensor data 142 may include attributes, which may be used by the prediction component 114 to perform signal processing and / or to obtain prediction data 168, possibly for performance of corrective actions. The prediction data 168 may be any data associated with the prediction system 110, for example, predicted metrology data such as a substrate, predicted properties of a substrate, predicted performance of a substrate or manufacturing equipment 124, and the like.
[0032] Each instance (e.g., set) of sensor data 142 may correspond to a product (e.g., a wafer), a set of manufacturing equipment, a type of substrate produced by the manufacturing equipment, combinations thereof, etc. Each instance of metrology data 160 and manufacturing parameters 150 may similarly correspond to a product, a set of manufacturing equipment, a type of substrate produced by the manufacturing equipment, combinations thereof, etc. The data store may further store information associating sets of different data types, e.g., information indicating that a set of sensor data, a set of metrology data, and / or a set of manufacturing data are all associated with the same product, manufacturing equipment, type of substrate, etc.
[0033] In some embodiments, the prediction system 110 may use machine learning to generate the predicted data 168 (e.g., a target output including data indicative of a manufacturing defect provided in the prediction system 110, etc.). In some embodiments, the prediction system 110 may use physics-based modeling to generate the predicted data 168. In some embodiments, the prediction system 110 may use statistical modeling to generate the predicted data 168. Two or more of these techniques may also be combined. The operation of the prediction system 110 is discussed in more detail with respect to FIGS. 2-3 and 4A-4D.
[0034] The client devices 120, manufacturing equipment 124, sensors 126, metrology equipment 128, prediction server 112, data store 140, server machine 170, and server machine 180 may be coupled to one another via network 130 to generate prediction data 168. The prediction data 168 may be used in implementing corrective action. The prediction data 168 may be used in model learning and / or displaying relationships between inputs (e.g., manufacturing parameters) and outputs (e.g., substrate characteristics, process chamber performance, etc.) of one or more models. The prediction data 168 may be used in model learning and / or presenting one or more indications of relationships between inputs and outputs of one or more models.
[0035] In some embodiments, network 130 is a public network that gives client device 120 access to prediction server 112, data store 140, and / or other publicly available computing devices. In some embodiments, network 130 is a private network that gives client device 120 access to manufacturing equipment 124, sensors 126, measurement equipment 128, data store 140, and / or other privately available computing devices. Network 130 may include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), personal area networks, routers, hubs, switches, server computers, cloud computing networks, and / or combinations thereof.
[0036] The client device 120 may include one or more computing devices, such as a personal computer (PC), laptop, mobile phone, smartphone, tablet computer, netbook computer, network-connected television ("smart TV"), network-connected media player (e.g., Blu-ray player), set-top box, over-the-top (OTT) streaming device, operator box, etc. The client device 120 may include a corrective action component 122. The corrective action component 122 may receive user input of indications associated with manufacturing equipment 124 (e.g., via a graphic user interface (GUI) displayed via the client device 120). In some embodiments, the corrective action component 122 sends indications to the prediction system 110, receives output (e.g., prediction data 168) from the prediction system 110, determines corrective actions based on the outputs, and implements the corrective actions. In some embodiments, the presentation component 115 of the client device 120 may perform one or more actions associated with providing information to a user indicative of model learning and / or visualization of input and output dependencies. The presentation component 115 may generate one or more plots, may perform calculations or have calculations performed by another device (e.g., prediction server 112), may receive user instructions via a GUI associated with presenting information to the user, etc.
[0037] In some embodiments, the prediction system 110 may further include a prediction component 114. The prediction component 114 may generate predicted data 168 using data retrieved from the model 190. In some embodiments, the prediction component 114 provides the predicted data 168 to the client device 120, which takes the predicted data 168 into account and takes corrective action via the corrective action component 122 (e.g., including displaying the predicted data 168 for a user). In some embodiments, the corrective action component 122 obtains an indication of the data included in the corrective action or presentation element, retrieves the data (e.g., from the data store 140, by providing one or more inputs from and receiving one or more outputs from the model 190, by providing instructions to the prediction component 114 or the prediction system 110 and receiving the outputs, etc.), and displays the data for a user. In some embodiments, the corrective action component 122 may store the data (e.g., store one or more plots, one or more configuration parameters, etc.), for example, via the data store 140 as analysis data 169. Analysis data 169 may include any data produced as output by any method described herein, for example, a method described in connection with FIG. 3 or FIGS. 4A-4D.
[0038] In some embodiments, the prediction server 112 may store the output of the trained model 190 (e.g., the prediction data 168) in the data store 140, and the client device 120 may retrieve the output from the data store 140. In some embodiments, the corrective action component 122 receives an indication of the corrective action from the prediction system 110 and implements the corrective action (e.g., displays the data to a user). Each client device 120 may include an operating system that enables a user to one or more of create, view, or edit data (e.g., an indication associated with the manufacturing equipment 124, a corrective action associated with the manufacturing equipment 124, etc.).
[0039] In some embodiments, metrology data 160 corresponds to historical characteristic data of a product (e.g., generated using historical sensor data and manufacturing parameters associated with historical manufacturing parameters), and predictive data 168 is associated with predicted characteristic data (e.g., of a product that will be or has been generated under conditions recorded by current sensor data 144 and / or current parameters 154). In some embodiments, predictive data 168 is predicted metrology data (e.g., virtual metrology data) of a product that will be or has been generated according to conditions recorded as current sensor data and / or current manufacturing parameters. In some embodiments, predictive data 168 is or includes an indication of an anomaly (e.g., an abnormal product, an abnormal component, an abnormal manufacturing equipment, an abnormal energy usage, etc.) and / or one or more causes of the anomaly. In some embodiments, predictive data 168 includes an indication of changes or fluctuations over time in some components, such as manufacturing equipment 124, sensors 126, measurement equipment 128, etc. In some embodiments, the forecast data 168 includes an indication of the lifespan of a component, such as manufacturing equipment 124 , a sensor 126 , or a measurement device 128 .
[0040] In some embodiments, one or more outputs of model 190 may be provided to a presentation component 115. The presentation component 115 may generate an indication of model learning, model input / output mapping and / or association, etc. The outputs of the presentation component 115 (e.g., plots, GUI elements, etc.) may be similar to those discussed in connection with FIGS. 5A-5C . In some embodiments, the presentation component 115 may generate an indication of the size of the output space of model 190. For example, simulated data 161 may be provided to model 190 as multiple data sets. The simulated data 161 may span the input space of model 190 or a portion of the input space. For example, the simulated data 161 may include data associated with multiple inputs of model 190. The simulated data 161 may include data sets having values of the inputs of model 190 ranging from a minimum input value (e.g., a minimum value of an input on which the model is trained, a minimum value of an input for which the model produces an output that meets a threshold confidence, etc.) to a maximum input value (e.g., subject to constraints similar to the minimum input value). In some embodiments, simulated data 161 may be generated to span the input space of model 190. For example, simulated data 161 may be generated in a grid spanning the dimensions of the input space (e.g., a first input may take one of several values, a second input may take one of a second list of values, and a third input may take one of a third list of values, all in combination to generate a multidimensional grid that spans the input space of model 190). A set of inputs (e.g., a set that approximately spans the input space of the model) may be provided to the model. The outputs of the model may be stored.
[0041] In some embodiments, the presentation component 115 may generate an indication of the size of the output space of the model 190. The model 190 may be provided with a set of inputs that approximately span the input space of the model 190 (e.g., the space of input values on which the model is trained, the full range of input values for which the model produces outputs that satisfy a threshold confidence condition, etc.). The model 190 may be provided with a randomly sampled set of inputs. For example, a distribution of training inputs may be generated (e.g., a distribution of values provided during training as a first input, a second distribution of values provided during training as a second input, etc.). A random distribution of simulated training inputs may be generated that matches the statistical scatter of the training data for the model. In some embodiments, a random sampling of input data, such as within the input space or within a portion of the input space, may be generated. In some embodiments, the random sampling of input data may be distributed according to a metric other than the distribution of inputs in the training data (e.g., a Gaussian, Lorentzian, or other distribution shape may be generated based on minimum and maximum input values for one or more inputs). Each of the sets of inputs (e.g., together approximately span the input space of the model) may be provided to the model. The model's outputs may be stored.
[0042] In some embodiments, a model (e.g., model 190) may generate a single output from a set of inputs. In some embodiments, a model (e.g., model 190) may generate multiple outputs from a set of inputs. In some embodiments, a model may generate several outputs targeted at a single feature from the set of inputs (e.g., thickness measurements from multiple locations on the substrate). In some embodiments, a model may generate several outputs targeted at multiple features from the set of inputs (e.g., thickness and resistance measurements from multiple locations on the substrate). A model will generally be described herein as a model that receives a set of inputs and generates multiple values for a single feature as output. Models that operate differently, such as models that produce a single output value as output, models that produce output values targeted at multiple features and / or properties of those features, are within the scope of this disclosure. In some embodiments, an ensemble model including multiple models may be utilized; for example, one set of inputs may be provided to a model, and two sets of outputs (e.g., each set of outputs targeted at a different feature) will be generated. No distinction is made herein between an ensemble model that generates an indication of multiple features and two separate models that generate an indication of multiple features; for example, providing a set of inputs to a model or receiving an output from a model may, in some applications, be interpreted as an operation associated with a sub-model of an ensemble model.
[0043] In some embodiments, one or more model outputs may be presented to a user. Each set of input conditions (e.g., input conditions that span approximately the input space, input conditions that span a portion of the input space, etc.) may correspond to one or more outputs. An indication of the output for each set of input conditions may be presented on a plot. The plot may represent approximately the full range of the output space of the model 190.
[0044] In some embodiments, each set of inputs (e.g., inputs that approximately span the input space of one or more models) may correspond to one plotted point on a scatter plot. In some embodiments, the axes of the plot (e.g., the horizontal (“x”) and vertical (“y”) axes of a two-dimensional scatter plot, two horizontal axes and one vertical axis of a three-dimensional scatter plot, etc.) may correspond to the outputs of one or more models. In some embodiments, the axes of the plot may correspond to characteristics of features associated with the outputs of one or more models. The features may include any feature of interest of the substrate. The features may include measurable metrology features of the substrate. The features may include thickness, resistance, sheet resistance, refractive index, extinction coefficient, critical dimension, linewidth, depth, sidewall height, etc. The characteristics may include statistical metrics of the features; for example, several measurements of the features (e.g., simulated measurements at various locations on the substrate, simulated metrology measurements, etc.) may be provided as output from the model 190, and the characteristics of the features may include the mean, median, standard deviation, uniformity, interquartile range, kurtosis, skew, or any other statistical metric that describes the spread of the feature values. In some embodiments, the features or characteristics of the features may include a subset of the values output by the model. For example, the values of a feature near the outer edge of the wafer, the values of a feature within an angular or radial range, or the values of a feature near the center of the wafer may be of interest. The features or characteristics of the features may include a spatial combination of the feature measurements, or a statistical combination of the feature measurements (e.g., the average of a portion of the measurements, such as the lower quartile, upper / lower median half, etc.), etc.
[0045] In some embodiments, a first axis of the output plot may correspond to a first characteristic of the first feature (as a graphical example, average thickness). A second axis of the plot may correspond to a second characteristic of the first feature (still as a graphical example, standard deviation of simulated thickness measurements). Each set of input conditions may generate a point on the scatter plot having a particular value corresponding to the first axis and a particular value corresponding to the second axis (still as a graphical example, each set of input conditions may generate a simulated substrate having an average thickness and standard deviation of thickness measurements, and points may be placed on the plot at appropriate locations for each set of inputs, thus visually mapping the magnitude of the model's input space into the two-dimensional output space represented by the plot). In some embodiments, a three-dimensional scatter plot may be generated that includes a third axis corresponding to a third characteristic of the first feature.
[0046] In some embodiments, a first axis of a plot of the output may correspond to a first characteristic of a first feature (as an illustrative example, average thickness). A second axis of the plot may correspond to a second characteristic of a second feature (still as an illustrative example, standard deviation of simulated resistance measurements). The second characteristic may be the same characteristic as the first characteristic or a different characteristic. In some embodiments, the characteristic associated with the second axis of the plot may be generated by a second model. Each set of input conditions may generate a point on the scatter plot having a particular value corresponding to the first axis and a particular value corresponding to the second axis (still as an illustrative example, each set of input conditions may generate an average thickness value and a resistance standard deviation value, and data points may be plotted on the plot at locations corresponding to the average thickness and resistance standard deviation). In some embodiments, a three-dimensional scatter plot including a third axis may be generated. The third axis may correspond to a third characteristic of the first feature, a third characteristic of the second feature, a third characteristic of the third feature, etc.
[0047] In some embodiments, outputs associated with a set of inputs that span the input space of one or more models (e.g., span approximately the input space, span a portion of the input space, etc.) may be displayed at sporadic points, e.g., on a background "cloud" (e.g., associated with the plotted feature or feature characteristic) that indicates the magnitude of the output space of one or more models in the dimensions of the plot.
[0048] In some embodiments, a set of inputs of interest (e.g., a reference set of inputs, a median set of inputs, etc.) may be provided to one or more models. The one or more models may generate outputs associated with the reference set of inputs. The outputs may be displayed on a plot, e.g., with a scatterplot cloud representing the full range of the output space of the one or more models.
[0049] In some embodiments, a sequential set of inputs of interest may be provided to one or more models. For example, the sequential set of inputs may include a reference set of inputs. The sequential set of inputs may further include a set of inputs in which one input is varied compared to the reference set of inputs. For example, a set of inputs may be provided to one or more models, where all inputs except one input are held at their reference values, and the value associated with the one input is varied among the set of inputs. The value associated with one input may vary from a minimum value to a maximum value. The minimum and maximum values may be associated with the size of the model's input space, the size of the input space measured orthogonally to the reference value, a confidence interval (e.g., the input space in which the output satisfies a threshold confidence metric value condition), a subset of one of these intervals, etc. The value associated with one input may vary systematically or randomly. The value associated with one input may vary in uniform increments, in increments according to a function (e.g., logarithmically spaced), increments according to a distribution (e.g., a Gaussian distribution, a distribution defined by a set of training data associated with one or more models, etc.), etc. The set of outputs may show the learning effect of varying one input value on an output metric associated with the plot (e.g., associated with two or three axes of the plot). The visual depiction of the outputs associated with a series of sets of inputs may include a visual distinction between the output associated with the minimum value of one input and the maximum value of one input. For example, the output point associated with the minimum input value may be colored green, and the output point associated with the maximum input value may be colored red, the output point associated with the minimum input value may be different (e.g., unique in the plot) from the output point associated with the maximum input value, and each output point may be presented with an arrowhead shape pointing toward the next output point (e.g., the next highest value of the input), etc. The outputs associated with a series of sets of input values may generate a curve on the plot that indicates the learning of the model, the effect of varying a variable on the model's output, etc.In some embodiments, the points of the curve may be distinguished from one another; for example, the presented data point associated with the lowest varied input value may be visually distinguished from the presented data point associated with the highest varied input value.
[0050] In some embodiments, multiple sequential sets of inputs may be provided to one or more models. For example, each sequential set of inputs may constrain all input values except those to a reference value, but vary one input value (e.g., one input value per sequential set) as described above. Each sequential set may be plotted in a visually different manner, for example, using a different color, shape, pattern, etc. of the visual representation of the output data. In some embodiments, each sequential set of inputs may be plotted to generate a curve showing model learning associated with the variation of one input associated with that sequential set. Each curve associated with a sequential set of inputs may meet and / or intersect at an output point associated with the reference set of inputs.
[0051] In some embodiments, the presentation component 115 may update the plot based on a user selection of, for example, a setting, a parameter, or a reference input condition. A user may be able to manipulate the input and / or output space via a GUI, for example, associated with the client device 120. The operation of the GUI, the operation of the presentation component 115, example plots, etc. are discussed in more detail in connection with FIGS. 5A-5C.
[0052] Performing a manufacturing process that results in a defective product is costly in terms of time, energy, product, components, manufacturing equipment 124, the cost of identifying the defects and discarding the defective product, etc. By inputting sensor data 142 (e.g., measurements of conditions in a processing chamber) and / or manufacturing parameters 150 (e.g., processing recipe parameters) into model 190, receiving output of predicted data 168, and implementing corrective action based on predicted data 168, system 100 can have the technical advantage of avoiding the costs of creating, identifying, and discarding defective product.
[0053] In some embodiments, the training of the model 190 (e.g., a machine learning model) (e.g., one or more aspects of the model's input / output mapping) may be subjected to additional testing, validation, etc. In some embodiments, one or more aspects of the association between the model's inputs and outputs may be displayed for review, inspection, validation, utilization, to update process parameters, or for designing new processing strategies, etc.
[0054] In some embodiments, a user may choose to implement a model; for example, the user may determine what actions will be taken based on the model's output. In some embodiments, the user may not have been involved in the development of the model (e.g., a customer of the model seller), may not be an expert in modeling techniques (e.g., an engineer or technician rather than an expert in model building or model training, etc.), may not have experience in using model outputs (e.g., may have developed an intuitive understanding, may rely on insider knowledge or past experience, etc.), etc. The user may be presented with model outputs in a conventional format and may be unable to conceptualize the results as an indication of connections, mappings, or learnings (e.g., input / output mappings) learned by the model (e.g., during model training). The user may not want to use a model that the user does not understand, may not trust a model that does not produce results in accordance with the user's understanding, etc. A model with associated learnings made clear to the user in an understandable manner may have the advantage of being trusted by the user, validated by the user, etc. Models trusted by the user may be more fully utilized and may enable more comprehensive improvements to the processing procedures or operation of the processing equipment.
[0055] Implementation of a manufacturing process that results in a failure of a component of manufacturing equipment 124 may result in costly downtime, product damage, equipment damage, ordering of replacement components, etc. By inputting sensor data 142 (e.g., measurements of conditions in a processing chamber) and / or manufacturing parameters 150 (e.g., recipe parameters of a processing recipe) into models 190 (e.g., machine learning models, physics-based models, statistical models, etc.), receiving output of predictive data 168, and performing corrective action (e.g., predictive operational maintenance such as replacing, treating, cleaning, etc.) based on the predictive data 168, system 100 may have the technical advantage of avoiding the costs of one or more of unexpected component failures, unscheduled downtime, lost productivity, unexpected equipment failures, product scrap, etc. Monitoring the performance of components, such as manufacturing equipment 124, sensors 126, metrology devices 128, etc., over time may provide an indication of a deteriorating component. Monitoring the performance of a component (e.g., a substrate support) over time may extend the operational life of the component if, for example, after a standard replacement interval has elapsed, measurements indicate that the component may still function well (performance above a threshold) for some time (e.g., until the next planned maintenance event).
[0056] In some embodiments, there may be a cost associated with implementing an action recommended by a model. For example, the model may suggest that a component is failing and / or being maintained before the component is scheduled to be replaced / maintained (e.g., the model may generate an output indicating that the component is replaced or maintained). Performing the recommended maintenance may be costly in terms of downtime, the cost of replacing the component, reduced active run time (e.g., green time), etc. A user may choose to avoid the cost of implementing an action if the user is not familiar with model training, correlation, mapping, etc. In some embodiments, avoiding the cost of implementing a corrective action may incur more cost; for example, the component may fail, which may result in lost product, costly unscheduled downtime, damage to other components in the system, etc. In some embodiments, the machine learning model may recommend (e.g., generate an output indicating this delay) a delay of maintenance (e.g., the component may be performing the scheduled forecast above). The system may be less costly to operate if the user follows the instructions (e.g., implements a corrective action such as updating the maintenance schedule) and chooses to delay the maintenance (e.g., by extending active processing time). Generating a model that the user trusts may be beneficial in terms of the cost of operating the manufacturing system by allowing the user to take action in response to predicted component lifespans of components in the manufacturing system.
[0057] Manufacturing parameters may be suboptimal for producing a product, which may have costly consequences such as increased consumption of resources (e.g., energy, coolant, gas, etc.), an increased amount of time to produce a product, increased component failures, an increased rate of defective products produced, etc. By inputting sensor data 142 into a trained model 190 (e.g., a machine learning model, a physics-based model, etc.), receiving the output of prediction data 168, and implementing corrective actions (e.g., based on the prediction data 168) to update manufacturing parameters (e.g., set optimal manufacturing parameters), system 100 may have the technical advantage of using optimal manufacturing parameters (e.g., hardware parameters, process parameters, optimal design) that avoid the costly consequences of suboptimal manufacturing parameters. In some embodiments, a user may choose to implement a recommended corrective action. A user's choice of whether to implement a corrective action may be based on the user's confidence in the model, the user's understanding of the model's behavior, the user's understanding of the model's learning, etc. The methods and systems disclosed herein may enable a user to gain a deeper understanding of the model's behavior and learning. The methods and systems disclosed herein may facilitate users to implement recommended corrective actions associated with a manufacturing system, e.g., updating manufacturing parameters, processing recipes, etc. This may increase the efficiency of the manufacturing system, reduce the cost of operation, etc.
[0058] The corrective action may be associated with one or more of computational process control (CPC), statistical process control (SPC) (e.g., SPC for electronic components that determine the process being controlled, SPC that predicts the useful life of a component, SPC that compares to a 3 sigma graph, etc.), advanced process control (APC), model-based process control, preventative operational maintenance, design optimization, manufacturing parameter updates, manufacturing recipe updates, feedback control, machine learning corrections, etc.
[0059] In some embodiments, the corrective action includes providing an alert (e.g., an alarm to stop or not perform a manufacturing process if the predictive data 168 indicates a predicted anomaly, such as an anomaly in a product, component, or manufacturing equipment 124). In some embodiments, the corrective action includes providing feedback control (e.g., modifying a manufacturing parameter in response to the predictive data 168 indicating an anomaly). In some embodiments, the corrective action includes updating a process recipe (e.g., modifying one or more manufacturing parameters based on the predictive data 168). In some embodiments, implementing the corrective action includes updating one or more manufacturing parameters. In some embodiments, implementing the corrective action includes updating one or more calibration tables and / or equipment constants (e.g., a set point provided to a component may be adjusted by a value across multiple process recipes; e.g., the voltage applied to a heater may be increased by 3% for all processes using the heater).
[0060] The manufacturing parameters may include hardware parameters (e.g., component replacement, indication that the manufacturing system is using a particular component, indication of processing updates such as processing chip replacement or firmware update), and / or process parameters (e.g., temperature, pressure, flow, volume, current, voltage, gas flow, lift speed, etc.). In some embodiments, the corrective action includes performing preventative operational maintenance (e.g., replacing, processing, cleaning, etc., components of the manufacturing equipment 124). In some embodiments, the corrective action includes performing design optimization (e.g., updating manufacturing parameters, manufacturing process, manufacturing equipment 124, etc., for an optimized product). In some embodiments, the corrective action includes updating a strategy (e.g., that the manufacturing equipment 124 is in idle mode, sleep mode, warm-up mode, etc.). In some embodiments, the corrective action (e.g., recommended by the model 190, implemented by the client device 120, implemented by the presentation component 115, etc.) may include providing an alert to a user (e.g., preparing data for presentation to a user).
[0061] Prediction server 112, server machine 170, and server machine 180 may each include one or more computing devices, such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc.
[0062] The prediction server 112 may include a prediction component 114. The prediction component 114 may be used to generate predictive data 168. In some embodiments, the prediction component 114 may receive (e.g., from the client device 120, retrieve from the data store 140) current sensor data 146 and / or current parameters 154 and generate output for implementing corrective actions associated with the manufacturing equipment 124 based on the current data.
[0063] Manufacturing equipment 124 may be associated with one or more machine learning models, physics-based models, statistical models, etc., such as, for example, model 190. Machine learning models and other models associated with manufacturing equipment 124 may perform many tasks, including process control, classification, performance prediction, etc. Model 190 may be trained using data associated with manufacturing equipment 124 or products processed by manufacturing equipment 124, such as, for example, sensor data 142 (e.g., collected by sensors 126), manufacturing parameters 150 (e.g., associated with process control of manufacturing equipment 124), metrology data 160 (e.g., generated by metrology equipment 128), etc.
[0064] Another type of machine learning model that can be used to perform some or all of the above tasks is an artificial neural network, such as a deep neural network. An artificial neural network generally includes a feature representation component with a classifier or recurrent layer that maps features to a desired output space. A convolutional neural network (CNN), for example, hosts multiple layers of convolutional filters. In the lower layers, pooling may be performed to address nonlinearities, and a multilayer perceptron is typically added above the lower layers to map the top-layer features extracted by the convolutional layers to a decision (e.g., a classification output).
[0065] A recurrent neural network (RNN) is another type of machine learning model. Recurrent neural network models are designed to interpret a series of inputs that are inherently related to each other, such as time trace data, sequential data, etc. The output of the RNN's recognition result is fed back as input to the recognition result to generate the next output.
[0066] Deep learning is a class of machine learning algorithms that uses a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Deep neural networks can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Deep neural networks include a hierarchy of layers, where different layers learn different representation levels corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into slightly more abstract and complex representations. In an image recognition application, for example, the raw input may be a matrix of pixels; the first representation layer may abstract the pixels and encode edges; the second layer may construct and encode the edge configuration; the third layer may encode higher-level shapes (recognizing substrate structures such as gates, masks, etc.); and the fourth layer may generate a classification output. In particular, the deep learning process is capable of learning which features are optimally placed at which levels. The "deep" in "deep learning" refers to the number of layers through which data is transformed. More precisely, deep learning systems have substantial credit assignment path (CAP) depth. A CAP is a chain of transformations from input to output. The CAP potentially describes the causal connections between input and output. For feedforward neural networks, the CAP depth may be the network depth, which may be the number of hidden layers + 1. In recurrent neural networks, where signals may propagate through layers more than once, the CAP depth is potentially unlimited.
[0067] Training a neural network can be accomplished in a supervised learning fashion, which involves feeding a training dataset consisting of labeled inputs through the network, observing its outputs, defining an error (by measuring the difference between the output and the label value), and using techniques such as deep gradient descent and backpropagation to adjust the network's weights across all its layers and nodes so that the error is minimized. In many applications, repeating this process across many labeled inputs in the training dataset results in a network that can produce correct outputs when presented with inputs that differ from those present in the training dataset.
[0068] In some embodiments, the prediction component 114 may use one or more models 190 to determine outputs for implementing corrective actions based on current data. The models 190 may be a single model, an ensemble model, or a collection of models used to process data. The models 190 may include one or more physics-based digital twin models, supervised machine learning models, unsupervised machine learning models, semi-supervised machine learning models, statistical models, etc.
[0069] In some embodiments, client device 120 may provide current sensor data 146 (e.g., sensor data of interest) to forecasting system 110. In some embodiments, presentation component 115 may provide simulated data 161 to forecasting system 110. Simulated data 161 may include data correlated with sensor data 142, manufacturing parameters 150, etc. Simulated data 161 may be or include user input (e.g., input via a GUI of client device 120). The simulated data may be provided as input, e.g., as input of interest, to model 190 to generate (e.g., via presentation component 115) a graphical representation of model learning, model knowledge, and model mappings / associations between inputs and outputs.
[0070] In some embodiments, data indicative of characteristics of substrates produced using the manufacturing system (e.g., predicted data) may be provided to a model (e.g., model 190), such as a trained machine learning model. The model may be trained to output data indicative of corrective actions to produce substrates with different characteristics. In some embodiments, data indicative of predicted characteristics of substrates produced using manufacturing equipment 124 and metrology data as the substrates are produced by their substrate supports are provided as inputs to a model (e.g., model 190). The model may predict root causes (e.g., manufacturing defects, component aging or variation, etc.) for differences between the predicted data and the measured data.
[0071] The historical sensor data 142, historical parameters 152, and / or metrology data 160 may be used in combination with the current sensor data 146 and current parameters 154 to detect fluctuations, changes, aging, etc., of components of the manufacturing equipment 124. The sensor data 142 monitored over time may include information indicative of changes in one or more components of the manufacturing equipment 124 due to, for example, aging, fluctuations, component failure, material deposition or removal, etc. The prediction component 114 may use a combination and comparison of these data types to generate prediction data 168. In some embodiments, the prediction data 168 includes data that predicts the lifespan of components of the manufacturing equipment 124, sensors 126, etc. The presentation component 115, used in combination with the prediction system 110, data store 140, etc., may provide a representation of the learning of the model 190.
[0072] In some embodiments, the predictive component 114 may receive data, such as sensor data 142, manufacturing parameters 150, and metrology data 160, and perform pre-processing, such as extracting patterns in the data or combining the data into new composite data. The predictive component 114 may then provide the data as input to the model 190. The model 190 may include a physics-based digital twin model that accepts the sensor data 142, manufacturing parameters 150, simulated data 161, etc. as input. The model may include a trained machine learning model, a statistical model, etc., configured to further process data associated with substrate characteristics, performance of the manufacturing equipment 124, etc. The predictive component 114 may receive predictive data from the model 190 indicative of substrate support performance, predicted substrate characteristics, manufacturing defects, component variations, etc. The predictive component 114 may then take corrective action (e.g., recommend a corrective action to a user). The corrective action may include sending an alert to the client device 120. The corrective action may also include updating manufacturing parameters of the manufacturing equipment 124. Corrective action may also include generating predictive data 168 indicative of chamber or instrument fluctuations, aging, or malfunctions.
[0073] In some embodiments, a model may be trained and utilized to generate recommended corrective actions and / or to execute corrective actions (e.g., the model may be operatively coupled to one or more components of manufacturing equipment 124, and the model's output may automatically update future processing parameters, etc.). In both the case of recommended actions and the case of automatic implementation of corrective actions, it may be beneficial to provide a record of the model learning to the user, e.g., to verify model stability, validate model mapping, etc. For example, the user may be presented with an alert (e.g., via the presentation component 115) describing the model learning. In some embodiments, the alert may include a graphical representation of the model output for a range of smoothly varying inputs. The user may predict that the smoothly varying input will produce a smoothly varying output. The user may verify that the model has learned input / output associations that the user trusts, and the user may continue to utilize the model, implement the recommended corrective actions, etc.
[0074] Data store 140 may be memory (e.g., random access memory), a drive (e.g., a hard drive, a flash drive), a database system, or another type of component or device capable of storing data. Data store 140 may include multiple storage components (e.g., multiple drives or multiple databases), which may span multiple computing devices (e.g., multiple server computers). Data store 140 may store sensor data 142, manufacturing parameters 150, analytical data 169, simulated data 161, metrology data 160, and forecast data 168. Sensor data may include time traces of sensor data over the duration of a manufacturing process, associations of data with physical sensors, preprocessed data such as averages and composite data, and data indicative of sensor performance over time (i.e., over many manufacturing processes). Manufacturing parameters 150 and metrology data 160 may include similar characteristics. Forecast data 168 may include data output by prediction system 110. The analysis data 169 may include data (e.g., plots, visualizations, etc.) output by the presentation component 115 (e.g., for further analysis). The simulated data 161 may include characteristics similar to one or more of the sensor data 142 and / or manufacturing parameters 150. The simulated data 161 may be provided to the model 190, for example, to generate an output for use by the presentation component 115. The simulated data 161 may include data ranges (e.g., ranges of values associated with input parameters to the model 190). The historical sensor data 144 and / or historical parameters 152 may be utilized to train the model 190. The metrology data 160 may be utilized to train the model 190, may include predicted metrology data output by the model 190, etc. The metrology data 160 may be metrology data of fabricated substrates, as well as sensor data, manufacturing data, and model data corresponding to those products. The metrology data 160 may be utilized to design processes for making additional substrates.The predictive data 168 may include predictions of metrology data resulting from operation of the substrate support, predictions of component variations, aging, or failure, predictions of component life spans, etc. The predictive data 168 may also include data indicative of components of the system 100 aging and failure over time.
[0075] In some embodiments, prediction system 110 further includes server machine 170 and server machine 180. Server machine 170 includes dataset generator 172, which can generate datasets (e.g., a set of data inputs and a set of target outputs) for training, validating, and / or testing model 190. Some operations of dataset generator 172 are described in detail below with respect to FIGS. 2 and 4A. In some embodiments, dataset generator 172 may divide historical data (e.g., historical sensor data 144, historical measurement data of measurement data 160, historical parameters 152, etc.) into a training set (e.g., 60% of the data), a validation set (e.g., 20% of the data), and a test set (e.g., 20% of the data). In some embodiments, prediction system 110 generates (e.g., via prediction component 114) multiple sets of attributes (e.g., feature vectors, vectors, etc.). For example, the first set of attributes may correspond to a first set of types of sensor data (e.g., from a first set of sensors, a first combination of values from the first set of sensors, a first pattern in values from the first set of sensors) corresponding to each of the datasets (e.g., a training set, a validation set, and a test set), and the second set of attributes may correspond to a second set of types of sensor data (e.g., from a second set of sensors different from the first set of sensors, a second combination of values different from the first combination, a second pattern different from the first pattern) corresponding to each of the datasets.
[0076] In some embodiments, server machine 180 includes a training engine 182, a validation engine 184, a selection engine 185, and / or a test engine 186. The engines (e.g., training engine 182, validation engine 184, selection engine 185, and test engine 186) may refer to hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (e.g., instructions executed on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. Training engine 182 may be capable of training model 190 using one or more sets of attributes associated with a training set from dataset generator 172. Training engine 182 may generate multiple trained machine learning models 190, where each trained model 190 corresponds to a distinct set of attributes of the training set (e.g., sensor data from a distinct set of sensors). For example, a first trained machine learning model may have been trained using all attributes (e.g., X1-X5), a second trained machine learning model may have been trained using a first subset of attributes (e.g., X1, X2, and X4), and a third trained machine learning model may have been trained using a second subset of attributes (e.g., X1, X3, X4, and X5) that may overlap with the first subset of attributes. Dataset generator 172 may receive the output of the trained models (e.g., 190), collect the data into training, validation, and test datasets, and use the datasets to train a second model. Some or all of the operations of server machine 180 may be used to train various types of models, including physics-based models, supervised machine learning models, unsupervised machine learning models, etc.
[0077] The validation engine 184 may be capable of validating the trained machine learning models 190 using a corresponding set of features of the validation set from the dataset generator 172. For example, a first trained model 190 trained using a first set of attributes of the training set may be validated using a first set of attributes of the validation set. The validation engine 184 may determine the accuracy of each of the trained models 190 based on the corresponding set of features of the validation set. The validation engine 184 may discard trained models 190 with an accuracy that does not meet a threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting one or more trained models 190 with an accuracy that meets the threshold accuracy. In some embodiments, the selection engine 185 may be capable of selecting the trained model 190 with the highest accuracy of the trained models 190.
[0078] The testing engine 186 may be capable of testing the trained model 190 using a corresponding set of attributes of a test set from the dataset generator 172. For example, a first trained model 190 trained using a first set of attributes of the training set may be tested using the first set of attributes of the test set. The testing engine 186 may determine the trained model 190 with the highest accuracy from all of the trained models based on the test set.
[0079] Model 190 may refer to a machine learning model, which may be a model artifact created by training engine 182 using a training set that includes data inputs and corresponding target outputs (correct answers for each training input). Model 190 may also or alternatively refer to a statistical model or a physics-based model. Patterns in a data set that map data inputs to target outputs (correct answers) may be found, and model 190 is provided with a mapping that captures these patterns. In some embodiments, model 190 may predict substrate properties. In some embodiments, model 190 may predict failure modes of fabrication chamber components.
[0080] Model 190 may refer to a trained physics-based model. A trained physics-based model may be configured to find solutions to one or more equations that describe physical quantities in the processing chamber, such as mass flow (e.g., gas flow), heat transfer equations, fluid dynamics equations, etc. In some embodiments, the assumptions used to generate the physics-based model may not be entirely accurate (e.g., due to inaccurate measurements, manufacturing or material defects, mismatched component manufacturing tolerances, component aging, variation, or behavior that differs from expectations, etc.). Training the physics-based model may correct for one or more of these assumptions that introduce error into the physics-based model, for example, by changing one or more parameters of the model to allow a better fit to the training data.
[0081] The prediction component 114 may provide input data to the trained model 190 and may run the trained model 190 on the input to obtain one or more outputs. The prediction component 114 may be able to determine (e.g., extract) prediction data 168 from the output of the model 190 and may determine (e.g., extract) confidence data from the output indicating a level of confidence that the prediction data 168 is an accurate predictor of a process associated with the input data for a product that has been or will be made, or an accurate predictor of a component of the manufacturing equipment 124. The prediction component 114 may be able to determine the prediction data 168, including predictions for finished substrate properties and predictions of the useful life of components of the manufacturing equipment 124, sensors 126, or metrology equipment 128, based on the output of the model 190. The prediction component 114 or the corrective action component 122 may use the confidence data to determine whether to take a corrective action associated with the manufacturing equipment 124 based on the prediction data 168. The presentation component 115 may utilize the confidence data, for example, in visually presenting some regions of the uncertain model's output space (e.g., by presenting the data in different colors, shading, shapes, sizes, transparency, etc. to indicate modal confidence).
[0082] The confidence data may include or indicate a level of confidence. As an example, the predicted data 168 may indicate characteristics of a finished wafer given a set of manufacturing inputs (e.g., current parameters 154), including the use of manufacturing equipment 124. The confidence data may indicate that the predicted data 168 is an accurate prediction of a product associated with at least a portion of the input data. In one example, the level of confidence is a real number between 0 and 1, inclusive, where 0 indicates no confidence that the predicted data 168 is an accurate prediction of a product processed according to the input data, and 1 indicates absolute confidence that the predicted data 168 accurately predicts the characteristics of a product processed according to the input data. In response to the confidence data indicating a level of confidence below a threshold level for a predetermined number of instances (e.g., a percentage of instances, a frequency of instances, a total number of instances, etc.), the prediction component 114 may retrain the model 190 (e.g., based on the current sensor data 146, the current manufacturing parameters 150, etc.).
[0083] For purposes of explanation and not limitation, aspects of the present disclosure describe using historical data to train one or more models 190 and inputting current data into the one or more trained models 190 to determine predicted data 168. In other implementations, heuristic or rule-based models are used to determine predicted data (e.g., without using a trained machine learning model). The prediction component 114 may monitor historical data and measurement data 160. Any of the information described with respect to data input 210 in FIG. 2 may be monitored or otherwise used in a heuristic or rule-based model.
[0084] In some embodiments, the functionality of client device 120, prediction server 112, server machine 170, and server machine 180 may be provided by fewer machines. For example, in some embodiments, server machines 170 and 180 may be combined into a single machine, and in other embodiments, server machine 170, server machine 180, and prediction server 112 may be combined into a single machine. In some embodiments, client device 120 and prediction server 112 may be combined into a single machine.
[0085] In general, functionality described in one embodiment as being performed by client device 120, prediction server 112, server machine 170, and server machine 180 may be performed by prediction server 112 in other embodiments, where appropriate. Additionally, functionality attributed to a particular component may be performed by different or multiple components operating together. For example, in some embodiments, prediction server 112 may determine corrective actions based on prediction data 168. In another example, client device 120 may determine prediction data 168 based on output from model 190 (e.g., a trained machine learning model or a physics-based digital twin model).
[0086] Additionally, the functionality of a particular component may be performed by different or multiple components working together. One or more of prediction server 112, server machine 170, or server machine 180 may be accessed as a service offered to other systems or devices through an appropriate application programming interface (API).
[0087] In some embodiments, a "user" may be described as a single individual. However, other embodiments of the present disclosure encompass a "user" that is an entity controlled by multiple users and / or automated sources. For example, a collection of individual users working together as a group of administrators may be considered a "user."
[0088] Embodiments of the present disclosure may be applied to data quality assessment, feature enhancement, model evaluation, virtual metrology (VM), predictive maintenance (PdM), constraint optimization, etc. Embodiments of the present disclosure may be applied to any trained modeling system and may provide model evaluation, model validation, representation of model learning, etc. for any machine learning model, any machine learning model associated with manufacturing and / or processing of a product, any machine learning model that predicts metrology of processed wafers, etc.
[0089] FIG. 2 is a block diagram of an example dataset generator 272 (e.g., dataset generator 172 of FIG. 1 ) used to create a dataset for a model (e.g., model 190 of FIG. 1 ) according to some embodiments. The dataset generator 272 may be part of the server machine 170 of FIG. 1 . In some embodiments, the system 100 of FIG. 1 includes multiple machine learning models. In such cases, each model may have a separate dataset generator, or the models may share a dataset generator. For example, a separate model may be used for each feature of interest of a wafer associated with a set of input data. FIG. 2 illustrates a dataset generator associated with a machine learning model configured to receive as input sensor data (e.g., current sensor data 146 of FIG. 1 ) and provide output predicted metrology (e.g., predicted data 168 of FIG. 1 ) for the substrate. A model (e.g., a machine learning model, a physics-based model, a statistical model, etc.) may be configured to perform one or more of many different tasks. For example, the model may receive sensor data and generate as output a feedback control signal to adjust process conditions. The model may receive processing parameters and predict substrate performance (e.g., predict metrology of the substrate resulting from a processing recipe). The model may receive an image (e.g., a block diagram of a target substrate design) and generate an associated image (e.g., an actual image of the substrate to be simulated). The model may receive an indication of a target product design and generate as an output a predicted process recipe to create a product of that design. The model may receive measurements of one or more components of a manufacturing system and generate as an output a predicted performance of the system. Any of these or many other specific use cases of models (e.g., any machine learning model that maps a set of inputs to one or more outputs, a model associated with manufacturing, a model associated with substrate processing, a model associated with semiconductor wafers, etc.) can benefit from the methods and systems described herein, for example, to display one or more indications of model learning.
[0090] Referring to FIG. 2 , a system 200 including a dataset generator 272 (e.g., dataset generator 172 of FIG. 1 ) creates a dataset for a machine learning model (e.g., model 190 of FIG. 1 ). The dataset generator 272 may create the dataset using sensor data, e.g., historical sensor data 144 of FIG. 1 . The dataset generator 272 may create the dataset using manufacturing parameters, e.g., historical parameters 152. The dataset generator 272 may create the dataset using measurement data, e.g., measurement data 160. In some embodiments, the dataset generator 272 may create a dataset for the model using data output by another model, function, processing tool, etc. (e.g., predicted data 168, analyzed data 169, simulated data 161, etc.). Models that receive different types of data as input or generate different types of data as output may receive datasets created from the corresponding data types from the dataset generator 272. In some embodiments, the dataset generator 272 creates training inputs (e.g., data inputs 210) from sensor data associated with one or more processing steps (e.g., associated with one or more fabricated substrates) and / or process parameter data associated with one or more process recipes. The dataset generator 272 also generates target outputs 220 for training the machine learning model. The target outputs may include metrology data, such as one or more substrates processed under conditions indicated by the sensor data, one or more substrates processed according to one or more process recipes, etc. In some embodiments, the metrology data 230 may include multiple measurements of a simulated wafer, e.g., each measurement corresponding to a predicted value of a feature of interest at a different spatial location of the simulated wafer. The training input data 210 and the target output data 220 are provided to the machine learning model.For purposes of describing the operation of the dataset generator for training a model, the dataset generator 272 is described as training a machine learning model that accepts input data indicative of processing conditions for a substrate and generates predicted metrology data for the substrate as output, although any other model (e.g., machine learning model) configuration may benefit from aspects of the present disclosure.
[0091] Within the scope of the present disclosure, the training inputs 210 and target outputs 220 may be represented in a variety of different ways. A two-dimensional map of substrate properties, a function that replicates that map, or other data indicative of substrate performance data may be used as the target output 220. The data set may include processed data, smoothed data, cleaned data (e.g., outliers removed, etc.), combined data, data that has been reduced to data attributes (e.g., vectors, feature vectors, etc.), etc.
[0092] In some embodiments, the dataset generator 272 generates a dataset (e.g., a training set, a validation set, a test set) that includes one or more data inputs 210 (e.g., training inputs, validation inputs, test inputs) and may include one or more target outputs 220 corresponding to the data inputs 210. The dataset may also include mapping data that maps the data inputs 210 to the target outputs 220. The data inputs 210 may also be referred to as “features,” “attributes,” or “information.” In some embodiments, the dataset generator 272 may provide a dataset to the training engine 182, the validation engine 184, or the test engine 186 of FIG. 1, where the dataset is used to train, validate, or test the model 190 of FIG. 1. Some embodiments of generating a training set may be further described with respect to FIG. 4A.
[0093] In some embodiments, the data set generator 272 may generate a first data input corresponding to a first set of sensor data 244A and / or a first set of manufacturing parameter data 252A for training, validating, or testing a first model, and the data set generator 272 may generate a second data input corresponding to a second set of sensor data 244B and a second set of manufacturing parameter data 252B for training, validating, or testing a second model.
[0094] In some embodiments, the dataset generator 272 may perform operations on one or more of the data inputs 210 and the target outputs 220. The dataset generator 272 may extract patterns from the data (such as slope, curvature), combine the data (such as averages, feature creation), or separate the data into groups (e.g., train a model on a subset of the predicted performance data) and use the groups to train separate models.
[0095] The data inputs 210 and target outputs 220 for training, validating, or testing a model may include information about a particular substrate processing recipe (e.g., a particular substrate design). The data inputs 210 and target outputs 220 may include information about a particular substrate processing system (e.g., about a particular set of manufacturing equipment). The data inputs 210 and target outputs 220 may include information about a particular process type, target substrate design, target substrate characteristics, or may be grouped together in another manner.
[0096] In some embodiments, the data set generator 272 may generate a set of target outputs 220 that includes the metrology data 230. The target outputs 220 may be divided into sets that correspond to sets of input data. Different sets of target outputs 220 may be used in conjunction with similarly defined sets of data inputs 210, including using different sets for training different models, training, validation, and testing, etc.
[0097] In some embodiments, a model may be trained without a target output 220 (e.g., an unsupervised or semi-supervised model). A model that is not provided with a target output may, for example, be trained to recognize significant differences (e.g., outside an error threshold) between predicted and measured performance data.
[0098] In some embodiments, the information used to train the model may be from a particular type of manufacturing equipment at a manufacturing facility having particular characteristics (e.g., manufacturing equipment 124 in FIG. 1 ), allowing the trained machine learning model to determine an outcome for a particular group of manufacturing equipment 124 based on inputs of predicted performance data and measured performance data associated with one or more components that share the characteristics of the particular group. In some embodiments, the information used to train the model may be about components from more than one manufacturing facility, allowing the model to determine an outcome for a component based on inputs from one manufacturing facility.
[0099] In some embodiments, after generating a dataset and using the dataset to train, validate, or test a model, the model may be further trained, validated, or tested, or may be adjusted. For example, additional data may be provided to the model as retraining data, revalidation data, retest data, etc., from substrates processed after the model has been trained, validated, and tested.
[0100] The dataset generator 272 may generate datasets for training, validating, and / or testing a model. Training a model may include generating a model mapping 273 that the model utilizes to connect input data to output data (e.g., to generate output from a set of provided inputs). In some embodiments, the model mapping 273 may include model weights and biases, e.g., weights and biases connecting nodes in a layer of a machine learning model.
[0101] In some embodiments, a dataset generator performing functionality similar to dataset generator 272 may be utilized to train a physics-based model. The physics-based model may be configured to generate outputs based on a physical understanding of the system, based on physical assumptions about the operation of the system, based on one or more numerical solutions of one or more physical equations (e.g., heat transfer equations, mass balance equations, fluid dynamics equations, etc.), etc. The physics-based model may be trained in a manner similar to a machine learning model. The physics-based model may be provided with training inputs and target outputs, and one or more parameters, weights, biases, etc. may be adjusted to result in better alignment between the model output and the target output.
[0102] In some embodiments, the physics-based model may receive a set of inputs (e.g., indicative of processing conditions for a substrate). The physics-based model may generate an output (e.g., predicted metrology data for the substrate) based on the set of inputs. The physics-based model may be provided with a target output (e.g., measured metrology data for the substrate). The physics-based model may adjust one or more parameters of the model to generate an output (e.g., predicted metrology data) that is more similar to the target output than it was before the adjustments were made.
[0103] In some embodiments, dataset generator 272 facilitates the generation of models for improving manufacturing systems. In some embodiments, a record of the learning of a trained model is generated. For example, a visual representation of tasks, associations, mappings, etc. that the model has learned (e.g., learned by providing the model with a dataset created by dataset generator 272) may be generated. Operation of dataset generator 272 may facilitate the generation of model mappings 273 (e.g., weights and biases for machine learning models, values for adjustable parameters for physics-based models, etc.). Visualization of the model learning may include visualization of model mappings 273, e.g., changes to the model's output when inputs are varied.
[0104] 3 is a block diagram illustrating a system 300 for generating output data for model training, association, and mapping analysis (e.g., generating analysis data 169 of FIG. 1 ), according to some embodiments. System 300 may be used to train a model (e.g., a machine learning model) and to generate data that can be utilized to describe the associations learned by the model. Some or all of the operations of system 300 may be used to generate output data for a machine learning model, e.g., predicted data 168 of FIG. 1 . Some or all of the operations of system 300 may be used to generate output data for a physics-based model, e.g., predicted data 168 of FIG. 1 .
[0105] 3, in block 310, system 300 (e.g., a component of forecasting system 110 of FIG. 1) performs data partitioning (e.g., via data set generator 172 in server machine 170 of FIG. 1) of historical data 364 (e.g., historical sensor data, historical manufacturing parameter data, historical metrology data, etc.) to generate training set 302, validation set 304, and test set 306. For example, the training set may be 60% of the historical data, the validation set may be 20% of the historical data, and the test set may be 20% of the historical data.
[0106] At block 312, the system 300 performs model training (e.g., via training engine 182 of FIG. 1 ) using the training set 302. The system 300 may train one model or multiple models using multiple sets of attributes (e.g., feature vectors) of the training set 302 (e.g., a first set of attributes including a subset of the historical data of the training set 302, a second set of attributes including a different subset of the historical data of the training set 302, etc.). For example, the system 300 may train machine learning models to generate a first trained machine learning model using the first set of attributes in the training set and to generate a second trained machine learning model using the second set of attributes in the training set (e.g., data different from the data used to train the first machine learning model). In some embodiments, the first trained machine learning model and the second trained machine learning model may be combined to generate a third trained machine learning model (e.g., which may be a better predictor than the first or second trained machine learning models themselves). In some embodiments, the sets of attributes used to compare the models may overlap (e.g., one model may be trained with performance data indicative of film thickness, another model may be trained with performance data indicative of both film thickness and film stress, different models may be trained with data from different locations on the substrate, models may be trained with inputs from different sets of sensors or manufacturing parameters, etc.). In some embodiments, hundreds of models may be generated, including models with various permutations of model attributes and combinations.
[0107] In block 314, the system 300 performs model validation (e.g., via the validation engine 184 of FIG. 1 ) using the validation set 304. The system 300 may validate each of the trained models using a corresponding set of attributes in the validation set 304. For example, the validation set 304 may use a subset of historical data types that are the same as those used in the training set 302, but for different input conditions (e.g., associated with performance features on the same wafer, the same sensors, the same input parameters, etc.). In some embodiments, the system 300 may validate hundreds of models (e.g., models with various permutations of attributes, combinations of models, etc.) generated in block 312. In block 314, the system 300 may determine the accuracy of each of the one or more trained models (e.g., by model validation) and may determine whether one or more of the trained models have an accuracy that meets a threshold accuracy. In response to determining that none of the trained models have an accuracy that meets the threshold accuracy, flow returns to block 312, and the system 300 performs model training using a different set of attributes from the training set. In response to determining that one or more of the trained models have an accuracy that meets the threshold accuracy, flow continues to block 316. The system 300 may discard trained machine learning models that have an accuracy below the threshold accuracy (e.g., based on a validation set).
[0108] In block 316, the system 300 may perform model selection (e.g., via the selection engine 185 of FIG. 1 ) to determine which of the one or more trained models that meet the threshold accuracy has the highest accuracy (e.g., selected model 308 based on validation of block 314). If only a single model is trained, the operations of block 316 may be skipped. In response to determining that two or more of the trained models that meet the threshold accuracy have the same accuracy, flow may return to block 312, where the system 300 performs model training using a further refined training set (e.g., corresponding to a further refined set of attributes) to determine the trained model with the highest accuracy.
[0109] At block 318, the system 300 performs model testing using the test set 306 (e.g., via the test engine 186 of FIG. 1 ) to test the selected model 308. The system 300 may test the first trained machine learning model using the first set of attributes of the test set and determine that the first trained machine learning model meets a threshold accuracy (e.g., based on the first set of attributes of the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 is overfitted to the training set 302 and / or the validation set 304 and is not applicable to other datasets, such as the test set 306), the flow proceeds to block 312, where the system 300 performs model training (e.g., retraining) using a different training set, possibly corresponding to a different set of attributes, or performs a reorganization of the substrate (e.g., dataset) divided into training, validation, and test sets. In response to a determination that the selected model 308 has an accuracy that meets the threshold accuracy based on the test set 306, flow continues to block 320. At least in block 312, the model may learn patterns of simulated sensor data to make predictions, and in block 318, the system 300 may apply the model to the remaining data (e.g., the test set 306) to test the predictions.
[0110] In block 320, the system 300 uses the trained model (e.g., the selected model 308) to receive simulated data 354 (e.g., simulated data 161 of FIG. 1 , simulated data of interest, a user-selected set of criteria inputs, inputs spanning one or more dimensions of the input space or a portion thereof for mapping a portion spanning the input space to a portion of the model's output space, etc.), determine (e.g., extract) analytical data 369 (e.g., analytical data 169 of FIG. 1 ) from the output of the trained model, and perform an action (e.g., perform a corrective action associated with the manufacturing equipment 124 of FIG. 1 , provide an alert to the client device 120 of FIG. 1 , provide an alert to a user via the presentation element 115 of FIG. 1 , etc.).
[0111] In some embodiments, retraining of the machine learning model is performed by supplying additional data to further train the model. Current data 346 may be provided in block 312. The current data 346 may differ from the data originally used to train the model by incorporating combinations of input parameters that were not part of the original training and input parameters outside the parameter space spanned by the original training, or may be updated to reflect chamber-specific knowledge (e.g., differences from an ideal chamber due to manufacturing tolerance ranges, part aging, part variation, performed maintenance, etc.). The selected model 308 may be retrained based on this data.
[0112] In some embodiments, one or more of acts 310-320 may be performed in various orders and / or may involve other acts not presented or described herein. In some embodiments, one or more of acts 310-320 may not be performed. For example, in some embodiments, one or more of data partitioning of block 310, model validation of block 314, model selection of block 316, or model testing of block 318 may not be performed. A subset of these actions may be performed when training a physics-based digital twin model, for example, to take measurements of processing conditions as inputs and produce predicted performance data for a substrate as output.
[0113] 4A-4D are flow diagrams of methods 400A-D associated with describing and / or visualizing model learning, according to certain embodiments. Methods 400A-D may be implemented by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (e.g., instructions executed on a processing device, general-purpose computer system, or dedicated machine), firmware, microcode, or a combination thereof. In some embodiments, methods 400A-D may be performed in part by prediction system 110. Method 400A may be performed in part by prediction system 110 (e.g., server machine 170 and dataset generator 172 in FIG. 1 , dataset generator 272 in FIG. 2 ). Prediction system 110 may use method 400A to generate a dataset for at least one of training, validating, or testing a model, according to embodiments of the present disclosure. The model may be a physics-based (e.g., digital twin) model (e.g., for generating predicted performance data for a substrate), a machine learning model (e.g., for generating data indicative of corrective actions associated with a component of manufacturing equipment, etc., for generating predicted performance data for a wafer), a statistical model, or another model trained to receive inputs and generate outputs related to the manufacturing or processing of a substrate. Methods 400B-D may be performed by prediction server 112 (e.g., prediction component 114, etc.). Methods 400B-D may also be performed by other components of prediction system 110. Operations described as relating to methods 400B-D may be performed by server machine 180 (e.g., training engine 182). In some embodiments, a non-transitory storage medium stores instructions that, when executed by a processing device (e.g., of prediction system 110, of server machine 180, of prediction server 112, etc.), cause the processing device to perform one or more of methods 400A-D.
[0114] For ease of explanation, methods 400A-D are shown and described as a series of operations. However, operations in accordance with the present disclosure may occur in various orders and / or concurrently, and with other operations not presented or described herein. Moreover, not all illustrated operations may be performed to implement methods 400A-D in accordance with the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that methods 400A-D may alternatively be represented as a series of interrelated states, such as by a state diagram or events.
[0115] FIG. 4A is a flow diagram of a method 400A for generating a dataset for a model (e.g., a machine learning model) that generates output data (e.g., predicted data 168 of FIG. 1) according to some embodiments.
[0116] Referring to FIG. 4A, in some embodiments, at block 401, processing logic implementing method 400A initializes a training set T to an empty set.
[0117] At block 402, processing logic generates a first data input (e.g., a first training input, a first validation input) that may include sensor data, manufacturing parameter data, measured substrate performance data, substrate metrology data (e.g., film properties such as thickness, material composition, optical properties, roughness, etc.), etc. In some embodiments, the first data input may include a first set of attributes related to the type of data, and the second data input may include a second set of attributes related to the type of data (e.g., as described with respect to FIG. 3).
[0118] At block 403, processing logic generates a first target output for one or more of the data inputs (e.g., the first data input). In some embodiments, the first target output is performance data for the substrate. In some embodiments, the first target output is data indicative of a corrective action. In some embodiments, the target output is not generated (e.g., to train an unsupervised machine learning model).
[0119] At block 404, processing logic optionally generates mapping data indicating an input / output mapping. The input / output mapping (or mapping data) may reference data inputs (e.g., one or more of the data inputs described herein), target outputs for the data inputs, and associations between the data inputs and target outputs. In some embodiments (e.g., embodiments without target output data), these operations may not be performed.
[0120] At block 405, processing logic, in some embodiments, adds the mapping data generated at block 404 to dataset T.
[0121] At block 406, processing logic branches based on whether dataset T is sufficient for at least one of training, validation, and / or testing of model 190 of FIG. 1. If so, execution proceeds to block 407; if not, execution returns to block 402. It should be noted that in some embodiments, the sufficiency of dataset T may be determined solely based on the number of inputs in the dataset, which in some embodiments are mapped to outputs; in other implementations, the sufficiency of dataset T may be determined based on one or more other criteria in addition to or instead of the number of inputs (e.g., a measure of diversity of data examples, accuracy, the total coverage of the input and / or output space, etc.).
[0122] At block 407, processing logic provides dataset T (e.g., to server machine 180 of FIG. 1 ) for training, validating, and / or testing model 190. In some embodiments, dataset T is a training set and is provided to training engine 182 of server machine 180 to perform training. In some embodiments, dataset T is a validation set and is provided to validation engine 184 of server machine 180 to perform validation. In some embodiments, dataset T is a test set and is provided to test engine 186 of server machine 180 to perform testing.
[0123] The operations of block 407 may generate a trained model, e.g., a model mapping between inputs and outputs (e.g., weight and bias values between nodes in layers of a machine learning model, values of adjustable parameters in a physics-based model, etc.). The model mapping may be described (e.g., an indication of the model's learning may be visualized) by a method of the present disclosure, e.g., any of methods 400A-D.
[0124] FIG. 4B is a flow diagram of a method 400B for generating a visual representation of the model's learning (e.g., a mapping from input to output), according to some embodiments. In some embodiments, the model comprises a machine learning model. In some embodiments, the model comprises a physics-based model. In some embodiments, the model comprises a statistical model. In some embodiments, the model is configured to receive as input one or more indications of substrate processing conditions (e.g., sensor data, manufacturing parameters, etc.) for processing the substrate. In some embodiments, the model is configured to generate as output predicted characteristics of the substrate processed according to the conditions indicated by the input. In some embodiments, the model is configured to generate an indication of values of one or more predicted characteristics of the substrate (e.g., a thickness of the substrate at one or more locations, a resistance of the substrate at one or more locations, a sheet resistance of the substrate, an optical property (e.g., an attenuation coefficient, a refractive index), an indication of a substrate geometry (e.g., a critical dimension, a sidewall height, a linewidth, a depth, etc.), etc.). In some embodiments, one or more characteristics (e.g., a statistical metric) of the predicted characteristics may be calculated. The characteristics may include mean, median, standard deviation, uniformity, skew, kurtosis, interquartile range, or any other statistical metric applicable to predicted feature values. In some embodiments, methods similar to method 400B may be utilized to generate a visual display of the training of a differently configured model (e.g., configured to receive different data as input and / or produce different data as output than the model described in method 400B).
[0125] In some embodiments, the model may be trained (e.g., before the start of method 400B). Training the model may include providing training inputs to the model and target outputs to the model. Training the model may include processing logic receiving a plurality of sets of input values (e.g., values indicative of process conditions for processing a plurality of substrates). The processing logic may receive target output data (e.g., metrology measurements of substrates processed at process conditions associated with the set of input values) for training the model. Training the model may include processing logic providing the plurality of sets of input values as training inputs and target output data as target outputs to the model.
[0126] At block 410, processing logic receives a first value. The first value is associated with a first input parameter of the model (e.g., the model may accept the first value as a first input parameter). The first input parameter is associated with a first processing condition of the substrate processing procedure (e.g., represents a value corresponding to this first processing condition). The first processing condition (and subsequent processing conditions of methods 400B-D) may be any processing condition of the substrate processing system that may affect the product, such as temperature, gas flow, gas identity and / or mixture (e.g., 20% reactive gas in non-reactive carrier, 15% reactive gas in non-reactive carrier, etc.), pressure, chuck power, RF power, processing time, etc.
[0127] At block 412, processing logic receives a first plurality of values. The first plurality of values ranges from a lowest value of the first plurality of values to a highest value of the first plurality of values (e.g., the values span a range of values). In some embodiments, the values span a range determined by a user. In some embodiments, the values span a range determined by a trained model (e.g., defined with respect to the model's input space). In some embodiments, the values span a range that is a portion of a range associated with the model's input space (e.g., a user may select 10% of the full range, 20% of the full range, 30% of the full range, or any subset of the full range). Each of the first plurality of values is associated with a second input parameter of the model. The second input parameter of the model is associated with a second processing condition of the substrate processing procedure.
[0128] At block 414, processing logic provides the first value and the first plurality of values to the model. The processing logic may provide the first value and the first plurality of values as inputs to the model. The model may be configured to generate outputs based on the inputs. The inputs may include additional values (e.g., values associated with third, fourth, etc. input parameters). The model may separately (e.g., sequentially) generate outputs associated with each set of inputs (e.g., an input set including the first value and a first value of the first plurality of values, an input set including the first value and a second value of the first plurality of values, etc.) to generate outputs corresponding to each set of input values.
[0129] In some embodiments, the processing logic may provide additional sets of inputs to the model. For example, the processing logic may receive a second value associated with a second input parameter of the model. The processing logic may further receive a second plurality of values, each of which is associated with a first input parameter of the model. The processing logic may provide additional sets of inputs to the model, e.g., a first set including the second value and a first value of the second plurality of values, a second set including the second value and a second value of the second plurality of values, etc. Similar operations may be performed for additional multiple values associated with additional input parameters (e.g., associated with additional processing conditions). In some embodiments, a series of sets of values may be provided to the model, each of which includes all values except a central value, a reference value, a value of interest, etc., with the other values allowed to vary throughout a range. The series of sets of values may include a series of sets, each of which includes a different varied input parameter. In this manner, a set of outputs, each related to a reference set of input values but varying one parameter, may be obtained from the model. In some embodiments, multiple values may be varied within a set, e.g., instead of a single reference set of values, multiple reference sets of values may be used, multiple values associated with input parameters may be varied within the same set (e.g., to show how varying two or more input parameters affects the output, or to demonstrate model learning in response to variations in two or more input parameters, etc.), etc.
[0130] In some embodiments, the processing logic may receive a second value associated with a first input parameter of the model. For example, the GUI may include an option for a user to change a reference value of the first input parameter. The system may perform similar operations (e.g., providing the second value and the first plurality of values to the model, providing the second value and the second plurality of values to the model, etc.) in response to receiving a user input, in response to receiving an instruction to perform an operation based on the second value, etc. In some embodiments, the user may change one or more reference inputs, one or more ranges associated with the plurality of values, one or more features to be investigated, one or more characteristics to be displayed, etc., and the system may perform the operations described herein in response to the user's selection to reveal model training, association, mapping, etc.
[0131] At block 416, processing logic receives a first plurality of outputs from the model. Each of the first plurality of outputs is associated with a first value (associated with a first input parameter of the model) and one of the first plurality of values (associated with a second input parameter of the model). Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates (e.g., a characteristic of the substrate that the model is configured to predict).
[0132] In some embodiments, further results may be received from the model. For example, the model may provide a second plurality of outputs, a third plurality of outputs, etc. The second plurality of outputs may be associated with a second value (e.g., a center or reference value associated with the second input parameter) and a second plurality of values (e.g., a plurality of values associated with the first input parameter). In this manner, the model may provide two or more sets of output data, one of which is associated with varying the first input parameter while holding the remaining parameters at a set of reference values, another of which is associated with varying the second input parameter while holding the remaining parameters at a set of reference values, etc. In the example of two sets of output data, visually, this may be represented as two arcs, curves, etc., each arc including a plurality of data points (e.g., each of the plurality of data points is associated with one of the plurality of inputs). The curves may intersect at a point (represented in output space) associated with a criteria set of input conditions (e.g., the output of the model received when the criteria set of inputs is provided as inputs to the model).
[0133] In some embodiments, more input parameters may be varied, which may result in more curves. In some embodiments, multiple inputs may be varied to generate a curve (e.g., a curve may be generated in which data points on the curve are associated with the output of a model generated by simultaneously adjusting two or more input parameters). Varying multiple inputs may allow one output feature, one characteristic of that feature, etc. to be changed without affecting the others. For example, changing two input parameters may have a similar effect on one characteristic, one characteristic, etc. (e.g., adjusting both input parameters to higher values may change the output characteristic to a higher value), but an opposite effect on another characteristic, another characteristic, etc. (e.g., adjusting a first input parameter to a higher value may cause the value of the output characteristic to change in the opposite direction as adjusting a second input parameter to a higher value). A set of input parameters may be provided in which all inputs except two are held at reference values and the two are varied, e.g., simultaneously, to maintain one output characteristic or feature while adjusting the other.
[0134] In some embodiments, further results received from the model may include a third plurality of outputs, each of which is associated with a second value associated with the first input parameter (e.g., a different set of criteria conditions) and a plurality of values associated with the second input parameter. In some embodiments, the plurality of values associated with the second input parameter may be the same as the values used in association with the first value associated with the first input parameter. In some embodiments, the plurality of values may be different from the plurality of values used in association with the first value of the first input parameter (e.g., one or more values may be different, the total number of values may be different, etc.). In this manner, a user may specify a different set of criteria input conditions. In some embodiments, the plurality of criteria conditions (e.g., criteria values associated with the second input parameter, criteria values associated with the third input parameter, etc.) may be changed.
[0135] At block 418, the first plurality of outputs are prepared for presentation by the processing logic. The first plurality of outputs (e.g., indicators of the model's outputs) are to be presented via a presentation element of a graphic user interface (GUI). In one embodiment, the presentation element includes two axes (e.g., two orthogonal axes). In some embodiments, the presentation element may include three axes (e.g., three orthogonal axes). In some embodiments, the presentation element may include a scatter plot, e.g., each simulated substrate may be associated with a value corresponding to a first axis (e.g., a first characteristic of a first feature) and a value corresponding to a second axis (e.g., a second characteristic of a second feature), and a data point may be displayed at a location indicating both of these values. The first of the two axes may correspond to a first characteristic of the first feature (e.g., a first statistical metric of the feature the model is configured to predict). The second of the two axes may correspond to a second characteristic of the first feature. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element. The graph shows a value of a first characteristic of a first feature (according to the position of the data point relative to a first axis) and a value of a second characteristic of the first feature (according to the position of the data point relative to a second axis) associated with each of a first plurality of outputs.
[0136] In some embodiments, data points associated with each of the first plurality of outputs may generate a curve in the presentation element. In some embodiments, further outputs may also be plotted, e.g., additional curves may be generated. The further outputs may be represented as additional curves, e.g., intersecting at data points associated with a reference set of inputs. In some embodiments, data associated with second values of one or more input parameters (e.g., a second set of reference conditions) may be presented. In some embodiments, in response to a change in the reference input parameters, processing logic may generate a new graphic including new data points (e.g., one or more curves of data points) associated with the new reference input conditions. In some embodiments, this graphic may include data points (e.g., one or more curves of data points) associated with two or more sets of reference conditions. In some embodiments, the graphic of the presentation element may include other features, such as an indication of the magnitude of the input space and / or output space of the model. In some embodiments, the graphic of the presentation element may include one or more visual indicators, for example, to distinguish between outputs associated with each of the multiple inputs. For example, the multiple input values may range from a lowest value to a highest value. Outputs associated with the lowest input values may be distinguished from outputs associated with the highest input values, for example, by color, shape, pattern, labeling, etc. Presentation elements and GUI features are discussed in further detail in conjunction with Figures 5A-5C.
[0137] 4C is a flow diagram of a method 400C for generating a visual representation of the training of multiple models, according to some embodiments. The models associated with method 400C may share one or more characteristics with the models described in connection with FIG. 4B. The operations associated with method 400C may share one or more characteristics with the operations of method 400B.
[0138] At block 420, processing logic receives a first value associated with a first input parameter of the first model. The first value is also associated with a first input parameter of the second model. The first input parameter is associated with a first processing condition of the substrate processing procedure (e.g., the first input parameter may be a measured and / or predicted value of a condition in a processing chamber, such as temperature, pressure, etc.). The first model may be configured to receive one or more inputs and generate as outputs one or more predictions of substrate performance (e.g., substrate metrology, substrate material properties, etc.). The first model may be configured to provide one or more indications associated with a characteristic of the simulated / predicted substrate (the characteristic may be thickness, another physical dimension, resistance, or any other metric of interest). In some embodiments, processing logic may further receive a second value also associated with the first input parameter, a third value associated with the second input parameter, etc.
[0139] At block 422, processing logic receives a first plurality of values. The first plurality of values ranges from a lowest value of the first plurality of values to a highest value of the first plurality of values. Each of the first plurality of values is associated with a second input parameter of the first model. Each of the first plurality of values is also associated with a second input parameter of the second model. The second input parameter is associated with a second processing condition of the substrate processing procedure. In some embodiments, processing logic may receive additional plurality of values. Processing logic may receive a second plurality of values associated with the first input parameter of the first model and the second model. Processing logic may receive a third plurality of values associated with a third input parameter of the first model and the second model. Processing logic may receive a fourth plurality of values associated with the first input parameter but which may include different values (e.g., a different range, a different interval, etc.) from the first plurality of values.
[0140] At block 424, processing logic provides the first value and the first plurality of values to the first model. Processing logic also provides the first value and the first plurality of values to the second model. In some embodiments, different and / or additional values are provided, e.g., a second value, a second plurality of values, a third plurality of values, etc.
[0141] At block 426, processing logic receives a first plurality of outputs from the first model. Each of the first plurality of outputs is associated with a first value and one of the first plurality of values. Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates (e.g., a characteristic of the simulated substrate that the first model is configured to predict). In some embodiments, different and / or additional outputs may be received, e.g., a second value, an output associated with the second plurality of values, etc.
[0142] At block 428, processing logic receives a second plurality of outputs from the second model. Each of the second plurality of outputs is associated with the first value and one of the first plurality of values. Each of the second plurality of outputs is associated with a second characteristic of one of the first plurality of simulated substrates (e.g., a characteristic of the simulated substrate that the second model is configured to predict). In some embodiments, different and / or additional outputs may be received, e.g., the second value, an output associated with the second plurality of values, etc.
[0143] At block 429, processing logic prepares the first plurality of outputs and the second plurality of outputs for presentation via a presentation element of the GUI. In some embodiments, additional outputs, e.g., outputs associated with different sets of inputs or multiple sets of inputs, may be further prepared for presentation. The presentation element includes two axes. A first of the two axes corresponds to a first characteristic of the first feature (e.g., a feature for which the first model is configured to predict one or more values). A second of the two axes corresponds to a second characteristic of the second feature (e.g., a feature for which the second model is configured to predict one or more values). Preparing the first plurality of outputs and the second plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element that indicates, for each simulated substrate of the first plurality of simulated substrates, a value of the first characteristic of the first feature and a value of the second characteristic of the second feature.
[0144] In some embodiments, the first characteristic and the second characteristic (e.g., a statistical metric of one or more feature predictions) may be the same characteristic (e.g., both may be average values of the feature predictions). In some embodiments, the first characteristic and the second characteristic may be different characteristics. In some embodiments, the first model comprises a machine learning model. In some embodiments, the second model comprises a machine learning model. In some embodiments, one or more of the models are statistical models, physics-based models, etc. In some embodiments, the generated figure may share one or more characteristics with the figure described in connection with FIG. 4B.
[0145] 4D is a flow diagram of a method 400D for generating a diagram that demonstrates the learning of a model, according to some embodiments. Method 400D may share one or more features with method 400B and / or method 400C.
[0146] At block 430, processing logic receives a first value associated with a first input parameter of the model. The first input parameter is associated with a processing recipe for processing the substrate. The first input parameter may be associated with, for example, a manufacturing parameter, a sensor measurement, etc.
[0147] At block 432, processing logic receives a first plurality of values, the first plurality of values ranging from a lowest value of the first plurality of values to a highest value of the first plurality of values, each of the first plurality of values associated with a second input parameter of the model, the second input parameter associated with the process strategy.
[0148] At block 434, processing logic provides the first value and the first plurality of values to a first model. The first model may be configured to generate one or more outputs that predict performance of the simulated substrate in response to receiving the set of input values.
[0149] At block 436, the processing logic receives a first plurality of outputs from the model. Each of the first plurality of outputs is associated with a first value and one of the first plurality of values (e.g., each output is an output generated by the model based on the first value being used as an input for a first parameter and one of the first plurality of values being used as an input for a second input parameter). Each of the first plurality of outputs is associated with a first characteristic of one of the first plurality of simulated substrates. In some embodiments, additional inputs may be provided to the model, and additional outputs may be received from the model. The additional inputs may differ from the initial inputs in one or more input parameters. The additional inputs may provide different multiple inputs, different single values (e.g., different reference conditions), etc. Each output (e.g., each predicted characteristic, each set of predicted characteristics, each characteristic of the set of predicted characteristics, etc.) may be associated with a set of inputs (e.g., one value for each input parameter). In some embodiments, an additional set of inputs may be provided to the model, which may generate outputs, e.g., an additional set of outputs associated with a second plurality of simulated substrates.
[0150] At block 438, processing logic prepares the first plurality of outputs for presentation via a presentation element of a GUI. The presentation element includes two axes (two independent axes, e.g., axes associated with independent values). A first of the two independent axes corresponds to a first characteristic of the first feature. Preparing the first plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element. The graphic indicates a value of the first characteristic of the first feature associated with an output of the first plurality of outputs.
[0151] FIG. 5A illustrates an example presentation element 500A that displays an indication of the training of one or more models (e.g., input / output mappings, associations, etc.) according to some embodiments. The presentation element 500A includes a first axis 502 and a second axis 504. The axes may be orthogonal, independent, etc. The first axis 502 may be associated with a first feature, a first characteristic of the first feature, etc. For example, the presentation element 500A may display data associated with one or more machine learning models to visualize the training of the one or more models (e.g., learned associations, learned input / output mappings, etc.). The first axis 502 may indicate a value of a statistical metric associated with a first characteristic of a first feature, e.g., one or more predictive metrics of a simulated substrate. The second axis 504 may indicate a value of a second characteristic of a second feature. In some embodiments, the first characteristic is the same as the second characteristic. In some embodiments, the first feature is the same as the second characteristic. In some embodiments, the first feature is associated with an output of a first model and the second feature is associated with an output of a second model.
[0152] The presentation element 500A includes an indication of the output space 506. The indication of the output space 506 may include multiple data points, as illustrated in FIG. 5A . In some embodiments, the data points of the indication of the output space 506 may be generated by providing multiple sets of input data (e.g., one set per output data point) to one or more models associated with the presentation element 500A. In some embodiments, the inputs associated with the output data points may span the model's input space, may span a portion of the model's input space, may be randomly selected from the model's input space (e.g., chosen to effectively / substantially span enough points in the model's input space), etc. In some embodiments, the indication of the output space 506 may be represented differently, for example, by coloring or shading an area of the presentation element 500A, by enclosing a portion of the presentation element 500A in a shape to indicate the output space of one or more models, etc.
[0153] The presentation element 500A may further include a first curve 508 and a second curve 510. The first curve 508 may include multiple data points (represented as white filled triangles in FIG. 5A ). The first curve 508 may include a visual distinction between outputs associated with a set of inputs where one or more input values have varied. For example, the first curve 508 may be associated with a series of sets of inputs, each of which maintains all inputs except the first input at a reference value. The value of the first input may vary from a minimum value to a maximum value. The shape of the data points representing the output values may indicate the progression of the associated input values, e.g., the shape may result from an output value associated with a set of inputs including the lowest value of the first input to an output value associated with a set of inputs including the highest value of the first input.
[0154] The second curve 510 may include a plurality of output data points (represented as black triangles). The second curve 510 may include output data points each associated with a set of input parameter values. The second curve 510 may include output data points associated with a series of sets of inputs, whereby each of the series of sets is associated with changing values of one or more second input parameters of the model. The first curve 508 and the second curve 510 may intersect in a region 512 (shown by a dashed circle) associated with a reference set of input parameter values.
[0155] In some embodiments, more than one input parameter may be varied, e.g., presentation element 500A may include two or more curves. Each curve may be associated with a different input parameter. In some embodiments, each curve may traverse region 512. In some embodiments, different curves may be associated with different criteria sets of inputs, e.g., curves may intersect in multiple regions of presentation element 500A. In some embodiments, only a single curve may be displayed (e.g., associated with variation of one input parameter).
[0156] 5B is an example presentation element 500B illustrating the training of one or more models, according to some embodiments. In some embodiments, a user may select a set of reference input values. Presentation element 500B includes a second curve 514 (e.g., including a series of data points, each data point associated with a set of input parameter values, each data point in the series including all inputs, except for the second input parameter, set to a reference value). Presentation element 500B may be associated with the same model(s) as presentation element 500A of FIG. 1. Presentation element 500B may be displayed in response to processing logic receiving adjustments to the set of reference input values. Presentation element 500B may be displayed upon a user entering an adjusted set of reference input values via the GUI.
[0157] First curve 508 and second curve 514 may intersect at region 516; e.g., region 516 may indicate one or more output values associated with an updated reference set of inputs. In some embodiments, the same region of input values may be associated with the curve of the adjustment of the reference conditions; e.g., curve 508 includes the same data points in presentation element 500A and presentation element 500B. In some embodiments, when the reference conditions are changed, the range of input values, the density of input values, the values of multiple inputs, etc. may change. In some embodiments, in response to the change in the reference conditions, the change in one or more input values, etc., one or more sets of input conditions may be provided to one or more models associated with presentation element 500B. Outputs from the one or more models may be displayed via presentation element 500B.
[0158] In some embodiments, one or more curves associated with multiple sets of input parameter values (e.g., varying one input parameter value from a minimum to a maximum value) may be shaped differently depending on the values of other input parameter values. For example, curve 510 may be shaped differently from curve 514. Presentation elements 500A and 500B may enable a user to concisely display such nonlinearities in the input / output mapping of one or more models associated with the presentation elements by allowing a user to change the set of reference values.
[0159] 5C is an example GUI 500C according to some embodiments. The GUI includes a presentation element 520, which may include, for example, one or more features of presentation element 500B of FIG. 5B. GUI 500C may also include additional elements, such as elements that provide additional information to the user, elements for receiving commands from the user, etc.
[0160] The GUI 500C includes a spatial settings element 522. The spatial settings element 522 may include one or more options related to the space displayed in the presentation element 520, e.g., an input space of one or more models, one or more output spaces, etc. The spatial settings element 522 may include a spatial edit element 524. The spatial edit element 524 may allow a user to adjust the input space of one or more parameters associated with one or more curves of the presentation element 520. For example, the spatial edit element 524 may open a window for editing one or more ranges of input values to generate one or more sets of data points (e.g., curves) for the presentation element 520. For example, minimum values, maximum values, the difference between adjacent values, the number of values provided, etc. may be adjusted by the spatial edit element 524.
[0161] The space setting element 522 may further include a space constraint element 526. In some embodiments, a partial region of the model space (e.g., input space, output space, etc.) may be of particular interest. The space constraint element 526 may allow a user to limit the space displayed in the presentation element 520. For example, the space constraint element 526 may allow a user to input a percentage value. The presentation element 520 may present a portion of the maximum range of one or more varying input values (e.g., corresponding to the entered percentage value). In some embodiments, the presentation element 520 may display the same number of data points (e.g., a user-selected number) in the constrained space and in the unconstrained space (e.g., the density of points increases when the input space is constrained). In some embodiments, the presentation element 520 may display a different number of data points in the constrained space than in the unconstrained space. In some embodiments, in response to a user selection via an element of the space setting element 522, one or more successive sets of inputs may be provided to one or more models.
[0162] GUI 500C may include a presentation settings element 528. The presentation settings element 528 may be utilized to adjust various visual settings associated with presentation element 520. Settings such as data point shape, how to visually distinguish between an output associated with the lowest input value varied among a plurality of input values and an output associated with the highest input value varied among a plurality of input values, color of data points and / or curves, spacing of presented data points (e.g., an option to present fewer data points than received from one or more models), toggling the display of a feature representing the size of the output space, etc. may be associated with presentation settings element 528.
[0163] GUI 500C may further include a criteria adjustment element 530. Criteria adjustment element 530 may be utilized by a user to adjust a set of criteria values, for example, to adjust a point in input space where an input parameter is adjusted, to adjust a point in output space where two or more curves of output data points intersect, etc. Criteria adjustment element 530 may allow values of one or more criteria input parameters to be changed.
[0164] In some embodiments, the criteria adjustment element 530 may include an input selection element 532. The input selection element may allow a user to select an input whose criteria input value will be adjusted. The input selection element 532 may include a list, a drop-down list, a series of icons, an input field, opening a window for input parameter selection, etc. The criteria adjustment element 530 may further include a value selection element 534. The value selection element may allow adjustment of the value of the input parameter indicated by the input selection element 532. The value selection element may include an input field, a slider, one or more buttons or icons, etc.
[0165] In some embodiments, a new graphic may be generated in response to a user adjusting a reference value associated with presentation element 520. The new graphic may include the new reference value, e.g., two or more curves may extend from the new reference value to illustrate changes in output based on varying inputs from the new reference value. In some embodiments, in response to a user adjusting the reference value, the new value is provided to one or more machine learning models associated with generating the data presented via presentation element 520.
[0166] GUI 500C may further include axis adjustment elements 536. Axis adjustment elements 536 may adjust the axes of presentation element 520. Axis adjustment elements 536 may allow a user to adjust each axis separately, for example, adjusting a first axis and adjusting a second axis. In some embodiments, presentation element 520 may illustrate a three-dimensional plot (e.g., a three-dimensional scatter plot of three dimensions of the output space of one or more models). The axis adjustment elements may allow a user to adjust a third axis.
[0167] Adjusting an axis may include selecting a feature (e.g., a predicted feature and / or a measured value of a simulated substrate) to be represented by the axis. In some embodiments, selecting a different feature may associate a different model (e.g., a machine learning model) with the presentation element 520. When a user selects one or more settings of an axis, one or more sets of input conditions may be provided to one or more models, and multiple outputs may be received from the one or more models to generate the graphic presented in the presentation element 520. In some embodiments, adjusting an axis may include selecting a characteristic (e.g., a statistical metric) of the feature to be represented by the axis. Selecting one or more features and / or characteristics may include selecting from a list, a set of icons or buttons, an input field, etc. In some embodiments, the GUI 500C may display multiple graphics, e.g., multiple charts with varying characteristics corresponding to a first axis and a second axis. Multiple charts may be utilized to simultaneously display additional model training, input / output mapping, etc.
[0168] 5D is an example presentation element 500D illustrating model learning associated with multiple varying input parameters, according to some embodiments. Presentation element 500D includes a first axis 540. First axis 540 may correspond to a first characteristic (e.g., a statistical metric) of a first feature (e.g., a model output) of a simulated substrate. Second axis 542 may correspond to a second characteristic of a second feature. The second characteristic may be the same as or different from the first characteristic. The second characteristic may be the same as or different from the first characteristic.
[0169] Presentation element 500D may include an indication of output space 544. The indication of output space 544 may be represented as a plurality of points (as shown), a shaded region, a bounded region, etc. The indication of output space 544 may be generated by providing randomly sampled input conditions to one or more machine learning models associated with the data in presentation element 500D. The indication of output space 544 may be generated by providing systematically sampled input conditions to one or more machine learning models associated with the data in presentation element 500D.
[0170] Presentation element 500D includes curves 546. Each curve may be a representation of varying one input parameter from a maximum value to a minimum value while holding all other input parameters at a set of reference values. The output when providing the reference set of inputs to one or more models associated with the data in presentation element 500D may be represented by the location where the curves intersect. Varying a first input while holding the other inputs at a set of reference values may generate a first curve 548, varying a second set of inputs may generate a second curve 550, varying a third set of inputs while holding the other inputs at a set of reference values may generate a third curve 552, and so on.
[0171] Presentation element 500D may visually display the learning of one or more models. The curves in presentation element 500D may show how the model's output changes as inputs are changed. For example, a user may aim to reduce a second characteristic of a second feature associated with second axis 542 (as indicated by the arrow) (e.g., in a new board manufacturing process, an updated board design, etc.). For example, the second characteristic of the second feature may be the average thickness of the board, and a thinner board may be the goal. Curve 546 visually indicates which input parameters affect the second feature in output space and suggests how strong the influence of the second feature on the second characteristic is in the vicinity of a set of reference values (e.g., near a reference input in input space).
[0172] A user may determine, based on presentation element 500D, that an input parameter associated with curve 552 does not have a strong influence on a second characteristic of a second feature (e.g., compared to other input parameters associated with other curves). For example, curve 552 may be associated with an input parameter such as gas pressure, which may not have a strong influence on a certain output result such as substrate thickness. The strength of the influence may be estimated, for example, by the length of the curve, the density of points along the curve, etc., depending on the display settings of presentation element 500D.
[0173] Based on presentation element 500D, a user may determine that an input parameter associated with curve 550 has a moderate impact on a second characteristic of a second feature (e.g., compared to other input parameters associated with other curves). For example, curve 550 may be associated with an input parameter such as temperature, and temperature may affect an output result such as the average thickness of the substrate (e.g., as learned by one or more machine learning models). Presentation element 500D provides an additional cue associated with first axis 540 that indicates that changing the input value associated with curve 550 does not have a strong impact on the second output characteristic. For example, the second characteristic of the second feature may be the average resistance of the substrate. Presentation element 500D may visually clarify that changing the input value associated with curve 550 (e.g., temperature) affects the second characteristic of the second feature (e.g., the thickness of the substrate indicated by second axis 542) but does not have a strong impact on the first characteristic of the first feature (e.g., resistance).
[0174] Based on presentation element 500D, a user may determine that an input parameter associated with curve 548 has a strong influence on a second characteristic of a second feature (e.g., compared to other input parameters associated with other curves). For example, curve 548 may be associated with an input parameter such as a process time, which may affect an output result such as an average thickness of a substrate.
[0175] The inputs associated with the curve 548 also strongly influence the first characteristic, e.g., the average resistance. A user can easily discern the impact of changing the input parameters associated with the curve 548 on both the characteristic associated with the first axis 540 and the characteristic associated with the second axis 542. The presentation element 500D may show nonlinear learning of one or more models, e.g., a change in the input parameters associated with the curve 548 causes a decrease and then an increase in the second characteristic. These nonlinearities are not captured in traditional model learning presentations (e.g., bar graphs), which may only display learning in the immediate vicinity of a reference set of conditions. Using the GUI associated with the presentation element 500D, the reference set of conditions may be changed, and patterns of the curve 546 associated with different sets of reference conditions (e.g., reference conditions around different portions of one of the curves displayed by the presentation element 500D) may be displayed. A user may use the displayed learning of one or more models to develop an intuition about, for example, how substrate processing conditions affect substrate properties. The displayed learning may be used to visually confirm corrective actions (e.g., to confirm the impact of updating a process strategy). The displayed learning may be used to communicate to a user that the model has generated logical associations between inputs and outputs (e.g., a user may be more likely to trust a model when they can see that the model matches their intuition, that the model is consistent with their experience, and that the model contains associations that do not immediately seem suspicious, as indicated, for example, by one or more discontinuities in the curve). The displayed learning may also be used to indicate the reliability of the model; for example, a rapidly oscillating or discontinuous curve may indicate an unreliable model, sampling bias or incompleteness in the training set, etc.
[0176] 6 is a block diagram illustrating a computer system 600 according to a particular embodiment. In some embodiments, computer system 600 may be connected to other computer systems (e.g., via a network such as a local area network (LAN), an intranet, an extranet, or the Internet). Computer system 600 may operate in the capacity of a server or a client computer in a client-server environment, or as a peer computer in a peer-to-peer or distributed network environment. Computer system 600 may be provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web appliance, a server, a network router, switch, or bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions that the device should take. Furthermore, the term “computer” is intended to include any collection of computers that, individually or jointly, execute a set (or sets) of instructions to perform one or more of the methodologies described herein.
[0177] In a further aspect, the computer system 600 may include a processing device 602, a volatile memory 604 (e.g., random access memory (RAM)), a non-volatile memory 606 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 618, which may communicate with each other via a bus 608.
[0178] The processing device 602 may be provided by one or more processors, such as a general-purpose processor (e.g., a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing another type of instruction set, or a combination of instruction set types), or a special-purpose processor (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[0179] Computer system 600 may further include a network interface device 622 (e.g., coupled to a network 674). Computer system 600 may also include a video display unit 610 (e.g., an LCD), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generating device 620.
[0180] In some implementations, the data storage device 618 may include a non-transitory computer-readable storage medium 624 (e.g., a non-transitory machine-readable storage medium) that may store instructions 626 for encoding any one or more of the methods or functions described herein, including the instruction-encoding components of FIG. 1 (e.g., the prediction component 114, the presentation component 115, the model 190, etc.), and for implementing the methods described herein.
[0181] The instructions 626 may also reside, completely or partially, within the volatile memory 604 and / or within the processing device 602 during execution by the computer system 600; thus, the volatile memory 604 and the processing device 602 may also constitute machine-readable storage media.
[0182] Although computer-readable storage medium 624 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" is intended to include single or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" also includes any tangible medium that can store or encode a set of instructions for execution by a computer, such a set of instructions causing a computer to perform one or more of the methodologies described herein. The term "computer-readable storage medium" is intended to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0183] The methods, components, and features described herein may be implemented by discrete hardware components or integrated into the functionality of other hardware components, such as an ASIC, FPGA, DSP, or similar device. In addition, the methods, components, and features may be implemented by firmware modules or functional circuitry within a hardware device. Furthermore, the methods, components, and features may be implemented as any combination of hardware devices and computer program components, or as a computer program.
[0184] Unless specifically indicated otherwise, terms such as "receiving," "performing," "providing," "obtaining," "causing," "accessing," "determining," "adding," "using," "training," "generating," "preparing," "training," "facilitating," and the like refer to actions and processes performed or implemented by a computer system that manipulate and convert data, which are similarly represented as physical (electronic) quantities in the computer system's registers and memory, into other data, which are similarly represented as physical quantities in the computer system's memory or registers, or in other such information storage, transmission, or display devices. Additionally, terms such as "first," "second," "third," and "fourth," as used herein, are intended to be labels for distinguishing between different elements and may not have any sequential meaning according to their numerical designation.
[0185] The examples described herein also relate to apparatus for performing the methods described herein. This apparatus may be specially constructed to perform the methods described herein or may comprise a general-purpose computer system selectively programmed by a computer program stored on the computer system. Such a computer program may be stored on a computer-readable tangible storage medium.
[0186] The example methods and diagrams described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized apparatus to perform the methods described herein and / or each of their individual functions, routines, subroutines, or operations. Examples of structures for these various systems are set forth above.
[0187] The above description is intended to be illustrative, not limiting. While the present disclosure has been described with reference to particular illustrative examples and implementations, it will be understood that the disclosure is not limited to the described examples and implementations. The scope of the present disclosure should be determined with reference to the following claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. A non-transitory machine-readable storage medium having stored thereon instructions that, when executed, cause a processing device to perform operations, the operations including: receiving, by the processing device, a first value associated with a first input parameter of a model, the first input parameter being associated with a first processing condition of a semiconductor wafer processing procedure; receiving, by the processing device, a first plurality of values ranging from a lowest value of the first plurality of values to a highest value of the first plurality of values, each of the first plurality of values associated with a second input parameter of the model, the second input parameter associated with a second processing condition of the semiconductor wafer processing procedure; providing, by the processing device, the first value and the first plurality of values to the model; receiving, by the processing device, a first plurality of outputs from the model, each of the first plurality of outputs associated with the first value and one value of the first plurality of values, and each of the first plurality of outputs associated with a first characteristic of one of a first plurality of simulated semiconductor wafers; preparing, by the processing device, the first plurality of outputs for presentation via a presentation element of a graphic user interface (GUI), the presentation element including two axes, a first of the two axes corresponding to a first characteristic of the first feature and a second of the two axes corresponding to a second characteristic of the first feature, wherein preparing the first plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element indicating a value of the first characteristic of the first feature and a value of the second characteristic of the first feature, associated with each of the first plurality of outputs; 1. A non-transitory machine-readable storage medium, comprising:
2. The non-transitory machine-readable storage medium of claim 1 , wherein the model comprises a machine learning model.
3. The operation is receiving a plurality of sets of input values; receiving a plurality of metrology measurements, each of the plurality of metrology measurements associated with one of the plurality of sets of input values; training the model to generate simulated metrology measurements of a simulated semiconductor wafer using the plurality of sets of input values and the plurality of metrology measurements, wherein training the model includes providing the plurality of sets of input values to the model as training inputs and providing the plurality of metrology measurements to the model as target outputs; The non-transitory machine-readable storage medium of claim 2 , further comprising:
4. The first feature is the thickness of the semiconductor wafer, Resistivity of semiconductor wafers, Semiconductor wafer sheet resistance, the refractive index of the semiconductor wafer, the attenuation coefficient of the semiconductor wafer, or Indication of semiconductor wafer geometry 10. The non-transitory machine-readable storage medium of claim 1, comprising at least one of:
5. The first characteristic of the first feature includes a statistical metric associated with simulated metrology measurements at a plurality of locations on the simulated semiconductor wafer, the statistical metric comprising: an average value of the first characteristic values; the median of the values of the first feature; the standard deviation of the values of the first feature; or uniformity of the values of the first feature 10. The non-transitory machine-readable storage medium of claim 1, comprising at least one of:
6. The operation is receiving, by the processing device, a second value associated with the second input parameter of the model; receiving, by the processing device, a second plurality of values, the second plurality of values ranging from a lowest value of the second plurality of values to a highest value of the second plurality of values, each of the second plurality of values associated with the first input parameter of the model; providing, by the processing device, the second value and the second plurality of values to the model; receiving, by the processing device, a second plurality of outputs from the model, each of the second plurality of outputs associated with the second value and one value of the second plurality of values, each of the second plurality of outputs associated with the first characteristic of one of a second plurality of simulated semiconductor wafers; preparing, by the processing device, the second plurality of outputs for presentation via the presentation element of the GUI, wherein preparing the second plurality of outputs for presentation includes facilitating generation of graphics for display in the presentation element that indicate the first characteristic of the first feature and the second characteristic of the first feature associated with each of the second plurality of outputs; 10. The non-transitory machine-readable storage medium of claim 1, further comprising:
7. 2. The non-transitory machine-readable storage medium of claim 1, wherein preparing the first plurality of outputs for presentation, by the processing device, further comprises facilitating generation of graphics for display in the presentation element that visually distinguish an output of the first plurality of outputs associated with the highest value of the first plurality of values from an output of the first plurality of outputs associated with the lowest value of the first plurality of values.
8. The operation is receiving, by the processing device, a second value associated with the first input parameter of the model; providing, by the processing device, the second value and the first plurality of values to the model; receiving, by the processing device, a second plurality of outputs from the model, each of the second plurality of outputs associated with the second value and one of the first plurality of values, and each of the second plurality of outputs associated with a first characteristic of one of a second plurality of simulated semiconductor wafers; preparing, by the processing device, the second plurality of outputs for display via the presentation element, wherein preparing the second plurality of outputs for presentation includes facilitating generation of graphics for display in the presentation element indicating a value of the first characteristic of the first feature and a value of the second characteristic of the first feature associated with each of the second plurality of outputs; 10. The non-transitory machine-readable storage medium of claim 1, further comprising:
9. A non-transitory machine-readable storage medium having stored thereon instructions that, when executed, cause a processing device to perform operations, the operations including: receiving, by the processing device, a first plurality of outputs from a first model, each of the first plurality of outputs associated with a first input value and one value of a first plurality of input values, and each of the first plurality of outputs associated with a first characteristic of one of a first plurality of simulated substrates; receiving, by the processing device, a second plurality of outputs from a second model, each of the second plurality of outputs associated with the first value and one of the first plurality of values, and each of the second plurality of outputs associated with a second characteristic of one of the first plurality of simulated substrates; preparing, by the processing device, the first plurality of outputs and the second plurality of outputs for presentation via a presentation element of a graphic user interface (GUI), the presentation element including two axes, a first of the two axes corresponding to a first characteristic of the first feature and a second of the two axes corresponding to a second characteristic of the second feature, wherein preparing the first plurality of outputs and the second plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element showing, for each simulated substrate of the first plurality of simulated substrates, a value of the first characteristic of the first feature and a value of the second characteristic of the second feature; 1. A non-transitory machine-readable storage medium, comprising:
10. 10. The non-transitory machine-readable storage medium of claim 9, wherein the first model comprises a first machine learning model and the second model comprises a second machine learning model.
11. The operation is receiving a first plurality of sets of input values; receiving a second plurality of sets of input values; receiving a first plurality of metrology measurements, each of the first plurality of metrology measurements associated with one of the first plurality of sets of input values; receiving a second plurality of metrology measurements, each of the second plurality of metrology measurements associated with one of the second plurality of sets of input values; training the first model to generate first simulated metrology measurements of a simulated substrate using the first plurality of sets of input values and the first plurality of metrology measurements, wherein training the first model comprises providing the first plurality of sets of input values to the first model as training inputs and providing the first plurality of metrology measurements to the first model as target outputs; training the second model to generate second simulated metrology measurements of the simulated substrate using the second plurality of sets of input values and the second plurality of metrology measurements, wherein training the second model comprises providing the second plurality of sets of input values to the second model as training inputs and providing the second plurality of metrology measurements to the second model as target outputs; 11. The non-transitory machine-readable storage medium of claim 10, further comprising:
12. The first feature is Thickness of the substrate, Resistance of the board, the sheet resistance of the substrate, the refractive index of the substrate, the damping coefficient of the substrate, or Indication of board geometry 10. The non-transitory machine-readable storage medium of claim 9, comprising at least one of:
13. The first characteristic of the first feature includes a statistical metric associated with simulated metrology measurements at a plurality of locations of one of the first plurality of simulated substrates, the statistical metric comprising: an average value of the first characteristic values; the median of the values of the first feature; the standard deviation of the values of the first feature; or uniformity of the values of the first feature 10. The non-transitory machine-readable storage medium of claim 9, comprising at least one of:
14. 10. The non-transitory machine-readable storage medium of claim 9, wherein preparing the first plurality of outputs and the second plurality of outputs for presentation by the processing device further comprises facilitating generation of a graphic for display in the presentation element that visually distinguishes between a presented data point associated with a first simulated substrate of the first plurality of simulated substrates associated with the highest value of the first plurality of values and a presented data point associated with a second simulated substrate of the first plurality of simulated substrates associated with the lowest value of the first plurality of values.
15. receiving, by the processing device, the first input value, the first input value comprising: the first model, and The second model receiving a first input value associated with a first input parameter of a substrate processing procedure, the first input parameter associated with a first processing condition of a substrate processing procedure; receiving, by the processing device, the first plurality of input values, the first plurality of input values ranging from a lowest value of the first plurality of values to a highest value of the first plurality of values, each of the first plurality of values comprising: the first model, and The second model receiving the first plurality of input values associated with a second input parameter of the substrate processing procedure, the second input parameter associated with a second processing condition of the substrate processing procedure; 10. The non-transitory machine-readable storage medium of claim 9, further comprising:
16. The operation is receiving, by the processing device, second values associated with the second input parameters of the first model and the second model; receiving, by the processing device, a second plurality of values ranging from a lowest value of the second plurality of values to a highest value of the second plurality of values, each of the second plurality of values associated with the first model and the first input parameter of the second model; providing, by the processing device, the second value and the second plurality of values to the first model; providing, by the processing device, the second value and the second plurality of values to the second model; receiving, by the processing device, a third plurality of outputs from the first model, each of the third plurality of outputs associated with the second value and one value of the second plurality of values, and each of the third plurality of outputs associated with the first characteristic of one of a second plurality of simulated substrates; receiving, by the processing device, a fourth plurality of outputs from the second model, each of the fourth plurality of outputs associated with the second value and one value of the second plurality of values, and each of the fourth plurality of outputs associated with the second characteristic of one of the second plurality of simulated substrates; preparing, by the processing device, the third plurality of outputs and the fourth plurality of outputs for presentation via the presentation element of the GUI, wherein preparing the third plurality of outputs and the fourth plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element that indicates, for each simulated substrate of the second plurality of simulated substrates, a value of the first property of the first feature and a value of the second property of the second feature; 16. The non-transitory machine-readable storage medium of claim 15, further comprising:
17. receiving, by the processing device, second values associated with the first input parameters of the first model and the second model; providing, by the processing device, the second value and the first plurality of values to the first model and the second model; receiving, by the processing device, a third plurality of outputs from the first model and a fourth plurality of outputs from the second model, each of the third plurality of outputs and each of the fourth plurality of outputs being associated with the second value and one of the first plurality of values, each of the third plurality of outputs being associated with the first characteristic of one of the second plurality of simulated substrates, and each of the fourth plurality of outputs being associated with the second characteristic of one of the second plurality of simulated substrates; preparing, by the processing device, the third plurality of outputs and the fourth plurality of outputs for presentation via the presentation element, wherein preparing the third plurality of outputs and the fourth plurality of outputs for presentation includes facilitating generation of a graphic for display in the presentation element that indicates, for each of the second plurality of simulated substrates, an associated value of the first property of the first feature and an associated value of the second property of the second feature; 16. The non-transitory machine-readable storage medium of claim 15, further comprising:
18. receiving, by one or more processors, a first value associated with a first input parameter of a first model, the first input parameter being associated with a process recipe for processing a substrate; receiving, by the one or more processors, a first plurality of values ranging from a lowest value of the first plurality of values to a highest value of the first plurality of values, each of the first plurality of values associated with a second input parameter of the first model, the second input parameter associated with the process strategy; providing, by the one or more processors, the first value and the first plurality of values to the first model; receiving, by the one or more processors, a first plurality of outputs from the first model, each of the first plurality of outputs associated with the first value and one of the first plurality of values, each of the first plurality of outputs associated with a first characteristic of a simulated substrate; preparing, by the one or more processors, the first plurality of outputs for presentation via a presentation element of a graphic user interface (GUI), the presentation element including two independent axes, a first of the two independent axes corresponding to a first characteristic of the first feature, and preparing the first plurality of outputs for presentation includes facilitating generation of a graphic in the presentation element that visually displays a relationship between the outputs of the first plurality of outputs and the first characteristic of the first feature; A method comprising:
19. providing, by the one or more processors, the first value and the first plurality of values to the second model; receiving, by the one or more processors, a second plurality of outputs from the second model, each of the second plurality of outputs associated with a second feature; preparing, by the one or more processors, the second plurality of outputs for presentation via the presentation element of the GUI, wherein a second axis of the two independent axes corresponds to a second characteristic of the second feature, and preparing the first plurality of outputs and the second plurality of outputs for presentation includes facilitating generation of a graphic in the presentation element that visually displays a relationship between the first plurality of outputs and the second plurality of outputs and the first characteristic of the first feature and the second characteristic of the second feature; 20. The method of claim 18, further comprising:
20. receiving, by the one or more processors, a second value associated with the second input parameter of the first model; receiving, by the one or more processors, a second plurality of values ranging from a lowest value of the second plurality of values to a highest value of the second plurality of values, each value of the second plurality of values associated with the first input parameter of the first model; providing, by the one or more processors, the second value and the second plurality of values to the first model; receiving, by the one or more processors, a second plurality of outputs from the first model, each of the second plurality of outputs associated with the second value and one of the second plurality of values, and each of the second plurality of outputs associated with the first feature; preparing, by the one or more processors, the second plurality of outputs for presentation via the presentation element of the GUI, wherein preparing the second plurality of outputs for presentation includes facilitating generation of a graphic in the presentation element that visually displays a relationship between the first plurality of outputs and the second plurality of outputs and the first characteristic of the first feature; 20. The method of claim 18, further comprising:
Citation Information
Patent Citations
Manufacturing design and process analysis system
JP2005518007A
Data display method
JP2014182605A
Method for analyzing data and method for displaying data
JP2016148988A
Search device, search method, and plasma processing device
JP2019165123A
Sensor metrology data intergration
US20200264335A1