Model Management for Non-Static Systems
By analyzing multivariate data using Gaussian hybrid model and time-coupled multimode hybrid model in IC manufacturing, identifying and adjusting outliers, the problem of quality and output improvement in non-static systems is solved, and higher production efficiency and product quality are achieved.
Patent Information
- Application Number
- CN202080059534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-16
- Filing Date
- 2020-10-14
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-10-14
AI Technical Summary
When manufacturing microelectronics and semiconductor devices such as ICs, it is difficult for the prior art to effectively detect and improve abnormalities in non-static systems, resulting in difficulty in improving unit processing quality and output.
The multivariate data was analyzed using Gaussian hybrid model (GMM) and time-coupled multimode hybrid model (TMM), outliers were used to identify outliers by calculating the anomaly score, and variables were adjusted based on these outliers to improve subsequent execution, and normal states and exception patterns were captured using sparse graphical models.
Improves the quality and output of microelectronics and semiconductor devices, and provides robust abnormality detection and diagnostic capabilities by automatically detecting and repairing abnormalities, adapting to system drift and transformation.
Smart Images

Figure CN114303235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the electrical, electronic and computer fields, and more particularly to the fabrication of microelectronics and / or semiconductor devices, such as integrated circuit (IC) fabrication. Background Art
[0002] Manufacturing microelectronics and / or semiconductor devices, such as IC manufacturing, typically involves multiple stages involving hundreds of unit processes. The overall quality and yield of microelectronic products depends on the yield and quality of each unit process, as well as on the successful integration of hundreds of unit processes. Furthermore, semiconductor manufacturing can involve non-static systems. However, anomaly detection in non-static systems is difficult using simple statistical methods. Therefore, there is a long-held but unmet need to improve the quality and yield of individual unit processes or small collections thereof.
[0003] Therefore, there is a need in the art to solve the above problems. Summary of the Invention
[0004] From a first aspect, the present invention provides a method for improving at least one of the quality and yield of a physical process, comprising: obtaining values of multiple variables associated with the physical process from corresponding executions of the physical process; determining at least one Gaussian mixture model (GMM) representing the values of the multiple variables for the executions of the physical process; calculating at least one anomaly score for at least one of the variables for at least one of the executions of the physical process based at least in part on the at least one Gaussian mixture model; identifying at least one of the executions of the physical process as an outlier based on the at least one anomaly score of the at least one of the variables; and modifying the at least one of the variables for use in one or more subsequent executions of the physical process based at least in part on the identification of the outlier so as to improve at least one of the quality and yield of the physical process.
[0005] From another aspect, the present invention provides an apparatus for improving at least one of the quality and yield of a physical process, the apparatus comprising: a memory; and at least one processor coupled to the memory, the processor operable to: obtain values of a plurality of variables associated with the physical process from corresponding executions of the physical process; determine at least one Gaussian mixture model (GMM) representing the values of the plurality of variables for the executions of the physical process; calculate at least one anomaly score for at least one of the variables for at least one of the executions of the physical process based at least in part on the at least one Gaussian mixture model; identify at least one of the executions of the physical process as an outlier based on the at least one anomaly score of the at least one of the variables; and modify the at least one of the variables for use in one or more subsequent executions of the physical process based at least in part on the outlier identification so as to improve at least one of the quality and yield of the physical process.
[0006] Viewed from another aspect, the present invention provides a computer program product for improving at least one of the quality and yield of a physical process, the computer program product comprising a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method for performing the steps of the present invention.
[0007] Viewed from another aspect, the invention provides a computer program stored on a computer readable medium and loadable into the internal memory of a digital computer, the computer program comprising software code portions for performing the steps of the invention when said program is run on a computer.
[0008] From another aspect, the present invention provides a computer program product comprising a non-transitory machine-readable storage medium having machine-readable program code embodied therein for improving at least one of the quality and yield of a physical process, the machine-readable program code comprising machine-readable program code configured to: obtain values of a plurality of variables associated with the physical process from corresponding executions of the physical process; determine at least one Gaussian mixture model (GMM) representing the values of the plurality of variables for the executions of the physical process; calculate at least one anomaly score for at least one of the variables for at least one of the executions of the physical process based at least in part on the at least one Gaussian mixture model; identify at least one of the executions of the physical process as an outlier based on the at least one anomaly score of the at least one of the variables; and modify the at least one of the variables for use in one or more subsequent executions of the physical process based at least in part on the outlier identification so as to improve at least one of the quality and yield of the physical process.
[0009] One aspect of the present invention is directed to a method for improving at least one of quality and yield of a physical process. The method includes: obtaining values of a plurality of variables associated with a physical process from respective executions of the physical process; determining at least one Gaussian mixture model (GMM) representing the values of the plurality of variables for the executions of the physical process; calculating at least one anomaly score for at least one of the variables for at least one of the executions of the physical process based at least in part on the at least one GMM; identifying at least one of the executions of the physical process as an outlier based on the at least one anomaly score for the at least one of the variables; and modifying the at least one of the variables for use in one or more subsequent executions of the physical process based at least in part on the outlier identification to improve at least one of the quality and yield of the physical process.
[0010] As used herein, "facilitating" an action includes performing the action, making the action easier, assisting in performing the action, or causing the action to be performed. Thus, by way of example and not limitation, instructions executing on one processor may facilitate an action performed by instructions executing on a remote processor by sending appropriate data or commands to cause or assist in performing the action. For the avoidance of doubt, where an actor facilitates an action by an action other than performing the action, the action is still performed by some entity or combination of entities.
[0011] One or more embodiments of the present invention, or elements thereof, can be implemented in the form of a computer program product comprising a computer-readable storage medium having computer-usable program code for performing the indicated method steps. In addition, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of a system (or device) (e.g., a computer) comprising a memory and at least one processor, the at least one processor being coupled to the memory and operable to perform the exemplary method steps. Still further, in another aspect, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of an apparatus for performing one or more of the method steps described herein; the apparatus can include (i) a hardware module, (ii) a software module stored in a computer-readable storage medium (or multiple such media) and implemented on a hardware processor, or (iii) a combination of (i) and (ii); any of (i)-(iii) implementing the specific techniques set forth herein.
[0012] These and other features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the invention, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The present invention will now be described, by way of example only, with reference to preferred embodiments as illustrated in the following drawings:
[0014] Figure 1 is a flow chart illustrating aspects of IC fabrication according to the prior art and in which preferred embodiments of the present invention may be implemented;
[0015] Figure 2 is a flowchart illustrating a control process according to an exemplary embodiment;
[0016] Figure 3 is a graph illustrating time series measurements according to the prior art and in which a preferred embodiment of the present invention may be implemented;
[0017] Figure 4A is a graph illustrating periodically correlated normal variables with outliers according to the prior art and in which a preferred embodiment of the present invention may be implemented;
[0018] Figure 4B is a diagram illustrating a multimodal normal variable with outliers according to the prior art and in which a preferred embodiment of the present invention may be implemented;
[0019] Figure 4C is a graph illustrating a drifting normal variable with an outlier value according to the prior art and in which a preferred embodiment of the present invention may be implemented;
[0020] Figure 5An anomaly detection method according to one aspect of the present invention is shown;
[0021] Figure 6 Describe aspects of Gaussian mixture models (GMMs) that may be used with aspects of the present invention;
[0022] Figure 7 depicts aspects of a temporally coupled multimode mixture model (TMM) according to aspects of the present invention;
[0023] Figure 8 An anomaly detection system according to one aspect of the present invention is shown;
[0024] Figure 9 A quality and yield improvement system (e.g., for semiconductor manufacturing) according to an aspect of the present invention is shown;
[0025] Figure 10 A system for observing and changing characteristics of an asset according to an aspect of the present invention is shown;
[0026] Figure 11 A Multiple Graphical Model (MGM) algorithm according to an aspect of the present invention is shown;
[0027] Figure 12 An Inverse Covariance Update (ICU) algorithm according to aspects of the present invention is shown;
[0028] Figure 13 A temporal order clustering (TOC) algorithm according to aspects of the present invention is shown;
[0029] Figure 14 A sparse weight selection algorithm (SWSA) according to aspects of the present invention is shown;
[0030] Figure 15A shows univariate trace feature data that can be used in embodiments of the present invention;
[0031] Figure 15B shows the experimental results of backward analysis of univariate trace feature data using GMM according to an embodiment of the present invention;
[0032] Figure 15C Shows the experimental results of backward analysis of univariate trace feature data using Z scores;
[0033] Figure 15D shows the experimental results of backward analysis of univariate trace feature data based on box plots;
[0034] Figure 16A and 16B shows multivariate trace feature data that may be used with embodiments of the present invention;
[0035] Figure 17A shows the experimental results of backward analysis of three-dimensional trace feature data using GMM according to an embodiment of the present invention;
[0036] Figure 17B Shown is a Hotelling T square (T 2 ) Experimental results of backward analysis of statistical three-dimensional trace feature data;
[0037] Figures 18A-18C shows multivariate trace feature data that may be used with embodiments of the present invention;
[0038] Figure 19A shows the experimental results of backward analysis of six-dimensional trace feature data using GMM according to an embodiment of the present invention;
[0039] Figure 19B Shown is a Hotelling T square (T 2 ) Experimental results of backward analysis of statistical six-dimensional trace feature data;
[0040] Figures 20A-20C shows continuous testing of time-varying scores according to an embodiment of the present invention;
[0041] Figure 21A shows univariate trace feature data that can be used in embodiments of the present invention;
[0042] Figure 21B Showing experimental results of forward projection of univariate trace feature data using GMM according to an embodiment of the present invention;
[0043] Figure 21C Shows experimental results of forward projection of univariate trace feature data using Z-scores;
[0044] Figure 22A shows univariate trace feature data that can be used in embodiments of the present invention;
[0045] Figure 22B Showing experimental results of forward projection of univariate trace feature data using GMM according to an embodiment of the present invention;
[0046] Figure 22C Shows experimental results of forward projection of univariate trace feature data using Z-scores;
[0047] Figure 22D shows experimental results of forward projection of univariate trace feature data according to box plots; and
[0048] Figure 23A computer system is shown that may be used to implement one or more aspects and / or elements of the present invention. DETAILED DESCRIPTION
[0049] Although embodiments of the present invention are described primarily with reference to manufacturing microelectronics and / or semiconductor devices, such as integrated circuit (IC) manufacturing, those skilled in the art will appreciate that aspects of the present invention may be used in many other applications. For example, in addition to semiconductor manufacturing, the principles of the present invention may be generally applicable to, for example, Internet of Things (IoT) technologies and solutions, big data, and / or analytics.
[0050] Figure 1 1 is a flow chart illustrating various aspects of an IC manufacturing process 100. A mask 110 may be fabricated based on the final physical layout of the IC. In some embodiments, the IC layout may be instantiated as a design structure comprising physical design data, and the design structure may be provided to manufacturing equipment to facilitate fabrication of physical integrated circuits according to the design structure. Wafer 120 is then processed in step 130, for example by photolithography and etching the wafer 120 using mask 140. Typically, during process 130, wafers 130 having multiple copies 151, 152, 153 of the final design are fabricated and cut (i.e., sliced) such that each die 151, 152, 153 is a copy of the integrated circuit. Once the wafer is sliced, testing and sorting of each die is performed at 160 to filter out any defective dies. Thus, each IC (die) is classified as good 170 (passing all tests) or bad 180 (failed one or more tests).
[0051] As previously mentioned, manufacturing microelectronic products and / or semiconductor devices, such as IC manufacturing, typically involves multiple stages including hundreds of unit processes. The overall quality and yield of microelectronic products depends on the quality of the yield of each unit process, as well as on the successful integration of hundreds of unit processes. Process (e.g., unit process) quality can be inferred from tool sensor time series measurements (reflecting tool conditions and recipes), incoming (partially completed) product characteristics, and other in-process measurements. Therefore, illustrative embodiments provide a method for improving product quality and / or yield (e.g., of a metallization process) by employing a predictive model for semiconductor manufacturing using tool group-related data, wafer data, and auxiliary data to detect and repair process anomalies.
[0052] Figure 2is a flow chart illustrating a control process 200 according to an exemplary embodiment. During each of the process steps s1 ... sN, in-process measurements are taken at 202 (e.g., to observe the current state of the asset), and post-process measurements are predicted at 204 based on those in-process measurements (e.g., learning a predictive model). A non-limiting example of a post-process measurement is wafer resistivity. Non-limiting examples of in-process measurements are plasma voltage, current, temperature, and pressure, as well as time elapsed during material deposition or etching. At 206, at least one controllable variable of the current process step is adjusted (e.g., a control action is taken to change a characteristic of the asset) in response to the prediction of the post-process measurement to reduce the error difference between the prediction and the target value of the post-process measurement.
[0053] Figure 3 is a graph showing time series measurement results of an exemplary IC manufacturing process. More specifically, Figure 3 The variables in columns 310-360 are shown, with rows corresponding to executions (runs) of a process (recipe). Column 310 shows the room ID. Figure 3 These variables are the same for all rows shown in . Although not in Figure 3 , embodiments may also include a wafer ID and / or tool ID. Column 320 shows a timestamp, and column 330 shows a recipe step. Column 340 shows voltage, column 350 shows time, and column 360 shows pressure. Figure 3 Each row in contains a unique value for each of these values, although those skilled in the art will appreciate that duplicate values are possible.
[0054] Figure 4A is a graph showing periodically related normal variables. Specifically, Figure 4A is a graph showing the average DC power supply voltage DCSrc.rVoltage of chamber 1 (CH1) in step 7 for performing processes on different days. Figure 4A Each square shown in represents the voltage value (y-axis) for an execution of the process on a given date (x-axis), with many dates having multiple executions and therefore multiple values for that date. Figure 4A It is shown that the voltage value is a normal variable that is periodically correlated: there is a repeating pattern of the variable with a sudden increase (e.g., at least starting around October 10 to above 640), then a gradual decrease (e.g., at least until around November 15 to below 620), then another sudden increase (e.g., at least starting around November 15 to above 650), followed by another gradual decrease (e.g., at least until around January 15 to below 610), and so on. Figure 4A Also included are values shown as circles rather than squares, which are outliers relative to the normal pattern of periodic correlation described above.
[0055] Figure 4B is a graph showing multimodal normal variables. Specifically, Figure 4B is a graph showing the median heating time for chamber 2 (CH2) in step 2 of the process performed on different days. Figure 4B Each square shown represents the heat time value (y-axis) for an execution of the process on a given date (x-axis), with many dates having multiple executions and therefore multiple values for that date. Figure 4B Showing that the thermal time values are a multimodal normal variable: the variable clusters in the first range during the first period (between 21.885 and 21.887, until December 21st, at least starting from December 10th), and then clusters in the second range during the second period (between 21.880 and 21.882, starting from December 22nd, at least until January 11th). Figure 4B Also included are values shown as circles rather than squares, which are outliers relative to the aforementioned multi-mode normal pattern.
[0056] Figure 4C is a graph showing drift normal variables. Specifically, Figure 4C is a graph showing the average actual position of the magnet lift motor (MagNetLift / Motor.rPos) for chamber 4 (CH4) in step 1 for executing this process on different days. Thus, Figure 4C Each square shown represents the average actual position (y-axis) of the process executions on a given date (x-axis), with many dates having multiple executions and therefore multiple values for that date. Figure 4C The average actual position value is shown to be a drifting normal variable: the variable gradually decreases at a steady rate (eg, a constant slope), for example, from approximately 37 on or before December 21 to approximately 20 on or after January 18. Figure 4C Also included are values shown as circles rather than squares; these are outliers relative to the normal pattern of drift described above.
[0057] Figure 5 1 shows an anomaly detection method according to one aspect of the present invention. The anomaly detection method can start with univariate data 510 or multivariate (e.g., multidimensional) data 520. If univariate data 510 is provided, it is converted into multivariate (e.g., three-dimensional) data 520 using time information. Each data point X can be considered in three-dimensional space. t , including the observation x t , timestamp t, and a function of the gap between the observed value and one or more other values (e.g., one or more values before and / or after the observed value), as described below with reference to Figure 8 Further discussion.
[0058] Minimum gap between previous and next values:
[0059] min(|x t -x t-1 |,|x t -x t+1 |)
[0060] Whether provided directly or indirectly through transformations of the univariate data 510, the multivariate data 520 is used to determine a sparse graphical model 530, which can be, for example, a Gaussian graphical model (GGM) and / or a Gaussian mixture model (GMM), as described below with reference to Figure 6 Further discussed. In some embodiments, the one or more sparse graphical models 530 may include one or more dual-sparse multi-task, multimodal Gaussian graphical models (MTL-MMGGMs) learned from data based on a Bayesian formulation. The dual sparsity may include sparsity in the dependency structure of the GGM and sparsity on the mixture components.
[0061] Within a given GGM graph (such as the graph represented by 530 in FIG4 ), nodes represent variables and lines represent non-zero (positive or negative) dependencies between variables. Some embodiments may use solid lines to represent positive dependencies and dashed lines to represent negative dependencies. Some embodiments may depict the strength of the dependencies by the thickness of the connecting lines, where thicker lines represent stronger dependencies. Dependency information is the correlation coefficient captured by the GGM model and provides an overview of the normal state of a particular system. The normal state of operation is a mixture of different dynamic conditions captured by the GGM graph. This dependency information is useful for understanding variable insights.
[0062] The aforementioned univariate data 510, multivariate data 520, and sparse graphical model 430 may include training data 570. In contrast, the test data 590 may include anomaly scores 540. For a new sample x, the anomaly score 540 is expressed as where ln represents the natural logarithm, and It is based on the expression The training data is 570 to learn the predictive distribution.
[0063] In some embodiments, anomaly scores 540 may be generated at predefined times (e.g., periodically, such as every 15 minutes). In some embodiments, anomaly scores 540 may be generated for the entire system and for each (or at least one or more) of the sensors within the system. For example, a current (online) sample of the most recent multivariate sensor data 520 may be received, and using the received data for the current time window, an overall anomaly score 540 for the system may be generated. Additionally, the most recent multivariate data 520 corresponding to the most recent time series from each (or at least one or more) of the system's sensors may be processed, and separate "per-variable" anomaly scores 540 may be generated to indicate the dynamic behavior or "health" of the system. As will be described, the generation of anomaly score(s) 540 may involve automated solving (updating) of model(s) 530 associated with the system.
[0064] As previously mentioned, the sparse graphical model 530 may include a Gaussian graphical model (GGM), and more specifically, a Gaussian mixture model (GMM). A GMM is a probability distribution p(x) formed as a weighted sum of K single-component Gaussian densities and / or distributions x:
[0065]
[0066] where π k is the mixing coefficient, and is the component, where μ k is the mean and Σ k is the covariance. The inverse of the covariance, Precision matrix.
[0067] Figure 6 Describes aspects of a Gaussian mixture model (GMM) that can be used with aspects of the present invention. The x-axis represents x, and the y-axis represents p(x), which can be a probability distribution. Figure 6 , K=3, and the dashed lines represent three single-component Gaussian distributions 602, 604, 606. The solid line indicates that the GMM 600 is composed of the weighted sum of the components 602, 604, 606.
[0068] Conventionally, when using a GGM, the log-likelihood is maximized:
[0069]
[0070] The mixture weights are updated:
[0071]
[0072] The covariance is updated:
[0073]
[0074] However, these techniques do not exploit structured learning (e.g., sparsity and / or correlation). Irrelevant components can be removed by using a sparse model. The sparse model may include sparse mixing weights π k and sparse inverse covariance Sparse mixing weight π k Automatic determination of the number of modes can be provided, while sparse inverse covariance and / or inverse precision matrix A sparse Gaussian graphical model (GGM) may be provided. Thus, the resulting model may be a multi-layer sparse mixture of sparse GGMs, which may include both sparsity in the dependency structure of the GGMs and sparsity on the mixture components.
[0075] As reference Figure 7 For further discussion, Illustrative embodiments of the present invention use the exact non-convex l0 norm, rather than the approximate convex sparsity-promoting l1 norm. Embodiments of the present invention utilize a temporally coupled multimode mixture model (TMM). This model recognizes that assets can operate in different modes, but share some similarities between them. Modes are gradually adjusted. The structure of the dependency graph should share some commonalities.
[0076] Figure 7 Aspects of a temporally coupled multimode mixture model (TMM) according to aspects of the present invention are depicted. Figure 7 In, as in Figure 6 In , the solid line represents a Gaussian mixture model (GMM) consisting of a weighted sum of three components represented by dashed lines. However, in Figure 7 In
[15] , the three components correspond to multivariate data consisting of the same set of variables during different (possibly at least partially overlapping) time windows. Thus, the Gaussian graphical model (GGM) representations of the three components of the GMM (here, TMM) each have the same set of nodes, the only difference being the connections between the nodes.
[0077] As used herein, |v i | represents the absolute value of the i-th element of vector v. If for small ∈>0 many elements satisfy |v i |≤∈, then the vector V is called ∈ sparse solution. ‖v‖ ∈ Represents the ∈ norm, counting the number of entries |v i |>∈. For example, ‖v‖0 represents the l0 norm (number of non-zero elements) of vector V. As discussed before, for Illustrative embodiments of the present invention utilize an exact non-convex sparsity to promote the l0 norm, as opposed to an approximate convex l1 norm.
[0078] Therefore, in a TMM according to an illustrative embodiment of the invention, the constrained regularized log-likelihood is maximized:
[0079]
[0080] Among them, sparsity is achieved by specifying the precision matrix The important components are directly constrained by the constraint ‖π‖ ∈ ≤κ ∈ Come to repair.
[0081] The dependency graph across components imposes some structural similarity. Only important components are restricted (e.g., only significant components are penalized), i.e., the mixture weights are significantly larger:
[0082]
[0083] Sometimes, the data has been slightly changed, but there is no access to the original data. Some domain knowledge about the exact matrix may be known, such as a good model has been constructed to obtain the matrix. This can be done by adding regularization to set Σ k Close to
[0084]
[0085] Figure 8 An anomaly detection system 800 according to one aspect of the present invention is shown. Input data 810 may include Figure 5 The input data 810 may be time series sensor data, which may be, for example, as described above. Figure 1 -4 is received in real time from one or more sensors within the IC manufacturing process system. The data processor 820 processes the input data 810 into, for example, one or more sparse graphical models, such as GGM, GMM and / or TMM, as described above with reference to Figure 5 530 and reference Figure 6 and 7 The data processor 820 may also include a method for converting the input data 810 from univariate data 510 to multivariate data 520, as discussed above with reference to FIG. Figure 5 As discussed in reference Figure 5 As discussed, the anomaly score calculator 830 calculates one or more anomaly scores 540 based on the input data 810 and based on the model determined in the data processor 820 .
[0086] Backward analysis 840 is a backward prediction tool that can be used, for example, for offline diagnosis. For backward analysis, the data point X can be considered in three-dimensional space. t, including the observation x t , timestamp t, and the minimum gap between the immediately preceding and following values: min(|x t -x t-1 |,|x t -x t+1 |). The backward analysis 840 may include a learned sparse mixture of sparse GMMs from all datasets and / or a calculated anomaly score for each historical sample. The backward analysis may include observing the current state of the asset, learning a predictive model for the asset, and taking control actions to change the characteristics of the asset. Thus, the backward analysis 840 may include modifying the input data 810, as described above with reference to Figure 2 Discussion and references below Figure 9 and Figure 10 Further Discussion. The backward analysis 840 is further discussed below with reference to Figures 15-19.
[0087] Forward projection 850 is a forward prediction tool that can be used, for example, for online anomaly detection. Forward projection 850 can include serial testing, which is trained on n-1 data points (see Figure 5 ), and then train on the nth instance (see test data 590), and it may also include repeatedly computing the time-varying score for the nth point as more samples are received (e.g., from a sensor). In the forward analysis, a data point X may be considered in three-dimensional space. t , the data point includes the observation value x t , timestamp t, and the average interval between the observed value and one or more (preferably several, for example three) previous values: x t -average(x t-1 ,x t-2 ,x t-3 ), which is a specific case (N=3) with the following more general formula:
[0088]
[0089] The forward projection 850 is discussed further below with reference to Figures 20-22.
[0090] Figure 9 A quality and yield improvement system 900 (e.g., for semiconductor manufacturing) according to one aspect of the present invention is shown. A time series sensor network 910 connects individual sensors at each asset to a central control center 920 (a computer system), which can receive / store time series data (e.g., variable 911, variable 912, variable 913) from each sensor. By way of example, variable 911 can be a periodically correlated normal variable, such as the one described above with reference to FIG. Figure 4AThe DC source voltage discussed above, variable 912 may be a multi-modal normal variable such as the one described above. Figure 4B The heating time discussed above, and variable 913 may be a drift normal variable, such as the one referenced above Figure 4C Motor location discussed.
[0091] Each asset may be composed of many different parts, and individual parts may be monitored with multiple sensors. Because different components in a system (asset) are not necessarily independent, the signals from each sensor must be analyzed in a multivariate manner. From the time series sensor network 910, multivariate time series data associated with each individual asset is input to a computer system 920, which provides: a model building framework configured with a model builder module 930, the model builder module 930 being configured to call instructions for building an anomaly detection model for the asset; an anomaly score calculator module 940, which is configured to call instructions for calculating one or more anomaly scores for the asset that may be indicative of abnormal / faulty operation; an anomaly score coordinator module 950, which is configured to call instructions for processing the calculated anomaly scores in order to implement and prioritize asset maintenance; and a process operation update module 960, which is configured to call instructions for processing updates to operations, as described above with reference to Figure 2 and Figure 8 And the following references Figure 10 discussed.
[0092] The model database or similar storage device 935 stores S anomaly detection models provided by the model builder module 930, where S is the number of assets (systems or tasks) in the fleet. In one embodiment, the learned anomaly detection model is a multi-task, multimodal Gaussian graphical model (MTL-MM GGM model) and / or a temporally coupled multimodal mixture model (TMM). More generally, the anomaly detection model can be a Gaussian mixture model (GMM) and / or a Gaussian graphical model (GGM).
[0093] In one embodiment, building a model by the model builder module 930 may include computing a combination of two model components: (1) a set of S sparse mixing weights and (2) a set of sparse GGMs 330. The former can be different between assets and thus represent the individuality of the assets in the fleet. The latter is shared with the S assets and thus represents the commonality across the S assets. The individual sparse mixing weights for an asset specify the importance of the GGMs. That is, the calculated weights serve as selectors for the GGMs in the database 935, and different sensors have different weights or signature weight distributions. These weights will typically have many zeros for robustness and interpretability and are automatically learned from the data, for example, based on a Bayesian formula. Thus, for each asset, the model builder module 930 can use mixing weight learning to optimally construct a learned model based on a representation that combines a common set of sparse GGMs as a basis set and individual sparse mixing weights that provide a sparse solution for the mixing weights, by which the number of sparse GGMs in the database 935 can be automatically determined. In the illustrative embodiment, using a semi-closed form solution and convex mixed integer programming formulation, the model builder module 930 is fast, accurate (provides a global solution), and simple because it does not use any "hidden" parameters to truncate the least contributing weights.
[0094] Given the learned model, during the "online" process, the anomaly score calculator module 940 receives the multivariate 911, 912, 913 and calculates anomaly scores at predefined times (e.g., periodically, such as once every 15 minutes). The anomaly score calculator module 940 can generate anomaly scores for the entire system and for each sensor. In one embodiment, the anomaly score calculator module 940 receives the current (online) sample of the latest multivariate sensor data from the time series sensor network 910 and uses the received data for the current time window to implement the above referenced Figure 5 The steps for generating an overall anomaly score 540 for the system are described. In addition, the model builder module 930 can further process the most recent multivariate data corresponding to the most recent time series data from each sensor of the system and generate a separate "per-variable" anomaly score for each variable to indicate the dynamic behavior or "health" of the system. As will be described, the generation of the score can involve the automated solution (updating) of the model 530, which can be a dual sparse mixture model associated with the system.
[0095] The anomaly score coordinator module 950 may implement functionality for ranking the overall anomaly score and the variable-by-variable anomaly score for each asset and for comparing each score to a set of thresholds provided by the model builder module 930. If certain anomaly scores are greater than a threshold, this may indicate a possible failure of that asset, and the anomaly score coordinator module 950 generates signals indicating those assets and their corresponding scores. In one embodiment, these signals may be automatically transmitted to the maintenance planner module 970, located on the same or an external computing system, which executes instructions for scheduling and prioritizing maintenance actions.
[0096] In one embodiment, when an asset's anomaly score is determined to exceed a threshold value derived from historical anomaly values, an output signal indicates the overall anomaly score for the individual asset. The individual "per-variable" anomaly scores determined for each sensor variable of that particular asset are further compared to threshold values for each of those corresponding sensor variables, determined, for example, based on the quantile values of each variable. A per-variable comparison exceeding a certain threshold value indicates that a particular sensor corresponding to that variable, or component, may be operating in an abnormal manner based on that variable's score. Thresholds for the overall anomaly score and per-variable anomaly scores can be determined by calculating scores against a dataset acquired under normal conditions. Specifically, the threshold value can be determined, for example, as the 95th percentile or, more simply, as the maximum value of the anomaly score under normal conditions. Thus, an output signal can be further generated to indicate the variables and / or sensors associated with the specific components most likely to cause a potential failure of the asset. Based on the asset data represented in the signal, it can be determined which components of the asset may require immediate maintenance.
[0097] The process operation update module 960 updates tool conditions and recipes, incoming (partially completed) product characteristics, and other data. The process operation update module 960 then passes the updated data to the maintenance planner module 970. The maintenance planner module 970 can execute instructions for prioritizing maintenance actions. The maintenance planner module 970 can determine the need to indicate or mark a service interruption to repair a specific potentially troublesome part of an asset (system) based on severity and resource availability, as well as values retrieved from the attribute database 975. For example, a new part, component, or sensor may need to be replaced in a specific asset to resolve a problem determined by a per-variable anomaly score, discussed later with reference to the anomaly score coordinator module 950. In one embodiment, the maintenance planner module 970 can automatically generate further signals embodying messages to mark or schedule a service interruption, repair, or other type of maintenance for the potentially troublesome asset.
[0098] Figure 10 A system for observing and changing the characteristics of an asset according to one aspect of the present invention is shown. The data processor 1020 (which may be generally similar to Figure 8820 in ) retrieves input data 810 from the attribute database 1015 (which may generally be similar to Figure 9 975 in ).
[0099] The statistical analysis engine 1030 may generally be similar to Figure 9 The model builder module 930 in FIG. 1 and the anomaly score calculator 1040 may generally be similar to Figure 9 940 and / or 950 in. The statistical analysis engine can be operated to detect outliers for a batch of wafers using sparse GMM or GPR (Gaussian process regression). The anomaly score calculator 1040 can be operated to calculate an anomaly score for historical data or an instantaneous sample using single or multiple sensor measurements with sparse GMM and predict the next wafer measurement based on the previous measurement.
[0100] At least a portion of the output generated by anomaly score calculator 1040 (e.g., reference Figure 9 At least a portion of the output discussed in 940 and / or 950 of FIG. 100 may be communicated to the user via a visual interface (e.g., a graphical user interface) and / or an email notification system 1050. Module 1060 operates similarly to the above referenced Figure 9 Whether the anomaly score is less than a threshold is determined in the manner discussed in 950. If the anomaly score is less than the threshold, then the controlled tool 1070 operates normally.
[0101] If the anomaly score exceeds a threshold, the control action module 1080 modifies one or more operating parameters of the controlled tool 1070 based, for example, at least in part on the actual demand 1091 and / or engineering domain knowledge 1092. The control action module may be substantially similar to Figure 9 960 and / or 970 in. Thus, module 1060 is operable to compare the prediction (from module 1040) with subsequent measurements using GPR to determine whether the actual measurement is abnormal. If an abnormal measurement is identified in 1060, module 1080 is operable to take remedial action.
[0102] Figure 11 The Multi-Model Graphical Model (MGM) algorithm according to one aspect of the present invention is shown, which calls at line 1111 Figure 12 The Inverse Covariance Update (ICU) algorithm is shown. At line 1202, the ICU algorithm calls Figure 13 The Temporal Ordering Clustering (TOC) algorithm shown. Figure 14The sparse weight selection algorithm (SWSA) according to one aspect of the present invention is shown. Figure 11 π in row 1110 of .
[0103] Illustrative embodiments of the present invention advantageously utilize unsupervised Gaussian mixture models (GMMs), and more specifically time-coupled multimode mixture models (TMMs), to provide quality improvements for non-stationary systems. The illustrative embodiments provide predictive models to detect outliers and anomalies in non-stationary systems with temporal information (e.g., one or more temporal predictor variables). The illustrative embodiments automatically capture multiple normal operating states, adapt the model to drift and shifts, and respect the temporal order of observations. The illustrative embodiments are also robust to noise and highly interpretable for diagnostic purposes.
[0104] The illustrative embodiments provide a novel multimodal prediction model for learning density functions of non-stationary systems using a graphical mixture model. The illustrative embodiments take into account the temporal order of samples and enforce structural similarity across different operating modes, e.g., dependency graphs of components. Sparsity for exact matrices can be handled via the l0 norm, and domain knowledge can be incorporated into the model. The illustrative embodiments also provide optimization algorithms for training models, as well as backward and forward prediction tools for offline diagnosis and online anomaly detection. These tools may include converting time series univariate data into multivariate (e.g., three-dimensional) data. These tools may additionally or alternatively include updating anomaly scores as more observations are obtained.
[0105] Figure 15A Univariate trace feature data that may be used in embodiments of the present invention is shown. Figure 15A The voltage values (y-axis) in step 7 are shown for sequentially processed wafers (x-axis), as shown in FIG. Figure 4A As discussed above, voltage values are periodically correlated normal variables. Figure 4A The square in the figure can correspond to Figure 15A A subset of the hollow circles in . Figure 4A Similarly, FIG15 includes outliers that are shown as solid circles rather than hollow circles (or squares). Specifically, Figure 15A Shown are 931 samples of the mean voltage of chamber 1 (AT / CH1 / DCSrc.rVoltage_mean) during Halo Paste, including 3 outliers.
[0106] Figures 15B-15D Shown for Figure 15A Experimental results for backward analysis of univariate trace data are shown. The experimental design aims to detect outliers, compare counts, and visualize differences. Figure 15B The experimental results of backward analysis of univariate trace feature data using GMM according to an embodiment of the present invention are shown. Figure 15B is a visualization of the variation fraction (y-axis) of consecutively processed wafers (x-axis) in step 7. Figure 15B As shown by the solid circle in FIG, the exemplary embodiment of the present invention correctly identifies Figure 15A When compared with conventional techniques (such as Z-score and box plot), the technique of the present invention using only GMM correctly identified Figure 15A Three outliers in the univariate trace features in .
[0107] Figure 15C The experimental results of a backward analysis of univariate trace feature data using Z-scores are shown. A z-score, or standard score, shows how many standard deviations a given data point is below or above the mean. Figure 15C is a visualization of the change scores (y-axis) in step 7 for consecutively processed wafers (x-axis). Figure 15C As shown by the solid circle in the figure, this traditional technique only identifies one outlier, not Figure 15A There are three outliers in the univariate trace feature data in .
[0108] Figure 15D The experimental results of backward analysis of univariate trace feature data based on box plots are shown. Figure 15A similar, Figure 15D The voltage values observed from step 7 (y-axis) for successively processed wafers (x-axis) are shown. Figure 15D In the box plot shown, the horizontal dashed lines 1514, 1515, and 1516 represent the first quartile (25th percentile), second quartile (50th percentile or mean), and third quartile (75th percentile) values within the data set, respectively. Figure 15D As shown by the solid circle in FIG, the conventional technique only identifies one outlier (wafer 647), instead of Figure 15A There are three outliers in the univariate trace feature data in .
[0109] Figure 16A and 16B Multivariate trace feature data that can be used with embodiments of the present invention is shown. More specifically, Figure 16A and 16B Shows that for Figure 15A The same sequentially processed wafers (x-axis) and the values of the additional variables in step 7 (y-axis). As described above, Figure 15A The voltage values (y-axis) of the wafers processed successively (x-axis) in step 7 are shown, as shown above with reference to FIG. Figure 4AAs shown, the continuously processed wafers are a periodically correlated normal variable (open circles) with three outliers (filled circles). Figure 16A The current value (y-axis) of the continuously processed wafers in step 7 is shown. Figure 15A Another periodically correlated normal variable (open circles) with outliers (filled circles) at the same three wafers. Figure 16B The power value (y-axis) of the wafers processed in succession in step 7 is shown. Figure 15A and Figure 16A The normal variable has an outlier value at 1 of the 3 wafers.
[0110] Figure 15A 、 16A and 16B each show univariate trace feature data. However, Figure 15A 、 16A 16B and 16C show the values of different variables (y-axis) for the same process step (step 7) for the same wafer (x-axis). Figure 15A 、 16A and 16B can be regarded as 3D multivariate trace feature data. Figure 15A 、 16A When and 16B, the resulting multivariate trace feature data has three outliers: wafers 267, 470, and 647. Each of wafers 267, 470, and 647 is an outlier for at least one of the three dimensions.
[0111] Figure 17A The experimental results of backward analysis of three-dimensional trace feature data using GMM according to an embodiment of the present invention are shown. The illustrative embodiment of the present invention correctly identifies Figure 15A 、 16A 16B show three outliers (wafers 267, 470, and 647) within the three-dimensional multivariate trace feature data.
[0112] Figure 17B Shown is a Hotelling T square (T 2 ) is the experimental result of the backward analysis of the three-dimensional trace feature data. This conventional technology only identifies the presence of Figure 15A 、 16A One outlier (wafer 647) among the three outliers (wafers 267, 470, and 647) in the three-dimensional multivariate trace feature data shown in FIG16B.
[0113] Figures 18A-18C Multivariate trace feature data that can be used with embodiments of the present invention is shown. More specifically, Figures 18A-18C Shows that for Figure 15A 、16A Same sequentially processed wafer as in 16B (x-axis), with values of the additional variables in step 7 (y-axis). Figure 18A The pressure values (y-axis) of the wafers (x-axis) processed successively in step 7 are shown, which are the wafers having an abnormal value at wafer 307 (at Figure 15A 、 16A Normal variables (open circles) for wafers that are not outliers (or wafers in 16B). Figure 18B The current (AT.CH1.EMCoil.BOM) values (y-axis) in step 7 are shown for successively processed wafers (x-axis), which are normal variables (open circles) without outliers (filled circles). Figure 18C The backside gas pressure values (y-axis) in step 7 for successively processed wafers (x-axis) are shown as normal variables (open circles) with no outliers (filled circles).
[0114] and Figure 15A 、 16A Similar to 16B, Figures 18A-18C Each shows univariate trace feature data. However, Figure 15A 、 16A , 16B and 18A-18C show the values of different variables (y-axis) for the same process step (step 7) for the same wafer (x-axis). Figure 15A 、 16A , 16B and 18A-18C can be regarded as 6-dimensional multivariate trace feature data. Figure 15A 、 16A , 16B, and 18A-18C, the generated multivariate trace feature data has four outliers: wafers 267, 307, 470, and 647. Each of wafers 267, 307, 470, and 647 is an outlier for at least one of the six dimensions.
[0115] Figure 19A The experimental results of backward analysis of six-dimensional trace feature data using GMM according to an embodiment of the present invention are shown. The illustrative embodiment of the present invention correctly identifies Figure 15A 、 16A , 16B and 18A-18C collectively show four outliers (wafers 267, 307, 470 and 647) within the six-dimensional multivariate trace feature data.
[0116] Figure 19B The Hotelling T square (T 2 ) Statistical experimental results of backward analysis of six-dimensional trace feature data. This conventional technology only identifies the presence of Figure 15A 、 16A, 16B and 18A-18C collectively show two outliers (wafers 307 and 647) among the above four outliers (wafers 267, 307, 470 and 647) within the three-dimensional multivariate trace feature data.
[0117] Figures 20A-20C Continuous testing of time-varying scores according to an embodiment of the present invention is shown. Specifically, Figure 20A Shown similar to the above reference Figure 15A The univariate trace data discussed above. Figure 8 As discussed, the forward projection 850 may include serial testing, which trains on n-1 data points and then tests on the nth instance. Figure 20A In this example, these n-1 data points are denoted as the training set 2017, which can generally correspond to Figure 7 The training data 570 in , and the nth instance is represented as the test point 2011, which can generally correspond to Figure 7 The test data in 590.
[0118] As mentioned above Figure 8 As discussed, forward projection 850 may also include repeatedly computing the time-varying score for the nth point as more samples are received (e.g., from a sensor). Figure 20B In the training set 2027 and the test point 2021 correspond to Figure 20A However, the score for test point 2021 is recalculated after receiving subsequent data points (e.g., observations and / or samples) 2022 and 2023, which have steadily decreasing values relative to 2021, thereby showing that test point 2021 is not an outlier, but rather similar to the above reference point. Figure 4A Part of the discussion of periodicity-related normal patterns.
[0119] exist Figure 20C In, as in Figure 20B In the training set 2037 and the test point 2031 correspond to Figure 20A 2017 and 2011 in , and recalculate the score for test point 2031 after receiving subsequent data points (e.g., observations and / or samples) 2032 and 2033. However, here, subsequent samples 2032 and 2033 have values similar to (and different from) the immediately preceding data point 2031, thus showing that test point 2031 is not part of a normal pattern of periodic correlation, but is instead an outlier.
[0120] Figure 21A 1 shows univariate trace feature data that can be used in embodiments of the present invention. Specifically, Figure 21A Shown with the above reference Figure 15A 、 Figure 20A and Figure 20B Here, the test point 2111 is similar to the data discussed above. Figure 20A Test points in 2011 and Figure 20B The test point 2021 in FIG. 2 corresponds to the test point 2021 in FIG. 2 , where the subsequent points steadily decrease relative to the test point 2111 and thus show that the test point 2111 is part of a periodically related normal pattern rather than an outlier. Figure 21A In the example, the test window 2119 includes some points before the test point 2111 and all points after it, which is the same as Figure 20A In contrast to the training set in 2017, Figure 20A The training set 2017 in includes all points before the test point 2011 and does not include all points after the test point 2011.
[0121] Figure 21B The experimental results of forward projection of univariate trace feature data using GMM according to an embodiment of the present invention are shown. More specifically, Figure 21B 1 shows how the change score (shown on the y-axis) for test point 2111 is recalculated after receiving a number of subsequent data points (shown on the x-axis). Test point 2111 was initially determined to be an outlier because it was significantly (e.g., greater than 1 standard deviation) above the mean: in fact, test point 2111 was greater than 2 standard deviations above the mean. Receiving a few additional test points slightly raised the mean, but test point 2111 is still almost 2 standard deviations above the (now slightly higher) mean and is therefore still considered an outlier because it is more than 1 standard deviation above the mean.
[0122] Figure 21C The experimental results of forward projection of univariate trace feature data using Z-score are shown. Figure 21B Same, Figure 21C 1 shows how the change score for test point 2111 (shown on the y-axis) is recalculated after receiving a number of subsequent data points (shown on the x-axis). Figure 21C In , test point 2111 is initially determined to be an outlier because it is significantly higher than the immediately preceding data point. Figure 21C As shown, once additional data points are received (which show a comparison Figure 20B More like Figure 20C ), the test point 2111 is no longer classified as an outlier.
[0123] Figure 22A 1 shows univariate trace feature data that can be used in embodiments of the present invention. Figure 15A 、 20Aand 21A, and in fact the test window 2219 is similar to Figure 21A 2119 in the test window. However, with Figure 20A and Figure 21A Different, in Figure 22A There are no test points shown in . Figure 22A There are 3 outliers within the test window 2219 in .
[0124] Figure 22B The experimental results of forward projection of univariate trace feature data using GMM according to an embodiment of the present invention are shown. Note that Figure 22B Include only Figure 22A The exemplary embodiment of the present invention correctly identifies the value within the test window 2219 of Figure 22A There are three outliers within the test window 2219 of the trace feature data shown in .
[0125] Figure 22C The experimental results of forward projection of univariate trace feature data using Z-score are shown. Figure 22B similar, Figure 22C Include only Figure 22A The dashed line 2238 indicates the criteria for characterizing a value as an outlier: a value above 2238 is more than 1 standard deviation above the mean and is therefore identified as an outlier. Figure 22C 17 data points in the were identified as outliers, instead of Figure 22A There are three outliers within the test window 2219 of the trace feature data shown in .
[0126] Figure 22D The experimental results of the forward projection of the univariate trace feature data according to the box plot are shown. Note that, unlike Figure 22B and Figure 22C different, Figure 22D include Figure 22A All values shown in , not just the values within the test window 2219. The values to the left of the vertical dashed line 2248 are the training data, while the values to the right of the vertical dashed line 2248 are the test data. Figure 22D In the box plot shown, the horizontal dashed lines 2244, 2245, and 2246 represent the first quartile (25th percentile), second quartile (50th percentile or mean), and third quartile (75th percentile) values within the data set, respectively. These horizontal dashed lines indicate the criteria for characterizing a value as an outlier: a value that is not between the 25th and 75th percentiles (e.g., not between dashed lines 2244 and 2246) is identified as an outlier. Because Figure 22D None of the data points shown in meet these criteria, so by this conventional method, Figure 22ANone of the three outliers present in the test window 2219 of the trace feature data shown in is identified as an outlier.
[0127] One or more embodiments of the present invention, or elements thereof, may be implemented at least in part in the form of an apparatus including a memory and at least one processor coupled to the memory and operable to perform the exemplary method steps.
[0128] One or more embodiments may utilize software running on a general purpose computer or workstation. Figure 23 Such an implementation may employ, for example, a processor 2302, a memory 2304, and an input / output interface formed, for example, by a display 2306 and a keyboard 2308. As used herein, the term "processor" is intended to include any processing device, for example, a processing device including a CPU (central processing unit) and / or other forms of processing circuitry. Furthermore, the term "processor" may refer to more than one individual processor. The term "memory" is intended to include memory associated with a processor or CPU, for example, RAM (random access memory), ROM (read-only memory), fixed storage devices (e.g., hard drives), removable storage devices (e.g., disks), flash memory, etc. Furthermore, as used herein, the phrase "input / output interface" is intended to include, for example, one or more mechanisms for inputting data to a processing unit (e.g., a mouse), and one or more mechanisms for providing results associated with the processing unit (e.g., a printer). Processor 2302, memory 2304, and input / output interfaces such as display 2306 and keyboard 2308 may be interconnected, for example, via a bus 2310 that is part of a data processing system 2312. Appropriate interconnections (e.g., via bus 2310) may also be provided to a network interface 2314 (such as a network card) and a media interface 2316 (such as a disk or CD-ROM drive), where the network interface 2314 may be provided for interfacing with a computer network and the media interface 2316 may be provided for interfacing with media 2318.
[0129] Thus, computer software comprising instructions or codes for performing the methods of the present invention as described herein may be stored in one or more associated memory devices (e.g., ROM, fixed or removable memory), and when ready to be used, partially or fully loaded (e.g., loaded into RAM) and implemented by the CPU. Such software may include, but is not limited to, firmware, resident software, microcode, etc.
[0130] A data processing system suitable for storing and / or executing program code will include at least one processor 2302 coupled directly or indirectly to a memory 2304 through a bus 2310. The memory elements may include local memory employed during actual implementation of the program code, bulk storage, and cache memories which provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during implementation.
[0131] Input / output or I / O devices (including but not limited to keyboard 908, display 906, pointing device, etc.) can be coupled to the system either directly (such as via bus 2310) or through intervening I / O controllers (omitted for clarity).
[0132] Network adapters such as network interface 2314 may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modem and Ethernet cards are just a few of the currently available types of network adapters.
[0133] As used herein (including in the claims), a "server" includes a physical data processing system (e.g., a Figure 23 ). It will be understood that such a physical server may or may not include a display and keyboard.
[0134] It should be noted that any of the methods described herein may include the additional step of providing a system comprising various software modules contained on a computer-readable storage medium; the modules may include, for example, any or all of the elements depicted in the block diagrams or other figures and / or described herein. The method steps may then be performed using the various software modules and / or submodules of the system described above executed on one or more hardware processors 2302. Further, a computer program product may include a computer-readable storage medium having code adapted to be implemented to perform one or more of the method steps described herein, including providing a system comprising various software modules.
[0135] Exemplary System and Article of Manufacture Details
[0136] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to perform various aspects of the present invention.
[0137] Computer readable storage medium can be a tangible device that can retain and store the instructions used by the instruction execution device.Computer readable storage medium can be, for example but not limited to, electronic storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or any suitable combination of the above.The non-exhaustive list of more specific examples of computer readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device such as punch card or the protrusion structure in the groove with the instruction recorded thereon and any suitable combination of the above.Computer readable storage medium as used herein should not be interpreted as temporary signal itself, such as radio wave or other free propagation electromagnetic wave, electromagnetic wave propagated by waveguide or other transmission media (for example, light pulse passing through fiber optic cable) or electric signal emitted by wire.
[0138] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or downloaded to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0139] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data of integrated circuit or source code or object code written in any combination of one or more programming languages, these programming languages include object-oriented programming languages (such as Smalltalk, C++ etc.) and process programming languages (such as " C " programming languages or similar programming languages). The computer-readable program instructions can be performed completely on the user's computer, partly on the user's computer, performed as an independent software package, partly on the user's computer, partly on a remote computer or fully on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer by any type of network (including local area network (LAN) or wide area network (WAN)), or can be connected to an external computer (for example, using an internet service provider through the internet). In certain embodiments, the electronic circuit comprising for example programmable logic circuit, field programmable gate array (FPGA) or programmable logic array (PLA) can make the electronic circuit personalized to perform computer-readable program instructions by utilizing the state information of computer-readable program instructions, so as to perform various aspects of the present invention.
[0140] The present invention is described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0141] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in the flowchart and / or block diagram or multiple blocks. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in the flowchart and / or block diagram or multiple blocks.
[0142] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in the flowchart and / or block diagram or multiple boxes.
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of possible implementations of the systems, methods and computer program products according to different embodiments of the present invention. To this end, each box in the flowchart or block diagram may represent a module, segment or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions annotated in the box may not occur in the order annotated in the figure. For example, depending on the functions involved, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the opposite order. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.
[0144] The description of various embodiments of the present invention has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
[0145] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms as well. It should also be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the features, wholes, steps, operations, elements and / or parts, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, parts and / or combinations thereof.
[0146] All means or steps plus corresponding structures, materials, acts, and equivalents of functional elements in the following claims are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the invention. The embodiments are chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the various embodiments of the invention with various modifications suitable for the particular use contemplated.
Claims
1. A method for improving at least one of the quality and yield of a physical process, comprising: obtaining values of a plurality of variables associated with a physical process from respective executions of the physical process, wherein the physical process comprises at least a portion of a semiconductor manufacturing process; determining at least one Gaussian mixture model (GMM) of said values of said plurality of variables representing performance of said physical process; calculating an anomaly score for each of the plurality of variables for at least one of the executions of the physical process based at least in part on the at least one Gaussian mixture model; identifying the at least one of the executions of the physical process as an outlier based on at least one anomaly score of at least one variable of the plurality of variables exceeding a specified threshold for the at least one of the executions of the physical process; and Based at least in part on the identification of the outlier, the at least one variable of the plurality of variables is modified for one or more subsequent executions of the physical process to improve the at least one of quality and yield of the physical process.
2. The method according to claim 1, wherein The plurality of variables associated with the physical process include at least one of voltage, current, power, and pressure.
3. The method according to claim 1, wherein The at least one GMM comprises at least one time-coupled multimode model TMM.
4. The method according to claim 3, wherein: Determining the at least one TMM includes: determining a plurality of Gaussian graphical models GGM, each Gaussian graphical model representing values of the plurality of variables during a respective subset of the executions of the physical process, such that each respective subset of the executions of the physical process includes the executions of the physical process during a respective time period; determining corresponding mixing weights of the GGM; and A weighted sum of the plurality of GGMs is determined according to the mixing weights.
5. The method according to claim 4, wherein At least a portion of a first time period corresponding to a first GGM of the plurality of GGMs overlaps with at least a portion of a second time period corresponding to a second GGM of the plurality of GGMs.
6. The method according to claim 3, wherein: Determining the at least one TMM includes maximizing a constrained regularized log-likelihood.
7. The method according to claim 3, wherein: Determining the at least one TMM includes using an exact convex l0-norm.
8. A method according to any one of the preceding claims, wherein Obtaining the value includes converting the univariate data into multivariate data based at least in part on time data associated with the univariate data.
9. The method according to claim 8, wherein The temporal data includes respective timestamps of values within the univariate data.
10. The method according to claim 8, wherein The temporal data includes at least one difference between immediately temporally adjacent values within the univariate data.
11. The method of claim 1 , further comprising recalculating the at least one anomaly score for at least one of the variables for at least one of the executions of the physical process based at least in part on a value for at least one of the one or more subsequent executions of the physical process.
12. The method according to claim 1, further comprising: training the model with values from one or more past executions of the physical process; as well as The model is tested using values from one or more subsequent executions of the physical process.
13. The method according to claim 1, wherein Identifying the outliers includes at least one of a past-performed backward analysis and the one or more subsequently-performed forward projections.
14. The method according to claim 1, wherein Calculating an anomaly score for a given variable for a given execution of the physical process includes calculating the minimum of the absolute values of: a difference between a value of the given variable for the given execution of the physical process and a value of the given variable for an execution of the physical process immediately preceding the given execution; as well as The difference between the value of the given variable for the given execution and the value of the given variable for an execution of the physical process that directly follows the given execution.
15. The method according to claim 1, wherein For a given execution of the physical process, calculating an anomaly score for the given variable for the given execution includes calculating a difference between a value of the given variable for the given execution and an average value of the given variable for the one or more subsequent executions of the physical process.
16. The method according to claim 1, wherein Determining the at least one GMM includes determining a multi-model graphical model (MGM) at least in part by performing inverse covariance update (ICU), time-ordered clustering (TOC), and a sparse weight selection algorithm (SWSA).
17. An apparatus for improving at least one of the quality and yield of a physical process, the apparatus comprising: Memory; as well as At least one processor, coupled to the memory, the processor being operable to: obtaining values of a plurality of variables associated with a physical process from respective executions of the physical process, wherein the physical process comprises at least a portion of a semiconductor manufacturing process; determining at least one Gaussian mixture model (GMM) of the values of the plurality of variables representing performance of the physical process; calculating an anomaly score for each of the plurality of variables for at least one of the executions of the physical process based at least in part on the at least one Gaussian mixture model; identifying the at least one of the executions of the physical process as an outlier based on at least one anomaly score of at least one variable of the plurality of variables exceeding a specified threshold for the at least one of the executions of the physical process; and Based at least in part on the identification of the outlier, the at least one variable of the plurality of variables is modified for one or more subsequent executions of the physical process to improve the at least one of quality and yield of the physical process.
18. A computer program product for improving at least one of quality and yield of a physical process, the computer program product comprising: A computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
State monitoring and fault diagnosis method for large-scale semiconductor manufacture process
CN102361014A
Deep-learning-based method for intermittent process fault detection
CN109116834A
Clustering based continuous performance prediction and monitoring for semiconductor manufacturing processes using nonparametric bayesian models
US20150012250A1