Prediction of Device Failure Modes from Process Traces

A predictive model using multivariate analysis and machine learning addresses the challenge of identifying device failure modes in semiconductor manufacturing by accurately classifying anomalies and providing corrective actions, enhancing process stability.

JP7703011B2Active Publication Date: 2025-07-04PDF SOLUTIONS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023504511
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-23
Filing Date
2021-07-22
Publication Date
2025-07-04
Estimated Expiration
2041-07-22

AI Technical Summary

Technical Problem

Existing semiconductor manufacturing processes face challenges in identifying the root cause of device failures due to insufficient multivariate analysis of sensor trace data, leading to inadequate detection and classification of device failure modes.

Method used

A predictive model using multivariate analysis and machine learning to detect anomalies in sensor traces, identify important features, and search a database of past data for similar anomalies, enabling classification and retrieval of root causes and corrective actions.

Benefits of technology

Enhances the capability to identify and address device failure modes by accurately classifying anomalies and providing corrective actions based on past data, improving process stability and reducing defective wafers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703011000001
    Figure 0007703011000001
  • Figure 0007703011000002
    Figure 0007703011000002
  • Figure 0007703011000003
    Figure 0007703011000003
Patent Text Reader

Abstract

A predictive model for equipment failure modes. An anomaly is detected in a collection of trace data, then key features are calculated. A search is performed for the same or similar anomaly with the same key features in a database of historical trace data. If the same anomaly has occurred previously and is in the database, the type of anomaly, its root cause, and action steps to correct it can be retrieved from the database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the use of process trace analysis for the detection and classification of semiconductor device failures, and more particularly to a machine-based method for predicting device failure modes.

Background Art

[0002] Effective process control for semiconductor manufacturing applications is important for improving reliability and reducing on-site failures. One approach to process control is fault detection and classification ("FDC"), which focuses on monitoring thousands of device sensors installed in process equipment as a means to quickly identify and correct process instabilities. However, one of the major challenges in the use of FDC techniques to drive a quick response to device problems is the identification of the root cause for detected process trace anomalies.

[0003] The detection of device failures by monitoring the time-series traces of device sensors has long been recognized but is a very difficult problem in semiconductor manufacturing. Typically, FDC methods begin by splitting complex traces into logical "windows" and then calculating statistics (often called metrics or key numbers) for the trace data within the windows. Metrics can be monitored using statistical process control ("SPC") techniques, mainly based on engineering knowledge, to identify anomalies, and the metrics can be used as inputs for predictive models and root cause analysis. The quality of the metrics determines the value of all subsequent analyses. High-quality metrics require high-quality windows. However, the analysis of metrics for anomaly detection is still mainly univariate, where anomalies are considered feature by feature, and is generally insufficient for identifying the device failure modes associated with the detected anomalies.

[0004] Therefore, it would be desirable to improve the capabilities of anomaly detection systems to identify device failure modes, for example, through multivariate analysis of trace data.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0006] As used herein, the term "sensor trace" refers to time-series data that periodically measures important physical quantities, such as sample values of physical sensors at each point in time, during the operation of a semiconductor processing apparatus. The sampling rate can vary, and the period between samples is not necessarily the same. The term "trace" or "device trace" refers to a set of sensor traces for all important sensors identified for a particular processing instance. The term "step" refers to a separate device processing period, e.g., one of the steps in a process recipe.

[0007] Disclosed herein is a predictive model for device failure modes. The model detects and identifies current anomalies in the trace data, calculates important features associated with the current anomalies, and searches a database of past trace data for anomalies having those important features. If the same or similar anomalies are found in the past trace data, a likelihood can be determined as to whether the current anomaly can be accurately classified according to those past anomalies, e.g., whether the current anomaly most closely resembles a previous anomaly in the past trace data. If so, the type of anomaly, its root cause, and the corrective action steps can likely be obtained from the database of past trace data. However, if not, the model returns an error, meaning that the anomaly has not been seen before. The anomaly and its features are nonetheless stored for future reference, and the database is updated if the root cause and corrective action are later determined.

[0008] Referring to FIG. 1, an exemplary graph 100 of trace data is shown, representing approximately 200 individual traces, i.e., time series values obtained from individual sensors taken during individual discrete steps of a semiconductor manufacturing process run for semiconductor wafer fabrication. Sensor values are plotted on the y-axis and time is measured in seconds on the x-axis. It should be recognized that process steps typically begin at a particular point in time, but the length of a process step can vary.

[0009] For most of the period of approximately 40 to 90 seconds, the normal process operation is expected to gradually decline, thus resulting in trace data that is relatively stable and consistent. However, in this case, between approximately 45 and 60 seconds, both the first set of traces 112 within the upper group 110 of the trace and the second set of traces 122 within the lower set 120 of the trace show sensor measurement values that suddenly spike, then decline, then rise back up, and then return to a gentle downward pattern. This trace behavior is unexpected and indicates some kind of problem related to the process. Therefore, in order to analyze the abnormal behavior, windows 115 and 125 are defined over these type I abnormal regions within the upper group 110 and lower group 120, respectively, of graph 100 where unexpected type I anomalies occur for some wafers.

[0010] Figure 2 shows the same graph 100 of the trace data, but focuses on a different portion of the trace data where type II anomalies occur only for the third set of traces 132 within the upper group 110. In this case, a portion of the trace appears to decline early, then recover, then decline again, and then recover again before dropping to nominal as expected. Therefore, in order to analyze this type of abnormal behavior, window 135 is defined over the first downward region within the upper group 110 of the trace, and window 136 is defined over the second downward region within the upper group.

[0011] Typically, technical staff manually establish a window for analyzing a specific region of trace data based simply on a visual review of the graph results to (i) determine where the trace data is consistent and / or (ii) roughly manually define a window for stable process operation where the rate of change is the same. Regions where the trace data changes abruptly in value or rate of change are considered transition windows and are generally located between a pair of stable windows. However, anomalies such as Type I and Type II anomalies described herein as examples can appear within otherwise normal stable windows of the trace data, as illustrated in FIGS. 1 and 2, and can be selected for processing and analysis through windowing of the relevant region(s).

[0012] A machine learning model is configured to detect anomalies using known methods that include the use of data from window analysis. For example, a combination of wafer attributes and trace position features can be provided as input to a simple multiclass machine learning model, such as a gradient boosting model, that is trained with respect to a dataset to detect anomalous behavior within the trace data. However, once an anomaly is detected, it is important to know whether the same anomaly has occurred previously and, if so, what caused it and what action steps should be taken to correct the problem.

[0013] After the definition of anomaly windows 115, 125, 135, and 136, metrics are calculated from the traces within each of the windows. The metrics are then stored, along with the selected wafer attributes and the location of the anomaly within the trace, as instances of the features associated with the windows and the trace data for those wafers. Feature engineering and selection can be performed to narrow the set of features to the important features determined to be most important for detecting and discriminating specific anomalies using the detection model.

[0014] Regarding Type I and Type II anomalies, respectively illustrated in FIGS. 1 and 2, the classifications predicted from the detection model (including normal wafers) are summarized in Table 200 of FIG. 3, where 5 wafers have anomalies identified as Type I anomalies, 7 wafers have anomalies identified as Type II anomalies, and 1 wafer has anomalies of both Type I and Type II. Despite the small number of detected anomalies, it is important to identify and characterize the anomalies for use in the training of the prediction model, particularly for monitoring trace data to minimize instances of process instability that could result in defective wafers.

[0015] For each type of anomaly, as important features as inputs to the model, the model is configured to (i) search a database of previous trace data for the same or similar anomalies and (ii) identify one or more previous anomalies as being most similar to the current anomaly or indicate that there are no anomalies like the current one in the database.

[0016] If the same or a similar anomaly is found in the database of past trace data, the root cause and the corrective steps taken to correct the anomalous behavior are likely also stored in the database and can be retrieved for comparison with the current anomaly. By comparing the characteristics and patterns of the anomalies, the model determines the likelihood that the current trace anomaly is most similar to one or more similar or identical anomalies observed in past traces. If the likelihood exceeds a threshold, the anomaly is classified and previous knowledge regarding the root cause and corrective actions is retrieved from the database.

[0017] This process is summarized in FIG. 4 using a chart, which illustrates one embodiment of process 300 for identifying root causes for anomalies. At step 302, trace data is received and processed within a prediction model. At step 304, at least one anomaly is detected within the received trace data and its location within the trace is identified. Then at step 306, a window is defined to include the portion of the trace containing the anomaly, and at step 308, features of the anomalous trace, including statistical measures, are calculated and stored at step 310. Then, a search is performed at step 312 within a database having past trace data for anomalies having the same features associated with the current anomaly. If at step 314 the same or similar anomalies are found within the past trace data, a likelihood can be determined at step 316 as to whether the current anomaly can be accurately classified according to those past anomalies. If so, at step 318, the type of anomaly, its root cause, and the corrective action steps can be retrieved from the database for the same or similar past occurrences, and appropriate corrective action is taken at step 320. However, if the likelihood is low, the model returns an error, meaning that the anomaly has not been seen before. The anomaly and its features are nevertheless stored for future reference, and the database is updated if the root cause and corrective action are determined subsequently.

[0018] FIG. 5 presents a more generalized approach in process 400. At step 402, trace data is received and processed within a prediction model. At step 404, an anomalous pattern is detected within the trace data. At step 406, features of the detected anomalous pattern are calculated and compared at step 408 to the features of previous anomalous patterns stored within a database of past trace data. At step 410, it is determined whether the features match, and then at step 412, information regarding the anomalous pattern from the past trace data is retrieved from the database, including one or more root causes for the anomaly and corrective actions for the root causes. At step 414, appropriate corrective action is taken.

[0019] The multivariate analysis of trace data has been facilitated by the emergence of parallel processing architectures and the development of machine learning algorithms, which enable users to use vast amounts of data at high speed to gain insights and make predictions, making such an approach appropriate and realistic. Machine learning is a field of artificial intelligence involving the construction and study of modeled systems that can learn from data. These types of ML algorithms, along with parallel processing capabilities, enable much larger datasets to be processed and are much better suited for involvement in multivariate analysis. Furthermore, effective machine learning approaches for abnormal trace detection and classification should facilitate active learning and continuously improve the accuracy of both fault detection and classification using the information obtained.

[0020] The creation and use of a processor-based model for trace analysis can be done on a desktop-based, i.e., stand-alone, or as part of a network system, but given the large amount of information to be processed and displayed with some interactivity, the processor capabilities (CPU, RAM, etc.) should be of current state-of-the-art technology to maximize effectiveness. In a semiconductor manufacturing plant environment, the Exensio® analysis platform is a useful option for constructing an interactive GUI template. In one embodiment, the coding of the processing routine can be done using Spotfire® analysis software version 7.11 or higher, which is compatible with the Python object-oriented programming language, which is mainly used for coding machine language models.

[0021] The foregoing description has been presented for purposes of illustration only - neither exhaustive nor intended to limit the present disclosure to the precise form described. Many modifications and variations are possible in light of the above teachings.

Claims

1. A method comprising: Receiving device trace data from a plurality of semiconductor device sensors into a computer-based machine learning model during a plurality of steps in a semiconductor process; Detecting, by the machine learning model, a first anomaly in the trace data, the first anomaly having a relevant location in the device trace data; Defining, by the machine learning model, a window that includes a period of device trace data that includes the first anomaly; Calculating, by the machine learning model, a statistic for a period of device trace data within the window; Storing, in memory, the statistic and relevant location of the first anomaly as a plurality of important features associated with the first anomaly; Searching a database of past trace data by providing the plurality of important features of the first anomaly as an input to a machine learning model configured to find similar past trace data having the plurality of important features; Determining, by the machine learning model, that an instance of the past trace data has the plurality of important features of the first anomaly; Identifying, by the machine learning model, a root cause for an instance of the past trace data having the plurality of important features of the first anomaly; Taking an action to correct the root cause in the semiconductor process; A method comprising the above.

2. Further comprising obtaining the root cause and a corrective action for the root cause from the database The method according to claim 1.

3. The determination step comprises: Determining, by the machine learning model, the likelihood that an instance of past trace data in the database has the plurality of important features of the first anomaly; Obtaining the root cause when the likelihood exceeds a threshold; The method according to claim 1.

4. The step of obtaining the root cause comprises: Further comprising obtaining a corrective action for the root cause from the database The method according to claim 3.

5. A method for predicting a failure of a semiconductor processing apparatus, comprising: By a processor including a machine learning model trained to detect anomalies within a set of trace data using multivariate analysis, during a plurality of steps in a semiconductor process, detecting a first abnormal pattern within a first set of traces obtained from a plurality of semiconductor device sensors; Identifying, by the processor, a period window within the data of the first set that includes the first abnormal pattern; Calculating, by the processor, a plurality of features from the first set of traces located within the window using multivariate analysis; Searching, by the processor, a database of past trace data; Identifying, by the processor, at least one set of past trace data within the database that has the plurality of features within a related abnormal pattern; Determining, by the processor, the likelihood that the related abnormal pattern of the at least one set of past trace data is the same as the first abnormal pattern using multivariate analysis of the plurality of features within the related abnormal pattern; Obtaining, by the processor, from the database a root cause for at least one previous abnormal pattern when the likelihood exceeds a threshold; Taking, in the semiconductor process, measures to correct the root cause; A method comprising. **Claim 6**: The step of identifying the period window further includes defining the period window as an area within the first set of traces where the values change rapidly. The method according to claim 5, further comprising. **Claim 7**: The step of identifying the period window further includes defining the period window as an area within the first set of traces where the rate of change of the values changes rapidly. The method according to claim 5, further comprising.

Citation Information

Patent Citations

  • Fault management device and fault management method

    JP2006099249A

  • Integrated monitoring operation system and method

    JP2017207894A

  • Semiconductor producing device, failure prediction method of semiconductor producing device, and failure prediction program of semiconductor producing device

    JP2018178157A

  • Long short-term memory anomaly detection for multi-sensor equipment monitoring

    US20200104639A1