A self-supervised enhanced semi-supervised wafer failure prediction method and system

Through self-supervised mask pre-training and data enhancement technology, wafer failure prediction is used to use labelless data and a small amount of labeled data to solve the dependence on a large number of labeled data in the prior art, and high accuracy prediction is achieved when data availability is limited.

CN119557571BActive Publication Date: 2025-05-20ZHEJIANG UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510132618.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-20
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The prior art relies on a large amount of label data in wafer failure prediction, and it is difficult to effectively utilize labelless data, resulting in insufficient prediction accuracy when data availability is limited.

Method used

A self-supervised and enhanced semi-supervised wafer failure prediction method is proposed. Through self-supervised mask pre-training and data enhancement technology, it uses labelless data and a small amount of labeled data for training to improve the prediction accuracy.

Benefits of technology

In the case of limited data availability, the accuracy and robustness of wafer failure prediction are improved, resource waste and cost are reduced, and the learning ability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557571B_ABST
    Figure CN119557571B_ABST
Patent Text Reader

Abstract

The present invention discloses a self-supervised enhanced semi-supervised wafer failure prediction method and system, belonging to the field of wafer failure detection. Obtain wafer acceptability test / probe test data to construct a training set; perform self-supervised pre-training on a wafer feature encoder; use the pre-trained wafer feature encoder to generate the original wafer coding features of each sample, and add strong / weak noise respectively, introduce data enhancement constraints when adding strong noise and calculate the loss; use the wafer coding features before and after adding strong / weak noise to predict whether the wafer fails respectively using a wafer failure classifier, and calculate the consistency loss; in addition, it is necessary to calculate the supervision loss of a small amount of labeled data; semi-supervised training is performed on the pre-trained wafer feature encoder and the wafer failure classifier by combining the above losses, and the model after semi-supervised training is used for actual wafer failure detection. The present invention makes full use of a large amount of wafer unlabeled data and a small amount of wafer labeled data to improve the prediction accuracy of wafer failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wafer failure detection, and particularly relates to a semi-supervised wafer failure prediction method and system with enhanced self-supervision. Background Art

[0002] The manufacturing of integrated circuits (ICs) involves complex processes, including thin film deposition, ion implantation, etching, and meticulous polishing. As the industry moves towards miniaturization, circuit components shrink, and designs become increasingly layered, subjecting each chip to a series of extensive processing steps. In this intricate choreography, even the slightest deviation can cause defects in the wafer manufacturing process, leading to wafer failure. A series of tests are conducted during and after manufacturing to ensure the proper functioning of the chips, such as wafer acceptability tests and wafer probe tests. However, during wafer acceptability tests and wafer probe tests, there is a time lag (1 - 2 weeks), and during this period, the wafer fab does not stop. Only after the results of the wafer probe tests are available can it be determined whether there are wafer failures in the manufacturing process. If there were problems in the previous batch of manufacturing, it would result in a large number of wafer failures, leading to significant waste of time and money, potentially up to millions. Therefore, timely prediction of wafer failure problems is crucial.

[0003] The cornerstone of wafer failure prediction has always been machine learning-centered model construction. For example, some studies have adopted methods such as logistic regression, XGBoost, and RF. However, although machine learning models have good effects in most industrial scenarios, they require a large amount of data sets as support. For wafer failure prediction, one sample matches one wafer, and the manufacturing of one wafer requires hundreds or thousands of steps, with very high manufacturing costs. Collecting a large amount of data sets means high costs. At the same time, there are many wafers waiting for testing in wafer manufacturing, which meet the criteria of unlabeled data, and machine learning techniques cannot use this unlabeled data for learning. In recent years, inspired by the success of self-supervised and semi-supervised learning in the fields of images and language tasks, some recent studies have explored the application of self-supervised and semi-supervised learning in tasks in the field of tabular data, but there is no relevant research in the scenario of wafer failure prediction. For example, the self-supervised and semi-supervised framework VIME introduced in 2020 proposed to use unlabeled data in the field of tabular data to improve the performance of the model. However, despite the breakthrough, the VIME model still faces limitations in the industrial environment, mainly because its simple network structure is difficult to adapt to the complexity of industrial data. Specifically, VIME uses a simple fully connected layer as the encoder, which is difficult to extract the relationships between complex data sets, resulting in a decline in network performance. In addition, the performance of the VIME model under small sample conditions is limited, which is related to the use of many perturbations for consistency in the semi-supervised training process of the VIME model. Using masked data for perturbations in small samples makes the model's learning of the relationships between data more blurred and the model performance declines. In addition, many other studies have put forward their own views on the methods of self-supervised training and fine-tuning in the field of tabular data, mainly migrating the self-supervised training and pre-training methods in the field of images to the field of tabular data, such as models like SAINT, SCARF, and Subtab. However, the application of these models to unlabeled data is insufficient, resulting in limited model performance and poor performance on industrial data sets. Generally speaking, the above methods are all aimed at prediction tasks for data such as housing prices, incomes, and stock markets. Currently, there is no excellent self-supervised and semi-supervised model for industrial data sets that can efficiently extract the complex correlations between industrial data sets.

[0004] Facing these obstacles, there is an urgent need for a method that not only needs to extract and identify the interrelationships between the data in the manufacturing process with higher accuracy but also needs to provide robust wafer failure prediction results under limited data availability. In addition, in order to use the data from failure tests to strengthen the performance of the model and improve the prediction accuracy, a new model for the field of wafer failure prediction must be developed. Summary of the Invention

[0005] In order to achieve wafer failure prediction using unlabeled data and a small amount of labeled data, the present invention proposes a self-supervised enhanced semi-supervised wafer failure prediction method and system, which alleviates the data volume requirements of the machine learning model and improves the accuracy of wafer failure prediction by using a large amount of unlabeled data and a small amount of labeled data tested by probes during the wafer manufacturing process.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention proposes a self-supervised enhanced semi-supervised wafer failure prediction method, comprising:

[0008] Obtain wafer acceptability test data and probe test data, preprocess and match the data, and obtain labeled wafer samples and unlabeled wafer samples as training sets;

[0009] Use the training set to perform self-supervised mask pre-training on the wafer feature encoder to obtain a pre-trained wafer feature encoder;

[0010] Use the pre-trained wafer feature encoder to generate the original wafer coding features of each sample in the training set, add strong noise and weak noise to the original wafer coding features respectively; introduce data enhancement constraints when adding strong noise to the original wafer coding features, and calculate the data enhancement constraint loss;

[0011] The original wafer coding features, the strongly enhanced wafer coding features after adding strong noise, and the weakly enhanced wafer manufacturing acceptability data coding features after adding weak noise are respectively used by the wafer failure classifier to predict whether the wafer fails, and the consistency loss is calculated based on the failure prediction results; in addition, the supervision loss needs to be calculated for the labeled wafer samples in the training set based on the failure prediction results;

[0012] Combining the data enhancement constraint loss, consistency loss and supervision loss, the pre-trained wafer feature encoder and wafer failure classifier are semi-supervised trained, and the pre-trained wafer feature encoder and wafer failure classifier after semi-supervised training are used to predict whether the wafer fails.

[0013] Furthermore, the preprocessing and matching of data includes:

[0014] Handle outliers and missing values ​​for wafer acceptability test data;

[0015] Match the pre-processed wafer acceptability test data and probe test data, and use the successfully matched data as labeled wafer samples, and the unmatched wafer acceptability test data as unlabeled wafer samples, and remove the unmatched probe test data;

[0016] The label of the labeled wafer sample is the result of probe test data.

[0017] Furthermore, the self-supervised masked pre-training of the wafer feature encoder using the training set includes:

[0018] Randomly mask the samples in the training set;

[0019] The masked samples use the wafer feature encoder to generate wafer encoded features. Based on the wafer encoded features, the masked positions and masked values are predicted respectively, and the parameters of the wafer feature encoder are iteratively updated according to the prediction results to complete the self-supervised pre-training.

[0020] Furthermore, the weak noise added to the original wafer encoded features is weak Gaussian noise.

[0021] Furthermore, the method of adding strong noise to the original wafer encoded features includes:

[0022] Input the original wafer encoded features of the samples into the initialized strong noise decoder, and generate strong enhanced noise samples after decoding;

[0023] Use the strong enhanced noise samples to generate strong enhanced wafer encoded features by the pre-trained wafer feature encoder, so as to add strong noise to the original wafer encoded features;

[0024] The strong noise decoder participates in the semi-supervised training together with the pre-trained wafer feature encoder and the wafer failure classifier.

[0025] Furthermore, the data augmentation constraint loss is as follows:

[0026]

[0027] Among them, represents the data augmentation constraint loss, represents the original sample i in the training set, represents the strong enhanced noise sample, represents the square of the L2 norm, represents the dimension of the encoded features generated by the pre-trained wafer feature encoder, respectively represent the variance and mean of the j-th dimension of the encoded features generated by the original sample i through the pre-trained wafer feature encoder, represents a preset hyperparameter, represents the number of original samples in the training set.

[0028] Further, the supervision loss refers to the binary cross-entropy loss between the predicted result of the wafer failure classifier using the original wafer coding features and the true label; the consistency loss refers to the binary cross-entropy loss between the predicted results of the strongly augmented wafer coding features and the weakly augmented wafer coding features using the wafer failure classifier respectively.

[0029] Further, when jointly performing semi-supervised training on the pre-trained wafer feature encoder and the wafer failure classifier with the data augmentation constraint loss, the consistency loss, and the supervision loss, the weight of the consistency loss gradually increases from 0 to a preset weight.

[0030] In a second aspect, the present invention proposes a self-supervised enhanced semi-supervised wafer failure prediction system for implementing the above-mentioned self-supervised enhanced semi-supervised wafer failure prediction method.

[0031] The beneficial effects of the present invention are as follows:

[0032] The present invention proposes a self-supervised enhanced semi-supervised model framework for wafer failure prediction, which combines self-supervised and semi-supervised training processes to model the wafer acceptability test data in the wafer manufacturing process. By calculating the attention weights between data through self-supervised training, the learning ability of the semi-supervised process is enhanced. Further, data augmentation methods with strong noise and weak noise are adopted and introduced into the training of the semi-supervised process through the consistency principle between pseudo-labels. Important and effective data expansion is achieved with very little labor cost, and the semi-supervised training process is enhanced through self-supervised training, further improving the performance of the model in predicting wafer failures.

[0033] The present invention makes full use of the wafer data that has passed the probe test and the wafer data that has not passed the probe test, avoiding data waste, while improving the stability of the model, reducing resource waste in the manufacturing process, saving costs, and helping to accelerate the root cause analysis of wafer failures and improve the yield. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flowchart of wafer failure prediction data collection shown in an embodiment of the present invention;

[0035] Figure 2 is a flowchart of self-supervised training;

[0036] Figure 3 is a flowchart of semi-supervised training;

[0037] Figure 4 is a schematic diagram of semi-supervised application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention.

[0039] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0040] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.

[0041] A self-supervised enhanced semi-supervised wafer failure prediction method proposed by the present invention mainly includes the following steps:

[0042] S1. Wafer data collection.

[0043] Collect relevant test data in the wafer manufacturing process and perform preprocessing. The wafer manufacturing process and its relevant test data include wafer acceptability test data and wafer probe test data; the wafer probe test data refers to the test results of whether the wafer fails.

[0044] The methods of data preprocessing include but are not limited to data merging, data cleaning, outlier processing, missing value processing, etc. The preprocessing process can be implemented according to the existing technologies in this field and will not be elaborated here.

[0045] In a specific implementation of the present invention, such as Figure 1As shown, there are hundreds of steps in the wafer manufacturing process, such as chemical vapor deposition process, etching process, photolithography process, chemical mechanical polishing process, cleaning process, ion implantation process, etc. To ensure the stability of the process and the integrity of the chip function, it is necessary to perform acceptability testing and probe testing on the manufactured wafers. Wafer Acceptance Testing (WAT) is a key quality control step in the semiconductor manufacturing process. It is used to evaluate the performance and quality of wafers during the manufacturing process. The purpose of WAT is to detect the electrical characteristics of wafers at different manufacturing stages to ensure that they meet the design specifications and quality standards. Usually includes resistance, capacitance, leakage current, threshold voltage, etc. These parameters help identify defects in the manufacturing process. Circuit Probing Testing is to measure the electrical performance of each chip through a probe station and an automatic test equipment during the semiconductor manufacturing process to screen out unqualified chips. It evaluates parameters such as voltage, current, and logic function, generates a yield report, and provides feedback for process optimization, thereby improving product quality and production efficiency. By collecting wafer acceptability test data and probe test data and preprocessing the data.

[0046] Match the preprocessed data. The wafer acceptability test data and probe test data that match successfully are used as labeled wafer samples. The label of the labeled wafer sample is the result of the probe test data, including failure and normal. The unmatched wafer acceptability test data is used as unlabeled wafer samples. Here, "matching successfully" means that both the wafer acceptability test data and the probe test data are collected for the same wafer.

[0047] S2. Construct a self-supervised enhanced semi-supervised wafer failure prediction model.

[0048] The self-supervised enhanced semi-supervised wafer failure prediction model consists of a self-supervised network and a semi-supervised network.

[0049] In a specific implementation of the present invention, the self-supervised network includes a mask generator, a wafer feature encoder, and a self-supervised prediction layer. The wafer feature encoder uses an FTTransformer encoder, which integrates the Transformer function and includes an attention mechanism. The mask generator is used to randomly mask the samples in the training set to generate proxy tasks. The prediction layer is composed of two different fully connected layers, which are used as a mask position predictor and a mask value predictor respectively for two self-supervised proxy tasks.

[0050] The semi-supervised network described above includes a weak noise generator, a wafer feature encoder, and a semi-supervised prediction layer. The wafer feature encoder is the same as the one in the self-supervised network and directly loads the model parameters trained by self-supervision. The semi-supervised prediction layer consists of two different fully connected layers, serving as a strong noise decoder and a wafer failure classifier respectively, which are used to add strong noise to the original wafer encoded features generated by the wafer feature encoder during semi-supervised training and to predict whether the wafer fails. The weak noise generator is used to add weak noise (weak Gaussian noise is adopted in this embodiment) to the original wafer encoded features generated by the wafer feature encoder during semi-supervised training.

[0051] The FTTransformer encoder adopted in the present invention includes a feature tokenizer, a multi-head self-attention calculation unit, an adder, a normalization unit, a residual network, a fully connected layer, and an activation unit, etc. Among them, the feature tokenizer is responsible for encoding the input wafer data to facilitate the multi-head self-attention calculation unit to calculate the attention weights between the data. Through the attention weights, the interaction information between the data can be effectively extracted, enhancing the robustness of the model. The residual network and the adder are used to retain the original information of the data to avoid information loss after passing through multiple layers of the network. The normalization unit is responsible for normalizing the data to improve the stability and generalization ability of the model and avoid the phenomenon of gradient explosion or vanishing gradient caused by the data being too large or too small. The activation function introduces non-linearity to increase the expressive power of the network. Generally, common activation functions include ReLU, Sigmoid, Tanh, etc. The fully connected layer flattens the extracted features and then connects them to the output layer to generate wafer encoded features. The FTTransformer encoder uses a residual network to add skip connections on the basis of the original network, which can solve the problems of gradient explosion and vanishing gradient caused by the increase in network depth. The specific calculation process of the FTTransformer encoder belongs to the common knowledge in the art and will not be elaborated here.

[0052] S3. Self-supervised pre-training.

[0053] The labeled wafer samples and unlabeled wafer samples together form a training set, and the wafer feature encoder is pre-trained by self-supervised masking using the training set to obtain a pre-trained wafer feature encoder.

[0054] In a specific implementation of the present invention, a mask m = [m1……md] ∈ {0, 1} is randomly sampled and generated by a mask generator B×D to mask the input sample data, where B is the batch size of the input samples and D is the dimension of the sample data. As Figure 2 shown, and Representing labeled wafer samples and unlabeled wafer samples respectively, the samples are masked by a mask generator, and there is only one masked position for each input sample. The masked samples are input into the wafer feature encoder for encoding to generate the wafer encoded features after masking.

[0055] Self-supervised pre-training uses two independent fully connected layers to implement the self-supervised pre-training task, and has activation functions and dropout layers. The self-supervised pre-training tasks include:

[0056] Task 1: Predicting the masked positions using the wafer encoded features after masking, which is embodied as a multi-class classification task, and using multi-class cross-entropy loss to calculate the loss function during the training process, denoted as :

[0057]

[0058] Among them, represents the masked position loss, N represents the number of samples, D represents the data dimension, represents the position label of whether the c-th dimension of the n-th sample is masked, represents that the c-th dimension of the n-th sample is masked, represents that the c-th dimension of the n-th sample is not masked; represents the predicted value of the c-th dimension of the n-th sample after masking.

[0059] Task 2: Predicting the numerical value of the masked position using the masked data, which is embodied as a regression task, and using the root mean square error as the loss function during the training process, denoted as :

[0060]

[0061] Among them, represents the masked value loss, represents the true value of the i-th sample being masked, represents the predicted value of the i-th sample being masked.

[0062] The total loss of the above self-supervised pre-training is defined as . The weight α between the masked position loss and the masked value loss can be set according to the experience of those skilled in the art.

[0063] Through the above self-supervised pre-training tasks, the wafer feature encoder can effectively learn the relationships and attention weights between the sample data, which helps the convergence of the semi-supervised learning process and the prediction performance. After the self-supervised pre-training is completed, the weight parameters of the wafer feature encoder are saved.

[0064] S4. Semi-supervised training and fine-tuning.

[0065] As Figure 3 shown, after the self-supervised pre-training is completed, first, the original wafer encoding features are calculated for the sample data in the training set using the pre-trained wafer feature encoder, and the original wafer encoding feature vector is denoted as , which can effectively extract the mutual relationships and global information between the sample data.

[0066] Secondly, strong noise and weak noise are respectively added to the original wafer encoding features through data augmentation, and a data augmentation constraint is introduced when adding strong noise to the original wafer encoding features, and the data augmentation constraint loss is calculated.

[0067] In a specific implementation of the present invention, the weak noise added to the original wafer encoding features is weak Gaussian noise, and the strong noise is added using a strong noise decoder, specifically:

[0068] The original wafer encoding features of the sample are input into the initialized strong noise decoder, and after decoding, a strongly augmented noise sample is generated;

[0069] The strongly augmented noise sample is used to generate strongly augmented wafer encoding features using the pre-trained wafer feature encoder, realizing the addition of strong noise to the original wafer encoding features;

[0070] The strong noise decoder participates in the semi-supervised training together with the pre-trained wafer feature encoder and the wafer failure classifier.

[0071] To ensure the stability of strong noise addition, the data augmentation constraint introduced when adding strong noise to the original wafer encoding features using the strong noise decoder is calculated using KL divergence, which can be expressed as:

[0072]

[0073] Among them, represents the data augmentation constraint loss, represents the original sample i in the training set, represents the strongly augmented noise sample, represents the square of the L2 norm, represents the dimension of the encoding features generated by the pre-trained wafer feature encoder, respectively represent the variance and mean of the j-th dimension of the encoding features generated by the original sample i passing through the pre-trained wafer feature encoder, represents a preset hyperparameter, represents the number of original samples in the training set.

[0074] After adding strong noise and weak noise to the original wafer coding features of each sample in the training set respectively, use the wafer failure classifier to predict whether the wafer fails for the original wafer coding features, the strongly enhanced wafer coding features after adding strong noise, and the weakly enhanced wafer coding features after adding weak noise respectively. Calculate the consistency loss and the supervised loss according to the failure prediction results, where the supervised loss needs to use the sample labels, that is, the said supervised loss is only for the labeled wafer samples in the training set.

[0075] In the semi-supervised training part of the present invention, consistency regularization is introduced, and consistency regularization calculation is performed on the wafer failure prediction results of the classifier for the strong noise data and the weak noise data, so that the prediction results are as consistent as possible. In this embodiment, binary cross-entropy loss is used for calculation, denoted as and its formula is defined as:

[0076]

[0077] where, is the predicted category corresponding to the weakly enhanced wafer coding feature, is the predicted category corresponding to the strongly enhanced wafer coding feature, and H is the binary cross-entropy loss function.

[0078] In order to make full use of the labels of the labeled wafer samples, the present invention introduces a supervised loss, that is, calculates the difference between the predicted value of the label data and the original label. In this embodiment, binary cross-entropy loss is used for calculation, denoted as and its calculation formula is:

[0079]

[0080] where, is the label category of the labeled wafer sample, is the predicted value output by the model, is the number of labeled wafer samples.

[0081] Jointly use the data augmentation constraint loss, consistency loss and supervised loss to perform semi-supervised training on the pre-trained wafer feature encoder and wafer failure classifier. The total loss of the above semi-supervised training is defined as L semi =L s +βL c +λL r where the value of β gradually rises from 0 to the specified value γ during the training process to avoid the instability of the classifier prediction value caused by the random weight initialization at the beginning of training.

[0082] S5. Use the pre-trained wafer feature encoder and wafer failure classifier after semi-supervised training to predict whether the wafer fails.

[0083] During the process of wafer failure prediction, wafer acceptability test data is collected, wafer coding features are generated using a wafer feature encoder, and then whether the wafer fails is predicted using a wafer failure classifier.

[0084] In a specific implementation of the present invention, in order to optimize the applicability of the trained wafer feature encoder and wafer failure classifier in the IC industrial environment, the model generated in the initial stage is used to predict new wafers, the prediction results are saved, the failure samples are collected to update the training set, and the wafer feature encoder and wafer failure classifier are regularly subjected to semi-supervised training based on self-supervised enhancement. By giving more labeled data and allowing the continuous expansion of the data set, the performance of the model can be continuously improved.

[0085] As Figure 4 shown, a complete process is as follows:

[0086] Obtain the initial labeled wafer samples and unlabeled wafer samples, train the model using the above self-supervised enhanced semi-supervised method and save it;

[0087] Collect the data to be detected and input it into the model to obtain the wafer failure prediction result, which can be used by engineers to further analyze this batch of wafers, accelerating the analysis of the root cause of wafer failure and improving the yield;

[0088] Use the online wafer failure prediction results for the expansion of the initial training set, and regularly perform semi-supervised training on the model based on self-supervised enhancement to continuously optimize the model.

[0089] In order to quantitatively evaluate the performance achieved by the present invention, a factory real data set is used for verification. 128 tested wafers are used as labeled data and more than 1000 untested wafers are used as unlabeled data. A comparison is made with the commonly used machine learning methods in wafer failure prediction. The results are shown in Table 1. The results show that the method proposed by the present invention has achieved the optimal performance in multiple indicators, and at the same time improved the performance of the basic encoder FTTransformer, proving the effectiveness of the method of the present invention.

[0090] Table 1 Comparison of the self-supervised enhanced semi-supervised method of the present invention with the existing fully supervised method

[0091]

[0092] In addition, Table 2 also shows the comparison of the method of the present invention with the current state-of-the-art self-supervised and semi-supervised models. The results prove that the method of the present invention is superior to all the state-of-the-art self-supervised and semi-supervised models in the industrial data set.

[0093] Table 2 Comparison of the self-supervised enhanced semi-supervised method of the present invention with the existing self-supervised and semi-supervised methods

[0094]

[0095] The present invention also provides a self-supervised enhanced semi-supervised wafer failure prediction system, which is used to implement the above embodiments. The terms "module", "unit", etc. used hereinafter can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0096] A self-supervised enhanced semi-supervised wafer failure prediction system provided in this embodiment includes:

[0097] A data acquisition module, which is used to obtain wafer acceptability test data and probe test data during the training phase, preprocess and match the two parts of data to obtain labeled wafer samples and unlabeled wafer samples as a training set; and, to collect wafer acceptability test data to be predicted during the prediction phase;

[0098] A self-supervised training module, which is used to perform self-supervised masked pre-training on the wafer feature encoder using the training set to obtain a pre-trained wafer feature encoder;

[0099] A semi-supervised training module, which is used to generate the original wafer encoding features of each sample in the training set using the pre-trained wafer feature encoder, and add strong noise and weak noise to the original wafer encoding features respectively; when adding strong noise to the original wafer encoding features, introduce data augmentation constraints and calculate the data augmentation constraint loss;

[0100] Use the original wafer encoding features, the strongly augmented wafer encoding features after adding strong noise, and the weakly augmented wafer encoding features after adding weak noise to predict whether the wafer fails using the wafer failure classifier, and calculate the consistency loss and the supervision loss according to the failure prediction results; the supervision loss is only for the labeled wafer samples in the training set;

[0101] Jointly perform semi-supervised training on the pre-trained wafer feature encoder and the wafer failure classifier using the data augmentation constraint loss, the consistency loss, and the supervision loss;

[0102] A wafer failure prediction module, which is used to predict whether the wafer fails using the pre-trained wafer feature encoder and the wafer failure classifier after semi-supervised training.

[0103] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments, and the implementation methods of the modules will not be elaborated here. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.

[0104] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0105] Obviously, the above-described embodiments and drawings are only some examples of the present application. For a person of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it can be understood that although the work done during the development process may be complex and time-consuming, for a person of ordinary skill in the art, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application. Without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A self-supervised enhanced semi-supervised wafer failure prediction method, characterized in that: include: Obtain wafer acceptability test data and probe test data, preprocess and match the data, and obtain labeled wafer samples and unlabeled wafer samples as training sets; Using the training set to perform self-supervised mask pre-training on the wafer feature encoder to obtain a pre-trained wafer feature encoder; The pre-trained wafer feature encoder is used to generate the original wafer coding features of each sample in the training set, and strong noise and weak noise are added to the original wafer coding features respectively; when adding strong noise to the original wafer coding features, a data enhancement constraint is introduced, and the data enhancement constraint loss is calculated; The data enhancement constraint loss is as follows: Among them, L r represents the data augmentation constraint loss, x i Represents the original sample i, x in the training set g represents a strongly enhanced noise sample, represents the square of the L2 norm, l represents the encoding feature dimension generated by the pre-trained wafer feature encoder, They respectively represent the variance and mean of the j-th dimension of the encoded feature generated by the pre-trained wafer feature encoder for the original sample i, α and β represent the preset hyperparameters, and N represents the number of original samples in the training set; The original wafer coding features, the strongly enhanced wafer coding features after adding strong noise, and the weakly enhanced wafer coding features after adding weak noise are respectively used by the wafer failure classifier to predict whether the wafer fails, and the consistency loss and the supervision loss are calculated according to the failure prediction results; the supervision loss is only for the labeled wafer samples in the training set; The data enhancement constraint loss, consistency loss and supervision loss are combined to perform semi-supervised training on the pre-trained wafer feature encoder and wafer failure classifier, and the pre-trained wafer feature encoder and wafer failure classifier after the semi-supervised training are used to predict whether the wafer fails.

2. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: The preprocessing and matching of data includes: Handle outliers and missing values ​​on wafer acceptability test data; Match the pre-processed wafer acceptability test data and probe test data, use a group of successfully matched data as labeled wafer samples, use unsuccessfully matched wafer acceptability test data as unlabeled wafer samples, and discard unsuccessfully matched probe test data; The labels of the labeled wafer samples are the results of probe test data, including failure and normal.

3. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: The method of using the training set to perform self-supervisory mask pre-training on the wafer feature encoder includes: Randomly mask the samples in the training set; The masked samples use the wafer feature encoder to generate wafer coding features, and the mask position and mask value are predicted based on the wafer coding features. The wafer feature encoder parameters are iteratively updated according to the prediction results to complete self-supervised pre-training.

4. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: The weak noise added to the original wafer coding feature is weak Gaussian noise.

5. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: Methods to add strong noise to the original wafer code features include: Inputting the original wafer coding feature of the sample into the initialized strong noise decoder, and generating a strong enhanced noise sample after decoding; The strongly enhanced noise samples are used to generate strongly enhanced wafer coding features using a pre-trained wafer feature encoder, thereby adding strong noise to the original wafer coding features; The strong noise decoder participates in semi-supervised training together with the pre-trained wafer feature encoder and the wafer failure classifier.

6. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: The supervised loss refers to the binary cross entropy loss between the original wafer coding features using the prediction results of the wafer failure classifier and the true label; the consistency loss refers to the binary cross entropy loss between the prediction results of the strongly enhanced wafer coding features and the weakly enhanced wafer coding features using the wafer failure classifier respectively.

7. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: When the data enhancement constraint loss, consistency loss and supervision loss are combined to perform semi-supervised training on the pre-trained wafer feature encoder and wafer failure classifier, the weight of the consistency loss is gradually increased from 0 to a preset weight.

8. The self-supervised enhanced semi-supervised wafer failure prediction method according to claim 1, characterized in that: Also includes: After predicting whether a wafer has failed using the wafer feature encoder and the wafer failure classifier, the prediction results are saved, failure samples are collected to update the training set, and semi-supervised training based on self-supervised reinforcement is performed regularly on the wafer feature encoder and the wafer failure classifier.

9. A self-supervised enhanced semi-supervised wafer failure prediction system, used to implement the self-supervised enhanced semi-supervised wafer failure prediction method according to any one of claims 1 to 8, characterized in that: The wafer failure prediction system comprises: A data acquisition module, which is used to obtain wafer acceptability test data and probe test data in the training phase, preprocess and match the two parts of data, and obtain labeled wafer samples and unlabeled wafer samples as training sets; and, used to collect wafer acceptability test data to be predicted in the prediction phase; A self-supervised training module, which is used to perform self-supervised mask pre-training on the wafer feature encoder using the training set to obtain a pre-trained wafer feature encoder; A semi-supervised training module is used to generate original wafer coding features of each sample in a training set using a pre-trained wafer feature encoder, and to add strong noise and weak noise to the original wafer coding features respectively; a data enhancement constraint is introduced when adding strong noise to the original wafer coding features, and a data enhancement constraint loss is calculated; The original wafer coding features, the strongly enhanced wafer coding features after adding strong noise, and the weakly enhanced wafer coding features after adding weak noise are respectively used by the wafer failure classifier to predict whether the wafer fails, and the consistency loss and the supervision loss are calculated according to the failure prediction results; the supervision loss is only for the labeled wafer samples in the training set; Semi-supervised training of the pre-trained wafer feature encoder and wafer failure classifier is performed by combining the data augmentation constraint loss, consistency loss and supervision loss; The wafer failure prediction module is used to predict whether a wafer fails by using a pre-trained wafer feature encoder and a wafer failure classifier after semi-supervised training.

Citation Information

Patent Citations

  • Self-supervised learning framework training method and device, equipment and storage medium

    CN116975617A

  • High-precision wafer defect detection method suitable for different density changes

    CN117036333A

  • Double-noise auto-encoder process fault classification method

    CN118940122A