Robust wafer defect detection method and system based on distillation learning

By extracting and purifying common features in distillation learning and processing abnormal samples during training, the problem of noise data affecting model performance is solved, and the robustness and accuracy of wafer defect detection are improved.

CN120163794APending Publication Date: 2025-06-17TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269636.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When the wafer defect detection method based on distillation learning exists noisy data in the training data, the model performance will degrade, resulting in inaccurate detection results.

Method used

By extracting common features and purifying them during distillation learning, the effects of non-common features and abnormal samples in model training are reduced. Specific methods include using a smoother and a weight counter to eliminate prominent independent features and identifying and reducing the training weight of the individual samples through an abnormal training sample removal module.

Benefits of technology

The robustness of the model and the reliability of the detection results are improved, and the detection accuracy can be maintained in the presence of noise in the training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163794A_ABST
    Figure CN120163794A_ABST
Patent Text Reader

Abstract

The invention provides a robust wafer defect detection method and system based on distillation learning, and relates to the technical field of wafer defect detection. According to the method and the device, aiming at the training samples in the distillation learning process, the examination is performed from the dimension of one batch, and the common characteristics in the plurality of training samples are extracted. The common features are common in the samples, can represent the features of the sample attributes, and are of great significance to learning and understanding of the model. Due to the fact that the existence proportion of the abnormal data in the training data is low, the abnormal features are often filtered out in the extraction process of the common features, and the effect of reducing the interference features is achieved. In the model training process, the model only receives supervision and guidance of the carefully extracted and purified common features through the design of the invention. Under the constraint of the supervision information, the model focuses on learning real valuable features, and disturbance introduced by noise samples is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of wafer defect detection, and particularly to a robust wafer defect detection method and system based on distillation learning. Background Art

[0002] In the semiconductor industry, industrial manufacturers usually use defect detection technology to identify defects in products. Defect detection solutions have two significant positive meanings. On the one hand, it can find non-compliant areas in the current product, provide sufficient guiding information for downstream processes, and avoid subsequent processing costs at the defect location. On the other hand, through the statistical data of defect detection, problems existing in the production process can be analyzed and discovered in a timely manner, and corresponding corrective measures can be taken, thereby fundamentally reducing the generation of defects and improving the yield and quality of the entire production line. For the defect detection of semiconductor compounds, especially the surface of wafers, target detection or unsupervised anomaly detection methods are usually adopted. Through deep learning technology, knowledge learning is carried out on samples to discover possible abnormal defects on the wafer surface.

[0003] Compared with the supervised learning-based target detection solution, unsupervised learning does not require prior knowledge of the category information of defect samples, has no high requirements for the quantity and balance of training samples, and is closer to the actual scenario of industrial quality inspection. Among them, the unsupervised anomaly detection model based on distillation learning has the characteristics of small volume and high efficiency, and is especially suitable for application in the quality inspection process of large-batch products such as wafer defect detection. The unsupervised anomaly detection method based on distillation can real-time identify positions on the wafer surface that are significantly different from normal samples, thereby preventing defective samples from flowing into the next process manufacturing link.

[0004] However, the current defect detection method based on the anomaly detection model is based on an assumption that all training samples are positive samples. This is difficult to guarantee in actual scenarios. In the industrial quality inspection process, although most samples are normal samples, it is not excluded that there are individual abnormal samples mixed in. During the collection process of the dataset for model training, it is impossible to check the normal nature of all samples one by one. In the case where the training data contains noise, the unsupervised anomaly detection method based on distillation learning will be greatly affected. The existence of noise data will increase the negative impact of abnormal information on model training, and thus lead to the failure of the entire method. Summary of the Invention

[0005] This application provides a robust wafer defect detection method and system based on distillation learning to at least partially solve the above problems.

[0006] In the first aspect of this application, a robust wafer defect detection method based on distillation learning is provided. The method includes: Obtain a training sample set; Input the training samples of the current batch in the training sample set into the teacher model and the student model, respectively extract the sample features corresponding to the training samples of the current batch, fix the model parameters of the teacher model, and train the student model with the goal of minimizing the difference between the sample features extracted by the teacher model and the sample features extracted by the student model to obtain a robust wafer defect detection model; during the training process, through the non-common feature removal module, eliminate the prominent independent features in the training samples of the current batch; Based on the wafer defect detection model, detect the wafer image to be detected.

[0007] Optionally, the non-common feature removal module includes: a smoother and a weight counter. During the training process, through the non-common feature removal module, eliminate the prominent independent features in the training samples of the current batch, including: The smoother performs an average operation on the training samples of the current batch in the batch dimension and outputs the average feature in the batch dimension; The weight calculator compares the variance information between each sample feature and the average feature, sorts each sample feature, selects a preset number of sample features with the largest degree of difference, and reduces their weights.

[0008] Optionally, the non-common feature removal module further includes: a selector. During the training process, through the non-common feature removal module, eliminate the prominent independent features in the training samples of the current batch, and further include: The selector uses the operation of difficult sample selection to screen the most difficult part of the sample features for the backpropagation of the model.

[0009] Optionally, during the training process, it further includes: Through the abnormal training sample removal module, identify and remove the individual samples in the training samples of the current batch whose difference is greater than the preset threshold.

[0010] Optionally, through the abnormal training sample removal module, identify and remove the individual samples in the training samples of the current batch whose difference is greater than the preset threshold, including: Calculate the Euclidean distance between each training sample of the current batch, and quantitatively measure the mutual distance of each training sample in the multi-dimensional feature space; Compare the average distance between each training sample and other training samples of the current batch, and determine the training sample with the largest average distance from other training samples of the current batch as the individual sample; Set the training weight of the individual sample to a preset small threshold.

[0011] Optionally, the method further includes: the non-common feature removal module obtains common loss information after removing prominent independent features in the training samples of the current batch, and the abnormal training sample removal module identifies and removes individual samples with a difference greater than a preset threshold in the training samples of the current batch to obtain purified loss information; obtaining comprehensive information according to the common loss information and the purified loss information, and constructing a loss function of the student model according to the comprehensive information.

[0012] The second aspect of the present application provides a robust wafer defect detection system based on distillation learning, and the robust wafer defect detection system based on distillation learning includes: An acquisition module, configured to acquire a training sample set; A training module, configured to input the training samples of the current batch in the training sample set into a teacher model and a student model, respectively extract the sample features corresponding to the training samples of the current batch, fix the model parameters of the teacher model, and minimize the difference between the sample features extracted by the teacher model and the sample features extracted by the student model as the goal to train the student model to obtain a robust wafer defect detection model; during the training process, through the non-common feature removal module, prominent independent features in the training samples of the current batch are removed; A detection module, configured to detect a to-be-detected wafer image based on the wafer defect detection model.

[0013] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes, it implements the robust wafer defect detection method based on distillation learning as described in the first aspect of the present application.

[0014] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the robust wafer defect detection method based on distillation learning as described in the first aspect of the present application.

[0015] The fifth aspect of the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps in the robust wafer defect detection method as described in the first aspect of the present application.

[0016] In this application, for the training samples in the distillation learning process, from the dimension of a batch, the common features in multiple training samples are extracted. These common features are the features that are prevalent in the samples and can represent the sample attributes, and are of great significance for the learning and understanding of the model. Since the proportion of abnormal data in the training data is relatively low, abnormal features are often filtered out during the extraction of common features, thus achieving the effect of reducing interfering features. This series of complex common feature extraction operations is similar to the purification process of substances in the chemical field, aiming to remove the possible interfering items in the sample features of the training data and ensure the purity and reliability of the samples participating in the model training process. After purification, the common features will be more stable and consistent, and can provide more accurate and effective supervision guidance for model training. During the model training process, the design of this application enables the model to only receive the supervision guidance of these carefully extracted and purified common features. Under the constraint of this supervision information, the model will focus on learning the truly valuable features and avoid the perturbations introduced by noisy samples. Brief Description of the Drawings

[0017] In order to more clearly illustrate the technical solution of this application, the drawings required for the description of this application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 is the step flowchart of the robust wafer defect detection method based on distillation learning provided by this application; Figure 2 is the structural schematic diagram of the non-common feature removal module in the robust wafer defect detection method based on distillation learning provided by this application; Figure 3 is the overall training process schematic diagram of the robust wafer defect detection model in the robust wafer defect detection method based on distillation learning provided by this application. Detailed Embodiments

[0019] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in combination with the drawings and specific embodiments.

[0020] For ease of understanding, first, the terms involved in this application are explained: Wafer: As the core material for semiconductor manufacturing, wafers are usually made of high-purity single crystal silicon, silicon carbide and other advanced materials. It is the basic material for related semiconductor components in integrated circuits. Wafers are round sheets with common sizes ranging from a few inches to dozens of inches in diameter. In the manufacturing process of semiconductor devices, wafers need to go through a series of precise process steps to ensure the excellent performance and high reliability of the final product.

[0021] Defect detection: Defect detection is an indispensable part of semiconductor manufacturing. Semiconductor component manufacturers often use a series of cutting-edge technologies and precision equipment to conduct comprehensive and in-depth inspections and quality assessments of generated products. The main goal is to capture potential flaws and defects, including but not limited to tiny scratches, imperceptible spots or tiny cracks on the surface. Defect detection plays a vital role in product manufacturing, which is directly related to the reliability and yield of subsequent products.

[0022] Distillation learning: Distillation learning is a technology commonly used in the process of deep learning model training for model compression and feature optimization. Its core idea is to use a well-trained large model as a teacher model and transfer knowledge to a smaller model, the student model, through knowledge transfer, so that the student model can learn the performance of the teacher model while maintaining a small capacity and low computational complexity. In defect detection technology, the idea of ​​distillation learning is also often used to train anomaly detection models to improve the defect detection method's ability to capture anomalies of rare defects.

[0023] In the semiconductor production process, the chip structure is complex and delicate, and the manufacturing process requires extremely high precision. In the precise manufacturing process, any tiny defect may cause the chip performance to decline or even fail. Therefore, in the overall chip production process, effective defect detection of the wafer is required to ensure the correct intervention of subsequent process links.

[0024] Among all wafer defect detection technologies, the method based on visual recognition is widely favored by manufacturers due to its non-destructive nature. The visual defect detection solution first uses a specific imaging module to scan the wafer surface to obtain an imaging picture under a specific optical path setting. After that, based on the feature information in the imaging picture, the possible defects are located and classified. Due to the large differences in wafer background information in different process links, the target detection method based on supervised learning is often easily affected by its background information, resulting in a large number of over-inspections and missed inspections. At the same time, the rare characteristics of many defect categories also determine that wafer defect detection solutions usually use unsupervised anomaly detection models to quickly identify possible anomalies in wafer imaging pictures.

[0025] In unsupervised anomaly detection methods, the technical paradigm based on distillation learning mainly utilizes the assumption of knowledge boundaries. In the model training phase, the student model will imitate the behavior of the teacher model, extract features from the input images, and make the output features as similar as possible to those of the teacher model. Since the model has not seen anomaly samples during training, for anomaly samples, the feature extraction ability of the student model is relatively weak. Therefore, in the testing phase, the probability that the input sample belongs to a normal sample can be judged by the similarity between the output features of the student model and the teacher model. The unsupervised anomaly detection method based on distillation learning does not need to rely on a specific comparison sample library, and its inference efficiency is much higher than that of other anomaly detection methods, showing strong practicality.

[0026] However, the unsupervised anomaly detection method based on distillation learning also has a problem, that is, it is very sensitive to the purity of the training samples. If anomaly samples are mixed in the training samples, that is, when the training set is noisy data, the performance of the trained anomaly detection model will drop significantly. This is because the anomaly information in the training samples provides incorrect guidance information for model training, enabling the model to have a strong feature extraction ability for anomaly samples, breaking the inherent principle of distillation learning for anomaly detection. In practical application scenarios, it is impossible to determine whether a large amount of training data contains a small number of anomaly samples, which poses a great challenge to the wafer defect detection method based on distillation learning and makes it impossible to accurately determine its reliability.

[0027] The main objective of the embodiments of this application is to provide a robust learning paradigm for the wafer defect detection method based on distillation, enabling it to cope with scenarios where the training data contains noise. Specifically, by extracting the mutual information of the training samples during the distillation learning process, reducing the degree of non-common features participating in model training, and removing potential anomaly samples that may exist, the correctness of the anomaly detection model training process is ensured.

[0028] First, briefly describe the formal definition of the noisy anomaly detection task. Under expected conditions, in the training phase, a training data set D1 containing only normal samples will be given for model training to obtain an anomaly detection model M; in the testing phase, the model M is used to judge the input image I and output whether the image is a normal sample or an anomaly sample. However, in the actual scenario, the training data set may contain some potential anomaly samples, which breaks the assumption of the training data set D1 in the conventional scheme. A real-world training data set D2 may contain most normal samples and a small number of anomaly samples, and this part of the anomaly samples hidden in the training data may mislead the model to make wrong judgments in the inference phase.

[0029] The overall solution of this application is to extract and purify the common features of the training samples in the distillation learning process, so that the model training process only receives the supervision and guidance of the common features, eliminating the harm and negative impact that may be introduced by individual noise samples. By adding the detection of abnormal training samples during the training of the anomaly detection model, the robustness of the method can be improved.

[0030] The core concept of the technical solution proposed in this application focuses on optimizing the distillation learning process to improve the efficiency and quality of the anomaly detection model training process.

[0031] Specifically, in ordinary distillation learning, all kinds of features contained in the training samples will participate in the model training process, including not only common features that are beneficial to model training, but also interference features introduced by some noise samples. These interference features will mislead the learning direction of the model and reduce the generalization ability and accuracy of the model.

[0032] In order to solve this problem, the present application proposes: for the training samples in the distillation learning process, the common features in multiple training samples are examined from the dimension of a batch. These common features are common features in the samples that can represent the attributes of the samples, which are of great significance for the learning and understanding of the model. Due to the low proportion of abnormal data in the training data, abnormal features are often filtered out during the extraction of common features, thereby reducing the effect of interfering features. This series of complex common feature extraction operations is similar to the material purification process in the chemical field. The purpose is to remove the interference items that may exist in the sample features of the training data, and ensure the purity and reliability of the samples involved in the model training process. The common features after purification will be more stable and consistent, and can provide more accurate and effective supervision and guidance for model training. During the model training process, the design of this application enables the model to only receive the supervision and guidance of these carefully extracted and purified common features. Under the constraints of this supervisory information, the model will focus on learning truly valuable features and avoid the disturbance introduced by other noise samples.

[0033] In addition, the present application also proposes to remove abnormal training samples. The presence of abnormal training samples may have an adverse effect on the training of the model and reduce the robustness of the model. In order to meet this challenge, the present application adds the detection link of abnormal training samples in the process of training the anomaly detection model. By introducing the idea of ​​anomaly detection into the process of model training, the training samples are monitored and analyzed in real time, and the abnormal samples are discovered and identified in time by mutual comparison. The model will remove the training gradients of these abnormal samples, so that the method only uses part of the valid samples in the face of complex samples, maintains robustness, and provides a strong guarantee for the stable operation of the model.

[0034] Specifically, the core concept of this application lies in: effectively controlling and managing training data. The correct utilization and error troubleshooting of training data are achieved through two core modules. Specifically, through the non-common feature removal module, prominent independent features in batch training samples are eliminated to ensure the purity of data during the backpropagation process of the input model; meanwhile, through the abnormal training sample removal module, individual samples with large differences in the batch are identified and mined, and they are excluded from the model training process to reduce the harm of their negative impacts.

[0035] Specifically, as Figure 1 shown, it shows the step flow chart of a robust wafer defect detection method based on distillation learning provided by this application. The method includes the following steps: S101, Obtain a training sample set.

[0036] In this application, most samples in the training sample set are normal wafer image samples, but inevitably, a small number of abnormal wafer image samples are mixed in. In practical applications, in the training sample set, it is inevitable that a small number of abnormal wafer image samples are mixed in. The abnormal information provided by the abnormal wafer image samples in the training samples provides incorrect guidance information for the training of the student model, resulting in the model being able to have strong feature extraction capabilities for abnormal samples, and the detection accuracy of the unsupervised anomaly detection method based on distillation learning decreases. Therefore, this application proposes to analyze from the batch dimension during the training process to eliminate prominent independent features, that is, abnormal features, in the training samples of each batch.

[0037] S102, Input the training samples of the current batch in the training sample set into the teacher model and the student model, respectively extract the sample features corresponding to the training samples of the current batch, fix the model parameters of the teacher model, and train the student model with the goal of minimizing the difference between the sample features extracted by the teacher model and the sample features extracted by the student model to obtain a robust wafer defect detection model.

[0038] Among them, during the training process, through the non-common feature removal module, prominent independent features in the training samples of the current batch are eliminated.

[0039] Specifically, during the training process of the anomaly detection model based on distillation learning, for a training sample x, the teacher model and the student model will respectively extract its features MT(x) and MS(x). The goal of model training is to minimize the difference between MT(x) and MS(x). This difference can be defined as Delta(MT(x) – MS(x)). Assume that x is a set of b samples in the training batch, then the non-common feature removal module is used to improve the features generated by x.

[0040] In this application, the main objective of the non-common feature removal module is to utilize the common features among multiple samples in a batch to reduce the participation of non-common features (i.e., abnormal features) in model training. Abnormal samples mixed in the training samples may directly provide false normal features, which can cause direct harm to the model. These abnormal samples will expand the boundaries of normal samples to ranges where they should not exist, so that the positive example distribution covered by the model may overlap with the abnormal samples in the test phase. To avoid this negative information, this application utilizes the mutual information among the training batch samples to make the model training process more cautious.

[0041] In this application, the non-common feature removal module includes: a smoother and a weight counter. During training, through the non-common feature removal module, prominent independent features in the training samples of the current batch are eliminated, including: S1, The smoother performs an average operation on the training samples of the current batch in the batch dimension and outputs the average feature in the batch dimension.

[0042] S2, The weight calculator compares the variance information between each sample feature and the average feature, sorts each sample feature, selects a preset number of sample features with the largest degree of difference, and reduces their weights.

[0043] The smoother is an average operation on a set of training samples x of a batch in the batch dimension, which is used to make the feature output of rare samples smoother. Assume that the size of the sample feature is W and H, and the dimension is C. The sample feature output by each training sample is in the dimension of C×W×H. The smoother averages them on the sample features in the batch dimension based on this dimension and outputs the average feature F in the batch dimension.

[0044] The weight calculator calculates the weight information between each feature element (i.e., sample feature), which represents the feature difference degree of each feature element. The weight calculator sorts the feature elements by comparing the variance information between each sample feature and the smoothed output feature (i.e., average feature). Select the top p percentile feature elements with the largest degree of difference and reduce their weights to w; for feature elements other than the top p percentile, their weights remain unchanged (default weight is 1).

[0045] In an alternative embodiment, the non-common feature removal module further includes: a selector. During training, through the non-common feature removal module, prominent independent features in the training samples of the current batch are eliminated, and it further includes: S3, The selector uses the operation of difficult sample selection to screen the most difficult partial sample features for the backpropagation of the model.

[0046] In this application, the selector utilizes the operation of difficult sample selection to screen the most difficult feature differences. Based on the above-mentioned smoother and weight calculator, the selector will select the most difficult part of the features for the back propagation of the model. Experience with existing solutions has shown that this calculation method can effectively improve the effect of distillation learning. In related technologies, due to the direct selection of difficult samples, noise data samples may be selected, resulting in poor training results; and this application selects the remaining difficult samples for training after removing abnormal features, ensuring that no noise is introduced while maintaining good training results.

[0047] Combined with the smoother, weight calculator and selector, the non-common feature removal module essentially reduces the degree of participation of unique feature elements in a specific range in training, selects the content that is most needed to participate in model training, purifies and filters the training sample information in model training, avoids the generation of some unnecessary gradient information, and ensures that the distillation learning process proceeds in a stable and effective direction. The structural diagram of the non-common feature removal module is shown in the figure below. Figure 2 Specifically, the present application analyzes all sample features obtained from a batch of training samples from the batch dimension, obtains the average feature based on the smoother, and then compares each sample feature with the average feature based on the weight calculator, selects a preset number of sample features with the largest difference, reduces their weights, and makes them less involved in model training, avoiding the introduction of abnormal information during the model training process, and further selects the most difficult part of the sample features (purified features) based on the selector for the back propagation of the model.

[0048] In this application, in a wafer defect detection method based on distillation learning, the mutual information differences between samples in batches are used to measure and find commonly occurring feature information, thereby reducing the probability of potential abnormal samples participating in model training by removing non-common features, thereby ensuring the purity and effectiveness of distillation learning knowledge.

[0049] In an optional implementation, the training process further includes: identifying and removing personality samples whose differences in the current batch of training samples are greater than a preset threshold through an abnormal training sample removal module.

[0050] Compared with the above-mentioned non-common feature removal module that purifies information from the feature perspective, the present application also proposes to filter information from the individual perspective through an abnormal training sample removal module. The abnormal training sample removal module uses the feature distances of the same batch of training samples to select the most prominent samples and regards them as out-of-boundary samples to reduce their participation in training.

[0051] In this application, in the wafer defect detection method based on distillation learning, the mutual distance relationship between samples in different batches is utilized to determine whether there are outliers. By adopting the strategy of removing abnormal training samples, the possibility of abnormal samples participating in model training is directly excluded, and the potential hazards and negative impacts that abnormal samples may cause to model training are excluded.

[0052] Specifically, through the abnormal training sample removal module, the individual samples in the training samples of the current batch with a difference greater than the preset threshold are identified and removed, including: S11, Calculate the Euclidean distance between each training sample in the current batch to quantitatively measure the mutual distance of each training sample in the multi-dimensional feature space.

[0053] S12, Compare the average distance between each training sample and other training samples in the current batch, and determine the training sample with the largest average distance from other training samples in the current batch as the individual sample.

[0054] S13, Set the training weight of the individual sample to a preset small threshold.

[0055] In this application, during the training process of the deep learning model, the existence of abnormal training samples may have an adverse impact on the model performance and generalization ability. To address this challenge, this application calculates the Euclidean distance between training samples in the same batch to quantitatively measure the mutual distance of each training sample in the multi-dimensional feature space. By comparing the average distance between each training sample and other training samples, the training sample with the largest average distance from the rest of the training samples can be found, and this training sample is likely to be a potential abnormal sample.

[0056] After locating the abnormal sample, this application does not directly remove it from the training sample set, but adopts a more flexible and refined processing method. This application sets the training weight of the abnormal sample to a smaller value. By reducing the training weight of the abnormal sample, its impact on model training can be directly reduced. The advantage of doing this is that it not only avoids the loss of training data information that may be caused by directly removing samples, but also effectively suppresses the interference of abnormal samples on the model training process.

[0057] By introducing this processing strategy for abnormal training samples, the training process of the model can be optimized. The model can be trained according to a classroom learning process similar to from easy to difficult. In the initial stage of training, the model mainly focuses on those relatively "normal" samples with larger weights. The feature differences between these samples are relatively small, and the model can more easily learn the rules and patterns they represent. As the training progresses, the model will also pay a certain amount of attention to abnormal samples with smaller weights, but this attention is carried out on the premise that the model already has a certain foundation. This gradual training method helps the model to more stably learn the effective information in the training data, avoiding instability in the training process due to being affected by abnormal samples too early or excessively. Thus, it fundamentally ensures the stability of the final effect of the model and improves the prediction accuracy and reliability of the model when facing unknown data.

[0058] In an alternative embodiment, during the training process, it further includes: the non-common feature removal module obtains common loss information after removing prominent independent features in the current batch of training samples, and the abnormal training sample removal module obtains purified loss information after identifying and removing individual samples in the current batch of training samples with a difference greater than a preset threshold; obtaining comprehensive information according to the common loss information and the purified loss information, and constructing the loss function of the student model according to the comprehensive information.

[0059] In this application, the overall training process of the robust wafer defect detection model is as Figure 3 shown. It can be seen that during the training process, the student model and the teacher model (pre-trained and with frozen model parameters) can respectively obtain corresponding sample features for a batch of training samples. For these sample features, on the one hand, through the non-common feature removal module, non-common features are removed to obtain common loss information. On the other hand, the abnormal training sample removal module is used to reduce the training weight of the training sample with the largest average distance from the rest of the training samples in the same batch to obtain purified loss information. Then, based on the common loss information and the purified loss information, comprehensive information is obtained, and the loss function of the student model is calculated based on this comprehensive information, thereby avoiding the influence of prominent independent features (non-common features) and abnormal samples in the training samples on the student model during the training process.

[0060] S103, detecting the wafer image to be detected based on the wafer defect detection model.

[0061] Thus, this application realizes the control of the knowledge transfer process of distillation learning by adding a non-common feature removal module and an abnormal training sample removal strategy during the training process of the anomaly detection model, improves the prudence and effectiveness of model training, thereby ensuring the robustness of the final wafer defect detection method and enhancing the reliability of the detection results.

[0062] This application conducts experiments on three standard industrial anomaly detection datasets, and the experimental results fully demonstrate the effectiveness of the method proposed in this application compared to existing methods. To simulate the situation where training samples may contain anomalies in the real-world scenario, this application contaminates the original training samples according to a certain noise ratio. Specifically, some anomaly samples are collected from the test set, and a part of them is randomly selected and injected into the training set. After the contamination operation is completed, the noisy training dataset will be larger than the original training dataset and there will be an overlap with the test set.

[0063] The experimental results for the MVTec LOCO dataset with 10% noisy data are shown in Table 1 below.

[0064] Table 1 Experimental Results for the MVTec LOCO Dataset with 10% Noisy Data

[0065] In Table 1, the metric for the method is AUROC. From the experimental results, it can be seen that the method proposed in this application can achieve effective performance improvement compared to the original method on the MVTec LOCO dataset with 10% noisy data. Whether it is the samples in the structurally abnormal part or the logically abnormal samples in the dataset, the accuracy can be improved by more than 5%.

[0066] Similarly, the experimental results for the MVTec AD dataset and VisA with 10% noisy data are shown in Table 2 below.

[0067] Table 2 Experimental Results for the MVTec AD Dataset and VisA with 10% Noisy Data

[0068] In Table 1, the metric for the method is AUROC. It can be seen that the comparison of the effects of the method proposed in this application and the original method on the MVTec AD dataset and VisA dataset verifies that the method proposed in the invention can improve the robustness of the wafer defect detection method in the scenario where the training data contains noisy data.

[0069] Based on the same inventive concept, this application also provides a robust wafer defect detection system based on distillation learning. The robust wafer defect detection system based on distillation learning includes: An acquisition module for acquiring a training sample set; A training module, which is used to input the training samples of the current batch in the training sample set into a teacher model and a student model, extract the sample features corresponding to the training samples of the current batch respectively, fix the model parameters of the teacher model, and train the student model with the goal of minimizing the difference between the sample features extracted by the teacher model and the sample features extracted by the student model to obtain a robust wafer defect detection model; during the training process, through a non-common feature removal module, the prominent independent features in the training samples of the current batch are eliminated; A detection module, which is used to detect the wafer image to be detected based on the wafer defect detection model.

[0070] Optionally, the non-common feature removal module includes: a smoother and a weight counter. During the training process, through the non-common feature removal module, the prominent independent features in the training samples of the current batch are eliminated, including: The smoother performs an averaging operation on the training samples of the current batch in the batch dimension and outputs the average feature in the batch dimension; The weight calculator compares the variance information between each sample feature and the average feature, sorts each sample feature, selects a preset number of sample features with the largest degree of difference, and reduces their weights.

[0071] Optionally, the non-common feature removal module further includes: a selector. During the training process, through the non-common feature removal module, the prominent independent features in the training samples of the current batch are eliminated, and further includes: The selector uses the operation of difficult sample selection to screen the most difficult part of the sample features for the backpropagation of the model.

[0072] Optionally, during the training process, it further includes: Through an abnormal training sample removal module, the individual samples in the training samples of the current batch with a difference greater than a preset threshold are identified and removed.

[0073] Optionally, through the abnormal training sample removal module, the individual samples in the training samples of the current batch with a difference greater than a preset threshold are identified and removed, including: Calculate the Euclidean distance between each training sample of the current batch to quantify the mutual distance of each training sample in the multi-dimensional feature space; Compare the average distance between each training sample and other training samples of the current batch, and determine the training sample with the largest average distance from other training samples of the current batch as the individual sample; Set the training weight of the individual sample to a preset small threshold.

[0074] Optionally, during the training process, it further includes: the non-common feature removal module obtains common loss information after removing prominent independent features in the training samples of the current batch, and the abnormal training sample removal module identifies and removes individual samples with a difference greater than a preset threshold in the training samples of the current batch to obtain purified loss information; obtaining comprehensive information according to the common loss information and the purified loss information, and constructing a loss function of the student model according to the comprehensive information.

[0075] Based on the same inventive concept, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes, it implements the steps in the robust wafer defect detection method based on distillation learning as described in any of the above embodiments.

[0076] Based on the same inventive concept, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the robust wafer defect detection method based on distillation learning as described in any of the above embodiments.

[0077] Based on the same inventive concept, the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps in the robust wafer defect detection method based on distillation learning as described in any of the above embodiments.

[0078] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0079] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0080] The present application is described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate for implementation in the processFigure 1 one or more processes and / or blocks Figure 1 a system for the functions specified in one or more blocks

[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks

[0082] These computer program instructions can also be loaded onto a computer or other programmable terminal device, such that a series of operation steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks

[0083] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0084] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0085] The above has introduced in detail a robust wafer defect detection method based on distillation learning. In this article, specific examples are used to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A robust wafer defect detection method based on distillation learning, characterized in that: The method comprises: Obtain a training sample set; Input the current batch of training samples in the training sample set into the teacher model and the student model, extract the sample features corresponding to the current batch of training samples respectively, fix the model parameters of the teacher model, train the student model with the goal of minimizing the difference between the sample features extracted by the teacher model and the sample features extracted by the student model, and obtain a robust wafer defect detection model; during the training process, eliminate the prominent independent features in the current batch of training samples through a non-common feature removal module; The wafer image to be inspected is inspected based on the wafer defect detection model.

2. The robust wafer defect detection method based on distillation learning according to claim 1, characterized in that: The non-common feature removal module includes: a smoother and a weight counter. During the training process, the non-common feature removal module is used to eliminate the prominent independent features in the current batch of training samples, including: The smoother performs an average operation on the batch dimension for the training samples of the current batch and outputs the average features on the batch dimension; The weight calculator compares the variance information between each sample feature and the average feature, sorts each sample feature, selects a preset number of sample features with the largest degree of difference, and reduces their weights.

3. The robust wafer defect detection method based on distillation learning according to claim 2, characterized in that: The non-common feature removal module further includes: a selector, which eliminates prominent independent features in the current batch of training samples through the non-common feature removal module during the training process, and further includes: The selector uses the operation of difficult sample selection to filter out the most difficult sample features for the back propagation of the model.

4. The robust wafer defect detection method based on distillation learning according to claim 1, characterized in that: The training process also includes: Through the abnormal training sample removal module, personality samples whose differences in the current batch of training samples are greater than the preset threshold are identified and removed.

5. The robust wafer defect detection method based on distillation learning according to claim 4, characterized in that: Through the abnormal training sample removal module, the personality samples whose differences in the current batch of training samples are greater than the preset threshold are identified and removed, including: Calculate the Euclidean distance between each training sample in the current batch, and quantitatively measure the mutual distance between each training sample in the multidimensional feature space; Compare the average distance between each training sample and other training samples in the current batch, and determine the training sample with the largest average distance to other training samples in the current batch as the personality sample; The training weight of the personality sample is set to a preset small threshold.

6. The robust wafer defect detection method based on distillation learning according to any one of claims 1 to 5, characterized in that: The method also includes: a non-common feature removal module obtains common loss information after eliminating prominent independent features in the current batch of training samples, and an abnormal training sample removal module obtains purification loss information after identifying and removing individual samples whose differences are greater than a preset threshold in the current batch of training samples; comprehensive information is obtained based on the common loss information and the purification loss information, and a loss function of the student model is constructed based on the comprehensive information.

7. A robust wafer defect detection system based on distillation learning, characterized in that: The robust wafer defect detection system based on distillation learning includes: An acquisition module is used to acquire a training sample set; A training module is used to input the current batch of training samples in the training sample set into the teacher model and the student model, respectively extract the sample features corresponding to the current batch of training samples, fix the model parameters of the teacher model, and train the student model with the goal of minimizing the difference between the sample features extracted by the teacher model and the sample features extracted by the student model to obtain a robust wafer defect detection model; during the training process, the prominent independent features in the current batch of training samples are eliminated through the non-common feature removal module; The detection module is used to detect the wafer image to be detected based on the wafer defect detection model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the robust wafer defect detection method based on distillation learning as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the robust wafer defect detection method based on distillation learning as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps in the robust wafer defect detection method based on distillation learning described in any one of claims 1 to 6 are implemented.