Conditional data enhanced wafer defect detection method and system

By using conditional data enhancement method in wafer defect detection, data enhancement is performed on the training sample, and loss information comparison and weight adjustment are performed during the model training process, the problem of unstable defect detection effect in the prior art is solved, and the robustness and detection accuracy of the model are improved.

CN120163793APending Publication Date: 2025-06-17TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269630.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing wafer defect detection methods tend to decline when changes in light conditions, imaging clarity and scanning position, resulting in unstable results in actual industrial scenarios.

Method used

The training sample is enhanced by using the conditional data enhancement method, and the enhanced sample is generated. During the model training process, by comparing the training loss information of the enhanced sample with the average loss value of all samples in the time window, it is determined whether it exceeds the preset range. If it exceeds the limit, the amplitude of its participation in the backpropagation loss calculation will be reduced.

Benefits of technology

The robustness of the wafer defect detection model is improved, so that it can detect defects stably under different imaging conditions, reducing the occurrence of false detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163793A_ABST
    Figure CN120163793A_ABST
Patent Text Reader

Abstract

The invention provides a conditional data enhanced wafer defect detection method and system, and relates to the technical field of wafer defect detection. According to the wafer defect detection method, in the training process of the wafer defect detection method based on the anomaly detection model, a conditional data enhancement strategy is applied to training data. A data enhancement strategy constructs an intermediate state model in a model training process, the model carries out rough quantitative evaluation on an enhanced sample according to context information at a training moment, the probability that the current enhanced sample accords with a normal sample is judged, and the training weight of the current enhanced sample participating in model parameter updating is controlled. Therefore, in the model training process, the effective enhanced samples are conditionally added, so that the robustness of the model can be improved on the premise of keeping the precision of the model, and the effectiveness and practicability of the model in an actual application scene are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of wafer defect detection, and particularly to a wafer defect detection method and system with conditional data augmentation. Background Art

[0002] In the semiconductor manufacturing process, equipment manufacturers usually use computer vision methods to detect defects in wafers at each process to determine the wafer quality and control the product yield.

[0003] Optical scanning imaging is a technology that uses optical principles to obtain images. It involves multiple processes such as light emission, reflection, absorption, and detection. Usually, a light beam emitted by a light source is used to irradiate the target object, and the light reflected or scattered from the object surface is captured for imaging.

[0004] In wafer defect detection, multiple optical scanning imaging technologies are usually adopted. These technologies use light sources with different wavelengths to irradiate the wafer surface, and identify defects by analyzing the reflected or scattered light signals. For example, most surface defects can be presented through the bright-field scanning channel. Defects such as particles and scratches on the sample surface will change the reflection and transmission of light, thus presenting a specific pattern of morphology in the imaging effect. Some defect categories need to be imaged and identified through dark-field scanning. The light source in the dark-field channel irradiates the sample surface at an oblique angle, and only the light scattered or reflected by the defects on the sample surface will be collected and imaged. In addition, the principle of photoluminescence is usually used in wafer defect detection. By irradiating the sample surface of the semiconductor material with a laser, imaging is achieved by exciting some defect substances to a high-energy state. During the process of scanning imaging, information such as the light source position, light source intensity, light source brightness, physical material, and optical path setting will all affect the final imaging effect.

[0005] Since wafer defect data is scarce and the possible defect categories on the wafer surface are difficult to predict and define in advance, unsupervised anomaly detection schemes are usually used for defect detection. The core idea of the unsupervised anomaly detection method is to use the distribution characteristics of normal data to build a model, and in the test stage, use the model to evaluate whether the sample to be tested conforms to the data distribution expressed by the model, and regard the data that does not conform to the distribution as abnormal data. However, conventional unsupervised anomaly detection methods are sensitive to the imaging effect of images. When information such as lighting conditions, imaging clarity, and scanning position changes, the defect detection effect often drops rapidly. However, in actual industrial scenarios, it is impossible to ensure that all factors of the imaging effect are stable within the expected range, which poses a great challenge to the effect of wafer defect detection methods in practice. Summary of the Invention

[0006] The present application provides a wafer defect detection method and system with conditional data augmentation to at least partially solve the above problems.

[0007] In a first aspect of the present application, a wafer defect detection method with conditional data augmentation is provided. The method includes: Receiving training samples, performing data augmentation on some of the training samples according to a preset probability to obtain augmented samples; Training a preset model based on the training samples to obtain a wafer defect detection model. During the training process, comparing the training loss information of the augmented samples with the training loss information of all training samples within a time window of it to determine whether the augmented samples exceed a preset range after the data augmentation transformation; in the case where it is determined that the augmented samples exceed the preset range after the data augmentation transformation, reducing the amplitude of the augmented samples participating in the calculation of the model's backpropagation loss; Detecting the wafer image to be detected based on the wafer defect detection model.

[0008] Optionally, during the training process, comparing the training loss information of the augmented samples with the training loss information of all training samples within a time window of it to determine whether the augmented samples exceed a preset range after the data augmentation transformation includes: Based on global variables, statistically calculating the training loss information of the training samples within the time window one by one to obtain a statistical value. The statistical value is iteratively updated in the form of a sliding window, and the mean value of the statistical value is obtained to obtain the average loss value of the training samples within the current time window of the model; for the augmented samples within the current time window, determining whether the loss value of the augmented sample under the current model calculation exceeds a preset interval compared with the average loss value.

[0009] Optionally, in the case where it is determined that the augmented samples exceed the preset range after the data augmentation transformation, reducing the amplitude of the augmented samples participating in the calculation of the model's backpropagation loss includes: In the case where the loss value of the augmented sample under the current model calculation exceeds the preset interval compared with the average loss value, setting a first weight factor F1 according to the ratio of the loss value of the augmented sample under the current model calculation to the average loss value; Based on the first weight factor F1, controlling the degree of the augmented sample participating in the calculation of the model's backpropagation loss.

[0010] Optionally, the method further includes: during the training process, obtaining a second weight factor F2 according to the value of the current iteration number of the model and the total iteration number of the model, and based on the second weight factor F2, controlling the degree of the augmented sample participating in the calculation of the model's backpropagation loss.

[0011] Optionally, the loss function of the preset model is: L = (1 - a)×L_ori + a × L_cond where L_ori represents all the loss functions of the samples without data augmentation, L_cond represents all the loss functions of the augmented samples after data augmentation, and a represents the proportion of the augmented samples participating in the training.

[0012] Optionally, the method further includes: Obtain a test sample set, and based on the test sample set, obtain a first metric of the wafer defect detection model for the test sample set; Perform data augmentation on all the samples in the test sample set to obtain an augmented sample set, and based on the augmented sample set, obtain a second metric of the wafer defect detection model for the augmented sample set; Obtain the augmentation deterioration score of the wafer defect detection model based on the difference between the second metric and the first metric; Evaluate the robustness of the wafer defect detection model based on the augmentation deterioration score.

[0013] In a second aspect of the present application, a wafer defect detection system with conditional data augmentation is provided. The wafer defect detection system with conditional data augmentation includes: A receiving module, configured to receive training samples, and perform data augmentation on some of the training samples according to a preset probability to obtain augmented samples; A training module, configured to train a preset model based on the training samples to obtain a wafer defect detection model. During the training process, compare the training loss information of the augmented samples with the training loss information of all the training samples within a time window thereof to determine whether the augmented samples exceed a preset range after the data augmentation transformation; in the case where it is determined that the augmented samples exceed the preset range after the data augmentation transformation, reduce the amplitude of the augmented samples participating in the calculation of the model backpropagation loss; A detection module, configured to detect a wafer image to be detected based on the wafer defect detection model.

[0014] In a third aspect of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the conditional data augmentation wafer defect detection method as described in the first aspect of the present application.

[0015] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the conditional data augmentation-based wafer defect detection method as described in the first aspect of the present application.

[0016] A fifth aspect of the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps in the conditional data augmentation-based wafer defect detection method as described in the first aspect of the present application.

[0017] The present application aims to improve the robustness of the defect detection method so that it is not affected by changes in imaging conditions. The ability to keep the defect detection method stable and effective under different scanning imaging conditions can accurately identify defects and will not result in incorrect detection results due to changes in imaging effects. Description of the Drawings

[0018] To more clearly illustrate the technical solutions of the present application, the drawings required for the description of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of the steps of the conditional data augmentation-based wafer defect detection method provided by the present application; Figure 2 is a schematic diagram of the visualization effect of an industrial product under different data augmentation types and data augmentation degrees in the conditional data augmentation-based wafer defect detection method provided by the present application; Figure 3 is a schematic diagram of the training framework of the wafer defect detection model in the conditional data augmentation-based wafer defect detection method provided by the present application. Detailed Embodiments

[0020] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0021] First, the terms related to the present application are explained: Wafer: A wafer is the basic material for semiconductor manufacturing, usually made of high-purity single crystal silicon, silicon carbide, or gallium nitride. It is the starting material for manufacturing integrated circuits, microprocessors, memories, and other semiconductor devices. A wafer is a thin slice cut from a silicon ingot, usually circular, with diameters ranging from a few inches to dozens of inches. The wafer needs to go through multiple processes in the semiconductor device manufacturing process, including photolithography, etching, ion implantation, chemical vapor deposition, physical vapor deposition, and chemical mechanical polishing, etc.

[0022] Defect Detection: Defect detection refers to the use of different technical means and equipment in the production and manufacturing process to conduct a comprehensive and systematic inspection and evaluation of the product to be tested, aiming to discover possible flaws or defects. These defects include but are not limited to surface flaws, scratches, spots, color differences, cracks, irregular lines, etc. In semiconductor manufacturing, defect detection is particularly important because it directly affects the reliability and yield of chips. With the continuous progress of integrated circuit technology, the sizes of wafers and semiconductor components are constantly shrinking, and the difficulty of defect detection technology is also increasing.

[0023] Data Augmentation: Data augmentation is a data processing technique in machine learning and deep learning, which is used to increase the diversity and quantity of data by effectively transforming the training dataset. Data augmentation methods are usually used in scenarios with insufficient data volume and data imbalance, which can effectively solve the overfitting problem in model training and improve the generalization ability and performance of the model.

[0024] A common idea of unsupervised anomaly detection methods is based on knowledge distillation, that is, using existing normal data to build a model, assuming that the model has specific construction capabilities for this part of the data. When the construction result for the input data does not meet the expectation, the input data is considered abnormal. For example, using a pre-trained teacher model to guide a student model to learn on existing normal data, and through the way of knowledge distillation, making the output content of the student model as similar as possible to that of the teacher model. After the model training is completed, for the given test data, if the distance between the output content of the teacher model and the student model is within a certain range, it is considered that the test data meets the distribution of normal data, otherwise it is determined as abnormal.

[0025] Although unsupervised anomaly detection methods based on knowledge distillation can achieve effective detection of anomaly information, they have relatively high requirements for the distribution consistency between training data and test data. In other words, when the distribution information of test data is quite different from that of training data, such anomaly detection methods are likely to fail. In the scenario of wafer defect detection, if there are significant differences in the imaging effects of the data used for model training and the data in actual applications, the trained defect detection model cannot effectively detect defects in the input data, resulting in serious defect omission or over-detection. In this case, to enhance the robustness of the defect detection method and enable it to stably perform the defect detection function under different scanning imaging effects, a feasible solution is to make the model learn images with different imaging effects during the training stage. Based on this idea, data augmentation techniques can be used to perform data augmentation of different degrees and categories on the existing training data, expanding the sample quantity and diversity of the training dataset, enabling the model to learn the possible distribution information of normal data under different imaging effects, ensuring its ability to handle images that may appear under different imaging effects, and preventing overfitting during the model training process. However, in the actual use process, since data augmentation techniques change the image content of the training data, data augmentation may also bring a series of negative impacts. For example, normal data may introduce anomaly information after data augmentation beyond the scope, causing conceptual confusion during the model training process and instead reducing the accuracy of the model.

[0026] Therefore, in current anomaly defect detection methods, less attention is paid to how to use data augmentation to improve the robustness of the model, and data augmentation techniques are usually not used in defect detection methods. Some solutions may cope with different scenarios by increasing the parameters of the model or using multiple model integrations, but these solutions do not fundamentally solve the impact brought by the variable imaging effects and will also reduce the execution efficiency of the model.

[0027] The unsupervised anomaly detection task generally refers to the identification of rare objects or events that deviate from the normal pattern or distribution in a dataset without prior knowledge or label information. Specifically, the goal of this task is to establish an anomaly detection model to distinguish normal samples and anomaly samples by learning the distribution characteristics of normal data. Since anomaly samples usually account for a very small proportion in the dataset, the solutions to the unsupervised anomaly detection task mainly rely on modeling normal data.

[0028] In deep learning methods, data augmentation and other methods are usually used to expand the sample size and diversity, so that the model can obtain stronger robustness and generalization ability. However, in current industrial anomaly detection methods, data augmentation operations are rarely introduced, which results in the model trained based on deep learning being sensitive to image styles and unable to correctly handle input pictures with different image styles. One possible reason is that data augmentation cannot guarantee the semantic invariance of normal samples, that is, some normal samples may unexpectedly become an abnormal sample due to image transformation operations, thus deteriorating the model's performance. Nevertheless, this application believes that data augmentation methods should still be taken seriously. Therefore, this application proposes a conditional data augmentation method for wafer defect detection, which performs data augmentation on wafer image data and executes a training scheme with conditional strategies to meet the robustness requirements of the model.

[0029] Specifically, this application proposes: during the training process of the wafer defect detection method based on the anomaly detection model, apply a conditional data augmentation strategy to the training data. The data augmentation operation is mainly used to expand the diversity of training samples, so that the model has learning experiences for data under different imaging conditions; the conditional strategy emphasizes the application method for the augmented samples to avoid introducing abnormal information during the model training process and causing confusion in semantic concepts. The data augmentation strategy will construct an intermediate model during the model training process. This model makes a rough quantitative evaluation of the augmented samples according to the context information at the training moment, judges the probability that the current augmented sample conforms to a normal sample, and controls the training weight for its participation in model parameter update. Thus, during the model training process, by conditionally adding effective augmented samples, the robustness of the model can be improved while maintaining the model's accuracy, thereby ensuring the effectiveness and practicality of the model in actual application scenarios.

[0030] As Figure 1 shown, it shows a flowchart of the steps of a conditional data augmentation method for wafer defect detection provided by this application. The method includes the following steps: S101, receive training samples, and perform data augmentation on some samples in the training samples according to a preset probability to obtain augmented samples.

[0031] In this application, the data augmentation methods involved include image blurring, image noise, image compression quality, contrast adjustment, and brightness adjustment. The effect intensity of each image augmentation method can be adjusted through hyperparameters.

[0032] Specifically, for the image blurring enhancement method: Common image blurring schemes include Gaussian blurring, motion blurring, focus blurring, etc., which can cause large-scale effect transformations to the image. In the actual industrial wafer defect detection scenario, due to the movement of the imaging device or the parameter configuration of the focusing module, the captured wafer image may be blurred, and the degree of this blurring will randomly fluctuate within a certain range. This phenomenon poses a great challenge to the defect anomaly detection model. This method selects to add simulated image blurring processing in the data enhancement method to generate diverse training data and ensure the compatibility of the defect detection model with the image blurring phenomenon.

[0033] For the image noise enhancement method: During the transmission of information, noise may accompany it due to reasons such as signal interference. Common noise types include Gaussian noise and salt-and-pepper noise. Gaussian noise is a random noise with zero mean and a specific variance, and its probability density function follows a normal distribution. Many noises in the process of optical imaging and image acquisition approximately conform to Gaussian noise. For ordinary wafer defect detection methods, if the input image contains Gaussian noise, it may interfere with the performance of signal processing and image processing algorithms, resulting in false detections or misclassifications.

[0034] For the image compression quality method: During the transmission of image data, it may need to be compressed due to data bandwidth limitations. The compressed image may lose some quality, resulting in image distortion and resolution degradation. JPEG is a common image compression algorithm, and different parameters can be set to obtain image information of different qualities. Ordinary wafer defect detection methods will not be able to make correct judgments if they receive low-quality compressed images. Therefore, this application also takes image compression quality as part of data enhancement.

[0035] For the contrast adjustment method: Image contrast is an important concept in image processing and computer vision, which describes the degree of brightness difference between different regions in the image. Images with high contrast have obvious brightness boundaries and more prominent details; while images with low contrast appear relatively dull and the details are not clear enough. In wafer scanning imaging pictures, sometimes the contrast of the image needs to be dynamically adjusted according to actual needs to highlight the defect information and facilitate subsequent algorithm detection. Therefore, the defect detection method needs to adapt to input pictures with different contrasts. This method realizes a data enhancement method through contrast adjustment.

[0036] Brightness adjustment method: Brightness adjustment is a basic operation in image processing, used to adjust the overall brightness level of an image, making the image appear brighter or darker. Brightness adjustment can improve the visual effect of the image, making it more suitable for different display devices and viewing conditions. During the wafer defect detection process, according to the different materials of the sample to be measured, it may be necessary to adjust the light source of different brightness for illumination, resulting in different brightness levels of the pictures received by the wafer defect detection algorithm. This method takes brightness adjustment as one of the types of data augmentation, which helps to improve the robustness of the defect detection model to image brightness.

[0037] The visualization effects of an industrial product under different data augmentation types and data augmentation degrees are as Figure 2 shown. Among them, from the first row to the last row are the product images obtained by using the image blur augmentation method, the image noise augmentation method, the image compression quality augmentation method, the contrast adjustment augmentation method, and the brightness adjustment augmentation method respectively. From left to right are the product images obtained by using different data augmentation degrees.

[0038] In this application, for the input training samples, some training samples can be augmented according to a preset probability to obtain a training sample set including the augmented samples and the original samples.

[0039] S102, training a preset model based on the training samples to obtain a wafer defect detection model.

[0040] Among them, during the training process, the training loss information of the augmented samples is compared with the training loss information of all training samples within a time window to determine whether the augmented samples exceed the preset range after the data augmentation transformation; in the case where it is determined that the augmented samples exceed the preset range after the data augmentation transformation, the amplitude of the augmented samples participating in the calculation of the model's backpropagation loss is reduced.

[0041] In this application, in order to alleviate the negative impact of data augmentation operations on model training, a conditional data augmentation strategy is proposed, and the context information is used to conditionally determine the data augmentation results during the model training process.

[0042] Specifically, during the training process of the model, the conditional data augmentation strategy sets the process of applying data augmentation as the execution result of conditional determination, and uses the context information during the model training process to determine whether the result of the current sample after the data augmentation transformation exceeds the expected range. This is a self-constrained learning process.

[0043] During the ordinary model training process, the obtained training loss information is directly discarded after backpropagation. In this application, the training loss information of each training sample within a time window is statistically used as context information.

[0044] In this application, the training loss information of the enhanced sample is compared with the training loss information of all training samples within a time window of it to determine whether the enhanced sample exceeds a preset range after data augmentation transformation, including: statistically counting the training loss information of the training samples within the time window one by one based on a global variable to obtain a statistical value, and the statistical value is iteratively updated in the form of a sliding window, calculating the mean of the statistical value to obtain the average loss value of the training samples within the current time window of the model; for the enhanced sample within the current time window, it is determined whether the loss value of the enhanced sample under the current model calculation exceeds a preset interval compared with the average loss value.

[0045] In the case where it is determined that the enhanced sample exceeds the preset range after data augmentation transformation, the amplitude of the enhanced sample participating in the model backpropagation loss calculation is reduced, including: In the case where the loss value of the enhanced sample under the current model calculation exceeds the preset interval compared with the average loss value, a first weight factor F1 is set according to the ratio of the loss value of the enhanced sample under the current model calculation to the average loss value; Based on the first weight factor F1, the degree of the enhanced sample participating in the model backpropagation loss calculation is controlled.

[0046] In this application, during the training process, a second weight factor F2 can also be obtained according to the value of the current iteration number of the model and the total iteration number of the model, and based on the second weight factor F2, the degree of the enhanced sample participating in the model backpropagation loss calculation is controlled.

[0047] Specifically, the conditional strategy of this application will determine conditional information based on two factors. First, during the model training process, the training loss information of each training sample will be statistically counted one by one according to the time window, and the loss value of the current enhanced sample will be compared with the statistical values of all loss information within a period of time window to determine whether the current enhanced sample is over-augmented. Second, the training progress is also a key factor. When the training starts, the amplitude of data augmentation allowed should be small to promote the stable training of the model; when the training gradually progresses to the later stage, a larger amplitude of data augmentation effect can be allowed to enable the model to see more diverse training samples.

[0048] In this application, when the preset model starts training, a global variable is used to statistically record the training loss information of each training sample. This part of the statistical value will be iteratively updated in the form of a sliding window. By taking the mean of this part of the statistical information, the average loss value of the model for each training sample within the current time window can be obtained. For an augmented sample after data augmentation transformation, if its loss value calculated by the current model exceeds a certain distance from the average loss value, a first weight factor F1 can be set according to the ratio of its loss to the average loss. The first weight factor F1 represents the amplitude of the augmented sample participating in the loss calculation of the model's backpropagation. When the loss value of the augmented sample calculated by the current model is much larger than the current average loss value, the larger the ratio, the smaller the calculated weight factor F1, and the lower the degree of the augmented sample participating in the loss calculation of the model's backpropagation in the future, which plays an inhibitory role on over-augmented samples.

[0049] On the other hand, during the model training process, a second weight factor F2 can be obtained according to the value of the current iteration number of the model and the total iteration number of the model. The value of F2 gradually increases as the model training progresses, indicating that the later the model training, the greater the tolerance for the data augmentation amplitude.

[0050] Multiply the two weight factors F1 and F2 to obtain the overall conditional training weight factor F = F1 × F2. This weight factor F is used to control the weights of all augmented samples after data augmentation transformation participating in the training.

[0051] In this application, the weight factor F will not be applied to the original training samples that have not undergone data augmentation transformation. Therefore, the training effect of the original training samples will not be affected.

[0052] In this application, the loss function L of the overall model training is divided into two parts, namely, for the samples that have not undergone data augmentation and the samples that have undergone data augmentation. Among them, the samples that have undergone data augmentation provide rich diversity and augmented samples for model training, but need to be controlled by a conditional data augmentation strategy to avoid harm to model training.

[0053] L = (1 - a)×L_ori + a × L_cond.

[0054] Among them, L_ori represents all the loss functions of the samples that have not undergone data augmentation, L_cond represents all the loss functions of the samples that have undergone data augmentation, and a is the proportion of the augmented samples that control data augmentation participating in the training. When a = 0, the model training process degenerates into the original defect detection training method. In practical applications, the value of a can be adjusted according to the specific situation of the actual application scenario of the model to balance the robustness and stability of the model.

[0055] As Figure 3 shown, it shows a schematic diagram of the training framework of the wafer defect detection model in the conditional data augmentation-based wafer defect detection method provided by the present application. Among them, x' represents the augmented sample, and the context loss value represents the average loss value of each training sample within a time window. Among them, the four small squares included in the left dashed box represent the sample features obtained based on the training samples. The second darker square shown in the figure corresponds to the augmented sample after excessive data augmentation. The four small squares included in the left dashed box represent the sample features obtained after the conditional augmentation strategy. The second hatched square shown in the figure indicates that the weight of the corresponding augmented sample after excessive data augmentation is reduced.

[0056] S103. Detect the wafer image to be detected based on the wafer defect detection model.

[0057] In the present application, during the model training process, by conditionally adding effective augmented samples, the robustness of the model can be improved while maintaining the model accuracy, thereby ensuring the effectiveness and practicality of the model in actual application scenarios.

[0058] In an optional implementation manner, the method further includes the following steps: S1. Obtain a test sample set, and based on the test sample set, obtain a first metric of the wafer defect detection model for the test sample set.

[0059] S2. Perform data augmentation on all samples in the test sample set to obtain an augmented sample set, and based on the augmented sample set, obtain a second metric of the wafer defect detection model for the augmented sample set.

[0060] S3. Obtain the augmentation deterioration score of the wafer defect detection model based on the difference between the second metric and the first metric.

[0061] S4. Evaluate the robustness of the wafer defect detection model based on the augmentation deterioration score.

[0062] In this application, a new evaluation metric is also proposed: the Enhancement Deterioration score (CD), which is used to evaluate the robustness of the wafer defect detection model under different data augmentation conditions. Specifically, for an existing wafer defect detection dataset X, corresponding data augmentation versions X-C can be generated for it by means of the data augmentation methods mentioned above. Different data augmentation versions can be generated for different data augmentation types and degrees. When it is necessary to comprehensively evaluate the robustness of the wafer defect detection model, the metrics on all augmented versions of a specific dataset can be evaluated, and the Enhancement Deterioration score CD is defined as the difference between the metric of the wafer defect detection model on the augmented version dataset X-C and its metric on the original dataset X.

[0063] CD = P1– P0; Where P1 is the metric of the wafer defect detection model on the data augmentation version X-C, and P0 is the metric of the method on the original version X of the dataset. The value of the Enhancement Deterioration score CD is usually negative because the performance of the wafer defect detection model is usually worse on the augmented version. The mean of the Enhancement Deterioration scores of the wafer defect detection model on all augmented version datasets can be regarded as the evaluation criterion for its robustness. The larger this value is, the stronger the robustness of the method.

[0064] In this application, the metric can be a commonly used model evaluation metric, such as AUROC, the area under the receiver operating characteristic curve.

[0065] In this application, the robustness of the wafer defect detection model can be evaluated based on the Enhancement Deterioration score. When the robustness of the wafer defect detection model meets the preset standard, it is determined that the wafer defect detection model is qualified. Based on this qualified wafer defect detection model, the wafer image to be detected can be detected to obtain more accurate detection results.

[0066] In this application, experiments were conducted on the general datasets MVTec AD and MVTec LOCO to verify the effectiveness of the method. For the datasets MVTec AD and MVTec LOCO, their corresponding data augmentation versions MVTec AD-C and MVTec LOCO-C were generated respectively, and the conditional data augmentation-based wafer defect detection method provided in this application was used for experiments. The experimental results are as follows: Table 1 Effects of the conditional data augmentation-based wafer defect detection method on MVTec AD and its augmented version

[0067] In Table 1, the experimental data for MVTec AD has the AUROC as the metric, and for the experimental data of MVTec AD-C, the two data separated by a slash are the AUROC and the enhanced deterioration score respectively. It can be seen from Table 1 that compared with the original method, the conditional data augmentation wafer defect detection method has almost the same accuracy as the original version, but the metrics on the data augmentation version are nearly 10% higher than the original method. At the same time, the enhanced deterioration score CD of the conditional data augmentation wafer defect detection method is larger than that of the original method, only -0.040, which means that the metrics obtained by this application in different scenarios are relatively stable and can adapt to various style changes that may occur in the scanning imaging during the wafer defect detection process.

[0068] Similarly, the experimental metrics (AUROC) of the conditional data augmentation wafer defect detection method on the MVTec LOCO dataset are shown in the following table: Table 2 Effects of the conditional data augmentation wafer defect detection method on MVTec AD and its enhanced version

[0069] In Table 2, the experimental data for MVTec LOCO has the AUROC as the metric, and for the experimental data of MVTec LOCO-C, the two data separated by a slash are the AUROC and the enhanced deterioration score respectively. It can be seen that similar to the MVTec AD dataset, the effects of the conditional data augmentation wafer defect detection method provided by this application on the MVTec LOCO dataset also prove that this application can use the conditional data augmentation strategy to obtain a more robust defect detection model.

[0070] Based on the same inventive concept, this application also provides a conditional data augmentation wafer defect detection system, and the conditional data augmentation wafer defect detection system includes: A receiving module, configured to receive training samples, perform data augmentation on some of the training samples according to a preset probability to obtain augmented samples; A training module, configured to train a preset model based on the training samples to obtain a wafer defect detection model. During the training process, compare the training loss information of the augmented samples with the training loss information of all training samples within a time window of it to determine whether the augmented samples exceed a preset range after the data augmentation transformation; in the case where it is determined that the augmented samples exceed the preset range after the data augmentation transformation, reduce the amplitude of the augmented samples participating in the model backpropagation loss calculation; A detection module, configured to detect a wafer image to be detected based on the wafer defect detection model.

[0071] Optionally, during the training process, compare the training loss information of the augmented sample with the training loss information of all training samples within a time window of it to determine whether the augmented sample exceeds a preset range after the data augmentation transformation, including: Based on global variables, count the training loss information of training samples within the time window one by one to obtain a statistical value. The statistical value is iteratively updated in the form of a sliding window, and the mean value of the statistical value is obtained to obtain the average loss value of the training samples within the current time window of the model. For the augmented sample within the current time window, determine whether the loss value of the augmented sample under the current model calculation exceeds a preset interval compared with the average loss value.

[0072] Optionally, in the case where it is determined that the augmented sample exceeds the preset range after the data augmentation transformation, reduce the amplitude of the augmented sample participating in the model backpropagation loss calculation, including: In the case where the loss value of the augmented sample under the current model calculation exceeds the preset interval compared with the average loss value, set a first weight factor F1 according to the ratio of the loss value of the augmented sample under the current model calculation to the average loss value; Based on the first weight factor F1, control the degree of the augmented sample participating in the model backpropagation loss calculation.

[0073] Optionally, during the training process, obtain a second weight factor F2 according to the value of the current iteration number of the model and the total iteration number of the model, and based on the second weight factor F2, control the degree of the augmented sample participating in the model backpropagation loss calculation.

[0074] Optionally, the loss function of the preset model is: L = (1 - a)×L_ori + a × L_cond where L_ori represents all loss functions of samples without data augmentation, L_cond represents all loss functions of augmented samples after data augmentation, and a represents the proportion of augmented samples participating in training.

[0075] Optionally, the device further includes: A first metric acquisition module, configured to acquire a test sample set, and based on the test sample set, obtain a first metric of the wafer defect detection model for the test sample set; A second metric acquisition module, configured to perform data augmentation on all samples in the test sample set to obtain an augmented sample set, and based on the augmented sample set, obtain a second metric of the wafer defect detection model for the augmented sample set; An enhanced deterioration score determination module, configured to obtain an enhanced deterioration score of the wafer defect detection model based on the difference between the second metric and the first metric; An evaluation module, configured to evaluate the robustness of the wafer defect detection model based on the enhanced deterioration score.

[0076] Based on the same inventive concept, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the conditional data augmentation-based wafer defect detection method according to any one of the above embodiments are implemented.

[0077] Based on the same inventive concept, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the conditional data augmentation-based wafer defect detection method according to any one of the above embodiments are implemented.

[0078] Based on the same inventive concept, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps in the conditional data augmentation-based wafer defect detection method according to any one of the above embodiments are implemented.

[0079] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0080] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0081] The present application is described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate a system for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 of one or more processes and / or blocks Figure 1 specified in the flowcharts or block diagrams.

[0083] These computer program instructions can also be loaded onto a computer or other programmable terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 of one or more processes and / or blocks Figure 1 specified in the flowcharts or block diagrams.

[0084] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0085] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0086] The above has introduced in detail a method for wafer defect detection with conditional data enhancement provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A wafer defect detection method with conditional data enhancement, characterized in that: The method comprises: Receiving training samples, and performing data enhancement on some samples in the training samples according to preset probabilities to obtain enhanced samples; Based on the training samples, a preset model is trained to obtain a wafer defect detection model. During the training process, the training loss information of the enhanced sample is compared with the training loss information of all training samples within a time window to determine whether the enhanced sample exceeds a preset range after the data enhancement transformation; if it is determined that the enhanced sample exceeds the preset range after the data enhancement transformation, the amplitude of the enhanced sample participating in the back propagation loss calculation of the model is reduced; The wafer image to be inspected is inspected based on the wafer defect detection model.

2. The wafer defect detection method with conditional data enhancement according to claim 1, characterized in that: During the training process, the training loss information of the enhanced sample is compared with the training loss information of all training samples in a time window to determine whether the enhanced sample exceeds a preset range after data enhancement transformation, including: Based on the global variables, the training loss information of the training samples in the time window is counted one by one to obtain the statistical value, which is iteratively updated in the form of a sliding window. The statistical value is averaged to obtain the average loss value of the training samples in the current time window of the model; for the enhanced samples in the current time window, it is determined whether the loss value of the enhanced sample calculated by the current model exceeds the preset interval compared with the average loss value.

3. The wafer defect detection method with conditional data enhancement according to claim 2, characterized in that: When it is determined that the enhanced sample exceeds a preset range after the data enhancement transformation, reducing the extent to which the enhanced sample participates in the back propagation loss calculation of the model, including: When the loss value of the enhanced sample calculated under the current model exceeds the preset interval compared with the average loss value, a first weight factor F1 is set according to the ratio of the loss value of the enhanced sample calculated under the current model to the average loss value; Based on the first weight factor F1, the degree to which the enhanced sample participates in the calculation of the model back propagation loss is controlled.

4. The wafer defect detection method with conditional data enhancement according to claim 1, characterized in that: The method also includes: during the training process, obtaining a second weight factor F2 according to the value of the current number of model iterations and the total number of model iterations, and based on the second weight factor F2, controlling the degree to which the enhanced samples participate in the calculation of the model back propagation loss.

5. The wafer defect detection method with conditional data enhancement according to claim 1, characterized in that: The loss function of the preset model is: L = (1-a)×L_ori + a × L_cond Among them, L_ori represents all loss functions of samples without data augmentation, L_cond represents all loss functions of enhanced samples after data augmentation, and a represents the proportion of enhanced samples participating in training.

6. The wafer defect detection method with conditional data enhancement according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquire a test sample set, and obtain a first metric of a wafer defect detection model for the test sample set based on the test sample set; Performing data enhancement on all samples in the test sample set to obtain an enhanced sample set, and obtaining a second metric of a wafer defect detection model for the enhanced sample set based on the enhanced sample set; Obtaining an enhanced deterioration score of the wafer defect detection model based on a difference between the second metric and the first metric; The robustness of the wafer defect detection model is evaluated based on the enhanced degradation score.

7. A wafer defect detection system with conditional data enhancement, characterized in that: The wafer defect detection system with conditional data enhancement includes: A receiving module is used to receive training samples, and perform data enhancement on some samples in the training samples according to a preset probability to obtain enhanced samples; A training module is used to train a preset model based on the training samples to obtain a wafer defect detection model. During the training process, the training loss information of the enhanced sample is compared with the training loss information of all training samples within a time window to determine whether the enhanced sample exceeds a preset range after the data enhancement transformation; if it is determined that the enhanced sample exceeds the preset range after the data enhancement transformation, the amplitude of the enhanced sample participating in the back propagation loss calculation of the model is reduced; The detection module is used to detect the wafer image to be detected based on the wafer defect detection model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the wafer defect detection method with conditional data enhancement described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the wafer defect detection method with conditional data enhancement described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps in the wafer defect detection method with conditional data enhancement described in any one of claims 1 to 6 are implemented.