Wafer defect detection method based on one-shot
Through the twin network and meta-learning strategy based on one-shot learning, combined with dynamic threshold optimization, the accuracy and efficiency of wafer defect detection under the condition of scarcity of samples is solved, and efficient and accurate defect detection and new task adaptation are achieved.
Patent Information
- Application Number
- CN202510480864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing wafer defect detection algorithms are difficult to achieve high-precision and efficient detection under scarcity of samples, and traditional methods are sensitive to light changes and are difficult to adapt to complex background environments.
The wafer defect detection method based on one-shot learning is adopted, through twin networks and meta-learning strategies, combined with dynamic threshold optimization, the dependence on large-scale annotation data is reduced, and the generalization ability and adaptability of the model are improved.
It realizes rapid extraction of features under a small number of samples, reduces error detection rates and missed detection rates, improves detection efficiency and accuracy, and adapts to new tasks quickly.
Smart Images

Figure CN120013929A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular relates to a wafer defect detection method based on one-shot learning. Background Art
[0002] Wafer defect detection is an extremely critical part of the chip manufacturing process, which directly affects the yield and reliability of the chip. In the actual production process, the wafer defect detection algorithm faces multiple technical challenges. First, the accuracy and robustness of the detection algorithm are crucial. Due to the small size of wafer defects, high-precision algorithms are required to ensure high recall rates and low false alarm rates; second, there are many types of wafer defects, complex manifestations, and they are greatly affected by factors such as process and materials. Traditional image processing methods are difficult to adapt to complex defects; finally, due to the particularity and complexity of wafer defect detection, defect samples are relatively scarce and difficult to obtain. Therefore, it is particularly important to design a defect detection algorithm with good adaptability under the condition of scarce samples.
[0003] The existing mainstream wafer surface defect detection methods can be roughly divided into traditional optical image processing methods and deep learning-based target detection algorithms. Traditional image processing methods include threshold segmentation, morphological processing, etc. Most of these methods rely on manually set threshold parameters and are sensitive to environmental changes such as lighting. These problems make them unable to adapt to complex background environments and complex defects and cannot be efficiently automated. The YOLO series of algorithms are the most popular deep learning-based target detection algorithms. These algorithms usually require a lot of data training to optimize model performance. However, due to the particularity and complexity of wafer defect detection, similar public data sets are very scarce, and it is also difficult to obtain a large number of defect samples for training in actual production processes. Summary of the invention
[0004] In order to make up for the shortcomings of the prior art, the present invention aims to provide a wafer defect detection method based on one-shot learning to reduce the algorithm's dependence on large-scale labeled data and complete feature extraction with only a small number of samples; and combines dynamic threshold optimization to enhance the model's ability to capture tiny defects and reduce the false detection rate of noise; in addition, a meta-learning method is introduced to enhance a small number of samples through synthetic data to improve the generalization ability of the model.
[0005] The technical problem solved by the present invention can be achieved through the following specific technical solutions: The one-shot wafer defect detection method comprises the following steps: Step 1: preprocess the data and use the preprocessed data as network input; Step 2: Construct a twin network dual-branch structure; Step 3: Meta-learning strategy training model; Step 4: Use dynamic threshold for detection.
[0006] Furthermore, in step 1, the data includes the following three categories: The first category is the reference sample, which contains a single defect-free wafer image, which is an RGB (red, green and blue three-channel) image with a resolution of 2048×2048 pixels; The second category is the sample to be tested, which includes a single image of the wafer to be tested from the same batch or the same process conditions as the reference sample; The third category is meta-training task data, which contains a small number of historical defect samples and is used for meta-learning training.
[0007] Furthermore, in step 1, the preprocessing includes standardization, image quality enhancement and data enhancement. The standardization adopts grayscale normalization. The purpose of image quality enhancement is to improve the image signal-to-noise ratio. The image sharpness is calculated to determine whether to perform filtering and noise reduction on the image. The sharpness calculation formula is as follows: (1) In the formula, is the Laplace operator, N Refers to the total number of pixels, x and y is the pixel horizontal and vertical coordinates, I ( x , y ) represents the pixel gray value, S Indicates the sharpness value, S The smaller the value, the lower the sharpness of the image and the blurrier the image. S <0.2, use the following filter function to process the image: (2) In the formula, I Represents the original image, I deblur represents the denoised image, F and F -1 denote Fourier transform and inverse transform respectively, F(I) Indicates Fourier transform of the original image; K is the bias constant, K =0.01; H(u,v) is the point spread function, which is directly generated by the mathematical model, yes H(u,v) The complex conjugate of .
[0008] Furthermore, a data enhancement method combining traditional image transformation and physical simulation defects is adopted. On the one hand, the diversity of data is enhanced by transforming the image; on the other hand, physical simulation defects based on morphological operations are generated based on reference samples.
[0009] Furthermore, in step 2, the ResNet-18 network is selected as the default backbone network of the dual-branch structure, the input image is processed into a grayscale image in the preprocessing stage, and the number of input channels is changed to a single channel; the output is a feature vector for calculating cosine similarity, the classification head of the original ResNet-18 is removed, replaced with a global average pooling, and a 256-dimensional feature vector is output using a fully connected layer; and a drop layer is set before the fully connected layer and the drop rate is set to 0.3. The cosine similarity calculation formula is: (3) In the formula, f a , f b Represent the feature vectors of the reference image and the image to be detected respectively, ||fa|| and ||fb|| represent the L2 norm of the two vectors respectively; the output range of cosine similarity S is [-1,1], which is mapped to [0,1] through Sigmoid (an activation function) as the final similarity score.
[0010] Furthermore, in step 3, the specific content of the meta-learning strategy training model is as follows: ① Construct tasks. The support set of each task is one defect-free image and one defective image, and the query set is 5 images; ② Randomly initialize the parameters of the backbone network ResNet-18, sample 32 tasks per batch, set the meta-learning rate to 0.001, and set the task internal learning rate to 0.01; ③ Perform internal loop training within each task to update network parameters, and perform external loop training between different tasks; ④ After each round of training, select difficult samples with a similarity between [0.4, 0.6] and increase their sampling weight in the next round of training. If the loss value does not decrease for 5 consecutive rounds, terminate the training.
[0011] Furthermore, the loss function used by the meta-learning strategy training model is for: (4) In the formula, L contrastive , L triplet They are contrast loss and triplet loss, contrast loss L contrastiveThe specific expression is as follows: (5) In the formula, x a , x b is an input sample pair, y i is the sample pair label, d ( x a , x b ) represents the feature distance of the sample pair, reflecting the similarity of the sample pair; m is the preset boundary value, N is the total number of sample pairs; triple loss L triplet The expression is as follows: (6) In the formula, x a represents the anchor point sample, x p represents a positive sample, x n represents negative samples, d ( a , b )express a and b Feature distance between; triplet loss L triplet The role of is to make the distance between the reference sample and the positive sample smaller than the distance between the negative sample by at least one boundary value m , usually set between 0.2 and 1.0.
[0012] Furthermore, in step 4, the specific contents of using the dynamic threshold for detection are as follows: The image to be detected and the reference image are input into the network to extract the feature vector, and then the cosine similarity is calculated. If the cosine similarity value is lower than the threshold, it means that a defect is detected. The similarity threshold changes dynamically with the image sharpness value. The calculation formula of the sharpness value S is shown in formula (1), and the similarity threshold T is determined by the following formula: (7) In the formula, T is the similarity threshold, T base is the basic threshold, generally taken T base =0.3, k is the adjustment coefficient, generally taken as k =0.1, S normRefers to the normalized sharpness value. When the image sharpness value is low, the image is blurry, and the similarity threshold should be set to a higher value to reduce noise interference; when the image sharpness value is high, the feature distinction is strong, and the threshold can be lowered to reduce false detection.
[0013] Compared with the prior art, the present invention has the following advantages: (1) Reduced algorithm dependence on large-scale labeled data: The present invention only requires one reference sample and a small number of defect samples in conjunction with data enhancement technology to quickly extract sample features and complete model training, which significantly increases detection efficiency while reducing the algorithm's dependence on large-scale labeled data.
[0014] (2) Adaptive similarity threshold: The present invention dynamically adjusts the similarity threshold by using the image sharpness value. Compared with the traditional method of manually setting the threshold, it has a higher degree of automation and can effectively reduce the false detection rate and missed detection rate.
[0015] (3) Good adaptability to new tasks: Through the meta-learning strategy, when encountering a new task, the present invention only needs to perform one internal loop to fine-tune the model to quickly adapt to the new task. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of the steps of the detection method of the present invention; Figure 2 A schematic diagram of physically simulating different types of defects under the same wafer background of the present invention; Figure 3 This is a schematic diagram of the dual-branch structure of the twin network of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] In the field of semiconductor manufacturing, improving the accuracy and efficiency of wafer defect detection algorithms is one of the key challenges that need to be solved urgently. The application background of the present invention focuses on achieving accurate and efficient detection of wafer defects under the condition of scarce samples or even single samples.
[0019] like Figure 1 As shown, a wafer defect detection method based on one-shot learning includes data input and preprocessing, twin network construction, model training and dynamic threshold reasoning. The present invention combines twin networks with meta-learning to achieve rapid adaptation of small samples and introduces a one-shot learning method in wafer defect detection. The specific contents are as follows: (A) Data input and preprocessing.
[0020] The input data can be divided into the following three categories: the first category is reference samples, which include 10 defect-free wafer images from the same process batch provided by semiconductor manufacturers with an image resolution of 2048×2048. After grayscale conversion, they are used as the benchmark reference library. During actual inspection, one of them is selected as the reference image for the current task; the second category is samples to be inspected, which includes 100 images of wafers from the same batch, of which 50 are defective images containing different types of defects and 50 are defect-free images; the third category is meta-training task data, which includes 30 defect samples of 5 categories collected historically, and the defect types include scratches, particles, bumps, line width abnormalities, and edge damage.
[0021] Preprocessing includes three aspects: standardization, image quality enhancement, and data enhancement. The goal of standardization is to reduce the impact of equipment differences and shooting conditions and unify the input feature space. In this invention, a simple grayscale normalization process is used. The purpose of image quality enhancement is to improve the image signal-to-noise ratio. In this invention, the sharpness is calculated to determine whether to filter and reduce noise on the image. The sharpness calculation formula is as follows: (1) In the formula, is the Laplace operator, N Refers to the total number of pixels, x and y is the pixel horizontal and vertical coordinates, I ( x , y ) represents the pixel gray value, S Indicates the sharpness value, S The smaller the value, the lower the sharpness of the image and the blurrier the image. S <0.2, use the following filter function to process the image: (2) In the formula, I Represents the original image, I deblur represents the denoised image, F and F -1 denote Fourier transform and inverse transform respectively, F(I) Indicates Fourier transform of the original image; K is the bias constant, K =0.01; H(u,v) is the point spread function, which is directly generated by the mathematical model, yes H(u,v) The complex conjugate of .
[0022] In order to meet the model generalization requirements under small sample conditions, the present invention adopts a data enhancement method that combines traditional image transformation with physical simulation defects. On the one hand, the existing defect images are randomly rotated (plus or minus 15 degrees), horizontally / vertically flipped, and scaled (0.8 to 1.2 times) and other image transformations are performed; on the other hand, based on the reference sample, physical simulation defects based on morphological operations are generated, including scratches of random length (5 to 50 pixels) and particle pollution of random diameter (3 to 20 pixels). The physical simulation defect effect is as follows: Figure 2 As shown, 500 enhanced samples are generated.
[0023] 2. Hardware environment The hardware configuration used for deep learning network training in the example of the present invention is: NVIDIA 4070SUPER GPU (12GB video memory), AMD 9700X CPU (6 cores and 12 threads), 32GB memory; the deep learning framework used is pytorch 2.6.0, CUDA 12.6.
[0024] 3. Twin network training.
[0025] The twin network in the present invention is the core feature extraction module, which aims to extract the feature vectors of the reference image and the image to be detected through a dual-branch structure, and then use the similarity comparison between the feature vectors to combine meta-learning to achieve small sample defect detection. The following is a detailed description: like Figure 3 The figure shows the overall structure of the network. Different from the traditional deep learning network, the network in the present invention has two input branches. The reference image branch inputs the defect-free wafer image, and the to-be-detected branch inputs the to-be-detected wafer image. It is worth noting that the two branches must use the same convolutional neural network structure and share all parameters to ensure the consistency of feature extraction.
[0026] This application selects ResNet-18 (a residual network) as the default backbone network of the dual-branch structure. The 18-layer structure has an appropriate depth, balances feature expression capabilities and computational efficiency, and is suitable for processing high-resolution wafer images; residual connections alleviate the gradient vanishing problem, improve training stability, and are suitable for sample scarcity problems. The input and output formats of the ResNet-18 network also need to be adjusted to adapt to the wafer defect detection task. Specifically, the input image is processed as a grayscale image in the preprocessing stage, so the number of input channels is changed to a single channel; the output should be a feature vector that is convenient for calculating cosine similarity, so the classification head of the original ResNet-18 is removed and replaced with global average pooling, and a 256-dimensional feature vector is output using a fully connected layer. In order to improve the generalization ability of the model, a dropout layer is set before the fully connected layer and the dropout rate is set to 0.3.
[0027] The cosine similarity calculation formula of the present invention is: (3) In the formula, f a , f b Represent the feature vectors of the reference image and the image to be detected respectively, ||fa|| and ||fb|| represent the L2 norm of the two vectors respectively; the output range of cosine similarity S is [-1,1], which is mapped to [0,1] through Sigmoid (an activation function) as the final similarity score.
[0028] 4. Meta-learning strategy training process.
[0029] The model training of the present invention combines the twin network and the meta-learning method to achieve high-precision defect detection under small sample conditions through multi-task optimization. The following is a detailed step-by-step description of the training process: (1) Construct tasks. The support set of each task is 1 defect-free image (positive sample) and 1 defect image (negative sample), and the query set is 5 images (3 positive samples + 2 negative samples) to evaluate the generalization ability of the task.
[0030] (2) The parameters of the backbone network (ResNet-18) were randomly initialized, 32 tasks were sampled in each batch, the meta-learning rate was set to 0.001, the intra-task learning rate was set to 0.01, and the training rounds were 200.
[0031] (3) Internal loop training is performed within each task to update network parameters, aiming to improve the model's ability to distinguish image similarities and to force the feature similarity values of positive sample pairs within the same task to be high and the feature similarity of negative sample pairs to be low; external loop training is performed between different tasks to improve the model's ability to quickly adapt to new tasks and reduce the number of samples required for adaptation to new tasks.
[0032] (4) After each round of training, difficult samples with a similarity between [0.4, 0.6] are selected, and the sampling weight is increased by 50% in the next round of training; if the loss value does not decrease for five consecutive rounds, the training is terminated.
[0033] The loss function is a key factor in network training. The loss function used in this invention is for: (4) Where, L contrastive , L triplet They are contrast loss and triplet loss respectively. The specific expression of contrast loss is as follows: (5) In the formula, xa , x b is an input sample pair, y i is the sample pair label, d ( x a , x b ) represents the feature distance of the sample pair, reflecting the similarity of the sample pair; m is the preset boundary value, N is the total number of sample pairs; the role of contrast loss is to shorten the distance between similar sample pairs (positive sample pairs) and to extend the distance between dissimilar samples (negative sample pairs). L triplet The expression is as follows: (6) In the formula, x a represents the anchor point sample, x p represents a positive sample, x n represents negative samples, d ( a , b )express a and b Feature distance between; triplet loss L triplet The role of is to make the distance between the reference sample and the positive sample smaller than the distance between the negative sample by at least one boundary value m , set to 0.5.
[0034] (V) Use dynamic threshold for detection.
[0035] After the network training is completed, the new task can be inferred. In the inference stage, the new task must first be adapted: input a defect-free reference image of the new process wafer and a defective sample, and then perform an internal update to fine-tune the model. After the new task adaptation is completed, the image to be detected and the reference image can be input into the network to extract the feature vector, and then the cosine similarity calculation is performed. If the cosine similarity value is lower than the threshold, it means that a defect has been detected. The similarity threshold changes dynamically with the image sharpness value. The calculation formula of the sharpness value S is shown in formula (1), and the similarity threshold T is determined by the following formula: (7) In the formula, is the similarity threshold, As the basic threshold, , is the adjustment coefficient, , Refers to the normalized sharpness value. When the image sharpness value is low, the image is blurry, and the similarity threshold will be set to a higher value by formula (7) to reduce noise interference; when the image sharpness value is high, the feature distinction is strong, and the threshold can be lowered to reduce false detection.
[0036] The performance of the present invention is compared with that of the traditional fixed threshold method (threshold 0.5) and YoloV5 (1000 defective samples) on 100 test samples. The results are as follows:
[0037] From the above data, it can be seen that the present invention achieves an accuracy of 92.3% with only 30 original defect samples combined with data enhancement, which reduces the labeled data by 97% compared to YoloV5. The dynamic threshold reduces the false detection rate to 7.7% and the missed detection rate to 12.4%. In summary, the feasibility and effectiveness of the technical solution are verified by the present invention through specific sample configuration, data enhancement parameters, training process and quantitative experiments.
[0038] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A one-shot wafer defect detection method, characterized in that: The following steps are involved: Step 1: preprocess the data and use the preprocessed data as network input; Step 2: Construct a twin network dual-branch structure; Step 3: Meta-learning strategy training model; Step 4: Use dynamic threshold for detection.
2. The one-shot wafer defect detection method according to claim 1, characterized in that: In step 1, the data includes the following three categories: The first category is the reference sample, which contains a single defect-free wafer image, which is an RGB image with a resolution of 2048×2048 pixels; The second category is the sample to be tested, which includes a single image of the wafer to be tested from the same batch or the same process conditions as the reference sample; The third category is meta-training task data, which contains a small number of historical defect samples and is used for meta-learning training.
3. The one-shot wafer defect detection method according to claim 1, characterized in that: In step 1, the preprocessing includes standardization, image quality enhancement and data enhancement. The standardization adopts grayscale normalization. The purpose of image quality enhancement is to improve the image signal-to-noise ratio. The image sharpness is calculated to determine whether to perform filtering and noise reduction on the image. The sharpness calculation formula is as follows: (1) In the formula, is the Laplace operator, N Refers to the total number of pixels, x and y is the pixel horizontal and vertical coordinates, I ( x , y ) represents the pixel gray value, S Indicates the sharpness value; if S <0.2, use the following filter function to process the image: (2) In the formula, I Represents the original image, I deblur represents the denoised image, F and F -1 denote Fourier transform and inverse transform respectively, F(I) Indicates Fourier transform of the original image; K is the bias constant, K =0.01; H(u,v) is the point spread function, which is directly generated by the mathematical model, yes H(u,v) The complex conjugate of .
4. The one-shot wafer defect detection method according to claim 3, characterized in that: A data enhancement method combining traditional image transformation and physical simulation defects is adopted. On the one hand, the diversity of data is enhanced by image transformation; on the other hand, physical simulation defects based on morphological operations are generated based on reference samples.
5. The one-shot wafer defect detection method according to claim 1, characterized in that: In step 2, the ResNet-18 network is selected as the default backbone network of the dual-branch structure, the input image is processed into a grayscale image in the preprocessing stage, and the number of input channels is changed to a single channel; the output is a feature vector for calculating cosine similarity, the classification head of the original ResNet-18 is removed, replaced with a global average pooling, and a 256-dimensional feature vector is output using a fully connected layer; and a drop layer is set before the fully connected layer and the drop rate is set to 0.
3. The cosine similarity calculation formula is: (3) in, f a , f b Represent the feature vectors of the reference image and the image to be detected respectively, ||fa|| and ||fb|| represent the L2 norm of the two vectors respectively; the output range of cosine similarity S is [-1,1], which is mapped to [0,1] by the activation function Sigmoid as the final similarity score.
6. The one-shot wafer defect detection method according to claim 1, characterized in that: In step 3, the specific content of the meta-learning strategy training model is as follows: ① Construct tasks. The support set of each task is one defect-free image and one defective image, and the query set is 5 images; ② Randomly initialize the parameters of the backbone network ResNet-18, sample 32 tasks per batch, set the meta-learning rate to 0.001, and set the task internal learning rate to 0.01; ③ Perform internal loop training within each task to update network parameters, and perform external loop training between different tasks; ④ After each round of training, select difficult samples with a similarity between [0.4, 0.6] and increase their sampling weight in the next round of training. If the loss value does not decrease for 5 consecutive rounds, terminate the training.
7. The one-shot wafer defect detection method according to claim 6, characterized in that: The loss function used by the meta-learning strategy training model for: (4) in, L contrastive , L triplet They are contrast loss and triplet loss, contrast loss L contrastive The specific expression is as follows: (5) In the formula, x a , x b is an input sample pair, y i is the sample pair label, d ( x a , x b ) represents the feature distance of the sample pair, reflecting the similarity of the sample pair; m is the preset boundary value, N is the total number of sample pairs; triple loss L triplet The expression is as follows: (6) In the formula, x a represents the anchor point sample, x p represents a positive sample, x n represents negative samples, d ( a , b )express a and b Feature distance between; triplet loss L triplet The role of is to make the distance between the reference sample and the positive sample smaller than the distance between the negative sample by at least one boundary value m , usually set between 0.2 and 1.
0.
8. The one-shot wafer defect detection method according to claim 1, characterized in that: In step 4, the specific contents of using the dynamic threshold for detection are as follows: The image to be detected and the reference image are input into the network to extract the feature vector, and then the cosine similarity is calculated. If the cosine similarity value is lower than the threshold, it means that a defect has been detected. The similarity threshold changes dynamically with the image sharpness value. The calculation formula of the sharpness value S is shown in formula (1). The similarity threshold T Determined by the following formula: (7) In the formula, T is the similarity threshold, T base As the basic threshold, T base =0.3; k is the adjustment coefficient, k =0.1, S norm Refers to the normalized sharpness value.
Citation Information
Patent Citations
Industrial product large defect detection method based on twin network
CN113902939A
Detection algorithm for few sample defects in QFN chip
CN114937005A
Small sample oil storage tank bottom plate defect detection method based on improved twin network
CN117576035A
Defect detection method based on deep contrast learning
CN118967690A
Microchip appearance defect detection method based on convolutional neural network
CN119693363A
Cited By
Circuit board visual quality inspection system and method based on AOI
CN120369744A
Printed circuit board defect detection method and system based on cross-modal prompt learning and visual guidance
CN120471929A
A printed circuit board defect detection method and system based on cross-modal prompt learning and visual guidance
CN120471929B