Wafer defect detection method based on multi-core attention fusion and dynamic large-core jumping
By using multi-core attention fusion and dynamic large-core jump methods in wafer defect detection, the existing detection methods are solved, and the problems of inefficiency and susceptibility to subjective factors are achieved, and defect detection with high accuracy and efficiency are achieved, and good generalization capabilities are provided.
Patent Information
- Application Number
- CN202510655963.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing wafer defect detection methods are inefficient, susceptible to subjective factors, and are time-consuming and labor-intensive when facing massive data, making it difficult to meet the needs of efficient production.
The wafer defect detection method based on multi-core attention fusion and dynamic large-core jump is adopted. Through the U-Net structure and multi-core channel segmentation attention perception fusion, the characteristics of different scales and positions are deeply mined and fused, and the dynamic attention nucleus jump connection convolution is combined to improve the accuracy and efficiency of defect detection.
It significantly improves the accuracy and efficiency of defect detection, reduces the error rate of mask generation, has good generalization ability, and can adapt to different types of wafer defect detection tasks.
Smart Images

Figure CN120182269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting wafer defects, specifically a method for detecting wafer defects based on multi-core attention fusion and dynamic large-core jumping, belonging to the technical field of semiconductor manufacturing. Background Art
[0002] In recent years, with the continuous progress of semiconductor manufacturing technology, the demand for wafer defect detection has become increasingly urgent. Defects on the wafer surface not only seriously affect the yield and performance of chips, but may also lead to quality fluctuations in the large-scale production process. Therefore, accurately and efficiently identifying these defects is crucial for ensuring production quality. Common defect types include edge defects, central defects, and randomly distributed defects, and their diversity and complexity pose great challenges to the detection work. At the same time, with the continuous refinement of manufacturing processes, the amount of wafer data is growing rapidly, which further increases the difficulty of defect detection.
[0003] Traditionally, wafer defect detection relied on engineers to analyze defects in images through visual inspection. This method is not only inefficient, easily interfered by subjective factors, resulting in missed or misjudged detections, but also time-consuming, laborious, and costly when faced with massive data, and can no longer meet the needs of high-efficiency production. Therefore, the industry has gradually started to explore more efficient and automated detection methods to improve the accuracy of defect identification and reduce operating costs.
[0004] In recent years, the rapid development of deep learning technology has brought new solutions to wafer defect detection. Systems based on convolutional neural networks can automatically learn complex features in images, achieve efficient classification of various defect types, and significantly improve the detection efficiency and accuracy. For example, on the one hand, the Residual Network (ResNet) uses residual blocks to deepen the network depth, effectively alleviating the gradient vanishing problem, and thus improving the ability to extract deep features; on the other hand, the Dense Connectivity Network (DenseNet) achieves maximum reuse of features through a dense connection structure, reducing redundant parameters. These methods have made important breakthroughs in improving detection performance, but at the same time, they also bring high computational resource consumption and long training time.
[0005] In addition, multi-scale feature extraction and attention mechanisms have also been introduced to enhance the model's ability to capture defect details. The multi-scale method helps to comprehensively capture the global and local features of defects by fusing feature information at different scales; while the attention mechanism improves the model's sensitivity to important defect regions by dynamically weighting key features. However, at present, these improved methods are often optimized independently, and fail to fully utilize lightweight network structures to reduce the amount of computation, and at the same time comprehensively use different scale information for detection. Summary of the Invention
[0006] Objective of the Invention: Aiming at the above problems, the objective of the present invention is to provide a wafer defect detection method based on multi-core attention fusion and dynamic large kernel jumping.
[0007] Technical Solution: The wafer defect detection method based on multi-core attention fusion and dynamic large kernel jumping of the present invention includes the following steps: Obtain the wafer image to be detected and perform image preprocessing; Input the preprocessed image into three feature extraction layers with the same structure in sequence for feature extraction to obtain the first feature map, the second feature map, and the third feature map respectively, and then input the extracted third feature map into the fourth feature extraction layer and the fifth feature extraction layer in sequence for feature extraction to obtain the fourth feature map and the fifth feature map; Input the first to fourth feature maps into the DLKConv-SC module respectively to obtain four feature maps with different sizes, which are respectively denoted as the sixth feature map, the seventh feature map, the eighth feature map, and the ninth feature map; Use the fifth to ninth feature maps for fusion decoding to obtain the defect segmentation result; Input the fifth feature map into the global average pooling layer and the activation function layer in sequence to obtain the defect classification result.
[0008] Further, the step of inputting the preprocessed image into three feature extraction layers with the same structure in sequence for feature extraction to obtain the first feature map, the second feature map, and the third feature map respectively, and then inputting the extracted third feature map into the fourth feature extraction layer and the fifth feature extraction layer in sequence for feature extraction to obtain the fourth feature map and the fifth feature map includes: Input the preprocessed image into the first MKCS-SE module for feature extraction, and then perform max pooling on the extracted feature map to obtain the first feature map; Input the first feature map into the second MKCS-SE module for feature extraction, and then perform max pooling on the extracted feature map to obtain the second feature map; Input the second feature map into the third MKCS-SE module for feature extraction, and then perform max pooling on the extracted feature map to obtain the third feature map; Input the third feature map into the fourth MKCS-SE module for feature extraction, and then perform average pooling on the extracted feature map to obtain the fourth feature map; Input the fourth feature map into the fifth MKCS-SE module for feature extraction, and then perform average pooling on the extracted feature map to obtain the fifth feature map.
[0009] Further, the step of inputting the first to fourth feature maps into the DLKConv-SC module respectively includes: Take one of the first feature map, the second feature map, the third feature map, or the fourth feature map as the input feature map, and input the input feature map into the first depthwise separable convolution and the second depthwise separable convolution respectively for large kernel feature extraction, obtaining the first large kernel feature map and the second large kernel feature map respectively. After fusing the first large kernel feature map and the second large kernel feature map, perform average pooling and max pooling respectively to obtain the third large kernel feature map and the fourth large kernel feature map; Input the third large kernel feature map and the fourth large kernel feature into the third depthwise separable convolution for channel fusion, and after activating the result of the channel fusion through the Sigmod function, obtain the fifth large kernel feature; Perform a Scale operation on the fifth large kernel feature map and the first large kernel feature map to obtain the seventh large kernel feature map; Perform a Scale operation on the fifth large kernel feature map and the second large kernel feature map to obtain the eighth large kernel feature map; Perform a merge operation on the seventh large kernel feature map, the eighth large kernel feature map, and the input feature map to obtain the feature map finally output by the DLKConv-SC module.
[0010] Further, the steps of using the fifth feature map to the ninth feature map for fusion decoding to obtain the defect segmentation result include: After fusing the ninth feature map and the fifth feature map, use the Transpose-Conv module for decoding to obtain the tenth feature map; After fusing the eighth feature map and the tenth feature map, use the Transpose-Conv module for decoding to obtain the eleventh feature map; After fusing the seventh feature map and the eleventh feature map, use the Transpose-Conv module for decoding to obtain the twelfth feature map; Fuse the sixth feature map and the twelfth feature map, and use the Transpose-Conv module for decoding to obtain the thirteenth feature map; Adjust the size of the thirteenth feature map through depthwise separable convolution, and use the Sigmod function to obtain the output segmentation result.
[0011] Further, the structures of the first MKCS-SE module to the fifth MKCS-SE module are the same, and the process of feature extraction includes: Perform preliminary feature extraction on the input feature map through depthwise separable convolution, and then adjust the channels of the obtained preliminary feature map through depthwise separable convolution to adjust the channels to a multiple of four; Then perform a four-way channel split on the adjusted feature map; Perform feature extraction at different scales on the feature maps obtained by the four-way channel split using four depthwise separable convolutions with different sizes; Merge the feature maps extracted at four different scales, perform a residual connection with the preliminary feature map, and input the feature map after the residual connection into the SE attention module; Input the feature map processed by the SE attention module into the depthwise separable convolution for channel adjustment, and adjust the channels to the preset number of channels for output.
[0012] Furthermore, the steps for image preprocessing include: Adjust the size of the defective image to a preset size, perform median filtering, and perform data augmentation for unbalanced categories.
[0013] Furthermore, data augmentation includes vertically or horizontally flipping the image, vertically or horizontally moving the image, and rotating or scaling the image.
[0014] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are: (1) The present invention adopts a U-Net structure. Through multi-core channel segmentation attention perception fusion, it deeply mines and fuses features at different scales and positions, extracts multi-scale image features, enabling the model to perform better when dealing with wafer defects of different scales and shapes, significantly improving the accuracy of defect detection. And introducing a dynamic attention large kernel skip connection convolution improves the accuracy of the model in the segmentation task and reduces the error rate of mask generation; The present invention has good generalization ability and can adapt to different types of wafer defect detection tasks, with strong practical value; (2) The present invention uses a lightweight architecture of the U-Net structure and a multi-core channel segmentation attention perception fusion structure, which can not only efficiently extract and process rich features in the wafer image, ensure high-quality feature extraction while reducing the computational complexity, but also can simultaneously perform multi-task collaboration of defect segmentation and classification of the wafer map; (3) By introducing a squeeze-and-excitation attention mechanism, the present invention can adaptively adjust the weights of each feature channel, enabling the model to pay more attention to key defect information and suppress irrelevant features, thereby enhancing the feature representation ability and detection performance. Description of the Drawings
[0015] Figure 1 is the overall structure diagram of the MK-DANet network in the specific implementation of the present invention; Figure 2 is the category picture of the wafer map defect dataset; Figure 3 is the result after median filtering of the category picture of the wafer map defect dataset; Figure 4 is the structural schematic diagram of the MKCS-SE feature extraction layer; Figure 5 It is a schematic structural diagram of the SE attention module; Figure 6 It is a specific structural diagram of the DALKConv-SC module. Specific implementation manner
[0016] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0017] The wafer defect detection method based on multi-core attention fusion and dynamic large kernel jump described in this embodiment has an overall block diagram as Figure 1 shown, and includes the following steps: Step 1, obtain the wafer image to be detected and perform image preprocessing.
[0018] Further, the steps of performing image preprocessing include: Adjust the size of the defect image to a preset size, such as it can be adjusted to , perform median filtering, and perform data augmentation for unbalanced categories.
[0019] Further, the data augmentation includes vertically or horizontally flipping the image, vertically or horizontally moving the image, and rotating or scaling the image.
[0020] In one example, obtain the detection wafer image and adjust the size, such as it can be adjusted to a three-channel RGB image, Figure 2 is a category picture of the wafer map defect dataset, which contains eight different wafer map defects. In the figure, Center represents the central defect, denoted as C; Donut represents the annular defect, denoted as D; Edge-Loc represents the edge local defect, denoted as EL; Edge-Ring represents the edge annular defect, denoted as ER; Local represents the local defect, denoted as L; Random represents the random defect, denoted as R; Scratch represents the scratch defect, denoted as S; Near-full represents the near-complete defect, denoted as NF. Then, perform median filtering on the input defect image, Figure 3 is the result after median filtering of the category picture of the wafer defect dataset. After filtering, perform data augmentation for unbalanced categories on the filtered image; the data augmentation includes vertically or horizontally flipping the image, vertically or horizontally moving the image, and slightly rotating or scaling the image. After preprocessing, it helps to have a high tolerance for changes in the size, position and direction of defects in the picture.
[0021] Step 2: The preprocessed image is successively input into three feature extraction layers with the same structure for feature extraction to obtain a first feature map, a second feature map, and a third feature map respectively. Then, the extracted third feature map is successively input into a fourth feature extraction layer and a fifth feature extraction layer for feature extraction to obtain a fourth feature map and a fifth feature map.
[0022] Further, the step of successively inputting the preprocessed image into three feature extraction layers with the same structure for feature extraction to obtain a first feature map, a second feature map, and a third feature map respectively, and then successively inputting the extracted third feature map into a fourth feature extraction layer and a fifth feature extraction layer for feature extraction to obtain a fourth feature map and a fifth feature map includes: Input the preprocessed image into the first MKCS-SE module for feature extraction. Then, after performing max pooling (MP) on the extracted feature map, obtain the first feature map. The number of filters f in the first MKCS-SE module is 8. Input the first feature map into the second MKCS-SE module for feature extraction. Then, after performing max pooling on the extracted feature map, obtain the second feature map. The number of filters f in the second MKCS-SE module is 16. Input the second feature map into the third MKCS-SE module for feature extraction. Then, after performing max pooling on the extracted feature map, obtain the third feature map. The number of filters f in the third MKCS-SE module is 16. Input the third feature map into the fourth MKCS-SE module for feature extraction. Then, after performing average pooling (AP) on the extracted feature map, obtain the fourth feature map. The number of filters f in the fourth MKCS-SE module is 32. Input the fourth feature map into the fifth MKCS-SE module for feature extraction. Then, after performing average pooling on the extracted feature map, obtain the fifth feature map. The number of filters f in the fifth MKCS-SE module is 64.
[0023] Further, the structures of the first MKCS-SE module to the fifth MKCS-SE module are the same. The process of feature extraction includes: Perform preliminary feature extraction on the input feature map through depthwise separable convolution. Then, adjust the channels of the obtained preliminary feature map through depthwise separable convolution to make the number of channels a multiple of four. Then, perform channel splitting of the adjusted feature map into four equal parts. Use depthwise separable convolutions with four different sizes to perform feature extraction at different scales on the feature maps obtained by channel splitting into four equal parts. Perform a merging operation on the feature maps extracted at four different scales, perform a residual connection with the preliminary feature map, and input the feature map after the residual connection into the SE attention module; Input the feature map processed by the SE attention module into the depthwise separable convolution for channel adjustment, and adjust the channels to the preset number of channels for output.
[0024] The MKCS-SE feature extraction layer is a multi-scale feature extraction module, as Figure 4 shown, including convolutional kernels of different sizes and a channel splitting layer. In the example, the defect image of size after preprocessing is input into the MKCS-SE feature extraction layer for feature extraction. The process is as follows: First, input the defect image into the depthwise separable convolution (DWConv) with a convolutional kernel size of , an input stride of 1, and a padding method of same for preliminary extraction. Then, perform a depthwise separable convolution with a convolutional kernel size of , an input stride of 1, and a padding method of same on the obtained feature map to adjust the channels to a multiple of 4; then perform a four-way channel split (Split) on the feature map with adjusted channels; then input the feature maps of the four-way channels into the depthwise separable convolution with a convolutional kernel size of , , , , an input stride of 1, and a padding method of same for feature extraction at different scales. Perform a merging operation on the four feature maps extracted at different scales and perform a residual connection with the feature map obtained by the depthwise separable convolution with a convolutional kernel size of , an input stride of 1, and a padding method of same; finally, input the result after the residual connection into the SE attention module (SE-Attention). The structure of the SE attention module is as Figure 5 shown. Process the input in two ways. The first way keeps the original structure unchanged; the second way first inputs it into the global average pooling layer (GAP) for pooling operation, then inputs the pooling result into the fully connected layer (FullyConnected) for full connection. Secondly, after activation by the ReLU function, perform another full connection layer for full connection, and then input the full connection result into the Sigmod function for operation to obtain a feature map of the same size as the original input, and its content is the specific weight of each channel. Perform a channel weighting (Scale) operation on the obtained feature map and the first-way input, and finally obtain the feature map processed by the SE attention layer. Finally, input the feature map processed by the SE attention module into the depthwise separable convolution with a convolutional kernel size of , an input stride of 1, and a padding method of same for channel adjustment, and adjust the channels to the preset number of channels for output.
[0025] Among them, the second path of the attention mechanism layer SE mainly includes two operations: Squeeze and Excitation. The following are the detailed definitions and formula descriptions: During the Squeeze process, for the input feature map with the shape of global average pooling is performed to obtain the global features of each channel. The specific formula for this process is: , where Z C represents the global feature of the C-th channel, and X C (i, j) represents the pixel value of the C-th channel in the input feature map at the position (i, j).
[0026] During the Excitation operation, a series of non-linear transformations are performed on the channel descriptor z to generate the attention weights for each channel. The specific steps are as follows: First, it enters a fully connected layer. Through a fully connected layer, the dimension of the channel descriptor is reduced to a fixed ratio of the original, usually , where r is a compression rate parameter, to obtain the attention weights for each channel, expressed as: , where is the attention weight, W1 and b1 are the weights and biases of the fully connected layer, represents the LeakyReLU activation function, and the formula is: , where X is the input data and leaky is a constant.
[0027] Then it enters another fully connected layer to restore the dimension to C, obtaining the attention weights for each channel, expressed as: , where W2 and b2 are the weights and biases of the fully connected layer.
[0028] The finally generated attention weights are used to weight each channel of the original feature map , that is, the final Scale operation, expressed as: , where is the weighted output feature map.
[0029] Among them, the Sigmoid function is one of the commonly used activation functions in machine learning and deep learning, especially in binary classification problems. It maps the input to the interval (0, 1), so it is often used in the output layer to predict probability values and is defined as follows: , where x is the input value, which can be any real number, and e is the base of the natural logarithm.
[0030] Step 3: Input the first to fourth feature maps into the DLKConv-SC module respectively to obtain four feature maps with different sizes, denoted as the sixth feature map, the seventh feature map, the eighth feature map, and the ninth feature map respectively.
[0031] Input the first feature map into the DLKConv-SC module to obtain the sixth feature map; input the second feature map into the DLKConv-SC module to obtain the seventh feature map; input the third feature map into the DLKConv-SC module to obtain the eighth feature map; input the fourth feature map into the DLKConv-SC module to obtain the ninth feature map.
[0032] Furthermore, the steps of inputting the first to fourth feature maps into the DLKConv-SC module respectively include: Take one of the first feature map, the second feature map, the third feature map, or the fourth feature map as the input feature map, input the input feature map into the first depthwise separable convolution and the second depthwise separable convolution respectively for large kernel feature extraction to obtain the first large kernel feature map and the second large kernel feature map respectively. After fusing the first large kernel feature map and the second large kernel feature map, perform average pooling and max pooling respectively to obtain the third large kernel feature map and the fourth large kernel feature map; Input the third large kernel feature map and the fourth large kernel feature into the third depthwise separable convolution for channel fusion, and after activating the result of channel fusion through the Sigmod function, obtain the fifth large kernel feature; Perform a Scale operation on the fifth large kernel feature map and the first large kernel feature map to obtain the seventh large kernel feature map; Perform a Scale operation on the fifth large kernel feature map and the second large kernel feature map to obtain the eighth large kernel feature map; Perform a merge operation on the seventh large kernel feature map, the eighth large kernel feature map, and the input feature map to obtain the feature map finally output by the DLKConv-SC module.
[0033] In the example, the structural schematic diagram of the DLKConv-SC module is as Figure 6 shown. First, input the input into the depthwise separable convolution with a convolution kernel size of , an input stride of 1, and a padding method of same, and a convolution kernel size of The depthwise separable convolution with an input stride of 1 and a padding method of same is used for large kernel feature extraction. Then, the outputs of the two are fused and then average pooling (AVG) and max pooling (MAP) are performed respectively. Finally, the two results obtained after average pooling and max pooling are input into a convolution kernel with a size of The depthwise separable convolution with an input stride of 1 and a padding method of same is used for channel fusion, and the result of channel fusion is activated by the Sigmod function and then combined with the result after passing through a convolution kernel with a size of The depthwise separable convolution with an input stride of 1 and a padding method of same and a convolution kernel with a size of The depthwise separable convolution with an input stride of 1 and a padding method of same is used to perform a Scale operation on the result to obtain two feature maps. Finally, the two feature maps obtained after the Scale operation are combined with the input to obtain the final feature map.
[0034] Step 4: Use the fifth to ninth feature maps for fusion decoding to obtain the defect segmentation result; Furthermore, the steps of using the fifth to ninth feature maps for fusion decoding to obtain the defect segmentation result include: After fusing the ninth feature map and the fifth feature map, use the Transpose-Conv module for decoding to obtain the tenth feature map; After fusing the eighth feature map and the tenth feature map, use the Transpose-Conv module for decoding to obtain the eleventh feature map; After fusing the seventh feature map and the eleventh feature map, use the Transpose-Conv module for decoding to obtain the twelfth feature map; Fuse the sixth feature map and the twelfth feature map, and use the Transpose-Conv module for decoding to obtain the thirteenth feature map; Resize the thirteenth feature map through a depthwise separable convolution, use the Sigmod function, and obtain the output segmentation result. The convolution kernel size of the depthwise separable convolution is with an input stride of 1 and a padding method of same. Through this depthwise separable convolution, the feature map is resized to The activation function selects the Sigmod function, and finally the output segmentation result is obtained.
[0035] In the above step 4, four decoding operations are performed in sequence, and the number of channels for the four times are 32, 16, 16, and 8 respectively. Finally, the output is a feature map with a size of .
[0036] The Transpose-Conv module is the transposed convolution (also known as deconvolution), which is the core operation module for upsampling feature maps in deep convolutional neural networks. This operation realizes the expansion of the spatial dimension in a learnable parameterized way and shows significant advantages in tasks such as image segmentation, generative adversarial networks, and super-resolution reconstruction. The transposed convolution is not the inverse process of the traditional convolution operation, but the spatial transformation form of the gradient calculation in the backward propagation process of the standard convolution. Specifically, for the input feature map and the convolution kernel, the transposed convolution maps the low-dimensional input to the high-dimensional space through the transpose operation of the kernel matrix. Its output size follows a specific rule: when the stride is s, s - 1 zero values are inserted between input units to achieve implicit upsampling; the output boundary effect can be precisely controlled by adjusting the padding parameter. It should be noted that the introduction of the output padding parameter effectively solves the problem of output size ambiguity that may occur when the stride is greater than 1.
[0037] At the actual implementation level, a three-stage calculation process is adopted: first, the input features are expanded in spacing (zero value insertion), then boundary padding is applied, and finally, the conventional convolution operation is performed. This design not only maintains the computational efficiency but also seamlessly connects with modules such as batch normalization and activation functions in the network. Compared with the traditional interpolation upsampling method, the core advantage of the transposed convolution lies in the learnability of its parameters. The kernel weights are automatically optimized through gradient descent to adapt to the spatial feature distribution of specific tasks.
[0038] Step 5: Input the fifth feature map into the global average pooling layer and the activation function layer in sequence to obtain the defect classification result.
[0039] Input the feature map that has passed through the feature extraction part into the global average pooling layer for processing to obtain a feature map; then input the obtained feature map into a fully connected layer or a linear layer (Dense Layer) with the activation function ReLU to obtain a feature map; finally, input the obtained result into a Dense layer with the activation function SoftMax for classification output to obtain the final classification result.
[0040] Among them, the SoftMax layer is an activation function layer commonly used in classification problems. It converts the output of the fully connected layer into a probability distribution. Its output value range is between 0 and 1, and the sum of all outputs is 1, representing the probabilities of various classes. Its formula can be expressed as: , where, is the output probability of the jth class, z jis the j-th output of the fully connected layer, and m is the number of output categories. The SoftMax function ensures that the output values form a probability distribution, enabling the interpretation of the model's output as the probability of each category, which is convenient for classification decisions.
[0041] Using the detection method described in the present invention to construct a detection model, before testing the detection effect of the wafer detection method described in the present invention, it is necessary to train the detection model using wafer detection defect images and their corresponding labels. To further illustrate the superiority of the method described in the present invention for wafer defect detection, a comparison is made with traditional models. The evaluation metrics include accuracy, precision, recall, F1-score, intersection over union (IOU), and Dice coefficient, and their expressions are as follows: , , , , , , Among them, TP is the number of positive samples predicted as positive samples actually; FP is the number of positive samples predicted as positive samples actually being negative samples; TN is the number of negative samples predicted as negative samples actually; FN is the number of negative samples predicted as negative samples actually being positive samples; |A| is the number of pixels in the predicted region; |B| is the number of pixels in the true label; |A∩B| is the number of intersection pixels between the prediction and the true label; |A∪B| is the number of union pixels between the prediction and the true label; the F1-score is a weighted average of accuracy and recall, and the larger its value means the better the model; IoU is an index used to measure the overlapping degree of two regions, widely used in tasks such as object detection and image segmentation. It evaluates the prediction accuracy of the model by calculating the ratio of the intersection to the union of the predicted region and the true label region; the Dice Coefficient is an index used to measure the similarity of two sets, widely used in tasks such as image segmentation, especially suitable for evaluating the overlapping degree of the segmentation result and the true label, and is one of the indexes for evaluating the prediction accuracy of the model.
[0042] The detection results are shown in Tables 1 to 6. The wafer defect detection model constructed by the method in the present invention is denoted as the MK-DANet model. Table 1 shows the precision, recall, and F1-score of the MK-DANet model on the WM-811K dataset. This table shows that for 9 wafer patterns, the model proposed in the present invention achieves very high precision, recall, and F1-score, with an average of 97% for all of them.
[0043] Table 2 shows the accuracy comparison of different models for each defect category in the WM-811K dataset. The comparison models used in this example are: Wafer Map Defect Pattern Recognition Model (WMDPI), Wafer Defect Detection Network Model Based on Separable Convolution and Attention Mechanism (WDD-SCA), Wafer Segmentation Classification Network Model (WSCN), and Wafer Map Classification Network Model Based on PeleeNet (WM-PeleeNet). It can be seen from Table 2 that the average classification accuracy of the model in the present invention on the WM-811K dataset is 97%, which is better than all other comparison models, indicating the advantage of the MK-DANet model in single-wafer defect classification.
[0044] Table 1
[0045] Table 2
[0046] Table 3
[0047] Table 4
[0048] Table 3 shows the precision, recall, and F1-score of the MK-DANet model on the MixedWM38 dataset. This table shows that for 38 wafer patterns, the model proposed in the present invention achieves very high precision, recall, and F1-score, with averages of 98%, 98%, and 97% respectively.
[0049] Table 5
[0050] Table 6
[0051] Table 4 shows the comparison of the model proposed in the present invention with the wafer defect pattern classification methods in recent years, including Residual Network (ResNet), Deep Cross Network (DCNet), WSCN, WM-PeleeNet. The comparison results of the average accuracy are shown in Table 5, indicating the good generalization of the method proposed in the present invention. Moreover, the detection accuracy of single defect, double defect and triple defect is higher than that of other methods. The detection accuracy of single defect is 99.1%, that of double defect is 98.3%, that of triple defect is 97.9%, and the detection accuracy of quadruple defect also reaches 96.5%. The average classification accuracy of the model proposed in the present invention on the MixedWM38 dataset is 98.2%, which is better than all other comparison models, indicating the advantage of the MK-DANet model in single wafer defect classification.
[0052] Table 6 is the comparison of the number of parameters of each model and the accuracy on the WM-811K dataset and the MixedWM38 dataset. The comparison models used include LeNet, AlexNet, VGG16, MobileNetV2, Residual Network (ResNet18), WM-PeleeNet, PeleeNet, Dense Connection Network (DenseNet121), Lightweight Compression Network (Squeezenet), WSCN. Table 6 shows that while improving the generalization ability and detection ability of the model proposed in the present invention, the number of parameters of the model is reduced as much as possible.
[0053] Table 7 is the comparison of the intersection over union and Dice coefficient of each model on the WM-811K dataset and the MixedWM38 dataset. The comparison models used include U-Net, DeepLabV3+ (based on the Residual Network-50 backbone) and WSCN. It can be seen from Table 7 that the MK-DANet model of the present invention has good defect segmentation ability compared with other models.
[0054] Table 7
Claims
1. A wafer defect detection method based on multi-core attention fusion and dynamic large core jumping, characterized in that: The following steps are involved: Acquire the wafer image to be inspected and perform image preprocessing; The preprocessed image is sequentially input into three feature extraction layers with the same structure for feature extraction, and a first feature map, a second feature map and a third feature map are obtained respectively; then the extracted third feature map is sequentially input into a fourth feature extraction layer and a fifth feature extraction layer for feature extraction, and a fourth feature map and a fifth feature map are obtained; Input the first to fourth feature maps into DLKConv-SC Module, four feature maps of different sizes are obtained, which are respectively recorded as the sixth feature map, the seventh feature map, the eighth feature map and the ninth feature map; The fifth to ninth feature maps are used for fusion decoding to obtain a defect segmentation result; The fifth feature map is input into the global average pooling layer and the activation function layer in sequence to obtain the defect classification result.
2. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to claim 1 is characterized in that: The preprocessed image is sequentially input into three feature extraction layers with the same structure for feature extraction, and a first feature map, a second feature map and a third feature map are obtained respectively. Then, the extracted third feature map is sequentially input into a fourth feature extraction layer and a fifth feature extraction layer for feature extraction, and the steps of obtaining the fourth feature map and the fifth feature map include: The preprocessed image is input into the first MKCS-SE The module performs feature extraction, and then performs maximum pooling on the extracted feature map to obtain the first feature map; Input the first feature map into the second MKCS-SE The module performs feature extraction, and then performs maximum pooling on the extracted feature map to obtain a second feature map; Input the second feature map into the third MKCS-SE The module performs feature extraction, and then performs maximum pooling on the extracted feature map to obtain the third feature map; Input the third feature map into the fourth MKCS-SE The module performs feature extraction, and then performs average pooling processing on the extracted feature map to obtain a fourth feature map; Input the fourth feature map into the fifth MKCS-SE The module performs feature extraction, and then performs average pooling processing on the extracted feature map to obtain the fifth feature map.
3. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to claim 1 is characterized in that: Input the first to fourth feature maps into DLKConv-SC The steps of the module include: Taking one of the first feature map, the second feature map, the third feature map or the fourth feature map as the input feature map, inputting the input feature map into the first depthwise separable convolution and the second depthwise separable convolution respectively to extract large core features, obtaining the first large core feature map and the second large core feature map respectively, fusing the first large core feature map and the second large core feature map, and then performing average pooling and maximum pooling respectively, to obtain the third large core feature map and the fourth large core feature map respectively; The third largest kernel feature map and the fourth largest kernel feature are input into the third depth separable convolution for channel fusion, and the result of channel fusion is passed through Sigmod After the function is activated, the fifth core feature is obtained; The fifth largest kernel feature map and the first largest kernel feature map are combined Scale Operation, get the seventh kernel feature map; The fifth largest kernel feature map and the second largest kernel feature map are combined Scale Operation, the eighth kernel feature map is obtained; The seventh and eighth kernel feature maps are combined with the input feature map to obtain DLKConv-SC The feature map finally output by the module.
4. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to claim 1 is characterized in that: The step of using the fifth to ninth feature maps for fusion decoding to obtain a defect segmentation result includes: After fusing the ninth feature map with the fifth feature map, use Transpose-Conv The module is decoded to obtain the tenth feature map; After fusing the eighth feature map and the tenth feature map, use Transpose-Conv The module is decoded to obtain the eleventh feature map; After fusing the seventh feature map and the eleventh feature map, use Transpose-Conv The module is decoded to obtain the twelfth feature map; The sixth feature map and the twelfth feature map are fused and used Transpose-Conv The module is decoded to obtain the thirteenth feature map; The thirteenth feature map is resized by depth-wise separable convolution and used Sigmod Function to get the output segmentation result.
5. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to claim 2 is characterized in that: First MKCS-SE Module to fifth MKCS-SE The structure of the modules is the same, and the process of feature extraction includes: The input feature map is subjected to a depthwise separable convolution for preliminary feature extraction, and then the obtained preliminary feature map is subjected to channel adjustment through a depthwise separable convolution, and the channel is adjusted to a multiple of four; Then divide the adjusted feature map into four equal channels; The feature maps obtained by dividing the channels into four equal parts are used to extract features of different scales using four depth-wise separable convolutions of different sizes; The feature maps extracted from four different scales are merged and residually connected with the preliminary feature map. The feature map after residual connection is input into SE Attention module; will pass SE The feature map processed by the attention module is input into the depthwise separable convolution for channel adjustment, and the channels are adjusted to the preset number of channels for output.
6. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to any one of claims 1 to 5, characterized in that: The steps of image preprocessing include: The defect image is resized to a preset size, median filtered, and data enhancement is performed on the imbalanced category.
7. The wafer defect detection method based on multi-core attention fusion and dynamic large core jumping according to claim 6 is characterized in that: Data augmentation includes flipping images vertically or horizontally, shifting images vertically or horizontally, and rotating or scaling images.
Citation Information
Patent Citations
Multi-target detection method for wafer graph defects
CN116385391A
Improved DeSTSeg-based unsupervised grain defect anomaly detection method
CN118469912A
Power line detection method based on multi-scale characteristics
CN119810694A