Wafer edge defect detection method based on deep learning

By improving the YOLOv11n model and combining the receptive field attention head module and the cross-stage cross-fusion network CIFNet, the problem of detecting small edge defects in wafer edge detection is solved. It achieves sub-pixel level accuracy and cross-fusion network interaction of multi-scale features, which solves the problems of limited edge defect detection capability and single feature fusion in existing wafer edge detection technologies, and improves detection accuracy and consistency.

CN120976133APending Publication Date: 2025-11-18JIAXING BAISHENG PHOTOELECTRIC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511066757.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies have limited ability to detect minute edge defects in wafer edge defect detection, and single feature fusion methods are prone to feature confusion, making it difficult to meet the requirements of sub-pixel level accuracy and decoupling of multiple defect features.

Method used

An improved YOLOv11n model is adopted, which decouples localization and classification through the setting of the receptive field attention head module RFAHead and the cross-stage cross-fusion network CIFNet, thereby enhancing the detection capability of small edge defects. Furthermore, the false positive rate of bubbles/burrs is reduced through four-step cross-level feature interaction and multi-level feature reuse.

Benefits of technology

It significantly improves the accuracy and consistency of wafer edge defect detection, enhances the ability to capture minute defects, reduces computational complexity and resource consumption, adapts to high-throughput detection requirements, and meets millisecond-level real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976133A_ABST
    Figure CN120976133A_ABST
Patent Text Reader

Abstract

The invention discloses a wafer edge defect detection method based on deep learning. The method comprises the following steps: selecting wafer edge defect images to generate a defect data set; the wafer defect data set is input into the improved YOLOv11n model to be trained, and the trained improved YOLOv11n model is obtained; obtaining a wafer image in the wafer production process, and inputting the trained improved YOLOv11n model to obtain the category and positioning information of the wafer edge defect image; the construction process of the improved YOLOv11n model comprises the following steps: replacing a mixed local channel attention module C2PSA of a backbone network with a three-convolution separable large kernel attention module C3LSKA; a neck network is replaced by a cross-stage cross convergence network CIFNet; and the detection head is replaced by a field attention head module RFAHead. The method improves the detection precision, enhances the multi-scale feature fusion capability, optimizes the calculation efficiency and resource consumption, enhances the real-time processing capability, improves the robustness and generalization capability of the model, and simplifies the model deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wafer inspection technology, and particularly relates to a wafer edge defect detection method based on deep learning. Background Technology

[0002] With advancements in technology, semiconductor process nodes are moving towards 2nm, the critical size of wafer edge defects has decreased to below 8μm, the 12-inch wafer capacity share has exceeded 82%, and the dynamic imaging distortion rate under high-speed rotation (≥120rpm) exceeds 30%, forcing detection technology to evolve towards multi-physics coupling modeling. Meanwhile, the heterojunction defect rate of third-generation semiconductor materials (SiC / GaN) is as high as 12.7%, driving detection algorithms to possess cross-material feature generalization capabilities.

[0003] Traditional machine vision systems rely on two-dimensional planar imaging (such as CCD line scan cameras), which lacks sufficient accuracy in detecting three-dimensional morphological defects (such as burrs and chamfering). Experimental data shows that when the defect height is <15μm, the false negative rate of two-dimensional imaging is as high as 12%. In rotating scanning systems, the high-speed rotation of the wafer (120rpm for a 12-inch wafer to meet a 30s cycle time) causes motion blur in the image. When using a global shutter camera, existing technologies show motion compensation errors exceeding ±3μm at a resolution of 10.45μm. Traditional algorithms employ a single detection model (such as edge gradient method for crack detection and morphological processing for edge chipping), which may lead to high false positive rates for dense bubbles and dirt, as well as feature coupling between different defect types.

[0004] Chinese patent document CN117422700A discloses a wafer surface defect detection method based on an improved YOLOv8, comprising the following steps: preprocessing and labeling wafer surface defect image data to generate a wafer surface defect image dataset; dividing the wafer surface defect image dataset into a training set, a validation set, and a test set according to a set ratio; introducing the ContextAggregation module and the GhostConv module to optimize the YOLOv8 target detection algorithm and constructing an improved YOLOv8 wafer surface defect detection model; training the above wafer surface defect detection model based on the training set and the validation set to obtain the optimal wafer surface defect detection model; inputting the test set into the optimal wafer surface defect detection model and outputting the wafer surface defect detection results.

[0005] The aforementioned patented solutions have two shortcomings: 1. Limited ability to detect minute edge defects: The GhostConv module used in these solutions reduces the number of parameters through feature stitching, but sacrifices the ability to capture local details. Wafer edge defects require sub-pixel accuracy, while GhostConv's linear operations are insufficient for extracting minute deformation features. 2. Unoptimized coupling of multiple defect features: The aforementioned patented solutions only add ContextAggregation to the Neck section, but do not design a cross-scale feature decoupling mechanism. Wafer edges often contain composite defects such as dense bubbles and burrs simultaneously, and a single feature fusion method can easily lead to feature confusion. Summary of the Invention

[0006] To overcome the limitations of existing wafer defect detection methods, such as limited detection capability for minute edge defects and the tendency for feature confusion due to single feature fusion methods, this invention aims to provide a deep learning-based wafer edge defect detection method. This method enhances the detection capability of minute edge defects by incorporating a receptive field attention head module (RFAHead) into the YOLOv11n model and decoupling localization and classification through a dual-task mechanism. Furthermore, a cross-stage cross-fusion network (CIFNet) is implemented, employing four-step cross-level feature interaction, multi-level feature reuse, and CSP module channel compression to effectively identify decoupled dense defect features, thereby reducing the false positive rate for bubbles / burrs.

[0007] To achieve the above objectives, the present invention employs the following technical solution: a wafer edge defect detection method based on deep learning, comprising the following steps:

[0008] Step 1: Select high-quality wafer edge defect images, manually label wafer surface defects, and generate a wafer edge defect dataset;

[0009] Step 2: Input the wafer edge defect dataset into the improved YOLOv11n model for training, and input the validation set during the training process. Verify the training effect through the test set to obtain the trained improved YOLOv11n model.

[0010] Step 3: Acquire wafer images during wafer manufacturing, input them into the trained improved YOLOv11n model, and obtain the category, location information, and confidence score of the wafer edge defect images;

[0011] The improved YOLOv11n model is constructed as follows: the hybrid local channel attention module C2PSA of the backbone network of the YOLOv11n model is replaced with the three-convolutional separable large kernel attention module C3LSKA; the neck network of the YOLOv11n model is replaced with the cross-stage cross-fusion network CIFNet; and the detection head of the YOLOv11n model is replaced with the receptive field attention head module RFAHead.

[0012] Furthermore, the three-convolutional separable large kernel attention module C3LSKA consists of three 1*1 convolutions and a large kernel separable convolutional attention module LSKA. The input feature map passes through the first 1*1 convolution, which reduces the number of channels to 2, dividing the channels of the input feature map into two parts. One part enters the large kernel separable convolutional attention module LSKA to extract attention features with a large receptive field; the other part enters the second 1*1 convolution to retain the original local features. Then, the output attention features and the local features are concatenated along the channel dimension, and finally, the third 1*1 convolution is used to restore the initial number of channels of the feature map to obtain the output feature map.

[0013] The three-convolutional separable large kernel attention module C3LSKA can effectively capture global information while retaining local features, thereby enhancing the feature representation capability of the backbone network.

[0014] Furthermore, the Large Kernel Separable Convolutional Attention Module (LSKA) consists of two-dimensional convolutions, depthwise separable convolutions, and depthwise separable dilated convolutions; the processing procedure is as follows:

[0015] The first step involves inputting the original feature map into the Large Kernel Separable Convolutional Attention Module (LSKA), which then processes it through the depthwise separable convolution to extract local features along the horizontal and vertical axes, respectively.

[0016] The second step is to further expand the receptive field by using the depth-separable dilatation convolution to extract the local features in the horizontal and vertical directions obtained in the first step.

[0017] The third step involves fusing the feature maps with a wide range of horizontal and vertical information obtained in the second step through two-dimensional convolution to generate attention weights.

[0018] The fourth step involves performing element-wise multiplication and weighting of the attention weights generated in the third step with the original feature map input from the skip connections in the first step to obtain the output feature map.

[0019] Furthermore, the cross-stage cross-convergence network CIFNet consists of multiple convolutional layers, a cross-stage partial connection module (CSP), an upsampling module (UpSample), and a concatenation module (Concat). The main function of CIFNet is to bridge the backbone network with the receptive field attention head module (RFAHead), inputting the optimized features into the RFAHead through four steps:

[0020] In the first step, the high-resolution feature map S3, the medium-resolution feature map S4 output by the backbone network, and the feature map S5 output by the three-convolutional separable large kernel attention C3LSKA module are sequentially subjected to 1*1 convolution to achieve feature weighting and adjustment; the resolution of the output feature map S5 is doubled to align its resolution with that of the medium-resolution feature map S4, thereby achieving cross-level detail fusion and completing the upsampling operation;

[0021] The second step involves inputting the output feature map S5 and the medium-resolution feature map S4 after upsampling in the first step into the concat stitching module to generate multi-scale fusion features. The generated result is then input into the cross-stage partial connection module CSP, where channel stitching and convolution operations are used to generate a feature map D4 with stronger feature representation capabilities.

[0022] The third step is to input feature map D4 into a 1*1 convolution. By fusing effective parameters of multi-scale features, the weights of the convolution kernel are reconstructed to achieve reparameterization. Then, feature map D4 is upsampled and concatenated with feature map S3, which has also undergone 1*1 convolution reparameterization, to generate a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to obtain feature map P3.

[0023] The fourth step is to generate a low-resolution feature map by sampling the feature map P3 through 3*3 convolution, and then concatenate it with the upsampled feature map D4 to obtain a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to generate feature map P4.

[0024] The fifth step involves using a 3*3 convolution downsampled feature map P4 and a fused feature map S5 that has been reparameterized by a 1*1 convolution. This fused feature map is then fed into the cross-stage partial connection module CSP to obtain feature map P5.

[0025] The cross-stage cross-fusion network CIFNet achieves resolution alignment by using cross-level up / down sampling. Combined with the channel compression mechanism and dynamic parameter sharing strategy of the cross-stage partial connection module CSP, it reduces the number of model parameters and computational complexity while ensuring the accuracy of multi-scale feature interaction, thus further realizing the lightweighting of the model.

[0026] Specifically, the cross-stage partial connection module (CSP) consists of depthwise separable convolution, depthwise separable dilated convolution, and ordinary convolution; the input feature map is divided into two channels by two 1*1 convolutions to generate auxiliary branches and a main branch; wherein, the auxiliary branch retains the spatial distribution characteristics of the original features to obtain F. aux The main branch feature map F extracts multi-scale contextual information through multiple basic blocks; the basic block includes 3*3 convolution and lightweight residual convolution.

[0027] Specifically, the processing flow of the main branch feature map F is as follows: After the input feature map enters the RepConv submodule via the lightweight residual convolution module, features are extracted by parallel 1*1 convolution and 3*3 convolution respectively, and the outputs of the two are added together at the feature level to generate the fused feature map F. fuse F fuse The fusion-stimulation feature map σ is obtained by using the Swish activation function. Ffuse , σ Ffuse Channel information is integrated through 3*3 convolution to generate intermediate features of the main path. Finally, the intermediate features are added to the original input feature map of the skip connection by residual addition to obtain the module output feature map. After the main branch and the auxiliary branch are processed, feature fusion is achieved by concatenating along the channel dimension, and the output feature map is generated by lightweight 1*1 convolution.

[0028] Furthermore, the receptive field attention head module RFAHead consists of two task branches: a bounding box regression task and a class prediction task. The bounding box regression task uses two receptive field attention convolutions RFAConv to extract key region information from the feature map, transforms the number of channels to 4*α through two-dimensional convolution, and finally calculates the perfect intersection-over-union ratio and distribution focus loss. The class prediction task uses depthwise separable convolution and 1*1 convolution to process the feature map, transforms the number of channels to β through two-dimensional convolution, and finally calculates the binary cross-entropy loss. Here, α represents the discretization degree of bounding box regression; β represents the number of classes.

[0029] To meet the differentiated requirements of bounding box regression and category prediction, the receptive field attention head module RFAHead achieves efficient collaborative optimization of localization and classification tasks in object detection through a dual-task decoupling structure and a differentiated allocation strategy for α and β channels.

[0030] Specifically, the receptive field attention convolution RFAConv adopts a dual-task hybrid convolution framework, and the technical process is divided into two parallel sub-tasks: attention weight calculation task and receptive field feature extraction task. In the attention weight calculation task, the input feature map has a shape of h*w*c. First, global average pooling is used to generate global features. Then, 1*1 group convolution combined with Softmax activation function is used to normalize and realize cross-channel interaction, and finally the attention weight map h*w*9c is output. In the receptive field feature extraction task, the input feature map is activated by group convolution with a feature map stride of 3 and a convolution kernel of 3 combined with ReLU activation function to generate multi-scale local feature map h*w*9c.

[0031] Specifically, the number of channels in the feature map output by the receptive field attention convolution RFAConv is 9 times the number of channels in the initial convolution. The 9 overlapping pixels at a single position are converted into non-overlapping 3*3 pixel blocks, and the overall feature map shape becomes 3h*3w*c. Finally, the results of the attention weight calculation task and the receptive field feature extraction task are multiplied together, and the fused feature map is downsampled using a 3*3 convolution to obtain an output feature map with shape h*w*c.

[0032] Furthermore, the improved YOLOv11n model feature extraction network adopts the CSPDarknet53 architecture, which has 11 layers in total, with the following structure: Conv layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2f layer → SPPF layer → C3LSKA layer; the first Conv layer has 16 output channels, and the number of output channels of each subsequent Conv layer is twice that of the previous Conv layer; the Conv layer contains four types of operations, namely convolution operation, batch normalization operation, SiLU activation processing operation, and downsampling operation, wherein the downsampling operation is only executed when the stride in the Conv layer is 2.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0034] 1. Improved detection accuracy: By introducing the C3LSKA triple convolutional separable large kernel attention module and the RFAHead receptive field attention head, it is possible to capture minute defects (such as micron-level edge chipping and edge cracks) and complex texture features at the wafer edge more precisely, significantly improving the accuracy of defect localization and classification.

[0035] 2. Enhanced multi-scale feature fusion capability: The cross-stage cross-fusion network CIFNet improves the ability to fuse and perceive defects of different scales (such as edge cracks and deformation) through multi-level feature interaction and dynamic weight allocation, ensuring detection consistency in complex scenarios.

[0036] 3. Optimize computational efficiency and resource consumption: The three-convolutional separable large kernel attention module C3LSKA adopts a separable convolution and large kernel attention mechanism, which significantly reduces the number of parameters while maintaining a large receptive field; the cross-stage cross-fusion network CIFNet reduces redundant computation through cross-stage feature reuse, thereby reducing the overall computational complexity and adapting to high-throughput detection requirements.

[0037] 4. Enhanced real-time processing capabilities: Lightweight design (such as C3LSKA's separated convolutional structure) and efficient feature fusion strategy (CIFNet) are optimized in synergy, significantly improving inference speed and meeting the millisecond-level real-time detection requirements of wafer production lines.

[0038] 5. Improved robustness and generalization ability of the model: The improved YOLOv11n model can better capture the complex relationships between feature maps, thus improving the model's robustness and generalization ability to different defect types and environmental changes.

[0039] 6. Simplified model deployment: The improved YOLOv11n model maintains high accuracy while reducing computing resource consumption. It can be efficiently deployed in edge computing devices (such as industrial GPUs / FPGAs), reducing enterprise hardware upgrade costs and adapting to actual production line integration needs. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the improved YOLOv11n model of the present invention;

[0041] Figure 2 This is a schematic diagram of the C3LSKA three-convolution separable large kernel attention module of the present invention;

[0042] Figure 3 This is a schematic diagram of the Large Kernel Separated Convolutional Attention Module (LSKA) of the present invention;

[0043] Figure 4 This is a schematic diagram of the cross-stage convergence network CIFNet of the present invention;

[0044] Figure 5 This is a schematic diagram of the cross-stage connection module CSP of the present invention;

[0045] Figure 6 This is a schematic diagram of the receptive field attention head (RFAHead) of the present invention;

[0046] Figure 7 This is a schematic diagram of the receptive field attention convolution RFAConv of the present invention;

[0047] Figure 8 This is a schematic diagram of the detection results of a detection system for water contamination in the prior art;

[0048] Figure 9This is a schematic diagram of the detection results of the improved YOLOv11n model for water contamination according to the present invention;

[0049] Figure 10 This is a schematic diagram of the detection results of scratches by a detection system in the prior art;

[0050] Figure 11 This is a schematic diagram of the scratch detection results of the improved YOLOv11n model of this invention. Detailed Implementation

[0051] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0052] In the description of this invention, it should be noted that directional terms such as "center," "lateral," "longitudinal," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise" indicate the orientation and positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. They should not be construed as limiting the specific protection scope of this invention.

[0053] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features. Thus, the use of "first" and "second" to define a feature may explicitly or implicitly include one or more of that feature, and in the description of this invention, "a number" means two or more, unless otherwise explicitly specified.

[0054] In this invention, unless otherwise explicitly specified and limited, terms such as "set" and "install" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can also refer to a mechanical connection; they can refer to a direct connection or a connection through an intermediate medium; or they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0055] See Figures 1-11 A deep learning-based method for detecting wafer edge defects includes the following steps:

[0056] Step 1: Select high-quality wafer edge defect images, manually label wafer surface defects, and generate a wafer edge defect dataset;

[0057] Step 2: Input the wafer edge defect dataset into the improved YOLOv11n model for training, and input the validation set during the training process. Verify the training effect through the test set to obtain the trained improved YOLOv11n model.

[0058] Step 3: Acquire wafer images during wafer manufacturing, input them into the trained improved YOLOv11n model, and obtain the category, location information, and confidence score of the wafer edge defect images;

[0059] The improved YOLOv11n model is constructed as follows: the hybrid local channel attention module C2PSA of the backbone network of the YOLOv11n model is replaced with the three-convolutional separable large kernel attention module C3LSKA; the neck network of the YOLOv11n model is replaced with the cross-stage cross-fusion network CIFNet; and the detection head of the YOLOv11n model is replaced with the receptive field attention head module RFAHead.

[0060] Data Acquisition Device Design: A high-resolution industrial camera is used. This camera can continuously capture images of a 2mm area around the edge of a 12-inch glass-based wafer at high speed and high resolution, with a capture cycle time of ≤30s, making it ideal for real-time inspection on production lines. The camera's high sensitivity and low noise characteristics ensure image quality. A ring-shaped LED light source provides uniform and stable illumination, effectively eliminating reflections and shadows, ensuring clarity and contrast of defects in the image. The ring-shaped LED light source is placed above the moving wafer material and coaxially mounted with the high-resolution industrial camera, ensuring that the light source's beam is aligned with the camera's line of sight for optimal illumination. As the wafer material passes through the inspection area of ​​the camera and light source, the camera continuously captures images at a high frame rate. Image data is transmitted to a computer for storage and processing via a high-speed GigE interface.

[0061] Data annotation: Data collected using data acquisition devices is categorized and stored according to defect type. Annotation tools (such as Labelimg) are used for manual annotation of images, marking the location and type of each defect and generating corresponding annotation files (TXT format) for deep learning model training. There are eight defect types: chamfer chipping, edge chipping, edge cracks, burrs, deformation, dull edges, dense bubbles, and dirt.

[0062] like Figure 1 As shown, the overall structure of the improved YOLOv11n model is based on the improved network structure of YOLOv11n (You Only Look Once version 11).

[0063] Feature Extraction Network: The feature extraction network of the deep learning model adopts the CSPDarknet53 architecture, which consists of 11 layers, structured as follows: Conv layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2f layer → SPPF layer → C3LSKA layer. The first Conv layer has 16 output channels, and each subsequent Conv layer has twice the number of output channels as the previous layer, for a total of five Conv layers, with the last Conv layer having 256 output channels. The Conv layer contains four operations: convolution, batch normalization, SiLU activation, and downsampling. The first three are regular operations, while the last operation is only executed when the stride in the Conv layer is 2.

[0064] like Figure 2 As shown, the three-convolutional separable large-kernel attention module C3LSKA includes three 1*1 convolutions and a large-kernel separable convolutional attention module LSKA. The input feature map is processed by the first 1*1 convolution, which reduces the number of channels to two, dividing the channels of the input feature map into two parts. One part enters the large-kernel separable convolutional attention module LSKA to extract attention features with a large receptive field; the other part enters the second 1*1 convolution to retain the original local features. Then, the output attention features and the local features are concatenated along the channel dimension. Finally, the third 1*1 convolution is used to restore the initial number of channels of the feature map, resulting in the output feature map. This module effectively captures global information while preserving local features, thereby enhancing the feature representation capability of the backbone network.

[0065] like Figure 3 As shown, the large kernel separable convolutional attention module LSKA consists of two-dimensional convolution (Conv), depthwise separable convolution (DWConv), and depthwise separable dilated convolution (DW-D-Conv). The processing procedure is as follows: First, the original feature map (input) is input into the large kernel separable convolutional attention module LSKA, and after processing by the depthwise separable convolution (DWConv), local features are extracted along the horizontal and vertical directions respectively.

[0066] The second step is to further expand the receptive field by using the depthwise separable dilated convolution (DW-D-Conv) to extract the local features in the horizontal and vertical directions obtained in the first step.

[0067] The third step involves fusing the feature maps with a wide range of horizontal and vertical information obtained in the second step through two-dimensional convolution (Conv) to generate attention weights.

[0068] The fourth step involves performing element-wise multiplication and weighting of the attention weights generated in the third step with the original feature map input from the skip connections in the first step to obtain the output feature map.

[0069] like Figure 4 As shown, the cross-stage cross-convergence network CIFNet consists of multiple convolutional layers, a cross-stage partial connection module (CSP), an upsampling module (UpSample), and a concatenation module (Concat). The main function of CIFNet is to bridge the backbone network with the receptive field attention head module (RFAHead), inputting the optimized features into the RFAHead through four steps:

[0070] In the first step, the high-resolution feature map S3, the medium-resolution feature map S4 output by the backbone network, and the feature map S5 output by the three-convolutional separable large kernel attention C3LSKA module are sequentially subjected to 1*1 convolution to achieve feature weighting and adjustment; the resolution of the output feature map S5 is doubled to align its resolution with that of the medium-resolution feature map S4, thereby achieving cross-level detail fusion and completing the upsampling operation;

[0071] The second step involves inputting the output feature map S5 and the medium-resolution feature map S4 after upsampling in the first step into the concat stitching module to generate multi-scale fusion features. The generated result is then input into the cross-stage partial connection module CSP, where channel stitching and convolution operations are used to generate a feature map D4 with stronger feature representation capabilities.

[0072] The third step is to input feature map D4 into a 1*1 convolution. By fusing effective parameters of multi-scale features, the weights of the convolution kernel are reconstructed to achieve reparameterization. Then, feature map D4 is upsampled and concatenated with feature map S3, which has also undergone 1*1 convolution reparameterization, to generate a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to obtain feature map P3.

[0073] The fourth step is to generate a low-resolution feature map by sampling the feature map P3 through 3*3 convolution, and then concatenate it with the upsampled feature map D4 to obtain a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to generate feature map P4.

[0074] The fifth step involves using a 3*3 convolution downsampled feature map P4 and a fused feature map S5 that has been reparameterized by a 1*1 convolution. This fused feature map is then fed into the cross-stage partial connection module CSP to obtain feature map P5.

[0075] The cross-stage cross-fusion network CIFNet achieves resolution alignment by using cross-level up / down sampling. Combined with the channel compression mechanism and dynamic parameter sharing strategy of the cross-stage partial connection module CSP, it reduces the number of model parameters and computational complexity while ensuring the accuracy of multi-scale feature interaction, thus further realizing the lightweighting of the model.

[0076] like Figure 5 As shown, the cross-stage partial connection module (CSP) consists of depthwise separable convolution, depthwise separable dilated convolution, and ordinary convolution; the input feature map is divided into two channels by two 1*1 convolutions to generate auxiliary branches and a main branch; wherein, the auxiliary branch retains the spatial distribution characteristics of the original features to obtain F. aux The main branch feature map F extracts multi-scale contextual information through n basic blocks. In this paper, n is set to 2, and each basic block consists of a 3*3 convolution and a lightweight residual convolution (RepConv).

[0077] The processing flow of the main branch feature map F is as follows: The feature map F input to the lightweight residual convolution (RepConv) module is obtained by adding the lightweight versions of 1*1 convolution and 3*3 convolution. fuse σ is then obtained through the Swish activation function. Ffuse The input is processed by a 3x3 convolution to organize the channel information, and finally concatenated with the residual of the main branch feature map F to obtain F. main After the main branch and auxiliary branches have finished processing, F will be... main With F aux Feature fusion is achieved by concatenating along the channel dimension, and F is generated by lightweight 1*1 convolution. out Output feature map. The cross-stage partial connection module (CSP) combines channel concatenation and convolution to achieve cross-scale feature fusion and computational efficiency optimization, thereby improving the model's feature capture capability.

[0078] F fuse =Conv 1×1 (F)+Conv 3×3 (F);

[0079] σ(x) = x·sigmoid(x);

[0080] F main =Conv 3×3 (σ(F fuse ))+F;

[0081] F out =Conv 1×1 (Concat(F main ,F aux )).

[0082] like Figure 6 As shown, the receptive field attention head module RFAHead consists of two task branches: a bounding box regression task and a class prediction task. The bounding box regression task uses two receptive field attention convolutions RFAConv to extract key region information from the feature map (input), transforms the number of channels to 4*α through two-dimensional convolution, and finally calculates the complete intersection over union (CIOU) and distribution focal loss (DFL). The class prediction task uses depthwise separable convolutions and 1*1 convolutions to process the feature map (input), transforms the number of channels to β through two-dimensional convolution, and finally calculates the binary cross-entropy loss (BCE). Here, α represents the discretization degree of bounding box regression (reg_max); β represents the number of classes (nc).

[0083] To meet the differentiated requirements of bounding box regression and category prediction, the receptive field attention head module RFAHead achieves efficient collaborative optimization of localization and classification tasks in object detection through a dual-task decoupling structure and a differentiated allocation strategy for α and β channels.

[0084] like Figure 7 As shown, the receptive field attention convolution RFAConv adopts a dual-task hybrid convolution framework. The technical process is divided into two parallel sub-tasks: attention weight calculation and receptive field feature extraction. In the attention weight calculation task, the input feature map (input) has a shape of h*w*c. First, global features are generated through global average pooling (AvgPool). Then, 1*1 group convolution (GroupConv) combined with the Softmax activation function is used to normalize and achieve cross-channel interaction, finally outputting an attention weight map h*w*9c. In the receptive field feature extraction task, the input feature map (input) is activated by group convolution with a feature map stride of 3 and a kernel of 3, combined with the ReLU activation function to generate a multi-scale local feature map h*w*9c.

[0085] Because grouped convolution is used, the number of channels in the feature map output from the dual tasks is nine times the initial number of channels. Spatial resampling is required to convert the nine overlapping pixels at a single location into non-overlapping 3x3 pixel blocks, resulting in an overall feature map shape of 3h*3w*c. Finally, the dual task results are multiplied, and a 3x3 convolution is used to downsample the fused feature map, yielding an output feature map of shape h*w*c. The receptive field attention convolution RFAConv utilizes spatial resampling to integrate local context and optimize channel redundancy, improving the model's adaptive focusing ability on key regions.

[0086] Figure 8 This is a schematic diagram illustrating the detection results of a prior art detection system for water contamination. Figure 9 This is a schematic diagram illustrating the detection results of the improved YOLOv11n model for water contamination according to the present invention. For the same wafer defect, the confidence level in the prior art is 0.42, while the confidence level of the improved YOLOv11n model of the present invention is 0.46. Similarly, Figure 10 This is a schematic diagram of the detection results of scratches by a detection system in the prior art; Figure 11 This diagram illustrates the scratch detection results of the improved YOLOv11n model of this invention. For the same wafer defect, the confidence level in the prior art is 0.61, while the confidence level of the improved YOLOv11n model of this invention is 0.65. Clearly, this invention offers higher detection accuracy and is more effective at detecting minute defects.

[0087] The confidence score represents the probability that the detection model is confident in predicting the presence of a target object and its category within the detection box. For example, 0.42 means that the model has a 42% confidence that the object in the box belongs to the predicted category. It is mainly used to filter out low-quality detections.

[0088] Loss Function: The loss function mainly consists of two parts: category classification loss and bounding box regression loss. Category classification calculates the difference loss between the category of the target contained in each predicted box and the true category, using a binary cross-entropy loss for N targets. Bounding box loss is used to calculate the difference loss between the predicted box and the true box, mainly composed of IOU loss and DFL loss.

[0089] Loss=λ coord ·Localization_Loss+λ cls • Classification_Loss;

[0090] λ coord and λ cls These are weighting coefficients used to balance the effects of different losses.

[0091] Post-processing: During inference, deep learning models generate a large number of predicted bounding boxes, each containing class probability, confidence score, and bounding box coordinates. To extract meaningful detection results, a series of post-processing steps are required, including threshold filtering, non-maximum suppression (NMS), bounding box adjustment, and multi-class classification.

[0092] 1. Threshold Filtering: Threshold filtering is used to filter the bounding boxes predicted by the model. By setting a confidence threshold, only predictions with a confidence score higher than that threshold are retained, thereby reducing irrelevant or false detections. The purpose of this step is to improve the accuracy and reliability of the detection results, ensuring that only those bounding boxes that are more likely to represent the true target are retained.

[0093] 2. Non-Maximum Suppression (NMS): After threshold filtering, the remaining predicted boxes may still contain a large number of redundant boxes, especially when the target object is large, multiple predicted boxes may overlap. Non-maximum suppression (NMS) is a post-processing technique used to remove redundant overlapping bounding boxes. NMS compares the confidence scores of overlapping bounding boxes, retaining the box with the highest confidence score and suppressing other boxes whose overlap exceeds a certain threshold. This reduces duplicate detections and improves the accuracy and clarity of the detection results.

[0094] 3. Bounding Box Refinement: After Non-Maximum Suppression (NMS), the remaining predicted bounding boxes need further refinement to improve localization accuracy. Bounding box refinement is a step that further refines the bounding boxes predicted by the model. By correcting and adjusting the initially predicted bounding boxes, the accuracy of target localization can be improved. This includes converting relative coordinates to absolute coordinates in the image, enabling the bounding boxes to accurately locate the target in the image.

[0095] 4. Multi-class processing: Multi-class processing refers to simultaneously processing prediction results from multiple categories in a detection task. The model needs to be able to distinguish and correctly classify targets from multiple categories. By setting category probability thresholds, the prediction result with the highest confidence in each category can be selected, ensuring that the final output category prediction is accurate and reliable. In defect sample validation, the false negative rate (number of false negatives / total number of defective products) ≤ 0.5%; in normal product validation, the pass rate (number of pass inspections / number of inputs, excluding contaminants) ≤ 2%.

[0096] Advantages of this invention compared to existing technologies

[0097] 1. Leading Capability for Detecting Minor Defects: The hierarchical attention mechanism of the C3LSKA three-convolutional separable large-kernel attention module, compared to context-weighted fusion methods, possesses a stronger ability to capture minor defects and has advantages in cross-scale feature integration. The LSKA large-kernel separable convolutional attention module achieves a breakthrough expansion of the receptive field while maintaining the model's lightweight nature through the synergy of depthwise convolution and dilated convolution, solving the pain point of easily missing minor edge defects.

[0098] 2. Multi-scale Feature Fusion: The innovative cross-level dynamic fusion architecture of the CIFNet cross-stage fusion network exhibits superior semantic information fidelity and sub-pixel alignment accuracy compared to single-path feature transfer. Through the synergistic optimization of reparameterization and channel compression, it significantly reduces computational redundancy while achieving superior multi-defect simultaneous recognition capabilities compared to the industry average.

[0099] 3. Optimized Inspection Tasks: The RFAHead dual-task decoupling design of the receptive field attention head module achieves precise positioning, classification, and collaborative optimization through spatial resampling and channel differentiation processing, outperforming unified inspection head structures. This design demonstrates superior ability to distinguish complex defects (such as burrs and cracks) at wafer edges, providing a viable solution to the long-standing industry problem of misjudging similar defects.

[0100] The above description is only a specific embodiment of the present invention, but the technical features of the present invention are not limited thereto. Any changes or modifications made by those skilled in the art within the scope of the present invention are covered by the patent scope of the present invention.

Claims

1. A wafer edge defect detection method based on deep learning, characterized in that: The steps include the following: Step 1: Select high-quality wafer edge defect images, manually label wafer surface defects, and generate a wafer edge defect dataset; Step 2: Input the wafer edge defect dataset into the improved YOLOv11n model for training, and input the validation set during the training process. Verify the training effect through the test set to obtain the trained improved YOLOv11n model. Step 3: Acquire wafer images during wafer manufacturing, input them into the trained improved YOLOv11n model, and obtain the category, location information, and confidence score of the wafer edge defect images; The improved YOLOv11n model is constructed as follows: the hybrid local channel attention module C2PSA of the backbone network of the YOLOv11n model is replaced with the three-convolutional separable large kernel attention module C3LSKA; the neck network of the YOLOv11n model is replaced with the cross-stage cross-fusion network CIFNet; and the detection head of the YOLOv11n model is replaced with the receptive field attention head module RFAHead.

2. The detection method as described in claim 1, characterized in that: The three-convolutional separable large kernel attention module C3LSKA consists of three 1*1 convolutions and a large kernel separable convolutional attention module LSKA. The input feature map is processed by the first 1*1 convolution, which reduces the number of channels to 2, dividing the channels of the input feature map into two parts. One part enters the large kernel separable convolutional attention module LSKA to extract attention features with a large receptive field. The other part enters the second 1*1 convolution to retain the original local features. Then, the output attention features and the local features are concatenated along the channel dimension. Finally, the third 1*1 convolution is used to restore the initial number of channels of the feature map to obtain the output feature map.

3. The detection method as described in claim 2, characterized in that: The Large Kernel Separable Convolutional Attention Module (LSKA) consists of two-dimensional convolutions, depthwise separable convolutions, and depthwise separable dilated convolutions; the processing procedure is as follows: The first step involves inputting the original feature map into the Large Kernel Separable Convolutional Attention Module (LSKA), which then processes it through the depthwise separable convolution to extract local features along the horizontal and vertical axes, respectively. The second step is to further expand the receptive field by using the depth-separable dilatation convolution to extract the local features in the horizontal and vertical directions obtained in the first step. The third step involves fusing the feature maps with a wide range of horizontal and vertical information obtained in the second step through two-dimensional convolution to generate attention weights. The fourth step involves performing element-wise multiplication and weighting of the attention weights generated in the third step with the original feature map input from the skip connections in the first step to obtain the output feature map.

4. The detection method according to any one of claims 1-3, characterized in that: The cross-stage cross-convergence network CIFNet consists of multiple convolutional layers, a cross-stage partial connection module (CSP), an upsampling module (UpSample), and a concatenation module (Concat). The main function of CIFNet is to bridge the backbone network with the receptive field attention head module (RFAHead), inputting optimized features into the RFAHead through four steps: In the first step, the high-resolution feature map S3, the medium-resolution feature map S4 output by the backbone network, and the feature map S5 output by the three-convolutional separable large kernel attention C3LSKA module are sequentially subjected to 1*1 convolution to achieve feature weighting and adjustment; the resolution of the output feature map S5 is doubled to align its resolution with that of the medium-resolution feature map S4, thereby achieving cross-level detail fusion and completing the upsampling operation; The second step involves inputting the output feature map S5 and the medium-resolution feature map S4 after upsampling in the first step into the concat stitching module to generate multi-scale fusion features. The generated result is then input into the cross-stage partial connection module CSP, where channel stitching and convolution operations are used to generate a feature map D4 with stronger feature representation capabilities. The third step is to input feature map D4 into a 1*1 convolution. By fusing effective parameters of multi-scale features, the weights of the convolution kernel are reconstructed to achieve reparameterization. Then, feature map D4 is upsampled and concatenated with feature map S3, which has also undergone 1*1 convolution reparameterization, to generate a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to obtain feature map P3. The fourth step is to generate a low-resolution feature map by sampling the feature map P3 through 3*3 convolution, and then concatenate it with the upsampled feature map D4 to obtain a fused feature map. The generated result is then input into the cross-stage partial connection module CSP to generate feature map P4. Fifth step: Use the 3*3 convolution downsampled feature map P4 and the feature map S5 that has been reparameterized by 1*1 convolution to generate a fused feature map, which is then sent to the cross-stage partial connection module CSP to obtain feature map P5.

5. The detection method as described in claim 4, characterized in that: The cross-stage partial connection module (CSP) consists of depthwise separable convolution, depthwise separable dilated convolution, and ordinary convolution; the input feature map is divided into channels by two 1*1 convolutions to generate auxiliary branches and a main branch; wherein, the auxiliary branch retains the spatial distribution characteristics of the original features to obtain F. aux The main branch feature map F extracts multi-scale contextual information through multiple basic blocks; the basic block includes 3*3 convolution and lightweight residual convolution.

6. The detection method as described in claim 5, characterized in that: The processing flow of the main branch feature map F is as follows: After the input feature map enters the RepConv submodule through the lightweight residual convolution module, features are extracted by parallel 1*1 convolution and 3*3 convolution respectively, and the outputs of the two are added together at the feature level to generate the fused feature map F. fuse F fuse The fusion-stimulation feature map σ is obtained by using the Swish activation function. Ffuse , σ Ffuse Channel information is integrated through 3*3 convolution to generate intermediate features of the main path. Finally, the intermediate features are added to the original input feature map of the skip connection by residual addition to obtain the module output feature map. After the main branch and the auxiliary branch are processed, feature fusion is achieved by concatenating along the channel dimension, and the output feature map is generated by lightweight 1*1 convolution.

7. The detection method according to any one of claims 1-3, characterized in that: The receptive field attention head module RFAHead consists of two task branches: a bounding box regression task and a class prediction task. The bounding box regression task uses two receptive field attention convolutions RFAConv to extract key region information from the feature map, transforms the number of channels to 4*α through two-dimensional convolution, and finally calculates the perfect intersection-over-union ratio and distribution focus loss. The class prediction task uses depthwise separable convolutions and 1*1 convolutions to process the feature map, transforms the number of channels to β through two-dimensional convolution, and finally calculates the binary cross-entropy loss. Here, α represents the discretization degree of bounding box regression; β represents the number of classes.

8. The detection method as described in claim 7, characterized in that: The receptive field attention convolution RFAConv adopts a dual-task hybrid convolution framework, and the technical process is divided into two parallel sub-tasks: attention weight calculation task and receptive field feature extraction task. In the attention weight calculation task, the input feature map has a shape of h*w*c. First, global features are generated through global average pooling. Then, 1*1 grouped convolutions combined with the Softmax activation function are used for normalization to achieve cross-channel interaction, and finally, the attention weight map h*w*9c is output. In the receptive field feature extraction task, the input feature map is activated by grouped convolutions with a stride of 3 and a kernel of 3, combined with the ReLU activation function to generate a multi-scale local feature map h*w*9c.

9. The detection method as described in claim 8, characterized in that: The number of channels in the feature map output by the receptive field attention convolution RFAConv is 9 times the number of channels in the initial convolution. The 9 overlapping pixels at a single position are converted into non-overlapping 3*3 pixel blocks, and the overall feature map shape becomes 3h*3w*c. Finally, the results of the attention weight calculation task and the receptive field feature extraction task are multiplied, and the fused feature map is downsampled using a 3*3 convolution to obtain an output feature map with shape h*w*c.

10. The detection method according to any one of claims 1-3, characterized in that: The improved YOLOv11n model feature extraction network adopts the CSPDarknet53 architecture, which has 11 layers in total, with the following structure: Conv layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2 layer → Conv layer → C3k2f layer → SPPF layer → C3LSKA layer. The first Conv layer has 16 output channels, and the number of output channels of each subsequent Conv layer is twice that of the previous Conv layer. The Conv layer contains four types of operations: convolution operation, batch normalization operation, SiLU activation processing operation, and downsampling operation. The downsampling operation is only executed when the stride in the Conv layer is 2.

Citation Information

Patent Citations

  • Wafer surface defect detection method based on YOLOv8 improvement

    CN117422700A