Wafer cleaning effect detection method based on improved YOLOv7
By introducing the Concat_attention module and SIOU loss function in the YOLOv7 algorithm, combined with data enhancement technology, the problem of insufficient detection accuracy and model generalization ability of tiny pollutants in wafer cleaning effect detection is solved, and efficient and accurate wafer cleaning effect detection is achieved.
Patent Information
- Application Number
- CN202510481059.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to efficiently detect tiny pollutants on the wafer surface, especially in complex backgrounds, which are prone to false detection and missed detection, and the model generalization ability is insufficient.
The Concat_attention feature fusion module is used to replace the Concat module of the YOLOv7 algorithm, and the SIOU loss function and label smoothing technology are introduced to optimize the matching accuracy of the prediction box and the real box, and the data set is expanded through data enhancement technology to enhance the adaptability of the model.
It significantly improves the detection accuracy of small-size and low-recognition pollutants, enhances the robustness and generalization capabilities of the model, and can effectively detect small targets and adapt to complex backgrounds.
Smart Images

Figure CN120355690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting the cleaning effect of wafers based on improved YOLOv7. Background Art
[0002] Wafer cleaning is an important step in semiconductor manufacturing. If various contaminants such as particles, metal ions, and organic residues on the wafer surface are not cleaned thoroughly, it will affect subsequent process steps and ultimately lead to a decline or even failure of chip performance. Therefore, it is necessary to ensure the cleanliness of the wafer surface through wafer cleaning.
[0003] In the field of semiconductor processing and manufacturing, high requirements are placed on the cleaning effect of wafers. Therefore, it is necessary to detect the surface after cleaning. Due to the characteristics of the visual mode such as non-contact, high precision, and high speed, it provides the possibility for wafer cleaning detection. Visual detection usually includes a high-resolution camera, image processing software, and algorithms to identify surface defects, contaminants, scratches, etc., to ensure that there are no residual contaminants and to check for other types of defects. However, the foreign objects on the wafer surface are small in size and low in distinguishability. The present invention proposes a method for detecting the cleaning effect of wafers based on improved YOLOv7, designs a Concat_attention feature fusion module, introduces an SIoU loss function, and uses label smoothing technology to enhance the generalization ability of the model and optimize the detection accuracy of foreign objects after wafer cleaning. Summary of the Invention
[0004] The present invention provides a method for detecting the cleaning effect of wafers based on improved YOLOv7 to solve the problems existing in the above-mentioned prior art.
[0005] The technical solutions adopted by the present invention are as follows:
[0006] A method for detecting the cleaning effect of wafers based on improved YOLOv7 includes the following steps:
[0007] S1) Collect a dataset of pictures of contaminants on the wafer surface and perform preprocessing, and then divide the dataset into a training set and a test set;
[0008] S2) Improve the network structure based on the YOLOv7 algorithm to obtain an improved YOLOv7 detection model. The improvement points are as follows:
[0009] 1) Replace the Concat module of the original YOLOv7 algorithm with a Concat_attention feature fusion module. The feature fusion process of the Concat_attention feature fusion module is as follows:
[0010] Input the low-level feature map and the high-level feature map into two parallel branches respectively. Each branch includes a convolutional layer and a BN layer;
[0011] The convolutional outputs of the two branches are added element-wise and then passed through an activation function to generate a fused feature map;
[0012] A spatial attention weight map is generated through the Sigmoid activation function;
[0013] The low-level feature map is multiplied by the spatial attention weight map to generate the final output feature map;
[0014] 2) By introducing the SIOU loss function that combines angle loss, distance loss, shape loss, and IoU loss to replace the CIOU function of the original YOLOv7 algorithm, the matching accuracy between the predicted bounding box and the ground truth bounding box is optimized;
[0015] 3) Label smoothing is introduced at the output end of the original YOLOv7 algorithm;
[0016] S3) The trained improved YOLOv7 detection model is applied to the detection after wafer cleaning, and the detection results are output.
[0017] Furthermore, the preprocessing includes:
[0018] The dataset is expanded through data augmentation techniques;
[0019] The expanded dataset is labeled using LabelImg software, and the labeled files are converted to the YOLO format for model training.
[0020] Furthermore, the angle loss formula is expressed as follows:
[0021]
[0022] where x, σ, c h are expressed as:
[0023]
[0024] The distance loss is based on the angle loss formula and is expressed as follows:
[0025]
[0026] where
[0027] The shape loss formula is defined as follows:
[0028]
[0029] where θ represents the attention degree of the shape loss, then ω w and ω h are expressed as:
[0030]
[0031] The formula of the SIOU loss function is defined as:
[0032]
[0033] In all formulas:
[0034] Λ: Angle loss, used to measure the angular difference between the predicted bounding box and the ground truth bounding box;
[0035] x: The ratio calculated from the center point coordinates of the ground truth bounding box and the predicted bounding box, c h is the vertical offset, and σ is the Euclidean distance between the center point of the ground truth bounding box and the center point of the predicted bounding box;
[0036] α: The angle between the line connecting the centers of the predicted bounding box and the ground truth bounding box and the horizontal axis;
[0037] Δ: Distance loss, used to measure the distance difference between the predicted bounding box and the ground truth bounding box;
[0038] ρx: The square of the relative offset in the horizontal direction;
[0039] ρy: The square of the relative offset in the vertical direction;
[0040] γ: The parameter related to the angle loss;
[0041] Ω: Shape loss, used to measure the differences in width and height between the predicted bounding box and the ground truth bounding box;
[0042] θ: The attention parameter of the shape loss;
[0043] ω w : The relative difference in the width direction;
[0044] ω h : The relative difference in the height direction;
[0045] LossSioU: The total loss of the SIOU loss function.
[0046] IoU: Intersection over Union, used to measure the overlapping degree between the predicted bounding box and the ground truth bounding box.
[0047] Furthermore, the specific method of the label smoothing is as follows:
[0048] Convert the original label to the smoothed label, and the conversion formula is:
[0049] q′(k∣z i )=(1 - ε)δk,t i +εU(k)
[0050] q′(k|z i ): The smoothed label distribution, representing the probability of class k given the input z i .
[0051] z i : The input feature vector or input sample;
[0052] ε: The smoothing coefficient, with a value range of [0, 1], used to control the degree of label smoothing;
[0053] δk,t i : The indicator function, which is 1 when k = t i and 0 otherwise; t i is the original label;
[0054] U(k): The uniform distribution, representing the probability distribution over class k, usually 1 / N, where N is the total number of classes. The present invention has the following beneficial effects:
[0055] 1) By introducing the Concat_attention feature fusion module, this method can more effectively combine low-level features and high-level features. The Concat_attention module dynamically allocates feature weights through the spatial attention mechanism, avoiding the information loss problem caused by directly concatenating features in the traditional Concat module, thus significantly improving the detection accuracy of small-sized and low-identifiability pollutants.
[0056] 2) The SIOU loss function combines the angular loss, distance loss, shape loss, and IoU loss, and can more comprehensively measure the matching degree between the predicted bounding box and the ground truth bounding box. In particular, the introduction of the angular loss solves the deficiency of the traditional CIOU function in the direction matching of the predicted bounding box, further improving the accuracy of the detection bounding box.
[0057] 3) By using data augmentation techniques (such as rotation, scaling, flipping, etc.) to expand the dataset, the adaptability of the model to different illuminations, angles, and background noises is enhanced. At the same time, using the LabelImg software for standardized annotation ensures the data quality and provides reliable input for model training.
[0058] 4) Introducing label smoothing at the output end can effectively solve the problems of incorrect or inaccurate data annotation, improving the fault tolerance rate and generalization ability of the model. Label smoothing avoids the model's over-reliance on incorrect annotations by smoothing the label distribution, thus showing higher robustness in practical applications.
[0059] 5) The Concat_attention module and SIOU loss function are embedded in the YOLOv7 network as substructures, which not only retains the high efficiency of YOLOv7, but also improves the overall performance through local optimization. This modular design allows the model to maintain high accuracy while not significantly increasing the computational complexity.
[0060] 6) Contaminants on the wafer surface are usually small in size and have low recognition. This method can effectively detect tiny targets through the optimization of the Concat_attention module and the SIOU loss function, thus solving the limitations of traditional methods in small target detection.
[0061] 7) There may be a variety of complex backgrounds on the circular surface (such as scratches, reflections, etc.). Through data enhancement and attention mechanism, this method can adapt to these complex backgrounds and avoid false detection and missed detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 Schematic diagram of the Concat module.
[0063] Figure 2 Schematic diagram of the Concat-attention feature fusion module.
[0064] Figure 3 This is a schematic diagram of SIOU. DETAILED DESCRIPTION
[0065] The present invention will be further described below in conjunction with the accompanying drawings.
[0066] The present invention provides a wafer cleaning effect detection method based on improved YOLOv7, comprising the following steps:
[0067] S1) collecting and preprocessing a wafer surface contaminant image dataset, and then dividing the dataset into a training set and a test set;
[0068] S2) An improved YOLOv7 detection model is obtained by improving the network structure based on the YOLOv7 algorithm, wherein the improvements are as follows:
[0069] 1) Use the Concat_attention feature fusion module to replace the Concat module of the original YOLOv7 algorithm;
[0070] 2) By introducing the SIOU loss function that combines angle loss, distance loss, shape loss and IoU loss to replace the CIOU function of the original YOLOv7 algorithm, the matching accuracy between the predicted box and the real box is optimized;
[0071] 3) Introducing label smoothing at the output of the original YOLOv7 algorithm;
[0072] S3) Apply the trained improved YOLOv7 detection model to the inspection after wafer cleaning and output the inspection results.
[0073] The present invention proposes a Concat_attention structure based on the Concat structure in the YOLOv7 network structure. The Concat structure and the Concat_attention structure of YOLOv7 are as Figure 1 and Figure 2 shown. This structure is similar to the spatial attention mechanism of the CBAM module. The optimized network algorithm is not a complete network algorithm structure, but a sub-structure that needs to be embedded into other detection or classification models, which can improve the efficiency and accuracy of object detection to a certain extent.
[0074] In YOLOv7, the Concat structure directly performs feature fusion on the low-level feature map and the high-level feature map and outputs. The dimensions of the high-level feature map and the low-level feature map are both B×C×W×H, where B represents the batch size, C represents the number of channels, and H and W represent the height and width.
[0075] The Concat_attention feature fusion module of the present invention includes the following three steps:
[0076] The first step is to input the high and low level feature maps into two identical parallel branches respectively, including a convolutional layer and a BN layer. After obtaining the convolutional outputs of the two branches, they are fused by element-wise addition, and a fused feature map is obtained through an activation function.
[0077] The second step is to pass the fused feature map through a convolutional layer and a BN layer again. Then, the feature map generates a weight map (B×1×H×W) through the Sigmoid activation function, which is used to weight the low-level feature map, that is, the spatial attention map at the width and height positions of the feature map.
[0078] The third step is to multiply the input low-level feature map by the spatial attention map output in the previous step and generate the final output feature map through the Concat operation.
[0079] Introduce the SIOU loss function
[0080] CIOU is used to calculate the prediction box loss in YOLOv7. Although CIOU comprehensively considers the matching degree of the prediction box and the ground truth box from aspects such as the overlapping area, the Euclidean distance between the center points, and the intersection over union, it does not consider whether the directions of the two match, which may cause the phenomenon of low convergence efficiency.
[0081] Based on this, the present invention uses the SIOU loss function to replace the CIOU function to further study the vector angle between the predicted box and the ground truth box, and readjusts the loss function structure. The schematic diagram of the SIOU parameters is as shown in Figure 3 Figure (the red box in the figure is the ground truth box, and the blue box is the predicted box).
[0082] The SIOU Loss loss function mainly includes four aspects: angle, distance, shape, and IoU, which are specifically expressed as Angle cost, Distance cost, Shape cost, and IoU cost.
[0083] The formula for the angle loss Angle cost is as follows:
[0084]
[0085] In the formula, x, σ, c h are expressed as follows:
[0086]
[0087] The distance loss Distance cost is based on the above angle loss formula, and the formula is as follows:
[0088]
[0089] Among them,
[0090] The shape loss equation is defined as follows:
[0091]
[0092] θ represents the attention degree of the shape loss Distance cost, then ω w and ω h are expressed as:
[0093]
[0094] Finally, from the above four parts of the formula, the SIOU loss function formula is defined as:
[0095]
[0096] In the above four parts of the formula: Λ: angle loss, used to measure the angle difference between the predicted box and the ground truth box;
[0097] x: the ratio calculated from the center point coordinates of the ground truth box and the predicted box, c h is the offset in the vertical direction, and σ is the Euclidean distance between the center point of the ground truth box and the center point of the predicted box;
[0098] α: The angle between the line connecting the centers of the predicted box and the ground truth box and the horizontal axis;
[0099] Δ: Distance loss, used to measure the distance difference between the predicted box and the ground truth box;
[0100] ρx: The square of the relative offset in the horizontal direction;
[0101] ρy: The square of the relative offset in the vertical direction;
[0102] γ: A parameter related to the angle loss;
[0103] Ω: Shape loss, used to measure the differences in width and height between the predicted box and the ground truth box;
[0104] θ: The attention parameter of the shape loss;
[0105] ω w : The relative difference in the width direction;
[0106] ω h : The relative difference in the height direction;
[0107] LossSioU: The total loss of the SIOU loss function.
[0108] When training with the YOLOv7 network model, it may be necessary to collect a large number of datasets, and the datasets themselves may contain categories of different sizes. Therefore, annotation errors may occur due to the differences in a very small part, which will affect the accuracy of the algorithm training, and further affect the recognition results and detection efficiency of the entire algorithm.
[0109] To solve the above problems, the present invention introduces LabelSmoothing at the output prediction module of the YOLOv7 network. It can improve the efficiency and accuracy of the algorithm, and finally achieve the effect of overfitting.
[0110] Assume that the training samples are (zi, ti). To improve the applicability of the algorithm of the present invention, if it is impossible to ensure the correct annotation of all datasets in actual situations, then label smoothing can be used to ensure a certain error tolerance. The calculation formula of the label smoothing method can be expressed as follows:
[0111] q′(k∣z i )=(1-ε)δk,t i +εU(k)
[0112] q′(k∣z i ):The smoothed label distribution, indicating the probability of class k given the input z i ;
[0113] zi : Input feature vector or input sample;
[0114] ε: Smoothing coefficient, with a value range of [0, 1], used to control the degree of label smoothing;
[0115] δk,t i : Indicator function, which is 1 when k = t i and 0 otherwise; t i is the original label;
[0116] U(k): Represents the probability distribution over class k, and N is the total number of classes.
[0117] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.
Claims
1. A wafer cleaning effect detection method based on improved YOLOv7, characterized in that: It includes the following steps: S1) Collect a dataset of wafer surface pollutant pictures and preprocess it, and then divide the dataset into a training set and a test set; S2) Improve the network structure based on the YOLOv7 algorithm to obtain an improved YOLOv7 detection model. The improvement points are as follows: 1) Replace the Concat module of the original YOLOv7 algorithm with a Concat_attention feature fusion module. The feature fusion process of the Concat_attention feature fusion module is as follows: Input the low-level feature map and the high-level feature map into two parallel branches respectively. Each branch includes a convolutional layer and a BN layer; Add the convolutional outputs of the two branches element-wise and generate a fused feature map through an activation function; Generate a spatial attention weight map through a Sigmoid activation function; Multiply the low-level feature map by the spatial attention weight map to generate the final output feature map; 2) Replace the CIOU function of the original YOLOv7 algorithm with a SIOU loss function that combines angle loss, distance loss, shape loss, and IoU loss to optimize the matching accuracy between the predicted box and the ground truth box; 3) Introduce label smoothing at the output end of the original YOLOv7 algorithm; S3) Apply the trained improved YOLOv7 detection model to the detection after wafer cleaning and output the detection results.
2. The detection method for wafer inspection after cleaning according to claim 1, wherein: The preprocessing includes: Expand the dataset through data augmentation techniques; Use LabelImg software to annotate the expanded dataset and convert the annotation files into the YOLO format for model training.
3. The detection method for after wafer cleaning according to claim 1, wherein: The angle loss formula is expressed as follows: where x, σ, c h are represented as follows: The distance loss is based on the angle loss formula and is expressed as follows: Among them, The shape loss formula is defined as follows: where θ represents the attention of the shape loss, then ω w and ω h are expressed as: The SIOU loss function formula is defined as: In all formulas: Λ: Angle loss, used to measure the angle difference between the predicted box and the ground truth box; x: The ratio calculated from the center point coordinates of the ground truth box and the predicted box, c h is the offset in the vertical direction, and σ is the Euclidean distance between the center point of the ground truth box and the center point of the predicted box; α: The angle between the line connecting the centers of the predicted box and the ground truth box and the horizontal axis; Δ: Distance loss, used to measure the distance difference between the predicted box and the ground truth box; ρx: The square of the relative offset in the horizontal direction; ρy: The square of the relative offset in the vertical direction; γ: A parameter related to the angle loss; Ω: Shape loss, used to measure the difference in width and height between the predicted box and the ground truth box; θ: The attention parameter of the shape loss; ω w : Relative difference in the width direction; ω h : Relative difference in the height direction; LossSioU: The total loss of the SIOU loss function.
4. The detection method for wafer cleaning as described in claim 1, characterized in that: The specific method of the label smoothing is as follows: Convert the original label to a smoothed label, and the conversion formula is: q′(k|z i ) = (1 - ε)δk,t i + εU(k) q′(k∣z i ):The smoothed label distribution, representing the probability of class k given the input z i ; z i : Input feature vector or input sample; ε: Smoothing coefficient, with a value range of [0,1], used to control the degree of label smoothing; δk,t i : An indicator function that is 1 when k = t i and 0 otherwise; t i is the original label; U(k): Represents the probability distribution on class k, and N is the total number of classes.