An Anchor-Free Surface Defect Detection Method Based on Multi-Scale Features
The multi-scale feature-based anchor-free defect detection method addresses anchor-based model limitations by using a Swin Transformer and FCOS detection heads with center sampling, improving precision and robustness for diverse defect shapes, especially small and elongated defects.
Patent Information
- Application Number
- CN202211686786.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-26
AI Technical Summary
When detecting surface defects in high-tech industries such as integrated circuits, existing anchor frame-based defect detection methods have problems such as rigid detection scale, limited generalization ability, unbalanced positive and negative samples, and large calculation amounts. It is especially difficult to efficiently detect bar defects with small sizes and extreme horizontal and vertical ratios.
The anchor-free frame surface defect detection method based on multi-scale features is adopted, combined with the Swin Transformer model and feature pyramid network, target detection is performed through the FCOS detection head, and the target frame is optimized using the central sampling strategy to realize pixel-by-pixel regression coding, reducing hyperparameter dependence, and improving detection accuracy and robustness.
It realizes efficient and adaptable target surface defect detection, balances the positive and negative samples in the target frame, improves the detection accuracy and robustness of small-size and extreme horizontal and aspect ratio defect targets, and is suitable for the detection of integrated circuit chips and industrial parts.
Smart Images

Figure CN115861281B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to a surface defect detection method without anchor boxes based on multi-scale features. Background Art
[0002] The rapid development of the manufacturing industry is inseparable from a new ecological model of high quality and digitization. In the production process of high-tech industries such as integrated circuits, surface defects will seriously affect the performance of electronic products and lead to potential safety hazards. Therefore, defect detection, as a key step in industrial production and quality control, timely detecting and identifying industrial defects can effectively ensure the production quality of products.
[0003] With the development of deep learning, more and more defect detection methods have been introduced for defect detection, which are generally divided into two types: one is based on anchor boxes, and the other is anchor-box-free. Since the introduction of anchor boxes often brings some adverse effects, for example, in the implementation of existing Anchor-based models, a set of target boxes needs to be preset for the detection network model based on prior knowledge, and then the model outputs the fine-tuning parameters of the target boxes. Finally, the target boxes output by the final model are calculated through the fine-tuning parameters and the preset parameters of the target boxes. Therefore, the scale and shape of the preset target boxes directly affect the model training effect. And because the scale and shape of the preset target boxes are related to the input data, the model design depends on the designer's understanding of prior knowledge, the detection scale is too rigid, and the generalization ability is limited. In addition, it will also result in unbalanced positive and negative samples and introduce more hyperparameters, greatly increasing the design difficulty and computational complexity. It greatly limits the accuracy and robustness of detecting problems such as linear proportion defects that frequently appear on the surface of microelectronic device integrated circuits (ICs). Therefore, considering the limitations of the traditional model based on the anchor-box method, such as slow detection speed, limited detection scale, and difficult parameter adjustment, the present invention proposes a surface defect detection method without anchor boxes based on multi-scale features. Summary of the Invention
[0004] The purpose of the present invention is to propose a surface defect detection method without anchor boxes with multi-scale features for the above problems, which can achieve efficient and multi-scale target surface defect detection, balance the positive and negative samples within the target boxes, and better represent targets with large shape variations, especially having good detection accuracy and robustness for small-size defect targets and bar-shaped defect targets with extreme aspect ratios.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A surface defect detection method without anchor boxes based on multi-scale features proposed by the present invention includes the following steps:
[0007] S1. Establish a surface defect detection model. The surface defect detection model includes a Swin Transformer model, a first feature extraction unit, an N-layer feature pyramid network, a second feature extraction unit, and N FCOS detection heads, where:
[0008] The first feature extraction unit includes multiple parallel first convolutional layers. Each stage of the Swin Transformer model is sequentially connected to the first N - 1 layers of the feature pyramid network through the first convolutional layer. The second feature extraction unit includes multiple parallel second convolutional layers. The feature pyramid network and the FCOS detection heads are connected to each other through the second convolutional layer. The feature pyramid network corresponding to the last stage of the Swin Transformer model is regarded as the reference layer. Starting from the reference layer, the input feature maps of the corresponding feature pyramid networks are fused from top to bottom through successive upsampling operations, and the output feature map of the reference layer is input into the Nth layer of the feature pyramid network after passing through the third convolutional layer;
[0009] The FCOS detection head includes a classification branch and a regression branch. The classification branch includes a class prediction branch and a center-ness branch. The regression and classification of the output feature map of the corresponding second convolutional layer are performed through each FCOS detection head. The FCOS detection head updates the target box based on the FCOS center sampling optimization method;
[0010] S2. Collect part images as a dataset to train the surface defect detection model;
[0011] S3. Input the part image to be detected into the trained surface defect detection model, predict the center point (c x , c y ) of the target box and the distances (l * , t * , r * , b * ) from this center point to the four sides of the ground truth box, calculate the upper left corner coordinates (c x - l * , c y - t * ) and the lower right corner coordinates (c x + r * , c y + b * ) of the predicted box as the surface defect detection result, where l * is the distance from the center point (c x , c y ) of the target box to the left border of the ground truth box, t * is the distance from the center point (c x , c y ) of the target box to the upper border of the ground truth box, r *is the distance from the center point (c x , c y ) of the target box to the right border of the ground truth box, and b * is the distance from the center point (c x , c y ) of the target box to the lower border of the ground truth box.
[0012] Preferably, the FCOS center sampling optimization method is as follows:
[0013] Take a circular circumscribed sub-rectangle smaller than the target box with the center point of the target box as the center, and update the target box to the sub-rectangle. Consider the sampling points falling within the sub-rectangle as positive samples and other sampling points as negative samples. The position parameters of the sub-rectangle are (c x -rs, c y -rs, c x +rs, c y +rs), where (c x , c y ) are the center point coordinates of the target box, s is the downsampling ratio of the input feature map of the corresponding FCOS detection head, and r is a hyperparameter.
[0014] Preferably, the Swin Transformer model includes a Patch Partition operation module, a first stage, a second stage, a third stage, and a fourth stage connected in sequence. The first stage includes a Linear Embedding operation module and multiple Swin transformer modules connected in sequence. The second stage, the third stage, and the fourth stage each include a PatchMerging operation module and multiple Swin transformer modules connected in sequence.
[0015] Preferably, there are 2 Swin transformer modules in the first stage, the second stage, and the fourth stage, and 6 Swin transformer modules in the third stage.
[0016] Preferably, the convolution kernel size of the first convolutional layer is 1×1 and the number of channels is 256.
[0017] Preferably, the convolution kernel size of the second convolutional layer is 3×3.
[0018] Preferably, the convolution kernel size of the third convolutional layer is 3×3 and the stride is 2.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] This method combines and applies the Swin Transformer model based on the self-attention mechanism in the vision field and the Feature Pyramid Network (FPN model) to encode, fuse, and represent image features. It redefines the object detection head using the anchor-free per-pixel regression encoding method FCOS to obtain the results of the downstream object detection tasks. At the same time, it adopts a center sampling strategy to optimize the FCOS detection head, achieving efficient and multi-scale adaptable object surface defect detection. It not only balances the positive and negative samples within the object bounding box but also can better represent objects with large shape variations, especially having good detection accuracy and robustness for small-sized defect objects and bar-shaped defect objects with extreme aspect ratios, and can be applied to the detection of fields including integrated circuit chips and industrial parts. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the anchor-free surface defect detection method based on multi-scale features of the present invention;
[0022] Figure 2 It is a schematic structural diagram of the surface defect detection model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that when a component is referred to as being "connected" to another component, it can be directly connected to the other component or there may also be an intermediate component. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0025] As Figure 1-2 shown, an anchor-free surface defect detection method based on multi-scale features includes the following steps:
[0026] S1. Establish a surface defect detection model, where the surface defect detection model includes a Swin Transformer model, a first feature extraction unit, N-layer Feature Pyramid Network, a second feature extraction unit, and N FCOS detection heads, where:
[0027] The first feature extraction unit includes multiple parallel first convolutional layers. Each stage of the Swin Transformer model is sequentially connected to the first N-1 layer feature pyramid networks through the first convolutional layers one by one. The second feature extraction unit includes multiple parallel second convolutional layers. The feature pyramid network and the FCOS detection head are connected one by one through the second convolutional layers. The feature pyramid network corresponding to the last stage of the Swin Transformer model is regarded as the reference layer. Starting from the reference layer, the input feature maps of the corresponding feature pyramid networks are fused from top to bottom through progressive upsampling operations, and the output feature map of the reference layer is input into the Nth layer feature pyramid network after passing through the third convolutional layer;
[0028] The FCOS detection head includes a classification branch and a regression branch. The classification branch includes a class prediction branch and a center-ness branch. Regression and classification of the output feature maps of the corresponding second convolutional layers are performed through each FCOS detection head. The FCOS detection head updates the target box based on the FCOS center sampling optimization method.
[0029] In one embodiment, the FCOS center sampling optimization method is as follows:
[0030] Take a circular circumscribed sub-rectangle smaller than the target box with the center point of the target box as the center, and update the target box to the sub-rectangle. The sampling points falling within the sub-rectangle are regarded as positive samples, and other sampling points are regarded as negative samples. The position parameters of the sub-rectangle are (c x -rs, c y -rs, c x +rs, c y +rs), where (c x , c y ) are the center point coordinates of the target box, s is the downsampling ratio of the input feature map of the corresponding FCOS detection head. For example, if the downsampling ratio stride is 2, r is a hyperparameter, such as r = 1.5.
[0031] In one embodiment, the Swin Transformer model includes a Patch Partition operation module, a first stage, a second stage, a third stage, and a fourth stage connected in sequence. The first stage includes a LinearEmbedding operation module and multiple Swin transformer modules connected in sequence. The second stage, the third stage, and the fourth stage all include a Patch Merging operation module and multiple Swin transformer modules connected in sequence.
[0032] In one embodiment, there are 2 Swin transformer modules in the first stage, the second stage, and the fourth stage, and 6 Swin transformer modules in the third stage.
[0033] In one embodiment, the convolution kernel size of the first convolutional layer is 1×1, and the number of channels is 256 dimensions.
[0034] In one embodiment, the convolution kernel size of the second convolutional layer is 3×3.
[0035] In one embodiment, the convolution kernel size of the third convolutional layer is 3×3, and the stride is 2.
[0036] As Figure 2 shown, the feature extraction backbone network adopted by this method is the Swin-Tiny version in the SwinTransformer model, which greatly improves the model detection speed. The first stage to the fourth stage correspond to Swintransformer Stage1 to Swin transformer Stage4 in sequence. 2, 2, 6, and 2 Swin transformer modules are used in the 4 stages respectively. The input image (Images) is divided into 8×8 windows, and the 7×7 pixel small image patches of each window are divided and input into the Swin transformer module for self-attention calculation. The parameters calculated by each window are shared. Therefore, the sequence length input into the Swin transformer module is uniformly 49. Specifically:
[0037] In the forward propagation process, first, the image (Images) with a size of H×W×3 is split into 4×4 sizes through the patchpartition operation, with a total of H / 4×W / 4 small image patches.
[0038] Immediately afterwards, it is encoded through a stacked Swin transformer module. In the first stage, the vector dimension is changed to the sequence length C (such as 96) accepted by the Swin transformer module through the patchembedding operation module and flattened.
[0039] When entering the first stage, the second stage, and the third stage, through the patch merging operation module, which is similar to pooling downsampling in a convolutional neural network, the features of each 2×2 adjacent small image patch are merged, and the length and width are halved while the number of channels is expanded to four times. At the same time, in order to be unified with the convolutional form, a 1×1 convolution is used to halve the number of channels to twice.
[0040] Finally, the backbone network outputs 4 sampled feature maps with downsampling factors of 4, 8, 16, and 32 times respectively and channel numbers of 96, 192, 384, and 768 dimensions, ending the feature representation learning process of the backbone network.
[0041] To address the situation of large-scale changes in the detection target scale, it is necessary to fuse feature information of different scales onto feature maps with different scaling ratios to adaptively predict the size information of the target bounding box in advance. Therefore, in this embodiment, a 5-layer Feature Pyramid Network (FPN) is used to connect the backbone network SwinTransformer model and the FCOS detection head before and after. The first to fifth layers correspond to Feature Pyramid P3 to Feature Pyramid P7 in sequence, specifically as follows:
[0042] First, 4 convolutions with 1×1 channels of 256 dimensions (the first convolutional layer) are constructed on the input side of the feature pyramid network to change the channel number of the sampled feature map. The 4 different hierarchical sampled feature maps extracted by the backbone network are respectively input into the corresponding feature pyramid network through the first convolutional layer.
[0043] The feature pyramid network corresponding to the fourth stage of the backbone network is regarded as the base layer. The 32× sampled feature map extracted in the fourth stage is fused with the 16×, 8×, and 4× sampled feature maps extracted from the third stage to the first stage through three successive upsampling operations from top to bottom. Moreover, the 32× sampled feature map extracted in the fourth stage passes through a 3×3 convolutional layer with a stride of 2 (the third convolutional layer) to obtain a smaller feature map, so as to generate stronger semantic information.
[0044] Through the above method, target bounding boxes of different sizes and scales can be naturally assigned to the 5 feature maps output by the above 5-layer feature pyramid network. Further, the FCOS detection head is used to regress and classify the targets in different feature maps. At the same time, the adopted FPN network and FCOS center sampling optimization can well adapt to the detection field with multi-scale transformations, and the difficulty of parameter tuning can also be reduced based on the anchor-free mechanism.
[0045] The FCOS detection head takes the output feature map of the corresponding second convolutional layer as input and is divided into two branches (classification branch and regression branch) to further regress and classify the corresponding feature map. In both branches, 4 3×3 convolutions are continued to be used for further encoding to extract classification and regression feature representations. During the convolutional encoding process, the length and width scales of the feature map are not changed, and the channel number of the feature map is maintained at 256. The 5 FCOS detection head feature levels are shared.
[0046] The first branch is the classification branch, including the class prediction branch and the center-ness branch. The class prediction branch outputs H×W×C, that is, a C-dimensional vector is predicted for each point on the feature map to determine its class. The center-ness branch outputs H×W×1, that is, only one value is predicted for each point on the feature map to represent its distance weight from the target center.
[0047] The second path is the regression branch, which is responsible for regressing the position of the target box. The regression branch outputs H×W×4, that is, for each point on the feature map, it is regarded as a detection point to regress the 4 distance parameters from this point to the target box.
[0048] The FCOS detection head does not require prior boxes, which greatly reduces the sample size and the number of parameters, thus reducing the computational load. Moreover, it also has good performance for small target detection, and can increase the number of predictable boxes through multi-scale detection.
[0049] Pre-screen and optimize the prediction points within the target box through Center Sample. The core idea of Center Sample is to pre-screen the prediction points that have a higher probability of falling on the true positive samples. First, find the center point within the target box, and take a smaller rectangle circumscribed by a circle centered at the center point. Then redefine the sub-rectangle as the positive sample region. Only the sampling points that fall within the sub-rectangle are positive samples, and other sampling points are regarded as negative samples. After screening, most of the sampling points on the inner edge of the target box but actually falling in the background part are reasonably removed, which optimizes the convergence speed and training results of the FCOS detection head. Specifically, the center sampling defines the center region as a sub-rectangle of the labeled box. The position parameters of the sub-rectangle are (c x -rs,c y -rs,c x +rs,c y +rs), where (c x ,c y ) are the center point coordinates of the target box, s is the downsampling ratio of the input feature map corresponding to the FCOS detection head. For example, if the downsampling ratio stride is 2, r is a hyperparameter, such as r = 1.5, or it can be adjusted according to actual needs.
[0050] S2. Collect part images as a dataset to train the surface defect detection model.
[0051] S3. Input the part image to be detected into the trained surface defect detection model, predict the center point (c x ,c y ) of the target box and the distances (l * ,t * ,r * ,b * ) from this center point to the four sides of the true box, and calculate the upper left corner coordinates (c x -l * ,c y -t * ) and the lower right corner coordinates (c x +r * ,cy +b * ) as the surface defect detection result, where l * is the distance from the center point (c x , c y ) of the target box to the left border of the ground truth box, and t * is the distance from the center point (c x , c y ) of the target box to the upper border of the ground truth box, and r * is the distance from the center point (c x , c y ) of the target box to the right border of the ground truth box, and b * is the distance from the center point (c x , c y ) of the target box to the lower border of the ground truth box. By using the anchor-free method, instead of outputting multiple candidate boxes and then performing post-processing to obtain the detection result, the coordinates of the predicted box are directly output to obtain the surface defect detection result.
[0052] This method fuses and applies the Swin Transformer model based on the self-attention mechanism in the vision field and the Feature Pyramid Network (FPN model) to encode, fuse, and represent image features, and redefines the object detection head using the anchor-free per-pixel regression encoding method FCOS to obtain the results of the downstream object detection tasks. At the same time, a center sampling strategy is adopted to optimize the FCOS detection head, realizing efficient and multi-scale object surface defect detection. It not only balances the positive and negative samples within the target box but also can better represent objects with large shape variations, especially having good detection accuracy and robustness for small-sized defect objects and bar-shaped defect objects with extreme aspect ratios, and can be applied to the detection in fields including integrated circuit chips, industrial parts, etc.
[0053] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0054] The above-described embodiments only express the embodiments of the present application that are described more specifically and in detail, but should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A surface defect detection method without anchor boxes based on multi-scale features, characterized in that: The surface defect detection method based on multi-scale features includes the following steps: S1. Establish a surface defect detection model, which includes a Swin Transformer model, a first feature extraction unit, an N-layer feature pyramid network, a second feature extraction unit, and N FCOS detection heads, where: The first feature extraction unit includes multiple parallel first convolutional layers. Each stage of the Swin Transformer model is sequentially connected to the first N-1 layers of the feature pyramid network through the first convolutional layer. The second feature extraction unit includes multiple parallel second convolutional layers. The feature pyramid network is connected to the FCOS detection heads through the second convolutional layer one by one. The feature pyramid network corresponding to the last stage of the Swin Transformer model is regarded as the base layer. Starting from the base layer, the input feature maps of the corresponding feature pyramid networks are fused from top to bottom through progressive upsampling operations, and the output feature map of the base layer is input into the Nth layer of the feature pyramid network after passing through the third convolutional layer. The convolutional kernel size of the first convolutional layer is 1×1, and the number of channels is 256. The convolutional kernel size of the second convolutional layer is 3×3, and the convolutional kernel size of the third convolutional layer is 3×3, with a stride of 2; The FCOS detection head includes a classification branch and a regression branch. The classification branch includes a class prediction branch and a center-ness branch. The regression and classification of the output feature map of the corresponding second convolutional layer are performed through each FCOS detection head. The FCOS detection head updates the target box based on the FCOS center sampling optimization method; S2. Collect part pictures as a dataset to train the surface defect detection model; S3. Input the image of the part to be detected into the trained surface defect detection model, and predict the center point (c x , c y ) of the target box and the distances (l * , t * , r * , b * ) from this center point to the four sides of the ground truth box. Calculate the upper left coordinate (c x - l * , c y - t * ) and the lower right coordinate (c x + r * , c y + b * ) of the predicted box as the surface defect detection result. Here, l * is the distance from the center point (c x , c y ) of the target box to the left side of the ground truth box, t * is the distance from the center point (c x , c y ) of the target box to the upper side of the ground truth box, r * is the distance from the center point (c x , c y ) of the target box to the right side of the ground truth box, and b * is the distance from the center point (c x , c y ) of the target box to the lower side of the ground truth box.
2. The method for surface defect detection without anchor boxes based on multi-scale features according to claim 1, wherein: The FCOS center sampling optimization method is specifically as follows: Take a circular circumscribed sub-rectangle smaller than the target box with the center point of the target box as the center of the circle, and update the target box to the sub-rectangle. Consider the sampling points falling within the sub-rectangle as positive samples and other sampling points as negative samples. The position parameters of the sub-rectangle are (c x -rs, c y -rs, c x +rs, c y +rs), where (c x , c y ) are the center point coordinates of the target box, s is the downsampling ratio of the input feature map of the corresponding FCOS detection head, and r is a hyperparameter.
3. The method for surface defect detection without anchor boxes based on multi-scale features according to claim 1, wherein: The Swin Transformer model includes a Patch Partition operation module, a first stage, a second stage, a third stage, and a fourth stage connected in sequence. The first stage includes a Linear Embedding operation module and multiple Swin transformer modules connected in sequence. The second stage, the third stage, and the fourth stage all include a PatchMerging operation module and multiple Swin transformer modules connected in sequence.
4. The method for surface defect detection without anchor boxes based on multi-scale features according to claim 3, wherein: There are 2 Swin transformer modules in the first stage, the second stage, and the fourth stage, and 6 Swin transformer modules in the third stage.
Citation Information
Patent Citations
Intelligent ship target detection method based on remote sensing image
CN114972851A
Method and apparatus for detecting metal surface defects
WO2022160170A1