A two-stage surface defect recognition method combining classification and detection
A two-stage network model with deep convolutional neural networks and attention mechanisms enhances defect detection by focusing on key features, improving robustness and efficiency in identifying both large and small defects.
Patent Information
- Application Number
- CN202210528242.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-16
AI Technical Summary
Existing deep learning-based automatic optical detection algorithms have shortcomings in model robustness and interpretability, and a single network structure leads to low overall defect identification and local defect positioning efficiency.
A two-stage network model is built that combines classification and detection. By introducing a deep convolutional neural network with attention mechanism, the defect classification and detection network are trained separately to achieve targeted strengthening of the feature map, and overall defect classification and local defect positioning are carried out through a phased process.
It improves the robustness and interpretability of the algorithm, and at the same time realizes efficient identification of overall defects and precise positioning of local defects, with high computing efficiency.
Smart Images

Figure CN114898153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a two-stage surface defect recognition method combining classification and detection. Background Art
[0002] Automated optical inspection is a machine vision-based method for detecting surface defects of inspected objects. In recent years, due to advantages such as non-contact, high efficiency, and good reliability, automated optical inspection has been increasingly widely used in industrial scenarios and has basically replaced manual visual inspection means in manufacturing fields such as electronic circuits, metal products, food textiles, etc. Most of the current mainstream automated optical inspection algorithms are still traditional machine vision algorithms, such as image processing and template matching. These traditional methods are characterized by strong pertinence, carefully designing algorithms for different defects to achieve a high detection rate, but there are also problems such as poor generalization and scalability, and high false detection rates.
[0003] Automated optical inspection algorithms based on deep learning have been a hot topic in the past two years. Their advantages such as automatic feature extraction, end-to-end detection, and high portability have promoted the rapid popularization of such algorithms in the industrial field and have impacted traditional machine vision algorithms. However, current deep learning surface defect detection algorithms still have some deficiencies, which are reflected in the following two aspects.
[0004] On the one hand, although current deep learning-based surface defect detection algorithms have achieved good image feature learning by using the automatic feature extraction of deep neural networks, they have not specifically learned the key feature channels and feature regions in the feature map, resulting in poor model robustness and poor algorithm interpretability.
[0005] On the other hand, most methods use a single image classification network or object detection network. A single image classification network focuses on the discrimination of overall defects of the inspected object and lacks the ability to identify and locate small defect targets. A single object detection network can identify and locate defects, but there is unnecessary waste of computing resources and efficiency loss in the recognition of overall defects.
[0006] In view of the above two problems, the present invention introduces an attention mechanism and constructs a two-stage network model of image classification + object detection, achieving targeted enhancement of key features in the feature map and organically combining the advantages of the classification network and the detection network through staged deep learning. Summary of the Invention
[0007] The object of the present invention is to address the deficiencies in existing deep - learning - based automatic optical detection algorithms, namely poor model robustness and a single network structure. By introducing a new feature learning method and constructing a two - stage network model for classification + detection, while improving the robustness and interpretability of the algorithm, it realizes the efficient recognition of overall defects and the precise localization of local defects, and has high computational efficiency. The present invention is mainly aimed at detection scenarios where there are two surface defect forms in the object to be inspected, namely overall defects with a large - scale distribution and local defects with a micro - scale, such as defect detection of surface - mounted components on a PCBA.
[0008] The overall technical solution of the present invention is as follows: First, construct a defect classification network and a defect detection network model. Then, construct a data set to train the defect classification network and the defect detection network respectively. Finally, organize the image data stream direction according to the two - stage defect recognition process, and ultimately realize the classification of large - scale overall defects and the localization and recognition of small - scale local defects.
[0009] Specifically, this technical solution includes the following content:
[0010] Step1: Construct a defect classification network model
[0011] The defect classification network is a deep convolutional neural network with a depth of 54, which contains 17 Bottleneck layers. The Bottleneck layer is a spindle - shaped residual connection structure that first expands the channels and then compresses the channels, which can effectively avoid excessive feature loss caused by depth - separable convolutions used in the network. A CBAM channel - spatial attention module is added before the first Bottleneck layer and after the last Bottleneck layer of the network respectively.
[0012] Step2: Train the defect classification network model
[0013] Collect a sufficient number of defective and normal samples for the object to be inspected, use image classification annotation software to label the defect categories and generate a label file. Divide the data set into a training set and a validation set according to a ratio of 4:1 to 6:1. During training, for images with a large aspect ratio of length to width, rotate them uniformly to the same direction, and then obtain input images of a fixed size through scaling and cropping. The SGD method is used during training, and the learning rate is halved every 20 epochs.
[0014] Step3: Construct a defect detection network model
[0015] The defect detection network is divided into three parts: the backbone network, Neck, and the detection head. The backbone network is constructed using CSPDarknet53 and is used to extract image features. Neck is a bidirectional feature pyramid structure of FPN+PAN, which is used for high and low-level feature fusion to improve the network's ability to detect small defect targets. The detection head adopts a decoupled detection head for classification and localization, which can effectively improve the performance of defect detection. The defect detection network no longer predicts multiple anchor boxes at each position of the feature map, but directly predicts the four parameter values of the box: x, y, w, and h. When predicting the target, SimOTA is used for optimal label assignment.
[0016] Step4: Train the defect detection network model
[0017] Collect a sufficient number of defect samples for the object to be inspected, use object detection annotation software to label the defect categories and positions and generate label files. Divide the data set into a training set and a validation set according to a ratio of 4:1 to 6:1. During training, use Mixup and Mosaic data augmentation and disable them in the last 15 epochs. The learning rate adopts the cosine annealing method to avoid premature algorithm convergence.
[0018] Step5: Construct a two-stage surface defect recognition algorithm model
[0019] After the defect classification network and the defect detection network are trained, the object to be inspected is input into the defect classification network for the classification of large-scale overall defects. If the object to be inspected is determined to have overall defects, the object to be inspected is directly judged as NG; if the object to be inspected has no overall defects, the object to be inspected is input into the defect detection network; in the defect detection network, if local defects are detected in the object to be inspected, the object to be inspected is judged as NG, and if no defects are detected in the object to be inspected, the object to be inspected is judged as OK.
[0020] Beneficial effects:
[0021] The beneficial effects of the present invention are that by introducing a new feature learning method and constructing a two-stage network model of classification + detection, not only the targeted enhancement of key features in the feature map is realized, the robustness and interpretability of the algorithm are improved, but also the advantages of both the image classification network and the object detection network are integrated, realizing the efficient recognition of overall defects and the precise localization of local defects, and having high computational efficiency. Brief description of the drawings
[0022] Figure 1 It is a schematic flow chart of the two-stage surface defect recognition method combining classification and detection according to the present invention;
[0023] Figure 2 It is a schematic structural diagram of the defect classification network constructed according to the present invention;
[0024] Figure 3Schematic diagram of the defect classification network structure introducing the attention mechanism in the present invention;
[0025] Figure 4 Schematic diagram of the defect detection network structure constructed according to the present invention. Detailed implementation manners
[0026] The present invention will be further described below in conjunction with the detailed implementation manners. Examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.
[0027] Embodiment 1
[0028] As Figures 1 to 4 shown, this embodiment provides a two-stage surface defect recognition method combining classification and detection, and the method includes the following steps:
[0029] Step1: Construct a defect classification network model
[0030] The structure of the network model is as Figure 2 shown. The depth of the neural network is 54, which includes 17 spindle-shaped Bottleneck layers. For a three-channel input image with a resolution of 224×224×3, it is downsampled into a feature map of 112×112×16 before entering the first Bottleneck layer. As Figure 3 shown, at this time, the CBAM attention block parallel to the feature stream will perform weighted calculation on the feature map, as shown in the following formula.
[0031]
[0032]
[0033] Where
[0034]
[0035]
[0036] The feature map after attention weighting enters the Bottleneck layer. Inside the Bottleneck layer, first, channel expansion is performed through 1×1 convolution, then depthwise separable convolution is executed, and finally, it is compressed back to the original number of channels. In this process, ReLU6 is used as the activation function in both the expansion and convolution, and no activation is performed during the compression process. When stacking the Bottleneck layers, only the first layer will perform 2-fold downsampling, and only the layers that have not been downsampled will use residual connections.
[0037] After the last Bottleneck layer, a CBAM attention block is also set up in parallel with the feature stream. Finally, the feature map passes through average pooling and a fully connected layer to output the prediction vector.
[0038] Step2: Train the defect classification network model
[0039] Collect a sufficient number of defect samples and normal samples for the object to be inspected, use image classification annotation software to label the defect categories and generate a label file. Divide the dataset into a training set and a validation set according to a ratio of 4:1. During training, for images with a large aspect ratio of length to width, rotate them to the same direction uniformly, then scale and crop them to a fixed size (224×224×3 or 90×190×3), and finally normalize the images to obtain the input images. The parameter settings during training are as follows: batch size 64, train for 100 epochs, use the cross-entropy loss function, adopt the stochastic gradient descent method to adjust the learning rate, and halve the learning rate every 20 epochs. After training, save the weight file for later use.
[0040] Step3: Construct the defect detection network model
[0041] As Figure 4 shown, the defect detection network is divided into three parts: the backbone network, Neck, and detection head. The backbone network is constructed using CSPDarknet53 for image feature extraction. The first unit of the backbone network is the Focus block, which is used to stack the spatial blocks of the image in the channel direction. The backbone network stacks several Conv-BN-Act blocks and CSP blocks (the stacking quantity depends on the specific situation), and the SPP block is used for multi-receptive field feature extraction.
[0042] Neck is a bidirectional feature pyramid structure of FPN+PAN, which is used for high and low layer feature fusion to improve the network's detection ability for small defect targets. The three inputs of FPN come from the second and third CSP blocks and the final output of the backbone network respectively, corresponding to feature maps with input sizes of 1 / 8, 1 / 16, and 1 / 32 of the image.
[0043] The detection head adopts a detection head with decoupled classification and localization. The feature map output by PAN first compresses the number of channels to 256 through a 1×1 convolution, then one branch completes the classification function through 3×3 and 1×1 convolutions, and the other branch branches again through a 3×3 convolution to complete coordinate regression and IoU regression respectively. The defect detection network no longer predicts multiple anchor boxes for each position of the feature map, but directly predicts the 4 parameter values of the box's x, y, w, and h. The 3×3 area of the center point is regarded as a positive sample. When predicting the target, SimOTA is used for optimal label assignment to complete the division of positive and negative samples.
[0044] Step4: Train the defect detection network model
[0045] Collect a sufficient number of defect samples for the object to be inspected, use the target detection annotation software to annotate the defect categories and locations and generate label files. Divide the dataset into a training set and a validation set according to a ratio of 4:1. When training, use Mixup and Mosaic data augmentation, set the batch size to 32, train for 200 generations, with a momentum factor of 0.9. The learning rate adopts the cosine annealing method to avoid premature algorithm convergence, and stop data augmentation in the last 15 generations. After training, save the weight file for future use.
[0046] Step5: Construct a two-stage surface defect recognition algorithm model
[0047] After the defect classification network and the defect detection network are trained, input the object to be inspected into the defect classification network for classifying overall defects on a large scale. If the object to be inspected is determined to have overall defects, directly judge the object to be inspected as NG; if the object to be inspected has no overall defects, input the object to be inspected into the defect detection network; in the defect detection network, if local defects are detected in the object to be inspected, judge the object to be inspected as NG, and if no defects are detected in the object to be inspected, judge the object to be inspected as OK.
[0048] By introducing a new feature learning method and constructing a two-stage network model of classification + detection, the present invention not only realizes the targeted enhancement of key features in the feature map, improves the robustness and interpretability of the algorithm, but also combines the advantages of both the image classification network and the target detection network, realizes the efficient recognition of overall defects and the accurate positioning of local defects, and has high computational efficiency.
[0049] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A two-stage surface defect recognition method combining classification and detection, characterized in that The method includes the following steps: S1. Construct a defect classification network model and a defect detection network model respectively; The method for constructing the defect classification network model is as follows: The defect classification network is a deep convolutional neural network with a depth of 54, which contains 17 Bottleneck layers; the Bottleneck layer is a spindle-shaped residual connection structure that first expands the channels and then compresses the channels, effectively avoiding excessive feature loss caused by the depthwise separable convolution used in the network; a CBAM channel-spatial attention module is added before the first Bottleneck layer and after the last Bottleneck layer of the defect classification network respectively; The method for constructing the defect detection network model is as follows: The defect detection network is divided into three parts: a backbone network, a Neck, and a detection head; the backbone network is constructed using CSPDarknet53 for extracting image features; the Neck is an FPN+PAN bidirectional feature pyramid structure for fusing high and low-level features to improve the network's detection ability for small defect targets; the detection head adopts a detection head with decoupled classification and localization, effectively improving the defect detection performance; the defect detection network no longer predicts multiple anchor boxes at each position of the feature map, but directly predicts the four parameter values of the box's x, y, w, and h; when predicting the target, SimOTA is used for optimal label assignment; S2. Construct a dataset according to S1 and train the defect classification network model and the defect detection network model respectively; S3. Construct a two-stage surface defect recognition algorithm model according to S2; The method for constructing the two-stage surface defect recognition algorithm model is as follows: After the defect classification network and the defect detection network are trained, the object to be inspected is input into the defect classification network for classifying large-range overall defects. If the object to be inspected is determined to have overall defects, the object to be inspected is directly judged as NG; if the object to be inspected has no overall defects, the object to be inspected is input into the defect detection network; in the defect detection network, if local defects are detected in the object to be inspected, the object to be inspected is judged as NG, and if no defects are detected in the object to be inspected, the object to be inspected is judged as OK; S4. Input the object to be inspected into the two-stage surface defect recognition algorithm model in S3 for classifying large-range overall defects and locating and recognizing minute local defects.
2. The two-stage surface defect recognition method combining classification and detection according to claim 1, characterized in that The method for training the defect classification network model is as follows: Collect a sufficient number of defect samples and normal samples for the object to be inspected, use image classification annotation software to annotate the defect categories and generate a label file; divide the dataset into a training set and a validation set according to a ratio of 4:1 to 6:1; during training, for images with a large aspect ratio of length to width, rotate them to the same direction uniformly, and then obtain input images with a fixed size through scaling and cropping; the SGD method is adopted during training, and the learning rate is halved every 20 generations.
3. The two-stage surface defect recognition method combining classification and detection according to claim 1, characterized in that, The method for training the defect detection network model is as follows: Collect a sufficient number of defect samples for the object to be inspected, use the target detection annotation software to annotate the defect categories and locations and generate a label file; divide the data set into a training set and a validation set according to a ratio of 4:1 to 6:1; during training, use Mixup and Mosaic data augmentation and deactivate them in the last 15 generations, and adopt the cosine annealing method for the learning rate to avoid premature algorithm convergence.