A metal surface defect detection method based on improved Faster RCNN algorithm
By improving the Faster RCNN algorithm, combining the ResNet101 backbone network and feature pyramid structure, the loss function and anchor point generation method are optimized, and the problem of insufficient position information accuracy and category accuracy in metal surface defect detection is solved, and more efficient metal surface defect detection is achieved.
Patent Information
- Application Number
- CN202210438440.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-25
AI Technical Summary
The prior art has problems with insufficient position information accuracy and category accuracy in metal surface defect detection, especially when the defect target and background are small, the scale changes are large, the types are diverse, and the contrast is low, the detection efficiency is low and the effect is poor.
Using the improved Faster RCNN algorithm, by obtaining metal surface defect image samples from the NEU-DET public database for pre-processing, ResNet101 is selected as the backbone network, the feature pyramid structure is designed for feature fusion, and deformable convolution is introduced to adapt to unknown changes. Optimize the loss function, adopt Rank&Sort Loss and GIOU loss, and improve the anchor point generation method to improve the matching degree of anchor boxes.
It significantly improves the position information accuracy and category accuracy of metal surface defect detection, improves the detection effect of small targets, solves the problem of category imbalance, and shows good detection performance on the metal surface defect data set.
Smart Images

Figure CN114897802B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision target detection, and in particular relates to a metal surface defect detection method based on an improved Faster RCNN algorithm, which can be used for metal surface defect detection. Background Art
[0002] For the problem of metal surface defect detection, the main methods currently used can be divided into two categories: one is the traditional surface defect detection method based on machine vision; the other is the target detection method based on deep learning. Traditional machine vision-based detection methods usually use conventional image processing algorithms or manually designed features plus classifiers. Due to the changing detection scenes in the process of industrial defect detection, it often faces challenges such as small differences between defect targets and backgrounds, large variations in defect scales, diverse types, low contrast, and a large amount of noise in defect images. Traditional machine vision-based methods have high application costs and low universality, and cannot better cope with these challenges, resulting in low detection efficiency and poor detection results.
[0003] With the rapid development of computer vision technology, target detection methods based on deep learning have also been more widely used in the industry. The general process is to first extract the features of the image, and then use the deep convolutional network model to identify and locate the target. In the defect detection algorithm based on deep learning, firstly, unreasonable anchor frame design will make the defect positioning inaccurate, resulting in missed detection and false detection; secondly, the defect detection effect on very small targets is not very ideal. Finally, due to the small number of industrial defect samples that can be provided in the real industrial environment, the network learning ability is insufficient, which easily leads to low model accuracy and weak robustness. Summary of the invention
[0004] In order to overcome the shortcomings and deficiencies in the above-mentioned prior art, the object of the present invention is to provide a metal surface defect detection method based on an improved Faster RCNN algorithm, which can effectively improve the position information accuracy and category accuracy of metal surface defect detection.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A metal surface defect detection method based on an improved Faster RCNN algorithm comprises the following steps:
[0007] Step 1: Obtain metal surface defect image samples from the NEU-DET public database collected by Northeastern University and preprocess them;
[0008] Step 2: Based on the Faster RCNN target detection model, select a backbone network with fewer parameters and better performance as the feature extraction network of the model, design a feature pyramid structure for feature fusion, and introduce deformable convolution with better adaptability to unknown changes;
[0009] Step 3: Optimize the loss function, simplify the complexity of model training, and improve model performance;
[0010] Step 4: Optimize the Region Proposal Networks (RPN) and improve the anchor point generation method to make the produced anchor frames more consistent with the defect target scale;
[0011] Step 5: Model training;
[0012] Step 6: Use the trained model to predict the test samples and output the category and location information of metal surface defects.
[0013] In step 1 described above, the specific steps are as follows: First, obtain metal surface defect image samples from the NEU-DET public database collected by Northeastern University, normalize the samples in the data set to 416×416 size, divide the training set and test set into a ratio of 8:2, and perform a series of data enhancement operations on the training samples, such as translation, flipping, adjusting saturation and contrast.
[0014] In the step 2 described above, the backbone network ResNet101 with fewer parameters and better performance is selected as the feature extraction network in the Faster RCNN target detection model. The ResNet101 network consists of 100 convolutional layers and one fully connected layer. The image scale is changed by the convolution step size. Its network structure is divided into five stages, and the image is downsampled by 2 times in each stage. In order to enhance the detection effect of small targets, a feature pyramid is introduced for feature fusion. All 3x3 traditional convolutions in the last three stages of ResNet101 are replaced with deformable convolutions.
[0015] In step 3 described above, the total loss of the model is the weighted sum of the classification loss and the regression loss. The specific improvement method is to replace the cross entropy loss in the classification loss with Rank&Sort (RS) Loss, and the regression loss uses GIOU loss. The weighted parameter of the total loss of the improved model is RS Loss divided by the regression loss. The calculation formula of RS Loss is as follows:
[0016]
[0017] Where: p is the set of positive samples; l RS (i) is the sum of the current rank error and the current sort error, that is, the current RS error; It is the sum of the target rank error and the target sort error, that is, the target RS error.
[0018] In the fourth step described above, the region proposal network is optimized and the anchor point generation method is improved. According to the aspect ratio and area size feature information of the defect target annotation box on each image sample in the database obtained by analysis, the aspect ratio and scaling ratio in the anchor point generation method are improved to make the generated anchor box and the defect target more matched.
[0019] In step 5 described above, the model training includes the following steps:
[0020] 5.1 Network parameter initialization;
[0021] 5.2 Set training parameters;
[0022] 5.3 Load training data;
[0023] 5.4 Iterative training.
[0024] In the above-mentioned step 5.1 network parameter initialization, the specific operation is: using the resnet101 model to extract the feature information of the input image.
[0025] In the above step 5.2, the training parameters are set as follows: the initial learning rate of the network is set to 0.0012, the learning momentum is set to 0.9, the weight decay coefficient is set to 0.0001, and the SGD optimizer is used.
[0026] In the iterative training of step 5.4 described above, the stochastic gradient descent algorithm is used to iteratively train the improved network structure, and the network parameters are saved every 621 iterations. After continuous iterations, the optimal solution of the network is obtained.
[0027] Beneficial effects of the present invention: Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1. High average accuracy: The backbone network of the present invention adopts the resnet101 model, which avoids the problem of gradient vanishing and degradation in the deep network, and uses the feature pyramid to fuse high-level semantic features and bottom-level definition features, thereby obtaining better feature representation and improving the detection effect of small targets. The deformable convolution structure is embedded in ResNet101 to better learn defect features;
[0029] 2. More accurate target location information: The present invention improves the anchor point generation method, making the generated anchor frame positioning more accurate and more in line with the defect target;
[0030] 3. Improved the class imbalance problem: The present invention adopts RS Loss in the loss function part, so that the model does not need additional auxiliary heads during training. RS Loss distinguishes positive and negative samples according to the classification scores, thereby solving the extreme class imbalance problem, simplifying the complexity of model training, and achieving better performance of the model;
[0031] 4. The present invention is compared with other commonly used target detection algorithms on metal surface defect data sets and also has good detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the process of the present invention;
[0033] Figure 2 It is a network structure diagram for feature extraction of the present invention;
[0034] Figure 3 This is a deformable convolution structure diagram. DETAILED DESCRIPTION
[0035] In order to deepen the understanding of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The embodiments are only used to explain the present invention and do not limit the protection scope of the present invention.
[0036] Embodiment: A metal defect detection method based on an improved Faster RCNN algorithm, such as Figure 1 As shown, the steps are as follows:
[0037] Step 1: Analysis and processing of metal surface defect dataset: Obtain metal surface defect image samples from the NEU-DET public database collected by Northeastern University. The data samples contain six types of defects: cracks, inclusions, patches, pits, rolled-in scale, and scratches.
[0038] The data samples are divided into training set and test set in a ratio of 8:2, the samples are normalized to 416×416 size, and a series of data enhancement operations such as translation, flipping, saturation adjustment and contrast are performed on the training samples.
[0039] Analyze the length-to-width ratio, area size and other information of the defect annotation box of each image in the data sample.
[0040] Step 2: Based on the Faster RCNN target detection model, select the backbone network model ResNet101 with fewer parameters and better performance as the feature extraction network of the model, design a feature pyramid structure for feature fusion, and introduce deformable convolution with better adaptability to unknown changes.
[0041] Among them, the ResNet101 network model consists of 100 convolutional layers and one fully connected layer, and the image scale is changed by the convolution step size, such as Figure 2 As shown in the figure, it is divided into five stages. After each stage, the image is downsampled by 2 times. The size of the feature maps output in the five stages is 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, respectively. The first stage is composed of a 7x7 convolution. The remaining four stages introduce 3, 4, 23, and 3 residual structures composed of 1x1, 3x3, and 1x1 convolutions, respectively, and set the convolution-related parameters in the residual block. The introduction of the residual structure can effectively avoid the problem of gradient disappearance and degradation in the deep network. The feature pyramid structure unifies the number of channels of the four feature maps output by ResNet101 in the last four stages through a 1x1 convolution kernel with a channel number of 256, and then unifies the size of the feature map by upsampling twice. The upper feature map and the lower feature map are fused by adding the corresponding position elements to generate four feature maps of different scales. This method of fusing shallow features with deep features increases the mapping resolution of small targets and effectively improves the detection effect of small targets. Since the defect shapes in metal defect detection are often very irregular, the traditional convolution kernel has poor adaptability to unknown changes. Therefore, the present invention adopts a deformable convolution with better performance. The implementation process of the deformable convolution is as follows: Figure 3 As shown in the figure, the deformable convolution adds an offset to the position of each sampling point in the convolution kernel, so that the convolution kernel can be changed to any shape during the training process, which can better extract the features of defects.
[0042] Step 3. Optimize the loss function: The total loss of the Faster RCNN model is the weighted sum of the classification loss and the regression loss. The specific improvement method is to replace the cross entropy loss in the classification loss with Rank&Sort (RS) Loss, and the regression loss uses GIOU loss. The weighted parameter of the total loss of the improved model is RS Loss divided by the regression loss. The calculation formula of RS Loss is as follows:
[0043]
[0044] Where: p is the set of positive samples; l RS (i) is the sum of the current rank error and the current sort error, that is, the current RS error; It is the sum of the target rank error and the target sort error, that is, the target RS error. RS The calculation formula of (i) is as follows:
[0045]
[0046] Where: x ij is logits s i and j The difference between j is a continuous label (such as intersection over union (IoU)). NFP(i) is the number of negative samples with a score greater than that of the positive sample. Rank(i) is the number of positive and negative samples with a score greater than or equal to that of the positive sample. H(x ij ) is the unit step function.
[0047] The calculation formula is as follows:
[0048]
[0049] Where: is the target loss, the value is 0; H(x ij ) is a unit step function; y * is a continuous label (such as Intersection over Union (IoU)).
[0050] RS Loss consists of two parts: Rank and Sort. Rank refers to distinguishing positive and negative samples based on classification scores, so that all positive samples are ranked before negative samples. Sort refers to sorting positive samples in descending order based on the values in the range of 0-1 obtained by continuous label IoU, so that different positive samples have different priorities during training. RS Loss not only sorts positive and negative samples, but also sorts between positive samples. This feature eliminates the need for additional quality assessment branches for target detection boxes during training, and can effectively solve extreme class imbalance problems without adding any sample balancing strategies. Since RS Loss balances subtasks such as classification and regression by considering loss values, there is no need to repeatedly adjust hyperparameters during model training. Only the learning rate needs to be adjusted to continuously improve model performance.
[0051] Step 4, optimize the region proposal network and improve the anchor point generation method: In Faster RCNN, by default, a group of anchor boxes are generated by combining three aspect ratios (0.5, 1, 2) and three scaling ratios (8, 16, 32), that is, a group of anchor boxes consists of 9 anchor boxes. However, since there are many defect targets with small scales and large aspect ratio differences in the metal surface defect data set studied by the present invention, the anchor boxes generated by the default anchor point generation method have a large difference in scale with the target defect and cannot match the target scale well. Therefore, the present invention uses a data analysis tool to analyze the aspect ratio and other features of the defect targets in the metal defect data set to obtain an anchor box that is closer to the target scale. Finally, it is determined that a group of anchor boxes (a group of anchor boxes consists of 10 anchor boxes) are generated by combining five aspect ratios (0.2, 0.5, 1.0, 2.0, 5) and two scaling ratios (2, 8).
[0052] Step 5: Model training: The present invention adopts the MMDetection deep learning target detection framework based on pyTorch, and performs single-card training on a GPU of model NVIDIA GeForce RTX 2080.
[0053] Setting of training parameters: The initial learning rate of the model is set to 0.0012, the learning momentum is set to 0.9, the weight decay coefficient is set to 0.0001, and the SGD optimizer is used. Due to single-card training, a large Batch Size will cause insufficient video memory, so the Batch Size is set to 2 and the Epoch is set to 36.
[0054] After the hyperparameters are set, the model is trained and the trained model is evaluated using the performance evaluation indicators commonly used in target detection algorithms, namely Average Precision (AP) and Mean Average Precision (mAP).
[0055] Step 6: Use the trained model to predict the test samples and output the category and location information of metal surface defects.
[0056] By using the improved Faster RCNN model designed by the present invention, after the user provides an image, the system can detect relevant information of metal defects according to the trained model.
[0057] In addition to the above embodiments, the present invention may also be implemented in other ways. Any technical solutions formed by equivalent replacement or equivalent transformation shall fall within the protection scope required by the present invention.
Claims
1. A metal surface defect detection method based on an improved Faster RCNN algorithm, characterized in that: The following steps are involved: Step 1: Obtain metal surface defect image samples from the NEU-DET public database collected by Northeastern University and preprocess them; Step 2: Based on the Faster RCNN target detection model, the backbone network ResNet101 with fewer parameters and better performance is selected as the feature extraction network in the Faster RCNN target detection model. The ResNet101 network consists of 100 convolutional layers and one fully connected layer. The image scale is changed by the convolution step size. Its network structure is divided into five stages, and the image is downsampled by 2 times in each stage. In order to enhance the detection effect of small targets, the feature pyramid is introduced for feature fusion. Replace all 3×3 traditional convolutions in the last three stages of ResNet101 with deformable convolutions; Step 3: Optimize the loss function, simplify the complexity of model training, and improve model performance. The total loss of the model is the weighted sum of the classification loss and the regression loss. The specific improvement method is to replace the cross entropy loss in the classification loss with Rank&Sort Loss, and the regression loss uses GIOU loss. The weighted parameter of the total loss of the improved model is RS Loss divided by the regression loss. The calculation formula of RS Loss is as follows: Where: p is the set of positive samples; l RS (i) is the sum of the current rank error and the current sort error, that is, the current RS error; It is the sum of the target rank error and the target sort error, that is, the target RS error; Step 4: Optimize the region proposal network and improve the anchor point generation method to make the produced anchor frame more consistent with the defect target scale; Step 5: Model training; Step 6: Use the trained model to predict the test samples and output the category and location information of metal surface defects.
2. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 1 is characterized in that: In the step 1, first, metal surface defect image samples are obtained from the NEU-DET public database collected by Northeastern University, the samples in the database are normalized to a size of 416×416, the training set and the test set are divided into a ratio of 8:2, and a series of data enhancement operations such as translation, flipping, adjusting saturation and contrast are performed on the training samples.
3. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 1, characterized in that: In the step 4, the region proposal network is optimized, and the aspect ratio and scaling ratio in the anchor point generation method are improved according to the aspect ratio of the defect target annotation box on each image sample in the database obtained by analysis, and the area size feature information, so that the generated anchor box and the defect target are more matched.
4. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 1, characterized in that: In step 5, the training of the model includes the following steps: 4.1 Network parameter initialization; 4.2 Set training parameters; 4.3 Load training data; 4.4 Iterative training.
5. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 4 is characterized in that: In the step 4.1 of initializing the network parameters, the specific operation is: using the resnet101 model to extract the feature information of the input image.
6. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 4 is characterized in that: In step 4.2, the training parameters are set as follows: the initial learning rate of the network is set to 0.0012, the learning momentum is set to 0.9, the weight decay coefficient is set to 0.0001, and the SGD optimizer is used.
7. The metal surface defect detection method based on the improved Faster RCNN algorithm according to claim 4, characterized in that: In the iterative training of step 4.4, the improved network structure is iteratively trained using a stochastic gradient descent algorithm, and the network parameters are saved every 621 iterations. After continuous iterations, the optimal solution of the network is obtained.
Citation Information
Patent Citations
Hotel scene picture target detection method based on Faster R-CNN-FFS model
CN113469272A
RS loss-based target detection model training method and device
CN114005009A