A strip surface defect detection method based on an improved RetinaNet algorithm
By improving the retinanet algorithm, combining data augmentation, context aggregation and multi-scale receptive field enhancement modules, the problems of low detection efficiency and low accuracy in steel surface defect detection are solved, and more efficient defect detection is achieved.
Patent Information
- Application Number
- CN202211390496.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-08
AI Technical Summary
The prior art has problems such as low detection efficiency, easy missed detection and poor robustness in the detection of steel surface defects, especially in deep learning, loss of feature information, limited spatial information and sample number, resulting in low detection accuracy.
The improved retinanet algorithm is adopted to optimize the loss function to improve detection accuracy through data augmentation, introduction of context aggregation module and multi-scale receptive field enhancement module, combined with a bidirectional weighted fusion network.
The detection effect of small targets is improved, and the detection effect is poor due to large changes in defect size, different shapes and background noise is solved, which is improved detection accuracy and robustness.
Smart Images

Figure CN115760734B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial steel surface defect detection, and particularly relates to a strip steel surface defect detection method based on an improved RetinaNet algorithm. Background Art
[0002] In industrial manufacturing production, surface defect detection is one of the important links to ensure product quality. As an important industrial raw material, the quality of rolled steel directly affects the quality of products. Most traditional defect detections rely on manual quality inspection, which has the problems of low detection efficiency and easy omission of detections.
[0003] Currently, there are mainly two methods for dealing with surface defect detection problems: object detection methods based on machine learning and deep learning. The methods based on machine learning mainly extract artificial features designed through experience. However, in the actual detection process, due to problems such as the inconsistent scales and large variation ranges of surface target defects, the easy confusion of defect categories, and the existence of a large amount of background and noise interference in images, this will lead to an increase in detection costs using machine learning detection methods and cannot well solve the above problems, ultimately resulting in poor detection effects. With the gradual maturity of deep learning technology, object detection algorithms based on deep learning have gradually replaced traditional methods based on machine learning in the industrial field. The general process is to first extract image features through convolution operations, and then obtain multi-scale features after passing through a feature fusion network and send them to the detection layer for defect localization and classification.
[0004] In the defect detection algorithm based on deep learning, first, after extracting features in the backbone, rich feature information may be lost due to channel number reduction. Second, continuous downsampling of the backbone network will lead to the loss of spatial information in deep features, resulting in poor detection effects in dealing with problems of large defect size changes and various shapes. Then, in the feature fusion process, the context information between non-adjacent feature maps cannot be effectively connected, resulting in low precision for small target defects. Finally, the number of samples collected in the industrial environment is limited, the network's learning of defects is insufficient, and the robustness is poor, resulting in low detection accuracy. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a strip steel surface defect detection method based on an improved RetinaNet algorithm, which solves the above problems.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A strip steel surface defect detection method based on an improved RetinaNet algorithm, including:
[0007] Obtain a defective image and load it to form a defective sample. Augment the defective sample by data augmentation methods to increase the dataset samples, where the dataset samples include training set samples, validation set samples, and test set samples;
[0008] Build a surface defect detection network based on the retinanet model. By initializing the parameters, setting hyperparameters, loading the training set samples, and setting the number of iterations for the surface defect detection network in sequence, achieve model training to obtain an optimal detection model;
[0009] Import the test set samples into the detection model for testing, and perform class classification and location regression on the test set samples to obtain the final detection result.
[0010] Based on the above technical solutions, the present invention also provides the following alternative technical solutions:
[0011] Further technical solution: The surface defect detection network includes a feature extraction network, a feature enhancement network, a feature fusion network, and an optimized loss function. The enhancement network includes a context aggregation module and a multi-scale receptive field module. The specific detection method of the surface defect detection network is as follows:
[0012] S1. Obtain an effective feature map P i , (i = 3, 4, 5), where P3 is a shallow feature map, P4 is a mid-high level feature map, and P5 is a deep feature map. The specific operation steps are to select resnet50 with 5 extraction stages as the feature extraction network of the model. Each stage of resnet50 generates feature maps with different resolution sizes through convolutional strides, and select the feature maps of the last three stages as the effective feature maps to obtain the effective feature map P i , (i = 3, 4, 5);
[0013] S2. Introduce the context aggregation module and process P3 through the context aggregation module to obtain P3 with rich small target information;
[0014] S3. Introduce a multi-scale receptive field module with four branches and perform feature enhancement on P5 through the multi-scale receptive field module. In the first three branches of the multi-scale receptive field module, the original feature input feature and the output feature of the previous branch are concatenated as the input of the next branch, that is, point 2;
[0015] S4. Use deformable convolution and 3×3 convolution to downsample the feature-enhanced P5 and generate deeper layer features P6, P7. Send P i , (i = 3, 4, 5, 6, 7) as the effective feature map into the feature fusion network for multi-scale feature fusion;
[0016] S5. Use the bidirectional weighted fusion network BiFPN, that is, the feature fusion network, to perform multi-scale feature fusion on the five effective feature maps P i , (i = 3, 4, 5, 6, 7);
[0017] S6. Optimize the localization loss. Replace the L1 loss in the retinanet model with the SmoothL1 loss function to optimize the localization loss. The expression of the SmoothL1 loss function is:
[0018]
[0019] Specifically, use SGD as the model optimizer to optimize the localization loss.
[0020] Further technical solution: The specific operation steps of S2 are:
[0021] S201. Process P4 and P5 through 1×1 convolution, BN, and LeakyReLU, and then perform upsampling by a factor of two and four;
[0022] S202. Use 3×3 depthwise separable convolution and 1×1 convolution to refine the shallow feature P3 and refine the edge information of small defect features;
[0023] S203. Use a weighted sum operation to multiply each feature map by the learning parameter w i , (i = 3, 4, 5), and further obtain a feature map and perform element-wise addition;
[0024] S204. Add the CA attention mechanism after weighted combination to weaken background information and strengthen the fusion between channels;
[0025] S205. Add the feature map obtained in S204 to the shallow feature map P3 element-wise to obtain P3 with rich small target information, that is, point 1.
[0026] Further technical solution: The steps for the multi-scale receptive field module to enhance the features of P5 are:
[0027] S301. The first branch of the multi-scale receptive field enhancement module uses asymmetric convolutional kernels 1×3 and 3×1 to extract features, and then uses a 3×3 convolution with a dilation rate of 3 to convolve the features and output;
[0028] S302. The second branch of the multi-scale receptive field enhancement module uses asymmetric convolutional kernels 1×5 and 5×1 to extract features, and then uses a 5×5 convolution with a dilation rate of 5 to convolve the features and output;
[0029] S303. The third branch of the multi-scale receptive field enhancement module uses an operation to perform parallel element-wise addition of three convolutional kernels of 1×3, 3×1, and 3×3, and then uses a 3×3 convolution with a dilation rate of 6 to convolve the features and output them;
[0030] S304. The fourth branch of the multi-scale receptive field enhancement module uses global average pooling, followed by upsampling and output;
[0031] S305. Concatenate the outputs of the four branches and then perform dimensionality reduction to obtain the final feature-enhanced P5.
[0032] Further technical solution: The specific operation steps of S5 are as follows:
[0033] S501. Perform feature-guided upsampling on the input features and then add them weighted with to generate Repeat the feature-guided upsampling operation until
[0034] S502. Use the obtained in S501 as the shallow output The output operations of other layers are as follows: After undergoes downsampling, it is weighted and fused with P i td and P i in to be used as the output P i out , where i = 4, 5, 6, is obtained by undergoing downsampling and then weighted and added with ;
[0035] S503. Add a top-down fusion path to further strengthen the fusion between information, and use the Ghost module for dimensionality reduction.
[0036] Further technical solution: The specific steps of the feature-guided upsampling in S501 are as follows:
[0037] S5011. Use bilinear interpolation to upsample and then perform element-wise addition with P i in and use a 3*3 convolution to change the number of channels to 1 to obtain the spatial weight;
[0038] S5012. Normalize the spatial weight obtained in S5011 through the softmax function and use the normalized spatial weight w i and Multiply in the channel dimension to obtain a feature map with more detailed semantics That is, point 3;
[0039] S5013. Repeat the steps of S5011 and S5012 to Perform feature-guided upsampling operation and set P i in 、 Generate P after weighted sum i td , where i = 3, 4, 5, 6.
[0040] Further technical solution: The specific process of S501 is expressed by the formula:
[0041] FGUpsample = Softmax(f 3 (Concat(Upsample(p i+1 ), p i ))) × p i+1
[0042]
[0043] where i = 6, 5, 4, 3, and f 3 (·) represents a convolution operation with a convolution kernel of 3, w1 and w2 are learning weights obtained through fast normalization, and ε is a fixed value of 0.0001.
[0044] Further technical solution: The process of S502 is expressed by the formula:
[0045]
[0046]
[0047]
[0048] where i = 4, 5, 6, w'1, w'2, and w'3 are learning weights obtained through fast normalization, and ε is a fixed value of 0.0001.
[0049] Further technical solution: The specific operation of S503 is expressed by the formula:
[0050]
[0051]
[0052] where i = 6, 5, 4, 3.
[0053] Beneficial effects
[0054] The present invention provides a strip surface defect detection method based on an improved RetinaNet algorithm, which has the following beneficial effects compared with the prior art:
[0055] 1. Detect steel surface defects based on the improved one-stage model RetinaNet. First, the backbone extracts feature maps of different resolutions, and a context feature aggregation module is introduced after the shallow features to obtain more detailed location information such as small defect textures and edges, thereby improving the detection effect of small targets. Secondly, a multi-scale receptive field enhancement module is introduced after the deep feature map to effectively solve the problem of poor detection effect caused by large variations in defect sizes, different shapes, and high background noise on the steel surface. Then, the improved BiFPN effectively connects the deep semantic information and the high-level semantic information across scales, further fusing the context information and improving the detection accuracy. Compared with other mainstream detection network models, the present invention has achieved good detection results in detecting steel surface defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a flowchart of defect detection based on RetinaNet of the present invention.
[0057] Figure 2 It is an overall framework diagram.
[0058] Figure 3 It is a schematic diagram of the multi-scale receptive field enhancement module.
[0059] Figure 4 It is a schematic diagram of the multi-scale feature fusion module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0061] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0062] 1. Data preprocessing stage
[0063] Obtain and analyze the steel surface defect dataset, divide the dataset into a training set, a validation set and a test set according to the ratio of 7:1:2, and before training, resize the pictures to 256×256 pixels and then perform preprocessing operations on the picture data enhancement, including random rotation, translation, cropping, etc.
[0064] Among them, the dataset is from the publicly available dataset NEU-DET of Northeastern University, with six types of collected defect categories, including rolled-in scale (Rs), plaque (Pa), crack (Cr), pitting (Ps), inclusion (In), and scratch (Sc). There are a total of 1800 grayscale images and corresponding annotation files. Each image has 300 samples, and the resolution size is 200*200 for all of them.
[0065] 2. Model training stage
[0066] Step 1: Build a surface defect detection network based on the retinanet model. The surface defect detection network includes a feature extraction network, a feature enhancement network, a feature fusion network, and an optimized loss function. The enhancement network includes a context aggregation module and a multi-scale receptive field module. The backbone resnet50 with strong feature extraction ability is selected as the feature extraction network. Among them, resnet50 has five stages, and the last four stages use a residual structure to retain gradient information. Image downsampling is performed once in each stage, and finally five feature maps C with different resolutions will be obtained i , (i = 1, 2, 3, 4, 5). Select C3, C4, and C5 as effective feature layers and send them into the feature fusion network. Use 1×1 convolution, BN, and LeakyReLU to set the number of channels of the feature maps to 256. The feature maps after dimensionality reduction are denoted as P i , (i = 3, 4, 5).
[0067] Step 2: The present invention proposes a context information aggregation module. First, P4 and P5 are processed by 1×1 convolution, BN, and LeakyReLU and then upsampled by a factor of two and four times. Secondly, use 3×3 depthwise separable convolution and 1×1 convolution to refine the shallow feature P3 and refine the edge information of small defect features. Thirdly, finally use a weighted sum operation, and a learnable parameter w will be multiplied in front of each feature map i , (i = 3, 4, 5), to further obtain a feature map and perform element-wise addition; Thirdly, after the weighted sum, add the CA attention mechanism to weaken the background information and strengthen the fusion between channels; Finally, perform element-wise addition on the enhanced feature map and the shallow feature P3 to obtain P3 with rich small target information (point 1).
[0068] Step 3: A multi-scale receptive field enhancement module is added after the deep feature map P5. This module is divided into four branches in total. The first branch uses asymmetric convolutional kernels 1×3 and 3×1 to extract features, and then uses a 3×3 convolution with a dilation rate of 3; the second branch uses asymmetric convolutional kernels 1×5 and 5×1 to extract features, and then uses a 5×5 convolution with a dilation rate of 5; the third branch uses the operation of performing element-wise addition in parallel on three convolutional kernels 1×3, 3×1, and 3×3, and then uses a 3×3 convolution with a dilation rate of 6; the fourth branch uses global average pooling followed by upsampling. Finally, the outputs of the four branches are concatenated and then dimensionally reduced to obtain the final enhanced feature P5. In the first three branches, the original input feature and the output feature of the previous branch are concatenated as the input of the next branch (point 2).
[0069] Step 4: Use deformable convolution and 3×3 convolution to downsample P5 to generate deeper features P6 and P7. Finally, P i in , (i = 3, 4, 5, 6, 7) are sent as effective feature maps into the feature fusion network for multi-scale feature fusion.
[0070] Step 5: To achieve fast multi-scale feature fusion, the present invention proposes an improved bidirectional weighted fusion network BiFPN (feature fusion network). Among them, five effective feature maps P i in , (i = 3, 4, 5, 6, 7) will be sent into the feature fusion network. The specific operation is as follows:
[0071] Step A: Perform feature-guided upsampling on the input feature , and then perform weighted addition with to generate
[0072] Among them, the specific steps of feature-guided upsampling are as follows: First, use bilinear interpolation to upsample , and then perform element-wise addition with P i in and use a 3*3 convolution to change the number of channels to 1 to obtain the spatial weight; secondly, perform normalization through the softmax function, and multiply the normalized spatial weight w'1 with in the channel dimension to obtain a more detailed semantic feature map (point 3); finally, repeat the above operation, perform feature-guided upsampling on , and perform weighted summation on P i in , to generate P i td . The above process is expressed by the formula as follows:
[0073] FGUpsample = Softmax(f 3 (Concat(Upsample(p i+1 ), p i ))) × p i+1
[0074]
[0075] where i = 6, 5, 4, 3, f 3 (·) represents a convolution operation with a convolution kernel of 3, w1 and w2 are learning weights obtained through fast normalization, ε is a fixed value of 0.0001 to avoid numerical instability during training.
[0076] Step B: Take the result obtained in Step A as the shallow output The output operations for other layers are as follows: Take after downsampling and perform weighted sum fusion with P i td and P i in and use the result as output P i out where i = 4, 5, 6, is obtained by after downsampling and through weighted sum. The above process is expressed by the formula as follows:
[0077]
[0078]
[0079]
[0080] where i = 4, 5, 6, w'1, w'2, w'3 are learning weights obtained through fast normalization, ε is a fixed value of 0.0001 to avoid numerical instability during training.
[0081] Step C: Add a top-down fusion path to further strengthen the fusion of information, use the Ghost module for dimensionality reduction, which reduces the computational cost while avoiding information loss. The specific operation is expressed by the following formula:
[0082]
[0083]
[0084] where i = 6, 5, 4, 3
[0085] Step 6: Replace the L1 loss with the SmoothL1 loss function to optimize the localization loss. The specific formula is as follows:
[0086]
[0087] Step 7: This model uses SGD as the optimizer, with an initial learning rate of 0.01, a momentum set to 1e-4, a batch size equal to 16, a maximum number of training epochs set to 24, and the learning rate is adjusted to 0.1 of the original learning rate at the 18th and 22nd epochs during training.
[0088] Specifically, the model evaluation metrics in the present invention use Average Precision (AP) and mean Average Precision (mAP).
[0089] 3. Model training stage
[0090] The test set samples divided in the data processing stage are sent into the optimal network detection model obtained after training in the model training stage for testing. Finally, the category classification and position regression of the test data are performed through the trained weight parameters to obtain the final detection result.
[0091] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A strip surface defect detection method based on an improved RetinaNet algorithm, characterized in that, Including: Obtain a defect image and load it to form a defect sample, and augment the defect sample by a data augmentation method to increase the dataset samples. The dataset samples include training set samples, validation set samples, and test set samples; Build a surface defect detection network according to the retinanet model. By initializing the parameters, setting hyperparameters, loading the training set samples, and setting the number of iterations for the surface defect detection network in sequence, model training is achieved to obtain an optimal detection model; Import the test set samples into the detection model for testing, and perform class classification and location regression on the test set samples to obtain the final detection result; The surface defect detection network includes a feature extraction network, a feature enhancement network, a feature fusion network, and an optimized loss function. The feature enhancement network includes a context aggregation module and a multi-scale receptive field module. The specific detection method of the surface defect detection network is as follows: S1. Obtain the effective feature map , where is the shallow feature map, is the middle and high-level feature map, is the deep feature map. The specific operation steps are as follows: Select resnet50 with 5 extraction stages as the feature extraction network of the model. Each stage of resnet50 generates feature maps with different resolution sizes through convolutional strides. Select the feature maps of the last three stages as the effective feature map to obtain the effective feature map ; S2. Introduce a context aggregation module and process through the context aggregation module to obtain rich in small target information; S3. Introduce a multi-scale receptive field module with four branches and perform feature enhancement through the multi-scale receptive field module, where in the first three branches of the multi-scale receptive field module, the original feature input feature and the output feature of the previous branch are concatenated as the input of the next branch; S4. Use deformable convolution and 3×3 convolution to perform downsampling on the enhanced features and generate deeper features and , and send as the effective feature map into the feature fusion network for multi-scale feature fusion; S5. Use the bidirectional weighted fusion network BiFPN, that is, the feature fusion network, to perform multi-scale feature fusion on the five effective feature maps ; S6. Optimize the localization loss. Replace the L1 loss in the retinanet model with the SmoothL1 loss function to optimize the localization loss. The expression of the SmoothL1 loss function is: Specifically, use SGD as the model optimizer to optimize the localization loss.
2. The strip surface defect detection method based on the improved RetinaNet algorithm according to claim 1, characterized in that The specific operation steps of S2 are: S201. Take and After processing through 1×1 convolution, BN, and LeakyReLU, perform upsampling by a factor of two and four. S202. Refine the shallow features using 3×3 depthwise separable convolution and 1×1 convolution , and refine the edge information of small defect features; S203. Use a weighted sum operation to multiply each feature map by a learned parameter , and further obtain a feature map and perform element-wise addition; S204. After weighted summation, add the CA attention mechanism to weaken the background information and strengthen the fusion between channels; S205. Add the feature map obtained in S204 and the shallow feature map element by element to obtain a feature map rich in small target information . 3. The strip surface defect detection method based on the improved retinanet algorithm according to claim 1, characterized in that, The steps for the multi-scale receptive field module to perform feature enhancement are as follows: S301. The first branch of the multi-scale receptive field enhancement module uses asymmetric convolutional kernels 1×3 and 3×1 to extract features, and then uses a 3×3 convolution with a dilation rate of 3 to convolve the features and output; S302. The second branch of the multi-scale receptive field enhancement module uses asymmetric convolutional kernels 1×5 and 5×1 to extract features, and then uses a 5×5 convolution with a dilation rate of 5 to convolve the features and output; S303. The operation of the third branch of the multi-scale receptive field enhancement module is to perform parallel element-wise addition of the three convolutional kernels 1×3, 3×1, and 3×3, and then use a 3×3 convolution with a dilation rate of 6 to convolve the features and output; S304. The fourth branch of the multi-scale receptive field enhancement module uses global average pooling and then upsampling and outputs; S305. After splicing the outputs of the four branches and performing dimensionality reduction to obtain the final feature enhancement .
4. The strip surface defect detection method based on the improved retinanet algorithm according to claim 1, characterized in that The specific operation steps of S5 are: S501. Perform a feature-guided upsampling operation on the input feature , and then perform weighted addition with to generate . Repeat the feature-guided upsampling operation until is generated; S502. Take the as the shallow output . The output operations for other layers are as follows: Take the after downsampling operation and perform weighted sum fusion with and and take the result as the output , where , is obtained by taking the after downsampling and performing weighted sum with ; S503. Add a top-down fusion path to further strengthen the fusion between information, and use the Ghost module for dimensionality reduction.
5. The strip surface defect detection method based on the improved RetinaNet algorithm according to claim 4, characterized in that, The specific steps of feature-guided upsampling in S501 are: S5011. Use bilinear interpolation to perform upsampling, and then add it element-wise to perform an element-wise addition operation and use a 3*3 convolution to change the number of channels to 1 to obtain the spatial weights; S5012. Normalize the spatial weights obtained in S5011 through the softmax function and multiply the normalized spatial weights with in the channel dimension to obtain a feature map with more detailed semantics , where ; Repeat the steps of S5011 and S5012, and perform feature-guided upsampling operation, and , generate after weighted sum, where .
6. The strip surface defect detection method based on the improved RetinaNet algorithm according to claim 4, characterized in that, The specific process of S501 is expressed by the formula: Among them, , represents a convolution operation with a convolution kernel of 3, , are learning weights obtained through fast normalization, is a fixed value and is 0.0001.
7. The strip surface defect detection method based on the improved RetinaNet algorithm according to claim 4, characterized in that, The process of S502 is expressed by the formula: Among them, , , , are learning weights obtained through fast normalization, is a fixed value of 0.0001.
8. The strip surface defect detection method based on the improved RetinaNet algorithm according to claim 4, characterized in that, The specific operation of S503 is expressed by the formula: Among them, .
Citation Information
Patent Citations
Strip steel surface defect detection method based on improved efficientNet-RCNN
CN112991267A