A Steel Surface Defect Detection Method Based on Improved Faster R-CNN Algorithm
By improving the Faster R-CNN algorithm, converting grayscale images into pseudo-color images and performing feature fusion, the problem of poor detection of subtle defects is solved, and a higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202510714831.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The traditional Faster R-CNN model is difficult to effectively extract the texture characteristics of defects such as fine cracks and oxidative pressing in industrial grayscale images, resulting in poor detection results.
Using the improved Faster R-CNN algorithm, defect feature extraction is enhanced by converting grayscale images into pseudo-color images and using attention mechanisms to perform feature fusion.
It significantly improves the accuracy of steel surface defect detection and improves the detection effect.
Smart Images

Figure CN120219397B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial defect detection, and particularly to a steel surface defect detection method based on an improved Faster R-CNN algorithm, which can significantly improve the detection accuracy of weak defects such as cracks, oxidation indentations, and spots, and is applicable to high-precision industrial quality inspection scenarios. Background Art
[0002] Machine vision technology is generally used in the field of industrial quality inspection for product surface defect detection. For example, currently, traditional image processing algorithms such as edge detection and threshold segmentation, as well as detection models based on deep learning such as Faster R-CNN (Faster Region-based Convolutional Neural Networks) and YOLO (You Only Look Once), are usually used to detect defects in automotive steel components. Among them, the Faster R-CNN model realizes defect localization and classification through feature extraction and region proposal network, and is widely applied. However, in industrial grayscale images, defects such as fine cracks and oxidation indentations have low contrast with the background, and the defect features are weak, which makes it difficult for the traditional Faster R-CNN model to extract effective texture features and results in poor detection effects. Summary of the Invention
[0003] To solve the above problems, the present invention provides a steel surface defect detection method based on an improved Faster R-CNN algorithm, which uses a dual-branch feature extraction backbone network to extract features. This backbone network performs pseudo-color conversion on the grayscale image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to achieve the purpose of enhancing defect features.
[0004] To achieve the above purpose, the solution provided by the present invention is as follows:
[0005] A steel surface defect detection method based on an improved Faster R-CNN algorithm, comprising the following steps:
[0006] S1. Collect surface images of various steel products with surface defects and label various defects;
[0007] S2. Construct an improved Faster R-CNN network, where the improved Faster R-CNN network includes a backbone network, a Region Proposal Network (RPN), a Region of Interest Pooling layer (RoIPooling), and a classifier. The backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate a final feature map;
[0008] S3. Use the steel product surface image collected in step S1 as a grayscale image to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model;
[0009] S4. Collect the steel product surface image to be predicted and input it into the steel product surface defect recognition model to identify the defects in the steel product surface image to be predicted.
[0010] As a specific implementation manner of the present invention, in the backbone network, first convert the grayscale image into a pseudo-color image and calculate the texture gradient of the grayscale image. Then, perform weighted fusion of each channel of the pseudo-color image with the texture gradient to generate a pseudo-color branch input map. Use the Resnet50 network to convert the pseudo-color branch input map and the grayscale image into a pseudo-color branch feature map and a grayscale branch feature map respectively. Then, use channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map. Finally, use spatial-level attention to process the feature map fused based on channel-level attention to obtain the final feature map.
[0011] In the present invention, the conversion of the grayscale image into a pseudo-color image can adopt the existing technology. For example, use the existing Jet color mapping algorithm to map the grayscale image into a pseudo-color image.
[0012] As a specific implementation manner of the present invention, calculating the texture gradient of the grayscale image includes the following steps:
[0013] S21. Normalize the grayscale image to generate a single-channel grayscale image;
[0014] S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image, and the calculation formula is as follows:
[0015] ;
[0016] ;
[0017] In the formula, TGM (x , y ) is the texture gradient of the pixel point ([ x , y ) in the single-channel grayscale image; I ([ x + i , y + j ) is the pixel value of the pixel point ([ x , y ) in the neighborhood of the pixel point ([ x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean of the neighborhood of the pixel point ([ x , y ); k is the radius of the window;
[0018] S23. Normalize the texture gradient using the maximum-minimum normalization method.
[0019] In the backbone network of the present invention, channel-level attention is used to fuse the color branch feature map and the grayscale branch feature map into a feature map, which specifically includes the following steps:
[0020] S311. Perform global average pooling on the color branch feature map F p and the grayscale branch feature map F g respectively to extract channel-level statistical features;
[0021] ;
[0022] ;
[0023] GAP ([ F g ) is the value after global average pooling of the grayscale branch feature map F g ; GAP ([ F p ) is the value after global average pooling of the color branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function;
[0024] S312. Generate channel weights through a multi-layer perceptron (MLP) to dynamically adjust the contributions of the two branch features;
[0025] ;
[0026] Sigmoid () is Sigmoid a function; ReLU () is ReLU the () function; W 1. W 2 is the parameter of the multi-layer perceptron; GAP ( F g ); GAP ( F p )] represents GAP ( F g ) and GAP ( F p ) concatenation;
[0027] S313. Feature fusion of two branches of features based on channel-level attention;
[0028] ;
[0029] F fusion is the feature fused based on channel-level attention; represents element-wise multiplication;
[0030] In the backbone network of the present invention, spatial-level attention is used to process the feature map fused based on channel-level attention, which specifically includes the following steps:
[0031] S321. By performing different linear projections on the feature F fusion fused based on channel-level attention, Q (Query), K (Key), and V (Value) are obtained:
[0032] ;
[0033] ;
[0034] ;
[0035] In the formula, W q , W k , W v are fully connected layers adapted to the dimensions of F g and F p The purpose is to transform the query feature Q and the image feature K , VLinearly project them respectively onto the same dimension;
[0036] S322. Calculate Q and K similarity (usually using dot product or cosine similarity) to generate an attention weight matrix A;
[0037] S323. Use A to V perform weighted summation to obtain the output feature:
[0038] Advantageous effects: The present invention improves the backbone network of the Faster R-CNN model. The backbone network performs pseudo-color conversion on grayscale images, then extracts the features of grayscale images and pseudo-color images respectively, and performs feature fusion through the attention mechanism to achieve the purpose of enhancing defect features, significantly improving the accuracy of steel surface defect detection. Description of the Drawings
[0039] Figure 1 is a schematic flow chart of improving the Faster R-CNN model in the embodiment of the present invention;
[0040] Figure 2 is Figure 1 a schematic flow chart of the backbone network in the improved Faster R-CNN model. Detailed Embodiments
[0041] The following combines embodiments and the drawings to further elaborate on the present invention in detail, but the implementation manners of the present invention are not limited thereto.
[0042] A steel surface defect detection method based on an improved Faster R-CNN algorithm includes the following steps:
[0043] S1. Collect surface images of various steel products with surface defects;
[0044] Specifically, in this embodiment, the object detection dataset (NEU-DET dataset) released by Northeastern University is used. This dataset has collected 6 types of strip steel surface defect images, with 300 images in each type, totaling 1800 images. The defect categories include crazing, Inclusion, Patches, Pitted Surface, Rolled-in Scale, and Scratches.
[0045] The original size of the images in the dataset is 200×200. To make the model adapt to different datasets, before feeding the images into the model, the shape of the images is adjusted to 600×600 uniformly. Assume the original image size is W old×H old The new image size is W new ×H new For the coordinates (xmin, ymin, xmax, ymax) of each bounding box in the label file, they are adjusted according to the following formula:
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] S2. Construct an improved Faster R-CNN network. As shown in Figure 1 The improved Faster R-CNN network of this embodiment includes:
[0051] An improved backbone network, also known as a feature extraction network, which is used to extract the features of the product surface image to generate a feature map; the backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate the final feature map;
[0052] A Region Proposal Network (RPN), which is used to generate candidate regions on the feature map;
[0053] A Region of Interest (RoI) Pooling layer, which maps the candidate regions of different sizes to the feature map and pools them into features of a fixed size;
[0054] A Classifier, which performs fine classification and bounding box regression on the features processed by the RoI Pooling layer;
[0055] In this embodiment, the Region Proposal Network, the RoI Pooling layer, and the Classifier are all prior arts and will not be elaborated here. Below will be combined with Figure 2 to elaborate on the backbone network in detail.
[0056] In the backbone network of this embodiment, first, the grayscale image is converted into a pseudo-color image and the texture gradient of the grayscale image is calculated. Then, each channel of the pseudo-color image is weighted and fused with the texture gradient to generate a pseudo-color branch input image. The Resnet50 network is used to convert the pseudo-color branch input image into a pseudo-color branch feature map (high-dimensional feature map); the Resnet50 network is used to convert the grayscale image into a grayscale branch feature map (high-dimensional feature map); then, channel-level attention is used to fuse the color branch feature map and the grayscale branch feature map into a feature map, and finally, spatial-level attention is used to process the feature map fused based on channel-level attention to obtain the final feature map. Specifically, it includes the following steps:
[0057] First, the Jet color mapping algorithm is used to map the grayscale image into a pseudo-color image, including normalizing the grayscale image. For a given normalized scalar value, its components on the R, G, and B channels are calculated respectively, and then the components are combined into an RGB image, which is the pseudo-color image.
[0058] Secondly, calculate the texture gradient of the converted grayscale image, including the following steps:
[0059] S21. Generate a single-channel grayscale image after normalizing the grayscale image;
[0060] S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image. The calculation formula is as follows:
[0061] ;
[0062] ;
[0063] In the formula, TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single-channel grayscale image; I ( x + i , y + j ) is the pixel value of the pixel point ( x , y ) in the neighborhood of the pixel point ( x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean value of the neighborhood of the pixel point ( x , y ); k is the radius of the window;
[0064] S23. Normalize the texture gradient using the maximum - minimum normalization method. The specific normalization formula is as follows:
[0065] ;
[0066] In the formula, TGM ( x , y ) is the normalized value of the texture gradient of the pixel point ( x , y ) in the single - channel grayscale image; TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single - channel grayscale image; min( TGM ) is the minimum texture gradient; max( TGM ) is the maximum texture gradient;
[0067] Next, each channel of the pseudo - color image is weighted and fused with the texture gradient map according to the same weight to obtain the pseudo - color branch input image P’ ;
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] In the formula, P’ is the pseudo - color branch input image, P’ is; P ’ R 、 P ’ G 、 P ’ B are the R 、 G 、 B channel values of the pseudo - color branch input image respectively; P R 、 P G 、 P B are the R 、 G 、 B channel values of the pseudo - color image respectively; TGM is the normalized pixel value in the texture gradient map.
[0073] Next, use the Resnet50 network to convert the pseudo-color input image P’ into a pseudo-color branch feature map F p (high-dimensional feature map), and convert the grayscale image I into a grayscale branch feature map F g (high-dimensional feature map);
[0074] ;
[0075] ;
[0076] Resnet() is the Resnet50 network;
[0077] S31. Use channel-level attention to perform feature fusion on the color branch feature map and the grayscale branch feature map;
[0078] S311. Respectively perform global average pooling on the color branch feature map F p and the grayscale branch feature map F g to extract channel-level statistical features;
[0079] ;
[0080] ;
[0081] GAP ( F g ) is the value after global average pooling of the grayscale branch feature map F g ; GAP ( F p ) is the value after global average pooling of the color branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function;
[0082] S312. Generate channel weights through a multi-layer perceptron (MLP) to dynamically adjust the contributions of the two-branch features;
[0083] ;
[0084] Sigmoid () is Sigmoid function; ReLU () is ReLU () function; W1. W 2 is the parameter of the multi - layer perceptron; GAP ( F g ); GAP ( F p )] represents GAP ( F g ) and GAP ( F p ) are concatenated;
[0085] S313. Feature fusion of two - branch features based on channel - level attention;
[0086] ;
[0087] F fusion is the feature fused based on channel - level attention; represents element - wise multiplication across channels;
[0088] S32. Process the feature map fused based on channel - level attention using spatial - level attention;
[0089] S321. By performing different linear projections on F fusion , obtain Q (Query), K (Key), and V (Value):
[0090] ;
[0091] ;
[0092] ;
[0093] In the formula, W q , W k , W v are fully - connected layers adapted to the dimensions of F g and F p The purpose is to linearly project the query feature Q and the image feature K , V onto the same dimension respectively;
[0094] S322. Calculate the similarity between Q and K (usually using dot - product or cosine similarity) to generate the attention weight matrix A ;
[0095] S323. Use A to V perform weighted summation to obtain an output feature:
[0096] ;
[0097] S3. Use the surface image of the steel product collected in step S1 to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model;
[0098] S4. Collect the surface image of the steel product to be predicted and input it into the steel product surface defect recognition model to identify the defects in the surface image of the steel product to be predicted.
[0099] In this embodiment, the Adam optimizer is used for continuation connection, the initial learning rate is 1e-4, and the batch size is set to 4.
[0100] In addition, in order to verify the influence of the improved backbone network on the model recognition performance, this embodiment also sets up a Faster R-CNN model with the Resnet50 network as the backbone network, and uses the same images for training and testing. The results are shown in Table 1.
[0101] Table 1 Average precision table of different models
[0102]
[0103] As can be seen from Table 1, the average precision of the model is significantly improved after using the improved backbone network.
[0104] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be covered by the protection scope of the present invention.
Claims
1. A steel surface defect detection method based on an improved Faster R-CNN algorithm, characterized in that, It includes the following steps: S1. Collect surface images of various steel products with surface defects and label various defects; S2. Construct an improved Faster R-CNN network, which includes a backbone network, a region proposal network, a region of interest pooling layer, and a classification and regression head. The backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate a final feature map; In the backbone network, first convert the grayscale image into a pseudo-color image and calculate the texture gradient of the grayscale image. Then, perform weighted fusion of each channel of the pseudo-color image with the texture gradient to generate a pseudo-color branch input map. Respectively use the Resnet50 network to convert the pseudo-color branch input map and the grayscale image into a pseudo-color branch feature map and a grayscale branch feature map. Then use channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map. Finally, use spatial-level attention to process the feature map fused based on channel-level attention to obtain the final feature map; S4. Use the surface image of the steel product as a grayscale image to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model; S5. Collect the surface image of the steel product to be predicted and input it into the steel product surface defect recognition model to identify the defects in the surface image of the steel product to be predicted.
2. A steel surface defect detection method based on an improved Faster R-CNN algorithm according to claim 1, characterized in that: Calculating the texture gradient of the grayscale image includes the following steps: S21. Normalize the grayscale image to generate a single-channel grayscale image; S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image. The calculation formula is as follows: ; ; Wherein, TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single-channel grayscale image; I ( x + i , y + j ) is the pixel value of the pixel point ( x , y ) in the neighborhood of the pixel point ( x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean of the neighborhood of the pixel point ( x , y ); k is the radius of the window; S23. Normalize the texture gradient by using the maximum-minimum normalization method.
3. A steel surface defect detection method based on an improved Faster R-CNN algorithm according to claim 1, characterized in that: In the backbone network, using channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map specifically includes the following steps: S311. Perform global average pooling on the color branch feature map and the grayscale branch feature map respectively to extract channel-level statistical features; ; ; GAP ( F g ) is the value after global average pooling of the grayscale branch feature map F g ; GAP ( F p ) is the value after global average pooling of the color branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function; S312. Generate channel weights through a multi-layer perceptron (MLP) , and dynamically adjust the contributions of the features of the two branches; ; Sigmoid () is Sigmoid a function; ReLU () is ReLU a () function; W 1. W 2 is a parameter of the multi-layer perceptron; GAP ( F g ) GAP ( F p )] represents GAP ( F g ) and GAP ( F p ) concatenation; S313. Feature fusion of the two branches based on channel-level attention; ; F fusion is the feature based on channel - level attention fusion; represents element - wise multiplication.
4. A steel surface defect detection method based on an improved Faster R-CNN algorithm according to claim 3, characterized in that: In the backbone network, using spatial-level attention to process the feature map fused based on channel-level attention specifically includes the following steps: S321. By performing different linear projections on the features fused based on channel-level attention, F fusion obtain Q (Query), K (Key), and V (Value): ; ; ; In the formula, W q , W k , W v are fully-connected layers adapted to the F g and F p dimensions, aiming to linearly project the query feature Q and the image feature K , V onto the same dimension respectively; S322. Calculate Q the similarity with K to generate an attention weight matrix A ; S323. Use the attention weight matrix A to V perform weighted summation to obtain the output feature.
Citation Information
Patent Citations
Defect detection method and device based on improved Faster-RCNN
CN116091496A