Steel surface defect detection method based on improved Faster R-CNN algorithm
By performing pseudo-color conversion and feature fusion in the backbone network of the Faster R-CNN model, the problem of poor effectiveness in detecting subtle defects is solved, and the accuracy of defect detection is significantly improved.
Patent Information
- Application Number
- CN202510714831.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-30
AI Technical Summary
When traditional Faster R-CNN models detect defects such as fine cracks, oxidative pressing, etc. in industrial grayscale images, it is difficult to extract effective texture features, resulting in poor detection results.
The Faster R-CNN algorithm is improved, and the grayscale image is converted into pseudo-color in the backbone network, and the features of the grayscale image and pseudo-color image are extracted respectively, and the features are fused through the attention mechanism to generate the final feature map.
It significantly improves the accuracy of steel surface defect detection, enhances the ability to extract defect characteristics, and is suitable for high-precision industrial quality inspection scenarios.
Smart Images

Figure CN120219397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial defect detection, and specifically provides a steel surface defect detection method based on an improved Faster R-CNN algorithm, which can significantly improve the detection accuracy of weak defects such as cracks, oxide inclusions, and spots, and is applicable to high-precision industrial quality inspection scenarios. Background Art
[0002] Machine vision technology is generally used in the field of industrial quality inspection for product surface defect detection. For example, currently, traditional image processing algorithms such as edge detection and threshold segmentation, as well as detection models based on deep learning such as Faster R-CNN (Faster Region-based Convolutional Neural Networks) and YOLO (You Only Look Once), are usually used to detect defects in automotive steel components. Among them, the Faster R-CNN model realizes defect localization and classification through feature extraction and region proposal network, and is widely applied. However, in industrial grayscale images, defects such as fine cracks and oxide inclusions have low contrast with the background, and the defect features are weak, which makes it difficult for the traditional Faster R-CNN model to extract effective texture features and results in poor detection effects. Summary of the Invention
[0003] To solve the above problems, the present invention provides a steel surface defect detection method based on an improved Faster R-CNN algorithm, which uses a dual-branch feature extraction backbone network to extract features. The backbone network performs pseudo-color conversion on the grayscale image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to enhance the defect features.
[0004] To achieve the above object, the solution provided by the present invention is as follows: A steel surface defect detection method based on an improved Faster R-CNN algorithm, comprising the following steps: S1. Collect surface images of various steel products with surface defects and label various defects; S2. Construct an improved Faster R-CNN network, where the improved Faster R-CNN network includes a backbone network, a region proposal network (RPN), a region of interest pooling layer (RoIPooling), and a classification and regression head. The backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate a final feature map; S3. Use the surface image of the steel product collected in step S1 as a grayscale image to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model; S4. Collect the surface image of the steel product to be predicted and input it into the steel product surface defect recognition model to identify the defects in the surface image of the steel product to be predicted.
[0005] As a specific implementation manner of the present invention, in the backbone network, first convert the grayscale image into a pseudo-color image and calculate the texture gradient of the grayscale image, then perform weighted fusion of each channel of the pseudo-color image with the texture gradient to generate a pseudo-color branch input map, use the Resnet50 network to convert the pseudo-color branch input map and the grayscale image into a pseudo-color branch feature map and a grayscale branch feature map respectively, then use channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map, and finally use spatial-level attention to process the feature map fused based on channel-level attention to obtain a final feature map.
[0006] In the present invention, the conversion of the grayscale image into a pseudo-color image can adopt the prior art. For example, the existing Jet color mapping algorithm is used to map the grayscale image into a pseudo-color image.
[0007] As a specific implementation manner of the present invention, calculating the texture gradient of the grayscale image includes the following steps: S21. Normalize the grayscale image to generate a single-channel grayscale image; S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image, and the calculation formula is as follows: ; ; In the formula, TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single-channel grayscale image; I ( x + i , y + j ) is the pixel value of the pixel point ( x , y ) in the neighborhood of the pixel point ( x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean value of the neighborhood of the pixel point ( x , y );k is the radius of the window; S23. Normalize the texture gradient using the maximum - minimum normalization method.
[0008] In the backbone network of the present invention, channel - level attention is used to fuse the color - branch feature map and the grayscale - branch feature map into one feature map, which specifically includes the following steps: S311. Respectively perform global average pooling on the color - branch feature map F p and the grayscale - branch feature map F g to extract channel - level statistical features; ; ; GAP ( F g ) is the value after global average pooling of the grayscale - branch feature map F g ; GAP ( F p ) is the value after global average pooling of the color - branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function; S312. Generate channel weights through a multi - layer perceptron (MLP) to dynamically adjust the contributions of the two - branch features; Sigmoid () is Sigmoid function; ReLU () is ReLU () function; W 1. W 2 are multi - layer perceptron parameters; GAP ( F g ); GAP ( F p )] represents GAP ( F g ) and GAP ( F p ) concatenated; S313. Feature fusion of the two - branch features based on channel - level attention; ; F fusionIs the feature based on channel - level attention fusion; Represents channel - by - channel multiplication; In the backbone network of the present invention, the spatial - level attention is used to process the feature map based on channel - level attention fusion, which specifically includes the following steps: S321. By performing different linear projections on the feature based on channel - level attention fusion F fusion to obtain Q (Query), K (Key) and V (Value): ; ; ; In the formula, W q , W k , W v are fully - connected layers adapted to the dimensions of F g and F p The purpose is to linearly project the query feature Q and the image feature K , V onto the same dimension respectively; S322. Calculate the similarity between Q and K (usually using dot - product or cosine similarity) to generate the attention weight matrix A;
[0009] S323. Use A to perform weighted summation on V to obtain the output feature: Beneficial effects: The present invention improves the backbone network of the Faster R - CNN model. The backbone network performs pseudo - color conversion on the grayscale image, then extracts the features of the grayscale image and the pseudo - color image respectively, and performs feature fusion through the attention mechanism to achieve the purpose of enhancing defect features, significantly improving the accuracy of steel surface defect detection. Brief Description of the Drawings
[0010] Figure 1 is a schematic flowchart of improving the Faster R - CNN model in an embodiment of the present invention; Figure 2 is Figure 1 a schematic flowchart of the backbone network in the improved Faster R - CNN model. Detailed Embodiments
[0011] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0012] A method for detecting steel surface defects based on an improved Faster R-CNN algorithm includes the following steps: S1. Collect surface images of various steel products with surface defects; Specifically, in this embodiment, the object detection dataset (NEU-DET dataset) released by Northeastern University is used. This dataset has collected 6 types of strip steel surface defect images, with 300 images in each type, for a total of 1800 images. The defect categories include crazing, Inclusion, Patches, Pitted Surface, Rolled-in Scale, and Scratches.
[0013] The original size of the images in the dataset is 200×200. To make the model adapt to different datasets, before feeding the images into the model, the shape of the images is adjusted to 600×600 uniformly. Assuming the original image size is W old ×H old and the new image size is W new ×H new For the coordinates (xmin, ymin, xmax, ymax) of each bounding box in the label file, they are adjusted according to the following formula: ; ; ; ; S2. Construct an improved Faster R-CNN network. As Figure 1 shown, the improved Faster R-CNN network in this embodiment includes: An improved backbone network, also known as a feature extraction network, is used to extract features of the product surface image to generate a feature map; the backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate a final feature map; A Region Proposal Network (RPN) is used to generate candidate regions on the feature map; A Region of Interest (RoI) Pooling layer maps the candidate regions of different sizes to the feature map and pools them into features of a fixed size; The classifier and regressor head performs fine classification and bounding box regression on the features processed by the region of interest pooling layer; In this embodiment, the region proposal network, the region of interest pooling layer, and the classifier and regressor head are all prior arts and will not be elaborated here. The backbone network will be described in detail below in conjunction with Figure 2 the detailed description of the backbone network.
[0014] In the backbone network of this embodiment, first, the grayscale image is converted into a pseudo-color image and the texture gradient of the grayscale image is calculated. Then, each channel of the pseudo-color image is weighted and fused with the texture gradient to generate a pseudo-color branch input image. The Resnet50 network is used to convert the pseudo-color branch input image into a pseudo-color branch feature map (high-dimensional feature map); the Resnet50 network is used to convert the grayscale image into a grayscale branch feature map (high-dimensional feature map); then, channel-level attention is used to fuse the color branch feature map and the grayscale branch feature map into a feature map, and finally, spatial-level attention is used to process the feature map fused based on channel-level attention to obtain the final feature map. Specifically, it includes the following steps: First, the Jet color mapping algorithm is used to map the grayscale image into a pseudo-color image, including normalizing the grayscale image. For a given normalization scalar value, its components on the R, G, and B channels are calculated respectively, and then the components are combined into an RGB image, which is the pseudo-color image.
[0015] Secondly, calculate the texture gradient of the grayscale image converted, including the following steps: S21. Generate a single-channel grayscale image after normalizing the grayscale image; S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image, and the calculation formula is as follows: ; ; In the formula, TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single-channel grayscale image; I ( x + i , y + j ) is the pixel value of the pixel point ( x , y ) in the neighborhood of the pixel point ( x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean of the neighborhood of the pixel point ( x , y );k is the radius of the window; S23. The texture gradient is normalized using the maximum - minimum normalization method. The specific normalization formula is as follows: ; In the formula, TGM ( x , y ) is the normalized value of the texture gradient of the pixel point ( x , y ) in the single - channel grayscale image; TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single - channel grayscale image; min( TGM ) is the minimum texture gradient; max( TGM ) is the maximum texture gradient; Next, each channel of the pseudo - color image is weighted and fused with the texture gradient map according to the same weight to obtain the pseudo - color branch input image P’ ; ; ; ; ; In the formula, P’ is the pseudo - color branch input image, P’ is; P ’ R , P ’ G , P ’ B are respectively the R , G , B channel values of the pseudo - color branch input image; P R , P G , P B are respectively the R , G , B channel values of the pseudo - color image; TGM is the normalized pixel value in the texture gradient map.
[0016] Next, the Resnet50 network is used to convert the pseudo - color input image P’ into the pseudo - color branch feature mapF p (High-dimensional feature map), convert the grayscale image I to a grayscale branch feature map F g (High-dimensional feature map); ; ; Resnet() is the Resnet50 network; S31. Use channel-level attention to perform feature fusion on the color branch feature map and the grayscale branch feature map; S311. Respectively perform global average pooling on the color branch feature map F p and the grayscale branch feature map F g to extract channel-level statistical features; ; ; GAP ( F g ) is the value after global average pooling of the grayscale branch feature map F g ; GAP ( F p ) is the value after global average pooling of the color branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function; S312. Generate channel weights through a multi-layer perceptron (MLP) , dynamically adjusting the contributions of the two-branch features; ; Sigmoid () is Sigmoid function; ReLU () is ReLU () function; W 1. W 2 is the multi-layer perceptron parameter; GAP ( F g ); GAP ( F p )] represents GAP ( F g ) and GAP ( F p ) concatenation; S313. Feature fusion of two branches based on channel-level attention; ; F fusion is the feature based on channel-level attention fusion; represents channel-wise multiplication; S32. Process the feature map based on channel-level attention fusion using spatial-level attention; S321. Through different linear projections on F fusion to obtain Q (Query), K (Key), and V (Value): ; ; ; In the formula, W q , W k , W v are fully connected layers adapted to the dimensions of F g and F p respectively, aiming to linearly project the query feature Q and the image feature K , V onto the same dimension respectively; S322. Calculate the similarity between Q and K (usually using dot product or cosine similarity) to generate the attention weight matrix A ; S323. Use A to weighted sum V to obtain the output feature: ; S3. Use the surface image of the steel product collected in step S1 to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model; S4. Collect the surface image of the steel product to be predicted and input it into the steel product surface defect recognition model to identify the defects in the surface image of the steel product to be predicted.
[0017] In this embodiment, the Adam optimizer is used for continuation, with an initial learning rate of 1e-4 and a batch size set to 4.
[0018] In addition, in order to verify the impact of improving the backbone network on the model recognition performance, this embodiment also sets up a Faster R-CNN model with a Resnet50 network as the backbone network, and uses the same images for training and testing. The results are shown in Table 1.
[0019] Table 1 Average accuracy of different models
[0020] It can be seen from Table 1 that the average accuracy of the model is significantly improved after using the improved backbone network.
[0021] The above are only preferred specific implementations of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed in the embodiments of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for detecting steel surface defects based on an improved Faster R-CNN algorithm, characterized in that, It includes the following steps: S1. Collect surface images of various steel products with surface defects and label various defects; S2. Construct an improved Faster R-CNN network, which includes a backbone network, a region proposal network, a region of interest pooling layer, and a classification and regression head. The backbone network performs pseudo-color conversion on the grayscale image to generate a pseudo-color image, then extracts the features of the grayscale image and the pseudo-color image respectively, and performs feature fusion through an attention mechanism to generate a final feature map; S3. Use the surface image of the steel product as a grayscale image to train the improved Faster R-CNN network to obtain a steel product surface defect recognition model; S4. Collect the surface image of the steel product to be predicted and input it into the steel product surface defect recognition model to identify the defects in the surface image of the steel product to be predicted.
2. The steel surface defect detection method based on the improved Faster R-CNN algorithm according to claim 1, wherein In the backbone network, first convert the grayscale image into a pseudo-color image and calculate the texture gradient of the grayscale image. Then, perform weighted fusion of each channel of the pseudo-color image with the texture gradient to generate a pseudo-color branch input map. Use the Resnet50 network to convert the pseudo-color branch input map and the grayscale image into a pseudo-color branch feature map and a grayscale branch feature map respectively. Then, use channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map. Finally, use spatial-level attention to process the feature map fused based on channel-level attention to obtain the final feature map.
3. A steel surface defect detection method based on an improved Faster R-CNN algorithm according to claim 2, characterized in that: Calculating the texture gradient of the grayscale image includes the following steps: S21. Normalize the grayscale image to generate a single-channel grayscale image; S22. Calculate the texture gradient of each pixel point in the single-channel grayscale image, and the calculation formula is as follows: ; ; In the formula, TGM ( x , y ) is the texture gradient of the pixel point ( x , y ) in the single-channel grayscale image; I ( x + i , y + j ) is the pixel value of the pixel point ( x , y ) in the neighborhood of the pixel point ( x + i , y + j ) in the single-channel grayscale image; µ x,y is the local mean of the neighborhood of the pixel point ( x , y ); k is the radius of the window; S23. Normalize the texture gradient by using the maximum-minimum normalization method.
4. The steel surface defect detection method based on the improved Faster R-CNN algorithm according to claim 2, characterized in that: In the backbone network, using channel-level attention to fuse the color branch feature map and the grayscale branch feature map into a feature map specifically includes the following steps: S311. Perform global average pooling on the color branch feature map and the grayscale branch feature map respectively to extract channel-level statistical features; ; ; GAP ( F g ) is the value after global average pooling of the grayscale branch feature map F g ; GAP ( F p ) is the value after global average pooling of the color branch feature map F p ; adaptive_avg_pool2d () is the global average pooling function; S312. Generate channel weights through a multi-layer perceptron (MLP) , and dynamically adjust the contributions of the features of the two branches; ; Sigmoid () is Sigmoid a function; ReLU () is ReLU a () function; W 1. W 2 is the parameter of the multi-layer perceptron; GAP ( F g ) GAP ( F p )] represents GAP ( F g ) concatenated with GAP ( F p ) S313. Feature fusion of the two branches based on channel-level attention; ; F fusion is the feature based on channel-level attention fusion; represents channel-wise multiplication.
5. The steel surface defect detection method based on the improved Faster R-CNN algorithm according to claim 4, characterized in that: In the backbone network, using spatial-level attention to process the feature map fused based on channel-level attention specifically includes the following steps: S321. By performing different linear projections on the features based on channel-level attention fusion, obtain Q (Query), K (Key), and V (Value): F fusion ; ; ; In the formula, W q , W k , W v are fully connected layers adapted to the F g and F p dimensions, aiming to linearly project the query feature Q and the image feature K , V onto the same dimension respectively; S322. Calculate Q and K to generate an attention weight matrix A; S323. Use the attention weight matrix A to V perform weighted summation to obtain the output feature.
Citation Information
Patent Citations
Welding seam defect identification method based on improved LeNet-5 model
CN108596892A
Expressway agglomerate fog early warning system based on deep fusion network
CN112419745A
CT image liver automatic segmentation method based on deep convolutional neural network
CN113470044A
Defect detection method and device based on improved Faster-RCNN
CN116091496A
Deep forgery detection method and device based on image diversification features
CN116311430A