Steel surface defect detection method based on YOLOv8 multi-scale convolution attention mechanism
By introducing a multi-scale convolutional attention mechanism and deformable convolution module in the YOLOv8 model, the steel surface defect detection model is optimized, and the speed problem of existing detection methods in environments with limited computing resources is solved, and efficient and accurate defect detection is achieved.
Patent Information
- Application Number
- CN202510003391.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
The existing steel surface defect detection methods are slow in industrial environments with limited computing resources, and large neural networks require high computing capabilities, making it difficult to effectively deploy on mobile or embedded devices.
Using the multi-scale convolution attention mechanism based on YOLOv8, the model is optimized to improve the detection capability of different size targets by introducing the MSCA module and the new deformable convolution C2f_DCNv2 module, and the Wise-IoU loss function is used to improve the detection accuracy.
The mAP value of the model is significantly improved, the detection accuracy and efficiency are improved, and the confidence and missed detection rate of the improved model are also significantly improved in actual detection.
Smart Images

Figure BDA0005226124100000061 
Figure BDA0005226124100000062 
Figure BDA0005226124100000063
Abstract
Description
Technical Field
[0001] The present invention relates to a steel surface defect detection method, and in particular to a steel surface defect detection method based on the YOLOv8 multi-scale convolutional attention mechanism. Technical Background
[0002] As an important basic material, steel plays a vital role in the development of the national economy. It is widely used in various fields, including construction, manufacturing, transportation, energy, etc. The quality of steel directly affects the safety and economic benefits of the project. However, surface defects (such as cracks, plaques, inclusions, pitting surface, rolling scale and scratches) generated during the production and transportation of steel are likely to cause safety hazards. Therefore, it is very important to accurately and efficiently detect surface defects of steel. The detection methods are mainly divided into three categories: manual detection, traditional photoelectric detection and advanced machine vision detection. Among them, although manual detection is intuitive, it is limited by factors such as high labor costs and obvious differences in subjective judgment, resulting in low accuracy and efficiency; traditional photoelectric detection methods include eddy current detection, magnetic flux leakage detection, infrared detection, laser scanning detection, etc. However, due to the high cost, these methods have not been widely adopted.
[0003] In recent years, with the development of technologies such as machine vision and deep learning, the detection of steel surface defects has developed towards automation and artificial intelligence. Machine vision and deep learning technologies are widely used in steel surface defect detection. The current mainstream deep learning methods are divided into one-stage and two-stage detection algorithms. The two-stage detection algorithm includes two steps: object positioning and image classification. The representatives of the two-stage detection algorithm are R-CNN, Fast R-CNN, Faster R-CNN and Mask R-CNN. These algorithms are used to generate candidate boxes and then classify each candidate box. The one-stage algorithm processes the entire image input at the input end at one time. The algorithm mainly extracts features through convolutional neural networks. The representatives of the one-stage detection algorithm are the YOLO (You Only Look Once) series and SSD (Single Shot MultiBox Detector). After the detection is completed, this type of algorithm will directly create a candidate box and directly generate the class probability and coordinate value of the object to obtain the corresponding detection result.
[0004] Xia, KW et al. proposed an improved YOLOv5s model, innovatively designed a large core C3 module that can be re-parameterized, and improved the model's feature perception and extraction capabilities in complex texture environments. However, in industrial environments with limited computing resources, large network architectures have slow computing speeds. Raj, GD et al. proposed the YOLOv7-csf model, which introduced a lightweight and low-cost coordinate attention mechanism in the head structure of YOLOv7. Huang, Y. et al. proposed replacing the original C2F module with a new CFN structure, which reduced the number of network parameters and GFLOPs, but it may be insufficient in global feature extraction. Although the above networks can achieve the task of detecting surface defects of steel while ensuring accuracy, these networks are large-scale and complex networks with very high requirements for computing power. For practical applications, the algorithms need to be deployed on mobile or embedded devices, but because the computing power of mobile terminals is not as good as that of large computers and other devices, they cannot meet the needs of large-scale neural networks. Summary of the invention
[0005] The purpose of the present invention is to propose a steel surface defect detection method based on the YOLOv8 multi-scale convolutional attention mechanism. The method uses the YOLOv8 target detection algorithm to process the steel surface defect image, extracts the image features and transmits the required information data in the image to the steel surface defect detection module. By designing the YOLO model, the problems of low accuracy of steel surface defect detection and false detection and missed detection caused by small defect size are solved.
[0006] The objective of the present invention is achieved through the following technical solutions:
[0007] The steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism has the following steps:
[0008] 1. Collect the open source dataset NEU-DET;
[0009] 2. Use the deep learning pytorch framework to configure the network environment and complete the yolov8 model building in this environment;
[0010] 3. Improve the original yolov8 model framework, introduce the multi-scale convolutional attention mechanism MSCA module in the backbone network and propose a new deformable convolution C2f_DCNv2 module to replace the original C2f module to optimize the model;
[0011] 4. Use the preprocessed data set as the input of the network and train it, load the yolov8 pre-trained weights, and use Wise-IoU (WIoU) as the loss function;
[0012] 5. Put the steel images in the dataset into the improved YOLOv8 network model for defect detection;
[0013] 6. The steel defect detection system organizes and classifies the detected images, which include six types of defect images: cracks, patches, inclusions, pitted surfaces, rolled-in scale, and scratches.
[0014] The method for detecting surface defects of steel materials based on the YOLOv8 multi-scale convolutional attention mechanism is as follows: first, a preprocessing operation is performed on a picture data set of an input model: the data set adopts the open source data set NEU-DET, the pictures in the data set are collected and sorted, and the pictures are converted into JPG format; secondly, the picture data is manually labeled using the labelimg tool, and then the label format is output in the xml format. Since the label format of YOLO adopts txt, it is necessary to convert the xml format label into txt format through a Python program; secondly, the data needs to be divided during network training, with 80% of the data set used as the training set of the network, 10% of the data set used as the test set, and the remaining 10% used as the verification set.
[0015] The steel surface defect detection method based on the YOLOv8 multi-scale convolutional attention mechanism, the preprocessing of the image data set completes the construction of the image processing defect detection model; the model architecture is divided into four parts: Input input end, Backbone backbone network, Neck network layer, Head output end, the input end processes the input image, and adopts Mixup data enhancement, adaptive anchor frame calculation, and adaptive image scaling to improve the model training speed and improve the network accuracy.
[0016] The method for detecting surface defects of steel materials based on the multi-scale convolutional attention mechanism of YOLOv8 is described. The image data set is preprocessed, and a defect detection model is built. In order to further improve the accuracy of steel surface defect detection, a multi-scale convolutional attention mechanism is introduced to optimize the network model. An attention mechanism MSCA module is added to the built YOLOv8 backbone network, and a method combining the self-attention mechanism and convolution is integrated into the network, so that the accuracy of steel surface defect detection is improved; secondly, a new deformable convolution C2f_DCNv2 module is proposed to replace the original C2f module, which enhances the model's ability to capture complex shapes and irregular target features and further improves the ability of feature extraction; finally, Head uses the Wise-IoU (WIoU) loss function to replace the original loss function, accurately measures the similarity between target frames, and improves the detection accuracy of the prediction box; the model loads the YOLOv8 pre-training weights, the initial learning rate uses 0.0005, the momentum is set to 0.937, the loss function uses WIOU, and the remaining parameters use the default values.
[0017] The method for detecting surface defects of steel materials based on the YOLOv8 multi-scale convolutional attention mechanism is described. The model is trained, and some parameters such as the number of data input into the network at one time, the number of training rounds, and the working threads are set to start training the model. After the model training is completed, whether the performance indicators of the model are reasonable are checked, the input data is tested, and the confidence of the prediction box is checked.
[0018] The steel surface defect detection method based on the YOLOv8 multi-scale convolutional attention mechanism is described. The model training is completed and tested. The steel defect detection system organizes and classifies the detected images. The system classifies the six types of defect images: crazing, patches, inclusions, pitted surfaces, rolled-in scale, and scratches.
[0019] The advantages and effects of the present invention are:
[0020] The present invention adopts a method based on target detection and adds image recognition to the steel surface defect detection system. Through model training and testing of steel surface defect images, the classification of steel surface defects is realized, which greatly improves the efficiency of industrial detection and provides an efficient and feasible solution for the field of metal surface defect detection.
[0021] The present invention introduces a multi-scale convolutional attention mechanism MSCA module into the YOLOv8 target detection algorithm and proposes a new deformable convolution C2f_DCNv2 module to replace the original C2f module to optimize the model, thereby improving the model's detection ability for targets of different sizes, thereby improving the accuracy of the network. The mAP value of the improved model is significantly improved by 3.5% based on the original model. At the same time, it can be intuitively seen that the confidence and missed detection of the improved model are also significantly improved in actual detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a structural diagram of steel surface defect detection of the present invention;
[0023] Figure 2 It is the YOLOv8 model architecture diagram of the present invention;
[0024] Figure 3 It is a diagram of the improved YOLOv8 model architecture of the present invention;
[0025] Figure 4 This is the MSCA network structure diagram of the multi-scale convolutional attention mechanism of the present invention;
[0026] Figure 5 This is the network structure diagram of the deformable convolution C2f_DCNv2 module of the present invention;
[0027] Figure 6 It is a flow chart of model training of the present invention;
[0028] Figure 7 It is the PR curve of the YOLOv8 algorithm of the present invention;
[0029] Figure 8 It is the PR curve of the improved YOLOv8 algorithm of the present invention. DETAILED DESCRIPTION
[0030] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are further described in detail below with reference to the accompanying drawings.
[0031] The steel surface defect detection method based on the YOLOv8 multi-scale convolutional attention mechanism of the present invention has a detection structure diagram as shown in FIG. Figure 1 As shown, the present invention takes the following steps:
[0032] 1. Collect the open source dataset NEU-DET from the homepage of Associate Professor Song Kechen of Northeastern University;
[0033] 2. Use the deep learning pytorch framework to configure the network environment and complete the yolov8 model building in this environment;
[0034] 3. Improve the original yolov8 model framework, introduce the multi-scale convolutional attention mechanism MSCA module in the backbone network and propose a new deformable convolution C2f_DCNv2 module to replace the original C2f module to optimize the model;
[0035] 4. Use the preprocessed data set as the input of the network and train it, load the yolov8 pre-trained weights, and use Wise-IoU (WIoU) as the loss function;
[0036] 5. Put the steel images in the dataset into the improved YOLOv8 network model for defect detection;
[0037] 6. The steel defect detection system organizes and classifies the detected images, which include six types of defect images: crazing, patches, inclusions, pitted surfaces, rolled-in scale, and scratches.
[0038] The specific implementation of step 1 is as follows:
[0039] (1) Collect and organize the images in the dataset and convert them into JPG format;
[0040] (2) Use labelimg tool to label the image data;
[0041] (3) Convert the annotated data into txt format;
[0042] (4) Divide the dataset into training set and test set.
[0043] The specific implementation of step 2 is:
[0044] (1) Configuring the experimental environment required by the present invention;
[0045] (2) Build the YOLOv8 model. The model architecture is divided into four parts: Input, Backbone, Neck, and Head. The YOLOv8 model architecture is shown in the figure below. Figure 2 shown.
[0046] The specific implementation of step 3 is as follows:
[0047] The original YOLOv8 model framework is improved, and the multi-scale convolution attention mechanism MSCA module is introduced into the backbone network. A new deformable convolution C2f_DCNv2 module is proposed to replace the original C2f module to optimize the model. The improved model structure is shown in the figure below. Figure 3 shown.
[0048] MSCA is a multi-scale convolutional attention mechanism, which consists of three parts: deep convolution to collect local information, multi-branch deep strip convolution to capture multi-scale context, and 1×1 convolution to build the relationship between different channels. The MSCA model is shown in the figure Figure 4 As shown. Since the strip convolution of each branch is lightweight, in each branch, two depth strip convolutions are used as standard depth convolutions of large kernels. The specific mathematical formula is as follows:
[0049] M=Conv 5×5 (Input) (1)
[0050] M 1 =DWConv 7×1 (DWConv 1×7 (M)) (2)
[0051] M 2 =DWConv 11×1 (DWConv 1×11 (M)) (3)
[0052] M 3 =DWConv 21×1 (DWConv 1×21 (M)) (4)
[0053]
[0054] Among them, Input represents input, DWConv 1×i and DWConv i×1 Represents deep convolution, M1, M2, and M3 represent three branches, and each branch uses convolution kernels of different sizes to perform convolution processing on the input. In each branch, two deep convolutions are used to approximate the standard deep convolution with a large kernel, where the kernel sizes of the three branches are set to 7, 11, and 21, respectively. Finally, the results of M and M1, M2, and M3 are processed by a 1×1 convolution kernel to establish connections between different channels. The processing results are used as weights to weight the Input and output the Output. After the convolution attention mechanism is introduced, the model can adaptively adjust the attention to different areas, thereby improving the ability to capture key information and reducing the impact of background interference.
[0055] This paper proposes a new deformable convolution C2f_DCNv2 module to replace the original C2f module to optimize the model. The difference between the deformable convolution and the traditional convolution is that the convolution position is not fixed and can be adjusted adaptively. It introduces a learnable offset on the basis of the traditional convolution. Since the offset can be fractional, bilinear interpolation is used to calculate the pixel position and obtain the corresponding eigenvalue. The structure of the deformable convolution is shown in the figure. Figure 5 As shown. The output formula of deformable convolution is as follows:
[0056]
[0057] Where w(p n ) indicates p n The weight of the position convolution kernel, x(p 0 +p n +Δp n ) indicates p 0 +p n +Δp n The eigenvalues of the position feature map, expressed in p 0 +p n The offset to add to the position.
[0058] Using deformable convolution can cover objects of different scales, enhance the feature representation ability of the model, and improve detection performance. However, the introduction of offset may cause irrelevant areas to be covered, thereby interfering with feature extraction and reducing overall performance. To solve this problem, the present invention introduces the deformable convolution C2f_DCNv2 module. DCNv2 introduces a modulation mechanism by adding a modulation parameter Δm k ∈[0,1] to learn the weights of the sampling points. For areas of no interest, a weight coefficient Δm is assigned k A small value.
[0059]
[0060] After the introduction of DCNv2, the model can more accurately capture detailed information about the boundaries and complex shapes of target objects, and is especially suitable for detecting objects with varied shapes.
[0061] The specific implementation of step 4 is as follows:
[0062] Load the pre-trained weights of YOLOv8 to improve the speed and accuracy of network training;
[0063] The preprocessed data set is divided into training set and test set and sent to the network;
[0064] Set bath size to 16, epochs to 300, input size to 640, and hyperparameters to default settings;
[0065] Precision and recall are used as indicators to measure the model. The model training flow chart is as follows: Figure 6 As shown, the formulas are:
[0066]
[0067]
[0068] Among them, TP is a positive sample that is correctly identified, FP is a negative sample that is identified as a positive sample, and FN is a negative sample that is correctly identified; in the YOLOv8 model, DFL loss and CIoU loss are used as loss functions. In order to solve the problem that the above two loss functions are affected by the differences in target size and shape of various defect types in the steel surface defect dataset and the differences in image distribution, WIOU is used as the loss function. WIoU deletes the aspect ratio penalty term and balances the impact of high-quality and ordinary quality anchor boxes on model regression, thereby improving the generalization ability and overall performance of the model. The specific formula is as follows:
[0069] L WIoU =rR WIoU L IoU ,R WIoU ∈[1,e),L IoU ∈[0,1] (10)
[0070] The specific implementation of step 5 is as follows:
[0071] Test the training model’s performance indicators such as Precision, Recall, map, and loss function; input steel surface defect images into the network to test the confidence of the network’s recognition.
[0072] The specific implementation of step 6 is as follows:
[0073] The steel surface defect system detection and identification module organizes the defect conditions detected by the system and obtains the defect type conditions. These data will be transmitted to the data management module, which organizes the data and transmits them to the corresponding data modules for classification. The mAP value of the result is significantly improved by 3.5% on the basis of the original model. The PR accuracy of the original algorithm and the improved algorithm is shown in the figure below. Figure 7 , Figure 8 shown.
Claims
1. A steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism, characterized in that: The method builds a network model for steel defect detection, and the steel defect detection and recognition module extracts features from the input image and performs detection and classification, specifically including the following steps: 1) Collect the open source dataset NEU-DET; 2) Use the deep learning pytorch framework to configure the network environment and complete the yolov8 model building in this environment; 3) Improve the original yolov8 model framework, introduce the multi-scale convolutional attention mechanism MSCA module in the backbone network and propose a new deformable convolution C2f_DCNv2 module to replace the original C2f module to optimize the model; 4) The preprocessed data set is used as the input of the network for training, the yolov8 pre-trained weights are loaded, and Wise-IoU (WIoU) is used as the loss function; 5) Put the steel images in the dataset into the improved YOLOv8 network model for defect detection; 6) The steel defect detection system organizes and classifies the detected images, which include six types of defect images: crazing, patches, inclusions, pitted surfaces, rolled-in scale, and scratches.
2. According to claim 1, a steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism is characterized in that: The method includes two main parts of the system. First, the image data set of the input model is preprocessed: the data set uses the open source data set NEU-DET, the images in the data set are collected and sorted, and the images are converted into JPG format. Secondly, the image data is manually labeled using the labelimg tool, and then the label format is output in xml format. Since the label format of YOLO uses txt, it is necessary to convert the xml format label into txt format through a Python program; secondly, when training the network, the data needs to be divided, 80% of the data set is used as the training set of the network, 10% of the data set is used as the test set, and the remaining 10% is used as the verification set.
3. According to claim 2, a steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism is characterized in that: The preprocessing of the image data set completes the construction of the image processing defect detection model; the model architecture is divided into four parts: Input input end, Backbone backbone network, Neck network layer, and Head output end. The input end processes the input image and uses Mixup data enhancement, adaptive anchor frame calculation, and adaptive image scaling to improve the model training speed and network accuracy.
4. According to claim 3, a steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism is characterized in that: The image data set is preprocessed and a defect detection model is built. In order to further improve the accuracy of steel surface defect detection, a multi-scale convolutional attention mechanism is introduced to optimize the network model. An attention mechanism MSCA module is added to the built YOLOv8 backbone network, and a method combining the self-attention mechanism and convolution is integrated into the network, so that the accuracy of steel surface defect detection is improved; secondly, a new deformable convolution C2f_DCNv2 module is proposed to replace the original C2f module, which enhances the model's ability to capture complex shapes and irregular target features and further improves the feature extraction capability; finally, Head uses the Wise-IoU (WIoU) loss function to replace the original loss function, accurately measures the similarity between target frames, and improves the detection accuracy of the prediction box; the model loads the YOLOv8 pre-trained weights, the initial learning rate is 0.0005, the momentum is set to 0.937, the loss function uses WIOU, and the remaining parameters use the default values.
5. According to claim 4, a steel surface defect detection method based on YOLOv8 multi-scale convolutional attention mechanism is characterized in that: The model is trained by setting some parameters such as the number of data input to the network at one time, the number of training rounds, and working threads, and starting to train the model; after the model training is completed, check whether the performance indicators of the model are reasonable, detect the input data, and check the confidence of the prediction box.
6. After the model training is completed and tested according to claim 5, the steel defect detection system organizes and classifies the detected images. The system classifies them according to the six types of defect images: crazing, patches, inclusions, pitted surfaces, rolled-in scale and scratches.
Citation Information
Cited By
Multi-mode teenager idiopathic scoliosis screening method based on back RGB-D image
CN120543912A
Unmanned missile loading vehicle road defect detection method and related device
CN121032954A
Commercial vehicle chassis visual detection method, system and device based on neural network and medium
CN121304603A