A steel surface defect target detection method based on improved YOLOv8
By improving the YOLOv8 algorithm, using Mosaic data enhancement, C2f_MSBlock module, CARAFE upsampling operator and Wise-Inner-ShapeIoU loss function, the YOLOv8-ST detection network was built, solving the problem of insufficient accuracy when YOLOv8 detects elongated and irregular small-size defect targets, and achieving higher detection accuracy and feasible deployment.
Patent Information
- Application Number
- CN202410517653.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-04-28
AI Technical Summary
The YOLOv8 algorithm lacks detection accuracy when detecting slender and irregular small-size defect targets.
Based on the improved YOLOv8 steel surface defect object detection method, the YOLOv8-ST detection network is constructed by using Mosaic data augmentation, C2f_MSBlock module, CARAFE upsampling operator and Wise-Inner-ShapeIoU loss function.
It improves the detection accuracy of slender and irregular small-size defect targets, improves the detection accuracy of mAP@50, while maintaining high frame rates, meeting the feasibility of detection accuracy and edge device deployment.
Smart Images

Figure CN118823299B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision target detection, and in particular to a steel surface defect target detection method based on improved YOLOv8. Background Art
[0002] Steel surface defect target detection is a key technology widely used in the industrial field. Its main goal is to accurately locate and identify various defects on the steel surface, such as scratches, cracks, oxidation and dents, through automated methods. These defects may reduce the quality of the product and may cause safety hazards during use.
[0003] With the rapid development of computer vision and deep learning, steel surface defect target detection is increasingly using automated methods based on image processing and machine learning. There are two main types of deep learning-based methods: one is the one-stage algorithm represented by the YOLO series and SSD, which does not require the process of generating region candidate frames, but directly generates detection categories and positions, and can obtain the final detection results in one stage. The other is the two-stage algorithm represented by RCNN and Faster-RCNN, which requires the generation of region candidate frames first and then sample classification.
[0004] The YOLO series of algorithms have been widely used in the industry due to their good detection performance and speed. Compared with previous versions, YOLOv8 has some new features, including a C2f structure with richer gradient flow, mainstream decoupling heads, and a change from Anchor-Based to Anchor-Free. However, the detection accuracy of the YOLOv8 algorithm is still insufficient for slender and irregular small-sized defect targets. Summary of the invention
[0005] The present invention is based on the original YOLOv8 model and provides a steel surface defect target detection method based on improved YOLOv8, which aims to improve the detection accuracy of slender hidden defects by appropriately increasing the number of parameters.
[0006] The present invention is specifically implemented by the following technical scheme. According to a steel surface defect target detection method based on improved YOLOv8 proposed by the present invention, the steps are as follows:
[0007] 1. Obtain the public steel surface defect dataset NEU-DET, which contains 1,800 defect images of six categories: pressed iron oxide scale, cracks, pits, plaques, scratches, and inclusions;
[0008] 2. Divide the training set, test set, and validation set into a ratio of 8:1:1;
[0009] 3. Perform Mosaic data enhancement processing on image samples;
[0010] 4. Build the steel surface defect target detection network YOLOv8-ST based on improved YOLOv8;
[0011] 5. Import the training set and validation set into the network for training to obtain a steel surface defect target detection model;
[0012] 6. Use the target detection model to detect steel surface defects on the test set.
[0013] The present invention provides a preferred solution for a steel surface defect target detection method based on improved YOLOv8, wherein the Mosaic data enhancement process includes:
[0014] Four pictures are randomly read from the data set each time, and the four pictures are randomly flipped, scaled, and color gamut changed. After the operations are completed, the original pictures are placed in four directions with the first picture at the upper left, the second picture at the lower left, the third picture at the lower right, and the fourth picture at the upper right. The fixed areas of the four pictures are cut out using a matrix, and then they are spliced together to form a new picture, which contains a series of contents such as a frame.
[0015] The present invention provides a preferred solution for a steel surface defect target detection method based on improved YOLOv8, and the C2f_MSBlock module includes:
[0016] The MSBlock layer in the C2f_MSBlock module divides the input into multiple branches. As the network goes deeper, the size of the convolution kernel is gradually increased. Small kernel convolution is used in the shallow layer to process high-resolution features, while large kernel convolution is used in the deep layer to capture a wide range of information. Each branch processes a different subset of features. Within the branch, feature conversion is performed through 1×1 convolution; then a k×k deep convolution (depthwise separable convolution) is applied; followed by a 1×1 convolution to reduce the number of parameters; finally, all branches are fused through a 1x1 convolution to integrate the features of each branch. The mathematical expression is as follows:
[0017] (1)
[0018] Except X 1 In addition, each branch must pass through a reverse bottleneck layer, that is, accept the reverse propagation of the previous branch, expressed as , where k represents the kernel size.
[0019] The present invention provides a preferred solution for a steel surface defect target detection method based on improved YOLOv8, wherein the lightweight upsampling operator CARAFE includes:
[0020] The CARAFE operator, the content-aware feature reconstruction operator, is a plug-and-play lightweight operator. Traditional upsampling usually uses techniques such as interpolation or deconvolution, which may introduce some artifacts or blurring effects during the upsampling process. CARAFE uses an adaptive convolution method for upsampling, which can better preserve detail information and reduce artifact blurring. CARAFE uses adaptive convolution to reduce the amount of calculation while maintaining accuracy and increase the speed of the algorithm.
[0021] The present invention provides a preferred solution of a steel surface defect target detection method based on improved YOLOv8, Wise-Inner-ShapeIoU including:
[0022] Keep the weight of WiseIoU, add and normalize the scores of InneIoU and ShapeIoU, and multiply the loss value by the weight of WiseIoU to form a new Wise-Inner-ShapeIoU.
[0023] The following are the mathematical expressions of WiseIoU and Wise-Inner-ShapeIoU:
[0024] (2)
[0025] (3)
[0026] (4)
[0027] (5)
[0028] (6)
[0029] Among them, scores is the score of each bounding box output by the target detection, weight is the sum of the scores of all foreground samples (i.e. positive samples), S and I are the scores calculated by the ShapeIoU and InnerIoU loss functions respectively. is the bounding box regression loss function of the loss function;
[0030] It is the intersection-over-union loss (IoU cost) between the real box and the bounding box: ,in is the predicted box area, is the real frame area; and are the width and height of the predicted box and the real box respectively, is the width and height of the minimum bounding rectangle of the real box and the predicted box, is the center coordinate of the real frame, is the center coordinate of the prediction box. , is a hyperparameter, is the gradient gain (weight of WiseIoU), is the outlier degree, is the dynamic average intersection-combination ratio with momentum m. The superscript * indicates that in order to prevent the generation of gradients that hinder convergence, In , In and Separated from gradient computation.
[0031] The present invention discloses a steel surface defect target detection method based on improved YOLOv8. Compared with the original technology, the present invention has the advantages of optimizing and improving the steel surface defect dataset NEU-DET, being able to perform real-time detection of slender and irregular small-sized defect targets, and improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A flow chart of a steel surface defect target detection method based on improved YOLOv8 provided by the present invention.
[0033] Figure 2 A schematic diagram of the detection network structure of a steel surface defect target detection method based on improved YOLOv8 provided by the present invention.
[0034] Figure 3 This is the schematic diagram of the C2f_MSBlock module structure.
[0035] Figure 4 This is the training effect diagram of the YOLOv8-ST steel surface defect detection network.
[0036] Figure 5 This is the detection effect diagram of the YOLOv8-ST steel surface defect detection network.
[0037] Figure 6 Schematic diagram of the confusion matrix of the YOLOv8-ST steel surface defect detection network. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of the present invention more clear, the specific implementation methods of the present invention are further described below in conjunction with the accompanying drawings.
[0039] Figure 1 The figure is a flow chart of the method of the present invention. The method for detecting steel surface defects based on improved YOLOv8 provided by the present invention specifically comprises the following steps:
[0040] Step 1: Obtain a public steel surface defect dataset and divide the target detection dataset required for training, testing, and validation sets.
[0041] The operation in step 1 is: take the pictures in the NEU-DET dataset and divide them into training set, test set and validation set in a ratio of 8:1:1. Write a python script file to create summary files train.txt and val.txt that store the absolute path of the pictures, label positions and categories by line, and finally put the label tags and jpg pictures in the training set and test set into the same directory.
[0042] Step 2: Perform Mosaic data enhancement processing on the selected images.
[0043] The operation in step 2 is: randomly select 4 pictures, perform augmentation operations on the 4 pictures respectively, and paste them to the corresponding positions of the mask equal to the size of the final output image respectively; the augmentation operation includes random cropping, scaling, arrangement and color gamut change; obtain new pictures and enrich the data set.
[0044] Random cropping: Randomly crop the image to make obstacles appear at different positions in the original image with different proportions.
[0045] Zoom: Zoom the original image to a different size.
[0046] Arrangement: Combination of original pictures and arrangement of frames.
[0047] Color gamut change: change the brightness, saturation and hue of the original image.
[0048] Step 3: Build the road sign detection model YOLOv8-ST based on the improved YOLOv8. The network structure is as follows: Figure 2 shown.
[0049] The operation in step 3 is: first, replace Bottleneck in C2f of the backbone part of the target detection network with MSBlock;
[0050] The MSBlock layer in the C2f_MSBlock module divides the input into multiple branches. As the network goes deeper, the size of the convolution kernel is gradually increased. Small kernel convolution is used in the shallow layer to process high-resolution features, while large kernel convolution is used in the deep layer to capture a wide range of information. Each branch processes a different subset of features. Within the branch, feature conversion is performed through 1×1 convolution; then a k×k deep convolution (depthwise separable convolution) is applied; followed by a 1×1 convolution to enhance features and reduce the number of parameters; finally, all branches are fused through a 1×1 convolution to integrate the features of each branch. The mathematical expression is as follows:
[0051] (1)
[0052] Except X 1 In addition, each branch must pass through a reverse bottleneck layer, that is, accept the reverse propagation of the previous branch, expressed as , where k represents the kernel size.
[0053] Then, the lightweight upsampling operator CARAFE is introduced into the feature fusion network;
[0054] The CARAFE operator, the content-aware feature reconstruction operator, is a plug-and-play lightweight operator. Traditional upsampling usually uses methods such as interpolation or deconvolution, which may introduce some artifacts or blurring effects during the upsampling process. CARAFE uses an adaptive convolution method for upsampling, which can better retain detail information and reduce artifact blurring. At the same time, traditional upsampling methods are usually fixed and cannot dynamically adjust the upsampling process according to the characteristics of the input image. CARAFE can dynamically adjust the upsampling process according to the characteristics of the input image through adaptive convolution, so as to better adapt to different target detection tasks. Traditional upsampling methods usually lead to an increase in the amount of calculation, which affects the speed of the algorithm. CARAFE can reduce the amount of calculation while maintaining accuracy and increase the speed of the algorithm through adaptive convolution. It can be seen that CARAFE can increase the accuracy of the algorithm, reduce the amount of calculation and increase the speed of the algorithm.
[0055] Finally, Wise-Inner-ShapeIoU is proposed to replace CIoU as the loss function of the model. Wise-Inner-ShapeIoU includes:
[0056] Keep the weight r of WiseIoU, add and normalize the scores of InnerIoU and ShapeIoU, and multiply the loss value by the weight of WiseIoU to form a new Wise-Inner-ShapeIoU.
[0057] The following are the mathematical expressions of WiseIoU and Wise-Inner-ShapeIoU:
[0058] (2)
[0059] (3)
[0060] (4)
[0061] (5)
[0062] (6)
[0063] Among them, scores is the score of each bounding box output by the target detection, weight is the sum of the scores of all foreground samples (i.e. positive samples), S and I are the scores calculated by the ShapeIoU and InnerIoU loss functions respectively. is the bounding box regression loss function of the loss function.
[0064] It is the intersection-over-union loss (IoU cost) between the real box and the bounding box: ,in is the predicted box area, is the real frame area; and are the width and height of the predicted box and the real box respectively, is the width and height of the minimum bounding rectangle of the real box and the predicted box, is the center coordinate of the real frame, is the center coordinate of the prediction box. , is a hyperparameter, is the gradient gain (weight of WiseIoU), is the outlier degree, is the dynamic average intersection-combination ratio with momentum m. The superscript * indicates that in order to prevent the generation of gradients that hinder convergence, In , In and Separated from gradient computation.
[0065] Step 4: Evaluate the training results of the steel surface defect target detection model YOLOv8-ST based on the improved YOLOv8.
[0066] The operations in step 4 are: statistical model parameter quantity, average accuracy mAP@50 on the test set, model size, and parameter quantity compared with the original YOLOv8 target detection method.
[0067] Compared with the original YOLOv8 network, the model of the present invention improves the mAP by 2.7% with minimal increase in the number of parameters and model size, while the FPS is still above 30, which meets the accuracy of detection and the feasibility of deployment on edge devices.
[0068] The model training environment of the present invention is: the CPU uses Intel(R) Xeon(R) W-2102@ 2.90GHz, the GPU uses GeForce RTX 2080Ti, the running memory is 64GB, the operating system is Ubuntu 18.04.3 LTS, and the deep learning framework is PyTorch.
Claims
1. A steel surface defect target detection method based on improved YOLOv8, characterized by comprising: 1.
1. Obtain the public steel surface defect dataset NEU-DET, which contains 1,800 defect images, 300 images of each defect, and six categories in total; 1.
2. Divide the dataset into training set, test set and validation set in a ratio of 8:1:1; 1.
3. Perform Mosaic enhancement on the dataset; 1.
4. Construct the steel surface defect target detection network YOLOv8-ST based on the improved YOLOv8, including: Replace Bottleneck in the C2f layer of the backbone part of the target detection network with MSBlock to form the C2f_MSBlock module; Introduce the lightweight upsampling operator CARAFE into the feature fusion network; Wise-Inner-ShapeIoU is proposed to replace CIoU as the loss function of the network, including: Keep part of the weight of WiseIoU, then add and normalize the scores calculated by InnerIoU and ShapeIoU, and multiply the loss value by the weight of WiseIoU. The mathematical expression is as follows: in, is the weight of the WiseIoU loss function, S and I are the ShapeIoU and InnerIoU loss functions respectively, weight is the sum of the scores of all positive samples, scores is the score of each bounding box output by the target detection, β is the outlier degree, δ and α are hyperparameters; 1.
5. Import the training set and validation set into the network for training to obtain a steel surface defect target detection model; 1.
6. Use the detection model to detect steel surface defects on the test set.
2. A steel surface defect target detection method based on improved YOLOv8 as claimed in claim 1, characterized in that The Mosaic data enhancement process in step 1.
2. includes: first, randomly select four images for data enhancement operations, including flipping, scaling and color adjustment; after the operation is completed, the original images are placed in four directions with the first image at the upper left, the second image at the lower left, the third image at the lower right, and the fourth image at the upper right; fixed areas of the four images are cut out and spliced into a new image, which contains the frame selection information.
3. A steel surface defect target detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: The C2f_MSBlock module includes: The MSBlock layer in the C2f_MSBlock module divides the input into multiple branches, each branch processes a different subset of features, and inside the branch, feature conversion is performed through a 1×1 convolution; then a k×k depthwise separable convolution is connected; then another 1×1 convolution is connected to reduce the number of parameters; finally, all branches are fused through a 1×1 convolution to integrate the features of each branch.
Citation Information
Patent Citations
Steel surface defect detection algorithm based on improved YOLOv8 model
CN117745697A