Steel surface defect target detection method based on DMFSO-YOLO network

By introducing the speed optimization accuracy module, dynamic multi-scale feature fusion module and Normalized Gaussian Wasserstein Distance loss function in the YOLOv8n algorithm, the DMFSO-YOLO network is built, which solves the problem of difficult balance of speed and accuracy in the detection of steel surface defects in the existing technology, and achieves efficient and accurate defect detection.

CN120088196APending Publication Date: 2025-06-03ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017997.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to balance speed and accuracy in steel surface defect detection, and they lack the ability to handle multi-scale features, so they cannot accurately detect small defects and background noise-intensive areas.

Method used

Based on the YOLOv8n algorithm, the speed optimization accuracy module (SOPM) and the dynamic multi-scale feature fusion module (DMFF) were added, and the Normalized Gaussian Wasserstein Distance (NWD) was introduced as the regression loss function to build the DMFSO-YOLO network.

Benefits of technology

It realizes high-precision steel surface defect detection, while ensuring good detection efficiency, accurately identifying tiny defects in complex backgrounds, and significantly optimizing the positioning accuracy of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088196A_ABST
    Figure CN120088196A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic multi-scale fusion and speed optimization network (DMFSO-YOLO) for steel surface defect detection on the basis of a YOLOv8n algorithm. Comprising the steps that S1, in a DMFSO-YOLO backbone network, a speed optimization precision module (SOPM) is provided, residual partial convolution and channel extrusion-excitation operation are used for transmitting information, and redundancy calculation is reduced while high precision is kept; and S2, designing a dynamic multi-scale feature fusion module (DMFF) in the neck network, and enabling the model to accurately recognize tiny defects in a complex background through multi-stage feature extraction, an attention mechanism and deep feature fusion. And S3, in the head network, introducing a normalized Gaussian-Warisstein distance (NWD) so as to more effectively quantify the difference between a prediction frame and a real frame, provide stable gradient feedback and enhance the positioning capability of the model to small defects. Compared with a traditional method, the real-time performance of the model can be ensured while the detection performance is remarkably improved. The method can be used for a steel surface defect detection task in a complex scene, and the detection precision is greatly improved compared with that of a traditional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, involving technologies such as deep learning, image processing, and target detection and recognition. Specifically, it refers to a method for detecting steel surface defect targets based on the DMFSO-YOLO network. Background Art

[0002] Due to its strength, durability, and versatility, steel has become a key material in multiple industrial sectors such as construction and automotive manufacturing. However, surface defects can significantly reduce the quality of steel, thereby affecting the service life and performance of products. Therefore, identifying and dealing with these defects at the early stage of the production process is crucial for maintaining high standards and ensuring product integrity.

[0003] The emergence of deep learning has provided a promising approach to address the shortcomings of these steel surface defect detections. Neural networks trained on large-scale image datasets have significantly reduced manual intervention, improving the detection speed and reliability. They have become increasingly popular and gradually replaced machine vision-based detection methods. Existing object detection algorithms, such as YOLO, Faster R-CNN, ViT, etc., have shown effective performance and good adaptability in dealing with complex surface defect detection tasks. However, these models still face challenges, especially in dealing with multi-scale features and accurately detecting small defects that are closely similar to background noise. These challenges highlight the need for models that can effectively integrate multi-dimensional features and provide precise error localization. Specifically, the main problems of existing algorithms are as follows: First, it is difficult for them to strike a balance between speed and accuracy. Second, their ability to process multi-scale features is insufficient, and they cannot accurately detect tiny defects in similar noise. Third, their ability to locate small defects is weak. Summary of the Invention

[0004] The technical problem solved by the present invention is to propose a method for detecting steel surface defect targets based on the DMFSO-YOLO network in view of the deficiencies of the prior art, which is an object detection algorithm applicable to the steel surface defect detection platform. The present invention adds a Speed Optimization Precision Module (SOPM) and a Dynamic Multi-Scale Feature Fusion Module (DMFF) on the basis of YOLOv8n, and uses NWD as the regression loss function. The present invention has high detection accuracy and at the same time ensures good detection efficiency, providing important support for the quality control of steel industrial production.

[0005] To achieve the above object, a method for detecting steel surface defect targets based on the DMFSO-YOLO network provided by the present invention is carried out according to the following steps:

[0006] Step 1: Obtain a publicly available steel surface defect image dataset and perform preprocessing;

[0007] Step 2: Configure the model training environment;

[0008] Step 3: Construct the DMFSO-YOLO network model. The DMFSO-YOLO network model is based on the YOLOv8n model. A Speed Optimization Precision Module (SOPM) is added to the backbone network of the model, a Dynamic Multi-Scale Feature Fusion (DMFF) module is added to the neck network of the model, and the Detection Head Network introduces a Normalized Gaussian Wasserstein Distance (NWD) loss function;

[0009] Step 4: The DMFSO-YOLO network model is deployed in the preset training environment, and its parameter file is appropriately adjusted. The input image resolution is 640×640, the number of training rounds is 200, the batch-size is 8, and the number of object categories to be detected is 6. Subsequently, the pre-divided dataset is used to train and validate the model to evaluate its performance;

[0010] Step 5: The trained and validated DMFSO-YOLO network model is applied to steel surface defect detection. By inputting the image to be detected, the recognition and positioning of defect targets are realized.

[0011] Furthermore, the said Step 1 includes the following sub-steps:

[0012] Step 1.1: The preprocessing method involves formatting the acquired dataset to meet the requirements of the DMFSO-YOLO network model;

[0013] Step 1.2: Divide the dataset formatted in Step 1.1 into a training set and a dataset that meet the requirements.

[0014] Furthermore, in the said Step 3, the configured training environment is: Central Processing Unit (CPU) Core TM i9-12900H, NVIDIA GeForce RTX 3070 GPU, cuda11.7, deep learning framework pytorch1.13.1.

[0015] Furthermore, the said Step 3 includes the following sub-steps:

[0016] Step 3.1: Set the input resolution to 640×640×3 to adapt to the detection requirements of small-sized defects;

[0017] Step 3.2: The backbone network extracts multi-level features through the CBS module and the SOPM module, and uses the SPPF module to fuse local and global features through max-pooling operations with different kernel sizes, outputting feature maps from different depths, namely P1(80×80×128), P2(40×40×256), and P3(20×20×384), enhancing the receptive field and enriching the feature representation;

[0018] Step 3.3: The neck network introduces a dynamic multi-scale feature fusion module (DMFF), dynamically adjusts the feature weights through the multi-head attention mechanism, and combines upsampling and feature concatenation operations to achieve effective interaction and enhancement of features at different scales;

[0019] Step 3.4: The detection head adopts an improved multi-branch structure, combines the BCE loss function for classification prediction and the NWD loss function for bounding box regression, where the NWD loss can accurately quantify the difference between the predicted box and the ground truth box, provide stable gradient feedback, and significantly optimize the localization accuracy of small objects.

[0020] The NWD loss function uses the Wasserstein distance to calculate the distribution distance. For two two-dimensional Gaussian distributions and μ 1 and μ 2 The second-order Wasserstein distance between them is defined as:.

[0021]

[0022] It can be simplified to:

[0023]

[0024] where ||·|| F represents the Frobenius norm.

[0025] In addition, for the Gaussian distributions a modeled by the bounding boxes A = (cx a , cy a , w a ) and B = (cx b , cy b , w b , h b ) and and The above formula can be further simplified to:

[0026]

[0027] It is normalized in its exponential form, and a new metric is obtained, called the NormalizedWasserstein Distance (NWD):

[0028]

[0029] where C is a constant closely related to the dataset.

[0030] The loss function based on NWD is as follows:

[0031] L NWD = 1 - NWD(N p , N g )

[0032] where is the Gaussian distribution model of the predicted bounding box P, is the Gaussian distribution model of the ground truth bounding box G.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the YOLOv8n algorithm, the present invention proposes a network with dynamic multi-scale fusion and speed optimization (DMFSO-YOLO). First, in the DMFSO-YOLO backbone network, a speed-optimized precision module (SOPM) is proposed. This compact module reduces redundant calculations while maintaining high precision. Compared with the original YOLOv8n backbone, the DMFSO-YOLO backbone achieves higher precision, a smaller algorithm scale, and less computational complexity. In addition, a dynamic multi-scale feature fusion module (DMFF) is designed in the neck network, enabling the model to accurately identify tiny defects in complex backgrounds through multi-level feature extraction, attention mechanisms, and deep feature fusion. Finally, the normalized Gaussian-Wasserstein distance (NWD) is introduced to more effectively quantify the difference between the predicted box and the ground truth box, provide stable gradient feedback, and enhance the model's ability to locate small defects. Compared with traditional methods, it can ensure the real-time performance of the model while significantly improving the detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions of the present invention, a brief introduction is made to the drawings required for the present invention.

[0035] Figure 1 is a schematic diagram of the existing YOLOv8n network structure;

[0036] Figure 2 is a schematic diagram of the overall network structure of the DMFSO-YOLO proposed by the present invention;

[0037] Figure 3 is a schematic diagram of the SOPM network module structure proposed by the present invention;

[0038] Figure 4 Schematic diagram of the DFMM network module structure proposed by the present invention;

[0039] Figure 5 Schematic diagram of the decoupled head structure with NWD regression loss;

[0040] Figure 6 Detection result diagram using the present invention. Specific implementation manners

[0041] To clarify the purpose, technical solutions and advantages of the embodiments of the present application, the following details its technical solutions in conjunction with the accompanying drawings. It should be noted that only some embodiments are discussed, not all. Based on the examples provided herein, all other embodiments developed by those skilled in the art without creative efforts fall within the protection scope of the present application.

[0042] The following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings.

[0043] The YOLOv8n network structure mainly consists of a backbone, a neck, and a head, as Figure 1 shown. YOLOv8n uses a modified CSPDarknet53 as the backbone network. The input features are downsampled five times to obtain five different-scale features in sequence. The CSP module of the original backbone network is replaced by the C2f module. The C2f module adopts a split connection, enriching the information flow of the feature extraction network while maintaining light weight. The CBS module performs a convolution operation on the input information, then performs batch normalization, and finally uses the SiLU activation function to obtain the output result.

[0044] The present invention proposes a steel surface defect target detection method based on the DMFSO-YOLO network, which is applicable to the target detection algorithm of the steel surface defect detection platform. The present invention adds a speed optimization precision module (SOPM) and a dynamic multi-scale feature fusion module (DMFF) on the basis of YOLOv8nn, and uses NWD as the regression loss function. The present invention has high detection accuracy and at the same time ensures good detection efficiency, providing an important support for the quality control of steel industrial production. The specific implementation steps are as follows:

[0045] Step 1: Extract data samples from the strip steel surface defect image dataset and divide them into a training set and a test set according to a specific ratio. Then, perform necessary preprocessing on the images in the training set.

[0046] The preprocessing includes data augmentation and normalization. The data augmentation methods include increasing brightness, decreasing brightness, and adding Gaussian noise; normalization standardizes the image pixel values to a fixed range.

[0047] Step 2: Improve based on the network structure of the existing YOLOv8n algorithm to obtain the DMFSO-YOLO object detection model after improvement.

[0048] The DMFSO-YOLO network model is based on the YOLOv8n model as the base network. A Speed Optimization Precision Module (SOPM) is added to the backbone network of the model, and a Dynamic Multi-Scale Feature Fusion (DMFF) module is added to the neck network of the model. The detection head network introduces the Normalized Gaussian Wasserstein Distance (NWD) loss function. The network structure diagram is as Figure 2 shown.

[0049] Set the input resolution of the preprocessed image in Step 1 to 640×640×3 to adapt to the detection requirements of small-sized defects.

[0050] The backbone network extracts multi-level features through the CBS module and the SOPM module. The structure of the Speed Optimization Precision Module (SOPM) is as Figure 3 shown. The specific operation is as follows: In this module, the number of input channels is directly reduced (from C to 0.5C) through 1×1 convolution operations, which can reduce the subsequent computational amount. After the branch, a part of the data passes through a simple 1×1 convolution, while another part of the data passes through a more complex 3×3 convolution. The parallelization of complex and simple operations maintains the feature extraction ability while reducing unnecessary computational costs.

[0051] Then, use the SPPF module to fuse local and global features through max-pooling operations with different kernel sizes, and output feature maps from different depths, namely P1 (80×80×128), P2 (40×40×256), and P3 (20×20×384), which enhances the receptive field and enriches the feature representation.

[0052] The neck network adds a Dynamic Multi-Scale Feature Fusion module (DMFF), and the network structure is as Figure 4 shown. Dynamically adjust the feature weights through the multi-head attention mechanism, and combine upsampling and feature splicing operations to achieve effective interaction and enhancement of features at different scales. The specific operation is as follows:

[0053] The input feature map first passes through a 1×1 convolutional layer to adjust the number of channels and prepare the feature map for subsequent processing. It is then normalized using batch normalization to ensure a more stable distribution of activation values ​​during training. The normalized feature map is activated using the SiLu activation function. This function combines the advantages of the sigmoid and ReLU functions to provide a smooth nonlinear transformation of the input. After activation, the features are processed through a 3×3 partial convolution to ensure efficient spatial extraction of features without significantly increasing the computational load. To further stabilize the learning process, layer normalization is applied after the 3×3 partial convolution to ensure that each feature map has a consistent scale across different layers. It then passes through three layers of Bi-Layer Routing Attention (BRA), which filters and routes the most relevant information from the feature map. This mechanism is designed to selectively focus on important features while ignoring less relevant features, thereby improving the overall efficiency and effectiveness of the model. Afterwards, an MLP module processes the filtered information. It allows the selected features to be combined and refined. The outputs of the BRA mechanism and the MLP module are then spliced ​​together to form a comprehensive representation of the input data. Finally, the concatenated features are projected back to the original number of channels using another 1×1 convolutional layer to ensure compatibility with the rest of the network.

[0054] The three sizes of tensors obtained after the neck layer processing are input into the detection head. The detection head is a decoupling head with NWD regression loss. The specific structure is as follows Figure 5 As shown in the figure. The detection head adopts an improved multi-branch structure, combining the BCE loss function for classification prediction and the NWD loss function for bounding box regression. The NWD loss can accurately quantify the difference between the predicted box and the true box, provide stable gradient feedback, and significantly optimize the positioning accuracy of small targets. The NWD loss function uses the Wasserstein distance to calculate the distribution distance. For two two-dimensional Gaussian distributions and μ 1 and μ 2 The second-order Wasserstein distance between is defined as:

[0055]

[0056] It can be simplified to:

[0057]

[0058] where ||·|| F Represents the Frobenius norm.

[0059] In addition, for the bounding box a=(cx a ,cy a ,w a ,h a ) and B=(cxb , cy b , w b , h b ) Gaussian distribution for modeling and The above formula can be further simplified as follows:

[0060]

[0061] Normalize it using its exponential form and obtain a new metric, called Normalized Wasserstein Distance (NWD):

[0062]

[0063] where C is a constant closely related to the dataset.

[0064] The loss function based on NWD is as follows:

[0065] L NWD = 1 - NWD(N p , N g )

[0066] In is the Gaussian distribution model of the predicted bounding box P, is the Gaussian distribution model of the ground truth bounding box G.

[0067] Step 3: Use the training set to train the steel surface defect detection model to obtain the optimal model; use this model to detect the test set images; evaluate the model performance through the mean average precision (mAP) and the frames per second (FPS).

[0068] The dataset is an image dataset for metal surface defect detection provided by Northeastern University in China. This dataset includes six common defect types: rolled-in scale (rs), patches (pa), cracks (cr), pitted surface (ps), inclusions (in), and scratches (sc). Each defect type usually includes hundreds of high-resolution grayscale images collected from real industrial environments, providing complexity and diversity. These pictures are divided into the training set and the test set in a ratio of 7:3 for input training.

[0069] The training environment of the algorithm of the present invention is: Central Processing Unit (CPU) Core TM i9 - 12900H, NVIDIA GeForce RTX 3070 GPU, cuda11.7, deep learning framework pytorch1.13.1, the number of training epochs is 200, batch-size is 8, and the number of object categories to be detected is 6. Some detection results are as Figure 6as shown

[0070] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention, rather than limiting their applications. Those skilled in the art can make appropriate modifications or equivalent substitutions to the technical solutions according to needs. As long as these changes do not deviate from the core of the present invention, they fall within the scope of the technical solutions of the present invention.

Claims

1. A steel surface defect target detection method based on DMFSO-YOLO network, characterized in that The steps include: Step 1: Obtain a publicly available dataset of steel surface defects and perform preprocessing; Step 2: Set up the hardware and software environment required for model training; Step 3: Construct a DMFSO-YOLO network model, wherein the DMFSO-YOLO network model is based on the YOLOv8n model, a speed optimized precision (SOPM) module is added to the backbone network of the model, a dynamic multi-scale feature fusion (DMFF) module is added to the neck network of the model, and the detection head network introduces the Normalized GaussianWasserstein Distance (NWD) loss function; Step 4: Deploy the constructed DMFSO-YOLO network in the preset training environment, adjust the parameters appropriately, and use the training set and validation set divided after preprocessing for training and performance evaluation; Step 5: Apply the trained and verified model to steel surface defect detection, and accurately identify and locate defect targets by inputting the image to be detected.

2. According to claim 1, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that Step 1 includes: Step 1.1: Format the acquired data set to meet the input requirements of the DMFSO-YOLO network model; Step 1.2: Divide the formatted dataset into training set and validation set.

3. According to claim 1, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that In step 2, the training environment is configured as follows: Central Processing Unit (CPU) Core TM i9-12900H, NVIDIA GeForce RTX 3070 GPU, cuda11.7, deep learning framework pytorch1.13.

1.

4. According to claim 1, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that: Step 3 includes: Step 3.1: Set the input resolution to 640×640×3; Step 3.2: The backbone network extracts multi-level features through the CBS module and the SOPM module, and uses the SPPF module to fuse local and global features. The output feature maps are P1 (80×80×128), P2 (40×40×256), and P3 (20×20×384). Step 3.3: Add the DMFF module to the neck network to achieve feature interaction and enhancement through multi-head attention mechanism, upsampling and feature concatenation operations; Step 3.4: The detection head combines the BCE loss function for classification prediction and the NWD loss function for bounding box regression.

5. According to claim 4, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that: In step S302, the SOPM module reduces the number of channels from C to 0.5C through 1×1 convolution to reduce the amount of calculation, and maintains the feature extraction capability through parallelization of simple and complex convolution operations.

6. According to claim 4, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that: In step S303, the specific operation of the DMFF module is as follows: first, the input feature map is subjected to 1×1 convolution to adjust the number of channels, and is processed by batch normalization and SiLU activation function; then, 3×3 partial convolution is used to extract spatial features, and layer normalization is applied to stabilize the learning process; then, the important features are filtered and routed by three double-layer routing attention (BRA) modules; finally, the filtered features are combined by the MLP module, and the output is spliced ​​and projected back to the original number of channels by 1×1 convolution.

7. According to claim 4, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that: In step 3.4, the NWD loss function is used to quantify the difference between the predicted box and the true box, and its specific formula is: L NWD =1-NWD(N p ,OF g ) in is the Gaussian distribution model for the predicted bounding box P, is the Gaussian distribution model of G of the ground-truth bounding box.

8. According to claim 1, a steel surface defect target detection method based on DMFSO-YOLO network is characterized in that: In step 3, the training method includes: dividing the preprocessed data set into a training set and a validation set, setting a 200-round training cycle, and inputting 8 images per batch. During the training, Wandb is used to monitor the training log, and the results are saved after the training is completed.

9. The method for detecting steel surface defects based on the DMFSO-YOLO network according to claim 1, characterized in that: In step 4, modifying the network model parameter file involves: initializing parameter settings, adjusting the input image size to 640×640, setting the number of training rounds to 200, the batch-size to 8, and the number of detection object categories to 6.