Steel surface defect detection method based on YOLO-V5
Through the steel surface defect detection method based on YOLO-V5, by utilizing image preprocessing, sample balancing and innovative detection head, the problems of low detection accuracy and high missed detection rate in traditional detection methods are solved, and efficient and real-time steel surface defect detection is achieved.
Patent Information
- Application Number
- CN202510846557.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional steel surface defect detection methods have the problems of low detection accuracy and high missed detection rate.
A steel surface defect detection method based on YOLO-V5 was adopted. Through image preprocessing, sample balancing, the introduction of the information aggregation and distribution mechanism (Gold-YOLO) and the independently innovative FRMHead detection head, a steel surface defect detection model was constructed for training and verification.
It has achieved efficient and real-time detection of steel surface defects, with the comprehensive detection accuracy of various defect types reaching more than 80%, improving detection accuracy and reducing missed detection rate.
Smart Images

Figure CN120672740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surface defect detection, and in particular to a method for detecting small defects on the surface of steel based on YOLO-V5. Background Art
[0002] Steel is an indispensable raw material in daily life, machinery manufacturing, and the defense industry. As a key industrial product, its production technology has made significant progress with the continuous development of the industry. The market has also placed high demands on the appearance and quality of steel products. However, during the steel production process, factors such as raw material quality, production equipment, and production conditions can lead to various surface defects of varying sizes and characteristics. Surface defects in steel not only reduce its strength, performance, and wear resistance, but also hinder its normal use and may even lead to serious consequences. Therefore, defect detection on the surface of steel is an indispensable step in ensuring its quality. On the production line, detecting surface defects in steel plays a key role in quality control.
[0003] To achieve integrated and high-performance processes, advanced industries require closer collaboration between manufacturing and defect detection. Traditional methods for detecting steel surface defects, such as manual inspection and laser ultrasonic testing, are labor-intensive and inefficient, requiring inspectors to perform a large amount of repetitive work. Therefore, efficient, real-time defect detection systems are crucial. Summary of the Invention
[0004] The present invention provides a steel surface defect detection method based on YOLO-V5, the main purpose of which is to solve the problems of low detection accuracy and high missed detection rate in traditional steel surface defect detection methods.
[0005] To achieve the above objectives, the present invention provides a steel surface defect detection method based on YOLO-V5, comprising the following steps: Step 1: Preprocess the image data of the steel surface; Step 2: Label the defect types of the image data and divide the dataset into training set, test set, and validation set; Step 3: Perform sample balancing, assign weights to defect types, and calculate the weights of image samples; Step 4: Use YOLO-V5 as the backbone network for feature extraction, introduce the information aggregation and distribution mechanism (Gold-YOLO) in the neck part of YOLO-V5, and add the independently innovated FRMHead detection head to the head part of YOLO-V5 to replace the original detection head of YOLO-V5 to build a steel surface defect detection model; Step 5: Use the dataset to train the steel surface defect detection model, test and verify the defect detection results of the model until they meet expectations.
[0006] Furthermore, the image data is derived from a steel surface defect dataset, preferably a NEU-Det dataset, which contains 1,800 images.
[0007] Furthermore, if the image data of the defective part is too small, the image is divided into N parts from left to right and from top to bottom according to the sliding window. M blocks and ensure that the two adjacent images overlap by 10%~20%.
[0008] Furthermore, the preprocessing includes brightness adjustment, contrast adjustment, sharpening and denoising of the image.
[0009] Furthermore, the defect types include RS (rolled-in scale), Pa (patches), Cr (crazing), Ps (pitted surface), In (inclusion), and Sc (scratches).
[0010] Furthermore, the training set, test set, and validation set are randomly divided into the NEU-Det dataset in a ratio of 7:2:1, where the training set accounts for 70%.
[0011] Furthermore, the sample balancing process includes assigning a frequency weight to each defect category and normalizing the weight. The specific steps are as follows: Step 3.1: Assign a frequency weight to each defect category and normalize the weight, where the formula is: in For the The number of samples of the defect class; is the total number of defect categories; is the total number of all defect samples; For the Frequency of class defects; in is the original frequency weight; is a very small constant to prevent division by zero errors; is the original frequency weight; in is the normalized frequency weight; is the sum of the original weights of all categories; Step 3.2: Calculate the category weights, where the formula is: in is the original category weight; For the Class average precision, from the previous round of validation set; in is the normalized category weight; is the sum of the original weights of all categories; Step 3.3: Calculate the image weight, where the formula is: in is the original weight of the image; The defect presence indicator function is: ; Step 3.4: Weight clipping, where the formula is: in is the lower threshold; is the upper threshold; When The weight of each image is between 0.1 and 0.9. There is no need to crop the image weight. If it is less than 0.1, it will be cropped to 0.1. If it is greater than 0.9, it will be cropped to 0.9.
[0012] Furthermore, the fourth step, Information Aggregation-Distribution (Gold-YOLO), is an advanced object detection model that improves information fusion efficiency through an innovative Gather-and-Distribute (GD) mechanism. This mechanism uses convolution and self-attention operations to process information from different layers of the network. In this way, Gold-YOLO can more effectively fuse multi-scale features, achieving an ideal balance between low latency and high accuracy.
[0013] Furthermore, the independently innovated FRMHead detection head in step 4 mainly involves three main modules: DFL, PCRC and FRM, which are: DFL (Distribution Focal Loss) is an integrated module for distributed focal loss. The implementation of DFL transforms the input features through a convolutional layer. The purpose of this process is to re-encode the features to better capture the characteristics of distributed data. At the same time, Focal Loss (Focal Loss for Dense Object Detection) is also applied to calculate the classification loss on the transformed features. in is the predicted probability of the target class, and It is an adjustment parameter used to control the weights of easy-to-classify samples and difficult-to-classify samples; The PCRC (Pooling, Convolution, and ReSampling Combination) module processes and merges information from the input data by combining convolutional layers, upsampling, and multiple sequential modules (including max pooling, average pooling, and additional convolutional layers). It is designed to enhance the model's feature extraction capabilities, especially when dealing with multi-scale features and different types of information, which can improve the model's feature expression ability and robustness. FRM (Feature Reassembling Module) is a feature reassembly module that aims to combine and process feature maps of different scales to improve the performance of image processing tasks. Through operations such as convolution, upsampling, downsampling, and softmax, FRM is able to integrate and reassemble features at different levels to provide rich and effective feature representations. The core idea of FRM is to take advantage of the advantages of multi-scale feature maps and improve the performance of the model in processing complex image tasks by reorganizing and integrating these features. Specifically, FRM achieves this goal through the following steps: First, for the input high-resolution feature map and low-resolution feature maps , we extract preliminary features through convolution operations and get and : in, and Represent the convolution kernel respectively. Next, the low-resolution feature map Upsampling is performed to make its resolution consistent with high-resolution features Figure 1 To: At the same time, for high-resolution feature maps Perform downsampling to reduce the resolution: Then, the upsampled and downsampled feature maps are fused, and the fusion operation is achieved by concatenation: In the fusion feature map We apply a convolutional layer to further process the fused features: Finally, the softmax operation is used to normalize the input feature map for further image processing tasks: Through the above series of operations, FRM can effectively combine feature maps of different scales to provide richer and more effective feature representation, thereby improving the performance of the model in complex image processing tasks.
[0014] Furthermore, the steel surface defect detection model has a detection process including: inputting image data, extracting defect area features through YOLO-V5, introducing an information aggregation-distribution mechanism (Gold-YOLO) in the Neck part of YOLO-V5, fusing features at different levels and injecting the fused global information into each level, adding an independently innovative FRMHead detection head to the head part of YOLO-V5, re-encoding the features to better capture the characteristics of distributed data, and finally outputting the defect type.
[0015] Furthermore, the defect detection effect meets expectations, specifically: the comprehensive detection accuracy of each defect type reaches more than 80%. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present drawings or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present drawings. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without any creative work.
[0017] Figure 1 is a flow chart of the steps of the present invention; Figure 2 This is a structural diagram of the surface defect detection model of the present invention; Figure 3 is an image data output diagram of the present invention; Figure 4 This is the detection accuracy of the surface defect detection model of the present invention for the NEU-Det dataset.
[0018] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments provided herein without inventive effort are intended to fall within the scope of protection of the present invention.
[0019] In this embodiment, Figure 1 As shown, a steel surface defect detection method based on YOLO-V5 includes the following steps: Step 1: Preprocess the image data of the steel surface; Step 2: Label the defect types of the image data and divide the dataset into training set, test set, and validation set; Step 3: Perform sample balancing, assign weights to defect types, and calculate the weights of image samples; Step 4: Use YOLO-V5 as the backbone network for feature extraction, introduce the information aggregation and distribution mechanism (Gold-YOLO) in the neck part of YOLO-V5, and add the independently innovated FRMHead detection head to the head part of YOLO-V5 to replace the original detection head of YOLO-V5 to build a steel surface defect detection model; Step 5: Use the dataset to train the steel surface defect detection model, test and verify the defect detection results of the model until they meet expectations.
[0020] Furthermore, the image data is derived from a steel surface defect dataset, preferably a NEU-Det dataset, which contains 1,800 images.
[0021] Furthermore, if the image data of the defective part is too small, the image is divided into N parts from left to right and from top to bottom according to the sliding window. M blocks and ensure that the two adjacent images overlap by 10%~20%.
[0022] Furthermore, the preprocessing includes brightness adjustment, contrast adjustment, sharpening and denoising of the image.
[0023] Furthermore, the defect types include RS (rolled-in scale), Pa (patches), Cr (crazing), Ps (pitted surface), In (inclusion), and Sc (scratches).
[0024] Furthermore, the training set, test set, and validation set are randomly divided into the NEU-Det dataset in a ratio of 7:2:1, where the training set accounts for 70%.
[0025] Furthermore, the sample balancing process includes assigning a frequency weight to each defect category and normalizing the weight. The specific steps are as follows: Step 3.1: Assign a frequency weight to each defect category and normalize the weight, where the formula is: in For the The number of samples of the defect class; is the total number of defect categories; is the total number of all defect samples; For the Frequency of class defects; in is the original frequency weight; is a very small constant to prevent division by zero errors; is the original frequency weight; in is the normalized frequency weight; is the sum of the original weights of all categories; Step 3.2: Calculate the category weights, where the formula is: in is the original category weight; For the Class average precision, from the previous round of validation set; in is the normalized category weight; is the sum of the original weights of all categories; Step 3.3: Calculate the image weight, where the formula is: in is the original weight of the image; The defect presence indicator function is: ; Step 3.4: Weight clipping, where the formula is: in is the lower threshold; is the upper threshold; When The weight of each image is between 0.1 and 0.9. There is no need to crop the image weight. If it is less than 0.1, it will be cropped to 0.1. If it is greater than 0.9, it will be cropped to 0.9.
[0026] Furthermore, the fourth step, Information Aggregation-Distribution (Gold-YOLO), is an advanced object detection model that improves information fusion efficiency through an innovative Gather-and-Distribute (GD) mechanism. This mechanism uses convolution and self-attention operations to process information from different layers of the network. In this way, Gold-YOLO can more effectively fuse multi-scale features, achieving an ideal balance between low latency and high accuracy.
[0027] Furthermore, the independently innovated FRMHead detection head in step 4 mainly involves three main modules: DFL, PCRC and FRM, which are: DFL (Distribution Focal Loss) is an integrated module for distributed focal loss. The implementation of DFL transforms the input features through a convolutional layer. The purpose of this process is to re-encode the features to better capture the characteristics of distributed data. At the same time, Focal Loss (Focal Loss for Dense Object Detection) is also applied to calculate the classification loss on the transformed features. in is the predicted probability of the target class, and It is an adjustment parameter used to control the weights of easy-to-classify samples and difficult-to-classify samples; The PCRC (Pooling, Convolution, and ReSampling Combination) module processes and merges information from the input data by combining convolutional layers, upsampling, and multiple sequential modules (including max pooling, average pooling, and additional convolutional layers). It is designed to enhance the model's feature extraction capabilities, especially when dealing with multi-scale features and different types of information, which can improve the model's feature expression ability and robustness. FRM (Feature Reassembling Module) is a feature reassembly module that aims to combine and process feature maps of different scales to improve the performance of image processing tasks. Through operations such as convolution, upsampling, downsampling, and softmax, FRM is able to integrate and reassemble features at different levels to provide rich and effective feature representations. The core idea of FRM is to take advantage of the advantages of multi-scale feature maps and improve the performance of the model in processing complex image tasks by reorganizing and integrating these features. Specifically, FRM achieves this goal through the following steps: First, for the input high-resolution feature map and low-resolution feature maps , we extract preliminary features through convolution operations and get and : in, and Represent the convolution kernel respectively. Next, the low-resolution feature map Upsampling is performed to make its resolution consistent with high-resolution features Figure 1 To: At the same time, for high-resolution feature maps Perform downsampling to reduce the resolution: Then, the upsampled and downsampled feature maps are fused, and the fusion operation is achieved by concatenation: In the fusion feature map We apply a convolutional layer to further process the fused features: Finally, the softmax operation is used to normalize the input feature map for further image processing tasks: Through the above series of operations, FRM can effectively combine feature maps of different scales to provide richer and more effective feature representation, thereby improving the performance of the model in complex image processing tasks.
[0028] Furthermore, the steel surface defect detection model has a detection process including: inputting image data, extracting defect area features through YOLO-V5, introducing an information aggregation-distribution mechanism (Gold-YOLO) in the Neck part of YOLO-V5, fusing features at different levels and injecting the fused global information into each level, adding an independently innovative FRMHead detection head to the head part of YOLO-V5, re-encoding the features to better capture the characteristics of distributed data, and finally outputting the defect type.
[0029] Furthermore, the defect detection effect meets expectations, specifically: the comprehensive detection accuracy of each defect type reaches more than 80%.
[0030] The Neck part of the backbone network YOLO-V5 used in this invention plays the role of bridging the Backbone and Head in network communication. Its main function is to integrate features from different scales and prepare for the final target detection. The neck of YOLO-V5 mainly adopts the PANet (Path Aggregation Network) structure, combined with FPN (Feature Pyramid Network). However, in the steel surface defect detection task, it was found that the model still has the problem of information fusion. Therefore, we improved the Neck part of YOLO-V5 and introduced a new mechanism - Information Aggregation-Distribution (Gold-YOLO). This mechanism achieves more efficient information interaction and fusion by fusing features at different levels and injecting the fused global information into each level. This method enhances the model's neck information fusion capability without significantly increasing latency, thereby improving the model's performance in detecting objects of different sizes.
[0031] The detection head of YOLO-V5 is responsible for converting the feature map output by Neck into the final detection result, including the category, location and confidence of the target. Although the detection head of YOLO-V5 itself has cooperated well with Neck and is sufficient for most simple target detection tasks, in the task of detecting surface defects in steel, since most defects are small, its own detection head may not be able to effectively capture details, resulting in low detection accuracy of small targets. This may be because in the process of downsampling the feature map, the features of small targets are overly reduced and difficult to be detected by subsequent layers. At the same time, targets with similar appearances are easily confused, resulting in classification errors. This may be because the resolution of the feature map and the depth of feature extraction are not enough to distinguish subtle differences between categories. Based on the above two problems, the independently innovative FRMHead detection head is introduced to replace the original detection head of YOLO-V5, providing richer and more effective feature representation, thereby improving the performance of the model in complex and small image processing tasks.
[0032] It should be noted that the present invention is not limited to the above-described embodiments. The above-described embodiments are merely illustrative, and any embodiments having substantially the same structure and achieving the same effects as the technical concept within the scope of the present invention are also included within the technical scope of the present invention. Furthermore, within the scope of the present invention, various modifications that can be conceived by those skilled in the art to the embodiments, as well as other configurations constructed by combining some of the components of the embodiments, are also included within the scope of the present invention.
Claims
1. A steel surface defect detection method based on YOLO-V5, characterized in that: The following steps are involved: Step 1: Preprocess the image data of the steel surface; Step 2: Label the defect types of the image data and divide the dataset into training set, test set, and validation set; Step 3: Perform sample balancing, assign weights to defect types, and calculate the weights of image samples; Step 4: Use YOLO-V5 as the backbone network for feature extraction, introduce the information aggregation and distribution mechanism (Gold-YOLO) in the neck part of YOLO-V5, and add the independently innovated FRMHead detection head to the head part of YOLO-V5 to replace the original detection head of YOLO-V5 to build a steel surface defect detection model; Step 5: Use the dataset to train the steel surface defect detection model, test and verify the defect detection results of the model until they meet expectations.
2. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: For the image data, if the image of the defective part is too small, the image is divided into N parts from left to right and from top to bottom according to the sliding window. M blocks and ensure that the two adjacent images overlap by 10%~20%.
3. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: The preprocessing includes brightness adjustment, contrast adjustment, sharpening and denoising of the image.
4. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: The training set, test set, and validation set are randomly divided into the NEU-Det dataset in a ratio of 7:2:1, with the training set accounting for 70%.
5. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: The step three specifically includes: Step 3.1 assigns a frequency weight to each defect category and normalizes the weight; Step 3.2: Calculate the category weight based on the frequency weight calculated in step 3.1; Step 3.3: Perform weighted calculation based on the defect category weights calculated in step 3.2 to obtain the image weight; Step 3.4: The image weights obtained in step 3.3 need to be cropped to prevent the weights from being too extreme.
6. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: The steel surface defect detection model has a detection process including: inputting image data, extracting defect area features through YOLO-V5, introducing an information aggregation and distribution mechanism (Gold-YOLO) in the neck part of YOLO-V5, fusing features at different levels and injecting the fused global information into each level, adding an independently innovated FRMHead detection head to the head part of YOLO-V5, re-encoding features to better capture the characteristics of distributed data, and finally outputting the defect type.
7. The steel surface defect detection method based on YOLO-V5 according to claim 1, characterized in that: The defect detection effect meets expectations, specifically: the comprehensive detection accuracy of each defect type reaches more than 80%.