Self-adaptive lightweight real-time strip steel surface defect detection method
By constructing a precision-enhanced feature pyramid network and an adaptive NWD loss function under the YOLOv10n framework, combined with progressive module frozen architecture search, the contradiction between accuracy and speed in strip surface defect detection is solved, and efficient and robust detection results are achieved.
Patent Information
- Application Number
- CN202510737167.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-16
AI Technical Summary
Existing strip surface defect detection models have difficulty accurately capturing defect edges in low-contrast and high-noise environments, and their lightweight design often leads to low computational efficiency, making them unable to meet the real-time needs of industrial production.
A precision-enhanced feature pyramid network is constructed, the AcuF module and adaptive NWD loss function are introduced, and a progressive module freezing architecture search strategy is adopted to optimize the network structure to improve detection accuracy and reduce computational complexity.
It significantly improves the positioning accuracy and robustness of defect edges, while reducing computational complexity and detection delay, achieving a balance between high precision and real-time performance, and is suitable for industrial-grade strip steel surface defect detection.
Smart Images

Figure CN120655592A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to an adaptive lightweight real-time strip steel surface defect detection method. Background Art
[0002] Industrial defect detection is a core component in ensuring product quality and production efficiency, and is widely used in the steel, automotive, aerospace, and other fields. Hot-rolled strip, a key material in the steel industry, is commonly used in processes such as cutting, bending, stamping, and welding. However, unstable process and storage conditions during the production process often lead to various defects on the strip surface. These defects not only affect the appearance quality but also significantly reduce the mechanical properties of the material, such as tensile strength, fatigue life, and ductility, directly affecting the quality and service life of the final product. Therefore, automatic defect detection technology based on the strip surface has become an important research direction for improving product quality and production efficiency.
[0003] With the rapid development of computer vision and deep learning technologies, automatic defect detection algorithms have shown great application potential in the field of defect identification and localization, and have important industrial value. Surface defect detection (SDD) aims to automatically identify and locate defective areas on the material surface through target detection algorithms. However, strip surface defect detection still faces several unique challenges, mainly reflected in the following aspects:
[0004] First, surface defects in steel strips typically exhibit low contrast, and the defective area contains relatively little high-frequency information, resulting in blurred edges and details. To address this issue, existing mainstream methods often enhance feature differentiation capabilities by designing complex deep convolutional neural networks (such as FPN and attention modules). However, complex network structures can lead to the loss of low-level features during transmission, while insufficient integration of global information can prevent the model from effectively fusing low-level details with high-level semantics, thus affecting the accurate capture of defect edges, especially in defective areas lacking high-frequency information.
[0005] Secondly, the loss function has a direct impact on the network's convergence speed, detection accuracy, and robustness. In traditional object detection models, commonly used loss functions, such as cross-entropy loss and smooth L1 loss, rely primarily on clear boundary information. However, in low-contrast or high-noise environments, these standard loss functions struggle to effectively process subtle target features, resulting in poor localization accuracy.
[0006] Finally, the high-speed operation of the strip production line requires the defect detection system to be able to process large amounts of image data in real time, which places extremely high demands on the algorithm's computational efficiency and latency. To meet the challenge of the huge amount of computation, existing research has mostly adopted lightweight structures, knowledge distillation, hardware acceleration, and other methods to improve detection speed. However, lightweight design often sacrifices feature expression capabilities and pattern recognition accuracy. To compensate for this shortcoming, many studies have improved accuracy by adding feature extraction layers or introducing complex learning strategies. Despite this, these strategies often lead to more complex model structures, increased computational workload and memory consumption, thereby reducing inference speed and increasing the hardware burden. Therefore, how to improve detection accuracy while maintaining lightweight design and ensuring good hardware adaptability has become a core issue that needs to be urgently addressed in the field of strip surface defect detection. Summary of the Invention
[0007] Purpose of the invention: The purpose of the present invention is to provide an adaptive lightweight real-time strip steel surface defect detection method to solve the accuracy and speed problems of existing strip steel surface defect detection models.
[0008] Technical solution: The present invention provides an adaptive lightweight real-time strip steel surface defect detection method, comprising the following steps:
[0009] Step 1: Build a detection model, construct a precision-enhanced feature pyramid network in the YOLOv10n structure, and introduce a shallow feature path and AcuF module;
[0010] Step 2: Introduce the adaptive NWD loss function into the detection model;
[0011] Step 3: Introduce progressive module freeze architecture search into the detection model to build the backbone architecture;
[0012] Step 4: Iteratively train the detection model using a public dataset of strip surface defects;
[0013] Step 5: Detect surface defects of the target strip based on the trained detection model.
[0014] Furthermore, step 1 is specifically as follows: the feature pyramid part introduces a shallow path from the trunk, replaces the specific module in the feature pyramid structure with the AcuF module, and adopts a dual detection head structure; the AcuF module consists of a multi-scale extractor and a step-by-step feature accumulation module. The multi-scale extractor selects an appropriate receptive field according to different feature map resolutions, and then fuses features of different scales; first, the input feature map is evenly divided into 4 groups [Y0, Y1, Y2, Y3], and then the accumulated features of the previous layer are spliced with the current branch input to obtain the accumulated features of the next layer, as shown in Equation 1:
[0015]
[0016] The output of the module is obtained after Bottleneck processing, as shown in Formula 2:
[0017] Y out =Conv2(Concat(A0,A1,A2,A3)) (2)
[0018] Furthermore, step 2 is specifically as follows: let the predicted box B1 and the real box B2, and the bounding box coordinates are shown in formula 3:
[0019]
[0020] in, They represent the center coordinates and width and height of the predicted box respectively, and x, y, w, and h represent the center coordinates and width and height of the real box respectively.
[0021] By diagonal length The center point coordinates and width and height of the predicted box are directly normalized to obtain the predicted box and the true box, as shown in Equation 4:
[0022]
[0023] We construct probability distribution models for the predicted and true boxes in terms of position and scale, respectively, to capture the fine-grained differences between them. We use the center points of the predicted and true boxes as the mean of the two-dimensional normal distribution, and use the normalized Euclidean distance to calculate the position difference, as shown in Equation 5:
[0024]
[0025] where ω x and ω y are the position weights in the x and y directions respectively;
[0026] After taking the logarithm of the aspect ratio, the model is built and the normalized logarithmic difference is calculated. At the same time, the aspect ratio difference is added, such as
[0027] As shown in formula 6:
[0028]
[0029] where r and are the aspect ratios of the real box and the predicted box, σ w and σ h are the standard deviations of the differences in the logarithms of width and height, and Υ is the weight of the aspect ratio difference;
[0030] The position and scale difference metrics are globally normalized and mapped to a unified scale range to achieve balance and global adaptability of the final NWD loss. The position difference normalization and scale difference normalization are shown in Equations 7 and 8, respectively:
[0031]
[0032] Where m and MAD are the median and median absolute deviation calculated in advance on the training set, respectively;
[0033] The optimized NWD loss function is shown in Equation 9:
[0034]
[0035] Among them, α and β represent weight coefficients, and C represents the scaling factor of the distance before number mapping.
[0036] Furthermore, step 3 is as follows: when searching for an architecture, the progressive module freezing search method constrains the optimization goal to reduce the detection delay while ensuring that the model accuracy is not lower than the established benchmark value; setting the benchmark accuracy mAP baseline With the delay threshold T, the mAP and average delay LAT(m) of a given model m are constrained, as shown in Equation 10:
[0037]
[0038] Under the premise of satisfying the accuracy constraint, the optimization goal is transformed into minimizing the normalized delay As shown in formula 11:
[0039]
[0040] Set the penalty term g(m) = mAP baseline -mAP(m), the final objective function is shown in Equation 12:
[0041]
[0042] The progressive module freezing search method is to freeze all modules in the trunk and unfreeze them one by one, searching for the optimal structure according to the constraints.
[0043] Furthermore, the public datasets of strip surface defects include two datasets: NEU-DET and GC10-DET.
[0044] The present invention further discloses a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.
[0045] The present invention further discloses a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the method of the present invention when the computer program / instruction is executed by a processor.
[0046] The present invention further discloses a computer program product, comprising a computer program / instruction, which implements the steps of the method of the present invention when executed by a processor.
[0047] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0048] The present invention achieves structural optimization by constructing a precision-enhanced feature pyramid network under the YOLOv10n framework, in which a shallow feature path and an AcuF module (composed of a multi-scale extractor and a step-by-step feature accumulation module) are introduced to achieve deep fusion of features of different resolutions, effectively alleviating the problem of insufficient extraction of low-contrast and small defect information; at the same time, the adaptive Normalized Wasserstein Distance loss function proposed in the patent significantly improves the accuracy and robustness of bounding box positioning by modeling the predicted box and the true box as two-dimensional normal distributions, and using normalized Euclidean distance and logarithmic difference to measure the fine-grained differences in position and scale; in addition, a progressive module freezing architecture search strategy is adopted, that is, all modules of the backbone network are first frozen, and then some modules are gradually unfrozen and targeted for optimization, thereby optimizing the network structure and reducing the computational complexity while ensuring that the detection accuracy is not lower than the preset benchmark.
[0049] After verification on the NEU-DET and GC10-DET datasets, the improved detection model has a more concentrated activation response in the defect area, can more accurately capture and locate various surface defects, and significantly reduce false detections and missed detections; at the same time, the network optimized using the progressive module freezing strategy has obvious improvements in computational complexity and detection latency, achieving a balance between high precision and real-time performance, thus providing a more efficient and robust solution for industrial-grade strip steel surface defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is an overall flow chart of the adaptive lightweight real-time strip steel surface defect detection method of the present invention;
[0051] Figure 2 This is the structure diagram of the precision-enhanced feature pyramid network;
[0052] Figure 3 This is the structure diagram of the multi-scale extractor;
[0053] Figure 4 It is a structure diagram of level-by-level feature accumulation;
[0054] Figure 5 Examples of dataset images; (A) is an example of NEU-DET dataset image; (B) is an example of GC10-DET dataset image;
[0055] Figure 6 Comparison of AP results of various classes in the dataset; (A) Comparison of AP results of various classes in the NEU-DET dataset; (B) Comparison of AP results of various classes in the GC10-DET dataset;
[0056] Figure 7 Visual comparison of heatmap and detection map; (A) Visual comparison of heatmap and detection map of NEU-DET dataset; (B) Visual comparison of heatmap and detection map of GC10-DET dataset DETAILED DESCRIPTION
[0057] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0058] like Figure 1 As shown in the figure, an adaptive lightweight real-time strip surface defect detection method is:
[0059] The steps include:
[0060] S1. Construct an accuracy-enhanced feature pyramid network in the YOLOv10n structure, introduce shallow feature paths and AcuF modules. Among them, the accuracy-enhanced feature pyramid network structure is as follows: Figure 2 As shown in the figure, the feature pyramid part introduces a shallow path from the backbone, replaces the specific module in the feature pyramid structure with the AcuF module, and adopts a dual detection head structure;
[0061] The AcuF module consists of a multi-scale extractor and a step-by-step feature accumulation module. The specific structure is as follows:
[0062] Multi-scale extractors such as Figure 3 As shown in Figure 1, appropriate receptive fields are selected according to different feature map resolutions, and then features of different scales are fused;
[0063] The level-by-level feature accumulation module is as follows Figure 4 As shown:
[0064] First, the input feature map is divided into 4 groups [Y0, Y1, Y2, Y3]. Then, the accumulated features of the previous layer are concatenated with the current branch input to obtain the accumulated features of the next layer, as shown in Equation 1:
[0065]
[0066] The output of the module is obtained after Bottleneck processing, as shown in Formula 2:
[0067] Y out =Conv2(Concat(A0,A1,A2,A3)) (2)
[0068] S2. Introduce the adaptive NWD loss function, specifically:
[0069] Assuming the predicted box B1 and the real box B2, the bounding box coordinates are shown in formula 3:
[0070]
[0071] By diagonal length The center point coordinates and width and height of the predicted box are directly normalized to obtain the predicted box and the true box, as shown in Equation 4:
[0072]
[0073] We construct probability distribution models for the predicted and true boxes in terms of position and scale, respectively, to capture the fine-grained differences between them. We use the center points of the predicted and true boxes as the mean of the two-dimensional normal distribution, and use the normalized Euclidean distance to calculate the position difference, as shown in Equation 5:
[0074]
[0075] where ω x and ω y They are the position weights in the x and y directions, respectively, used to adjust the importance of the direction.
[0076] After taking the logarithm of the aspect ratio, the model is built and the normalized logarithmic difference is calculated. At the same time, the aspect ratio difference is added, such as
[0077] As shown in formula 6:
[0078]
[0079] where r and are the aspect ratios of the real box and the predicted box, σ w and σ h are the standard deviations of the differences in the logarithms of width and height, and Υ is the weight of the aspect ratio difference.
[0080] The position and scale difference metrics are globally normalized and mapped to a unified scale range to achieve balance and global adaptability of the final NWD loss. The position difference normalization and scale difference normalization are shown in Equations 7 and 8 respectively:
[0081]
[0082] Here, m and MAD are the median and median absolute deviation calculated in advance on the training set, respectively.
[0083] The optimized NWD loss function is shown in Equation 9:
[0084]
[0085] S3. Introducing progressive module freezing architecture search, the backbone architecture is constructed by gradually freezing some modules and optimizing the remaining modules. In order to simplify the search process, the progressive module freezing search method constrains the optimization goal to minimize the detection delay while ensuring that the model accuracy is not lower than the established benchmark value. Set the benchmark accuracy mAP baseline With the delay threshold T, the mAP and average delay LAT(m) of a given model m are constrained, such as
[0086] As shown in formula 10:
[0087]
[0088] Under the premise of satisfying the accuracy constraint, the optimization goal is transformed into minimizing the normalized delay As shown in formula 11:
[0089]
[0090] Set the penalty term g(m) = mAP baseline -mAP(m), the final objective function is shown in Equation 12:
[0091]
[0092] The progressive module freezing search method is to freeze all modules in the trunk and unfreeze them one by one, searching for the optimal structure according to the constraints.
[0093] S4. Use public datasets to iteratively train the detection model.
[0094] Experiment and result analysis:
[0095] This example uses two datasets, NEU-DET and GC10-DET, to evaluate the performance of the algorithm.
[0096] The NEU-DET dataset contains six typical types of hot-rolled strip surface defects: Crazing (Cr), Inclusion (In), Patches (Pa), Pitted Surface (PS), Rolled-in Scale (RS), and Scratches (Sc). The dataset includes 1,800 images, 300 for each defect. GC10-DET is a surface defect dataset collected in real industry. The dataset includes 3,570 grayscale images (the current version is 2,300), including 10 defect types such as punching (Punching_Hole), welding (Welding_Line), crescent (Crescent_Gap), water spot (Water_Spot), oil spot (Oil_Wpot), silk spot (Silk_Spot), foreign body intrusion (Inclusion), indentation (Rolled_Pit), severe crease (Crease), and waist fold (Waist_Folding). The dataset image examples are as follows: Figure 5 As shown in (A) and (B).
[0097] All training processes in this example are performed on a server equipped with an AMD Ryzen 57500F 6-Core CPU, 32GB of RAM, an NVIDIA GeForce RTX 4070Super GPU, and Windows 10. The software environment is Torch 2.4.1 with CUDA 11.7.
[0098] During training, the input image size was set to 640, the total number of training epochs was 500, the batch size was 16, and the early stopping patience value was set to 50. The SGD optimizer was used with an initial learning rate of 0.01, a momentum of 0.937, and a weight decay coefficient of 0.0005. For data augmentation, the hue, saturation, and brightness of the HSV parameters were adjusted to 0.015, 0.7, and 0.4, respectively. A translation range of 0.1, a scale range of 0.5, a horizontal flip probability of 0.5, and a random erase probability of 0.4 were used. Mosaic data augmentation was also enabled.
[0099] In the ablation experiment of this example, the independent effects and combined effects of the three improved modules were evaluated on the two datasets of NEU-DET and GC10-DET. The AP results are as follows: Figure 6 As shown in (A) and (B).
[0100] This example provides a detailed visual comparison of the detection results of the two datasets NEU-DET and GC10-DET, as shown in the following example: Figure 7Each group of images contains four rows: the first row is the original image, the second row is the heat map of the baseline model, the third row is the heat map of the proposed model, and the fourth row is the prediction result with the detection box.
[0101] The heat map and prediction results show that the baseline model has problems such as blurred activation areas and insufficient response to defect morphology for different types of defects, resulting in the coverage of the prediction box being usually larger or smaller than the actual defect area, and even obvious missed detection. In contrast, the model proposed in the present invention has a more concentrated response to the defect area in the heat map, and the degree of activation is closer to the defect morphology itself, which can better correspond to the defect position and range in the original image. Especially in samples with complex background textures, the baseline model often causes erroneous activation due to noise interference, while the model proposed in the present invention maintains a more stable focusing effect and achieves accurate highlighting of the target defect. It can be seen from the result graph with detection boxes that the method of the present invention is more accurate in positioning the detection boxes of the two data sets, and rarely has large-scale misjudgments or missed detections, indicating that the improved feature extraction effectively suppresses background interference and enhances sensitivity to defect details.
Claims
1. An adaptive lightweight real-time strip surface defect detection method, characterized in that: The steps include: Step 1: Build a detection model, construct a precision-enhanced feature pyramid network in the YOLOv10n structure, and introduce a shallow feature path and AcuF module; Step 2: Introduce the adaptive NWD loss function into the detection model; Step 3: Introduce progressive module freeze architecture search into the detection model to build the backbone architecture; Step 4: Iteratively train the detection model using a public dataset of strip surface defects; Step 5: Detect surface defects of the target strip based on the trained detection model.
2. The adaptive lightweight real-time strip steel surface defect detection method according to claim 1, characterized in that: Step 1 is as follows: the feature pyramid part introduces a shallow path from the backbone, replaces the specific module in the feature pyramid structure with the AcuF module, and adopts a dual detection head structure; the AcuF module consists of a multi-scale extractor and a step-by-step feature accumulation module. The multi-scale extractor selects an appropriate receptive field according to the resolution of different feature maps, and then fuses features of different scales; first, the input feature map is divided into 4 groups [Y0, Y1, Y2, Y3], and then the accumulated features of the previous layer are concatenated with the current branch input to obtain the accumulated features of the next layer, as shown in Equation 1: The output of the module is obtained after Bottleneck processing, as shown in Formula 2: <h2 style=";text-align:left;direction:ltr">Y<h2 style=";text-align:left;direction:ltr"> out <h2 style=";text-align:left;direction:ltr"> =Conv2(Concat(A0,A1,A2,A3)) (2) 3. The adaptive lightweight real-time strip steel surface defect detection method according to claim 1, characterized in that: Step 2 is as follows: Let the predicted box B1 and the real box B2, and the bounding box coordinates are shown in formula 3: in, Represent the center point coordinates and width and height of the predicted box respectively, and x, y, w, and h represent the center point coordinates and width and height of the real box respectively; By diagonal length The center point coordinates and width and height of the predicted box are directly normalized to obtain the predicted box and the true box, as shown in Equation 4: We construct probability distribution models for the predicted and true boxes in terms of position and scale, respectively, to capture the fine-grained differences between them. We use the center points of the predicted and true boxes as the mean of the two-dimensional normal distribution, and use the normalized Euclidean distance to calculate the position difference, as shown in Equation 5: where ω x and ω y are the position weights in the x and y directions respectively; The logarithm of the aspect ratio is taken and the model is built. The normalized logarithmic difference is calculated and the aspect ratio difference is added, as shown in Equation 6: where r and are the aspect ratios of the real box and the predicted box, σ w and σ h are the standard deviations of the differences in the logarithms of width and height, and Υ is the weight of the aspect ratio difference; The position and scale difference metrics are globally normalized and mapped to a unified scale range to achieve balance and global adaptability of the final NWD loss. The position difference normalization and scale difference normalization are shown in Equations 7 and 8, respectively: Where m and MAD are the median and median absolute deviation calculated in advance on the training set, respectively; The optimized NWD loss function is shown in Equation 9: Among them, α and β represent weight coefficients, and C represents the scaling factor of the distance before number mapping.
4. The adaptive lightweight real-time strip steel surface defect detection method according to claim 1, characterized in that: Step 3 is as follows: When searching for an architecture, the progressive module freezing search method constrains the optimization goal to reduce the detection delay while ensuring that the model accuracy is not lower than the established benchmark value; set the benchmark accuracy mAP baseline With the delay threshold T, the mAP and average delay LAT(m) of a given model m are constrained, as shown in Equation 10: Under the premise of satisfying the accuracy constraint, the optimization goal is transformed into minimizing the normalized delay As shown in formula 11: Set the penalty term g(m) = mAP baseline -mAP(m), the final objective function is shown in Equation 12: The progressive module freezing search method is to freeze all modules in the trunk and unfreeze them one by one, searching for the optimal structure according to the constraints.
5. The adaptive lightweight real-time strip steel surface defect detection method according to claim 1, characterized in that: In step 4, the public datasets of strip surface defects include two datasets: NEU-DET and GC10-DET.
6. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to claim 1.
7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.