Steel surface defect detection method and system

By introducing data enhancement technology of the YOLOV9s-Bi FPN-DW-C3 model and diffusion hidden model into the YOLO series algorithm, the problem of high overfitting and calculation costs in steel surface defect detection in the prior art is solved, and efficient and real-time defect detection effect is achieved.

CN120125565APending Publication Date: 2025-06-10GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510283729.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing YOLO series algorithms have problems of overfitting and high computational costs in steel surface defect detection, resulting in low detection accuracy and difficult to meet the real-time detection requirements of industrial sites.

Method used

The YOLOV9s-Bi FPN-DW-C3 model is adopted. This model enhances the utilization of multi-scale feature information and reduces the computational amount of model by introducing Bi FPN feature fusion network and DW depth. At the same time, the image data of steel surface defects is enhanced through the diffusion hidden model, enriching the diversity of the training data and improving the generalization ability of the model.

Benefits of technology

On the premise of ensuring detection accuracy, the calculation amount of the model is significantly reduced, the detection efficiency is improved, and the real-time detection needs can be met in the industrial site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125565A_ABST
    Figure CN120125565A_ABST
Patent Text Reader

Abstract

The invention discloses a steel surface defect detection method and system, and belongs to the technical field of machine vision and industrial detection. Comprising the steps that steel surface image data are received and input into a pre-trained YOLOV9s-Bi FPN-DW-C3 model to detect the steel surface defects, the model comprises a backbone network layer, a neck network layer and a detection head layer, and the step of detecting the steel surface defects comprises the substeps that the steel surface image data are input into the backbone network layer for multi-scale feature extraction, generating a feature map set containing different levels of semantic information; inputting the feature map set into a neck network layer, executing up-sampling matching, normalization preprocessing, weighted fusion and multi-path integration operation of a cross-scale feature map, and outputting an optimized fusion feature map; inputting the fusion feature map into a detection head layer, traversing the fusion feature map based on a preset anchor frame, and predicting the defect category probability and bounding box offset of each anchor frame; and generating a defect detection result list according to the defect category probability and the bounding box offset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision and industrial inspection, and particularly to a method and system for detecting steel surface defects. Background Art

[0002] In the process of steel production and manufacturing, various defects such as cracks, holes, and scratches will inevitably occur on the steel surface. These defects will not only affect the appearance quality of the steel, but also seriously reduce the physical properties and service life of the steel, and further have a significant impact on the quality of industrial products using the steel subsequently. Therefore, accurately and efficiently detecting steel surface defects is crucial for ensuring steel quality and industrial production safety.

[0003] In recent years, object detection technology based on deep learning has been widely used in the field of industrial inspection. Among them, the YOLO (You Only Look Once) series of algorithms have shown excellent performance in object detection tasks due to their high real-time performance and simplicity and efficiency. However, in the scenario of steel surface defect detection, there are still some problems with the existing YOLO series of algorithms. On the one hand, the image data of steel surface defects is often limited, and the defect types are diverse and the features are complex, which makes the model prone to overfitting during the training process, resulting in low detection accuracy. On the other hand, the traditional YOLO model has a large number of parameters and high computational costs, and it is difficult to meet the requirements of rapid detection in the real-time detection environment of the industrial site.

[0004] In addition, the current object detection algorithms have deficiencies in feature fusion and information transmission, and cannot make full use of feature information at different scales, especially the detection effect for some tiny defects is poor. Therefore, how to reduce the computational amount of the model and improve the detection efficiency on the premise of ensuring the detection accuracy is an urgent problem to be solved in the current field of steel surface defect detection.

[0005] The disclosure of the above background art content is only used to assist in understanding the concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0006] This application provides a method and system for detecting steel surface defects, which can reduce the computational amount of the model and improve the detection efficiency on the premise of ensuring the detection accuracy.

[0007] To achieve the above object, the embodiments of this application disclose the following technical solutions:

[0008] In a first aspect, the embodiments of this application provide a method for detecting steel surface defects, including the following steps:

[0009] Receive the steel surface image data;

[0010] Input the received steel surface image data into the pre-trained YOLOV9s-Bi FPN-DW-C3 model to detect the defects on the steel surface. The YOLOV9s-Bi FPN-DW-C3 model includes a backbone network layer, a neck network layer, and a detection head layer. The steps for the YOLOV9s-Bi FPN-DW-C3 model to detect the defects on the steel surface are as follows:

[0011] Input the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a set of feature maps containing semantic information at different levels;

[0012] Input the set of feature maps into the neck network layer, perform upsampling matching, normalization preprocessing, weighted fusion, and multi-path integration operations on the cross-scale feature maps, and output the optimized fused feature maps;

[0013] Input the fused feature maps into the detection head layer, traverse the fused feature maps based on the preset anchor boxes and predict the defect class probabilities and bounding box offsets for each anchor box; generate a list of defect detection results containing defect type labels and bounding box coordinates according to the defect class probabilities and bounding box offsets.

[0014] In some possible implementation manners of the first aspect, the backbone network layer includes an initial convolutional downsampling module, an ELAN1 module, an AConv module, and a first C3 module. Among them, an SE attention mechanism is added to the first C3 module. The steps for multi-scale feature extraction are as follows:

[0015] Perform spatial compression and channel expansion on the steel surface image through the initial convolutional downsampling module to generate a first-scale feature map containing underlying texture features;

[0016] Input the first-scale feature map into the ELAN1 module, fuse the feature information of different receptive fields through multi-branch heterogeneous convolutional operations, and output a middle-layer feature map with enhanced expression ability;

[0017] Input the middle-layer feature map into the AConv module for secondary spatial downsampling and channel expansion to generate a second-scale feature map containing high-level semantic information;

[0018] Input the second-scale feature map into the first C3 module, and output a set of feature maps that fuse multi-scale semantic information through multi-path convolutional splitting, heterogeneous processing, and channel splicing integration.

[0019] In some possible embodiments of the first aspect, the neck network layer includes a Bi FPN module and a second C3 module, wherein an SE attention mechanism is added to the second C3 module, and the cross-scale feature map processing steps include:

[0020] Upsample the high-level feature map output by the backbone network layer to match the size of the middle-level feature map, and perform normalization preprocessing on the middle-level feature map;

[0021] Input the upsampled high-level feature map and the preprocessed middle-level feature map into the Bi FPN module, and perform cross-level semantic fusion on the input feature maps through a dynamic weight allocation mechanism. The dynamic weight allocation mechanism includes: allocating learnable weight parameters to each input feature map, performing weighted summation after non-linear activation and weight normalization processing to generate a semantically enhanced fused feature map;

[0022] Input the fused feature map into the second C3 module, and output a discriminatively optimized fused feature map through multi-path convolution splitting, heterogeneous branch processing, and channel dimension splicing integration. The optimized fused feature map contains detailed features and context semantic information of the steel surface defects.

[0023] In some possible embodiments of the first aspect, the detection head layer includes a Dual DDetect module, and the defect detection result generation steps include:

[0024] Input the optimized fused feature map output by the neck network layer into the Dual DDetect module, and traverse the multi-scale feature maps through preset anchor boxes to initially locate potential defect regions;

[0025] Perform the following operations synchronously in the Dual DDetect module:

[0026] Extract feature information through the first branch convolution network and predict the class probabilities of cracks, holes, or scratches existing in each anchor box;

[0027] Calculate the bounding box offset of each anchor box through the second branch convolution network to accurately locate the spatial coordinates of the defect;

[0028] Filter the class probabilities and bounding box offsets based on the non-maximum suppression algorithm to remove redundant detection boxes;

[0029] Map the filtered detection results to the original image coordinate system to generate a detection list containing defect type labels and their corresponding bounding box coordinates.

[0030] In some possible embodiments of the first aspect, the training set and test set used during the training of the YOLOV9s-BiFPN-DW-C3 model are generated through the following steps:

[0031] Input the original steel surface defect image into the pre-trained diffusion latent model, and the diffusion latent model generates a new defect image according to the probability distribution inside the original steel surface defect image;

[0032] Unify the original data and the enhanced data into the YOLO training format, and divide the training set and the test set according to a preset ratio.

[0033] In a second aspect, an embodiment of the present application provides a steel surface defect detection system, including:

[0034] A first receiving unit for receiving steel surface image data;

[0035] A first analysis unit for inputting the received steel surface image data into the pre-trained YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects. The YOLOV9s-BiFPN-DW-C3 model includes a backbone network layer, a neck network layer, and a detection head layer. The steps for the YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects include:

[0036] Input the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a set of feature maps containing semantic information at different levels;

[0037] Input the set of feature maps into the neck network layer, perform upsampling matching, normalization preprocessing, weighted fusion, and multi-path integration operations on the cross-scale feature maps, and output an optimized fused feature map;

[0038] Input the fused feature map into the detection head layer, traverse the fused feature map based on a preset anchor box and predict the defect class probability and bounding box offset of each anchor box; according to the defect class probability and the bounding box offset, generate a defect detection result list containing defect type labels and bounding box coordinates.

[0039] In some possible implementation manners of the second aspect, the backbone network layer includes an initial convolutional downsampling module, an ELAN1 module, an AConv module, and a first C3 module. Among them, an SE attention mechanism is added to the first C3 module. The first analysis unit is specifically used for:

[0040] Perform spatial compression and channel expansion on the steel surface image through the initial convolutional downsampling module to generate a first-scale feature map containing underlying texture features;

[0041] Input the first-scale feature map into the ELAN1 module, fuse the feature information of different receptive fields through multi-branch heterogeneous convolutional operations, and output a middle-layer feature map with enhanced expression ability;

[0042] Input the middle - level feature map into the AConv module for secondary spatial down - sampling and channel expansion to generate a second - scale feature map containing high - level semantic information;

[0043] Input the second - scale feature map into the first C3 module, and through multi - path convolution splitting, heterogeneous processing and channel splicing integration, output a set of feature maps that fuse multi - scale semantic information.

[0044] In some possible implementation manners of the second aspect, the neck network layer includes a BiFPN module and a second C3 module, where an SE attention mechanism is added to the second C3 module. The first analysis unit is specifically further configured to:

[0045] Upsample the high - level feature map output by the backbone network layer to match the size of the middle - level feature map, and perform normalization pre - processing on the middle - level feature map;

[0046] Input the upsampled high - level feature map and the pre - processed middle - level feature map into the BiFPN module, and perform cross - level semantic fusion on the input feature maps through a dynamic weight assignment mechanism. The dynamic weight assignment mechanism includes: assigning learnable weight parameters to each input feature map, performing weighted summation after non - linear activation and weight normalization processing to generate a fusion feature map with enhanced semantics;

[0047] Input the fusion feature map into the second C3 module, and through multi - path convolution splitting, heterogeneous branch processing and channel - dimension splicing integration, output a discriminatively optimized fusion feature map. The optimized fusion feature map contains the detailed features and context semantic information of the steel surface defects.

[0048] In some possible implementation manners of the second aspect, the detection head layer includes a Dual DDetect module. The first analysis unit is specifically further configured to:

[0049] Input the optimized fusion feature map output by the neck network layer into the Dual DDetect module, and traverse the multi - scale feature map through a preset anchor box to initially locate potential defect regions;

[0050] Synchronously perform the following operations in the Dual DDetect module:

[0051] Extract feature information through the first - branch convolutional network and predict the class probabilities of cracks, holes or scratches existing in each anchor box;

[0052] Calculate the bounding box offset of each anchor box through the second - branch convolutional network to accurately locate the spatial coordinates of the defect;

[0053] Based on the non - maximum suppression algorithm, screen the class probabilities and bounding box offsets to remove redundant detection boxes;

[0054] Map the filtered detection results to the original image coordinate system to generate a detection list containing defect type labels and their corresponding bounding box coordinates.

[0055] In some possible implementations of the second aspect, the training set and the test set used in the training of the YOLOV9s-BiFPN-DW-C3 model are generated through the following steps:

[0056] Input the original steel surface defect images into a pre-trained diffusion latent model, and the diffusion latent model generates new defect images according to the probability distribution inside the original steel surface defect images;

[0057] Unify the original data and the augmented data into the YOLO training format, and divide the training set and the test set according to a preset ratio.

[0058] In a third aspect, an embodiment of the present application provides an electronic device, including one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any one of the technical solutions in the first aspect.

[0059] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the technical solutions in the first aspect is implemented.

[0060] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the technical solutions in the first aspect is implemented.

[0061] One or more technical solutions provided in the first aspect of the embodiments of the present application have at least the following technical effects or advantages:

[0062] The YOLOv9-Bi FPN-DW-C3 model constructed by improving the YOLOv9 algorithm, introducing the Bi FPN feature fusion network enhances the utilization of multi-scale feature information, adopting the DW depthwise separable convolution reduces the computational complexity of the model, and improving the C3 module improves the model's detection ability for tiny defects.

[0063] By using the diffusion latent model to augment the steel surface defect image data, the problems of insufficient data and overfitting are effectively solved, the diversity of the training data is enriched, and the generalization ability of the model is improved. At the same time, while improving the detection accuracy of steel surface defects, the computational complexity of the model is significantly reduced, which can meet the requirements of real-time detection in industrial sites.

[0064] Among them, for the technical effects brought by any one of the design methods in the second to fifth aspects, reference can be made to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained by extending the provided drawings.

[0066] Figure 1 It is a flowchart of a method for detecting steel surface defects provided by some embodiments of the present application;

[0067] Figure 2 It is a flowchart of the YOLOv9-Bi FPN-DW-C3 model provided by some embodiments of the present application;

[0068] Figure 3 It is a schematic structural diagram of a steel surface defect detection system provided by some embodiments of the present application;

[0069] Figure 4 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] Now the specific implementation schemes of the present invention will be described in detail. Although the present invention is described in connection with these specific implementation schemes, it should be understood that it is not intended to limit the present invention to these specific implementation schemes. On the contrary, these implementation schemes are intended to cover alternative, modified, or equivalent implementation schemes that may be included within the spirit and scope of the invention defined by the claims. In the following description, a large number of specific details are set forth in order to provide a comprehensive understanding of the present invention. The present invention can be implemented without some or all of these specific details.

[0071] When used in conjunction with the terms "comprising", "the method comprises", or similar language in this specification and the appended claims, the singular forms "a", "an", "the" include plural references unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0072] Application Overview: During the steel production and manufacturing process, various defects such as cracks, holes, and scratches inevitably occur on the steel surface. These defects not only affect the appearance quality of the steel, but also seriously reduce the physical properties and service life of the steel, thereby having a significant impact on the quality of industrial products using the steel subsequently. Therefore, accurately and efficiently detecting steel surface defects is crucial for ensuring steel quality and industrial production safety.

[0073] In recent years, object detection technologies based on deep learning have been widely applied in the field of industrial inspection. Among them, the YOLO (You Only Look Once) series of algorithms have shown excellent performance in object detection tasks due to their high real-time performance and simplicity and efficiency. However, in the scenario of steel surface defect detection, there are still some problems with the existing YOLO series of algorithms. On the one hand, the image data of steel surface defects are often limited, and the defect types are diverse and the features are complex, which makes the model prone to overfitting during the training process, resulting in low detection accuracy. On the other hand, the traditional YOLO model has a large number of parameters and high computational costs, and it is difficult to meet the requirements of rapid detection in the real-time detection environment of the industrial site.

[0074] For the above technical problems, please refer to Figure 1 and Figure 2 In the embodiments of this application, a method for detecting steel surface defects is provided, including the following steps:

[0075] S101: Receive steel surface image data;

[0076] S102: Input the received steel surface image data into the pre-trained YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects.

[0077] Among them, the YOLOV9s-BiFPN-DW-C3 model includes a backbone network layer, a neck network layer, and a detection head layer. The steps for the YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects include:

[0078] The first step is to input the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a set of feature maps containing semantic information at different levels;

[0079] Specifically, in some embodiments, the backbone network layer includes an initial convolutional downsampling module, an ELAN1 module, an AConv module, and a first C3 module. Among them, an SE attention mechanism is added to the first C3 module. The steps for multi-scale feature extraction include:

[0080] The first sub-step is to perform spatial compression and channel expansion on the steel surface image through the initial convolutional downsampling module to generate a first-scale feature map containing underlying texture features;

[0081] The second sub-step is to input the first-scale feature map into the ELAN1 module, fuse the feature information of different receptive fields through multi-branch heterogeneous convolution operations, and output a middle-level feature map with enhanced expressive ability.

[0082] The third sub-step is to input the middle-level feature map into the AConv module for secondary spatial downsampling and channel expansion to generate a second-scale feature map containing high-level semantic information.

[0083] The fourth sub-step is to input the second-scale feature map into the first C3 module, and through multi-path convolution splitting, heterogeneous processing, and channel splicing integration, output a set of feature maps that fuse multi-scale semantic information. The set of feature maps includes but is not limited to P3 / 8, P4 / 16, and P5 / 32 level feature maps. Among them, the size of P3 / 8 is 160×160, the size of P4 / 16 is 80×80, and the size of P5 / 32 is 40×40.

[0084] The second step is to input the set of feature maps into the neck network layer, perform upsampling matching, normalization preprocessing, weighted fusion, and multi-path integration operations on the cross-scale feature maps, and output an optimized fused feature map.

[0085] Specifically, in some embodiments, the neck network layer includes a Bi FPN module and a second C3 module. Among them, an SE attention mechanism is added to the second C3 module. The cross-scale feature map processing steps include:

[0086] The first sub-step is to perform upsampling on the high-level feature map output by the backbone network layer to match the size of the middle-level feature map, and perform normalization preprocessing on the middle-level feature map.

[0087] Exemplarily, the high-level feature map is the P5 / 32 level feature map; the middle-level feature map is the P4 / 16 level feature map.

[0088] The second sub-step is to input the upsampled high-level feature map and the preprocessed middle-level feature map into the BiFPN module, and perform cross-level semantic fusion on the input feature maps through a dynamic weight allocation mechanism. The dynamic weight allocation mechanism includes: assigning learnable weight parameters to each input feature map, performing weighted summation after non-linear activation and weight normalization processing, and generating a fused feature map with enhanced semantics.

[0089] The third sub-step is to input the fused feature map into the second C3 module, and through multi-path convolution splitting, heterogeneous branch processing, and channel dimension splicing integration, output a discriminatively optimized fused feature map. The optimized fused feature map contains the detailed features and context semantic information of the steel surface defects.

[0090] In the third step, input the fused feature map into the detection head layer, traverse the fused feature map based on the preset anchor boxes, and predict the defect class probability and bounding box offset of each anchor box; according to the defect class probability and bounding box offset, generate a list of defect detection results including defect type labels and bounding box coordinates.

[0091] Specifically, in some embodiments, the detection head layer includes a Dual DDetect module, and the defect detection result generation steps include:

[0092] The first sub-step: Input the optimized fused feature map output by the neck network layer into the Dual DDetect module, traverse the multi-scale feature map through the preset anchor boxes, and initially locate the potential defect areas;

[0093] The second sub-step: Synchronously perform the following operations in the Dual DDetect module:

[0094] The third sub-step: Extract feature information through the first branch convolutional network, and predict the class probability of cracks, holes or scratches in each anchor box;

[0095] The fourth sub-step: Calculate the bounding box offset of each anchor box through the second branch convolutional network, and accurately locate the spatial coordinates of the defect;

[0096] The fifth sub-step: Based on the non-maximum suppression algorithm, screen the class probability and bounding box offset, and remove redundant detection boxes;

[0097] The sixth sub-step: Map the screened detection results to the original image coordinate system, and generate a detection list including defect type labels and their corresponding bounding box coordinates. Exemplarily, the bounding box coordinates are represented by the upper left vertex coordinates and width and height values.

[0098] Preferably, in some embodiments, the training set and test set used during the training of the YOLOV9s-BiFPN-DW-C3 model are generated through the following steps:

[0099] The first step: Input the original steel surface defect image into the pre-trained diffusion latent model, and the diffusion latent model generates new defect images according to the probability distribution inside the original steel surface defect image;

[0100] The second step: Uniformly convert the original data and the augmented data into the YOLO training format, and divide the training set and test set according to a preset ratio.

[0101] The specific training process is as follows: Initialize the model: Initialize the YOLOv9-BiFPN-DW-C3 model using the pre-trained YOLOv9 model parameters to accelerate the convergence speed of the model.

[0102] Set hyperparameters: Set the learning rate to 0.001, the batch size to 64, and the number of training epochs to 300. Use the Stochastic Gradient Descent (SGD) optimization algorithm to update the parameters of the model.

[0103] Training process: During the training process, batch the training set data and input it into the model. The model calculates the value of the loss function based on the input images and annotation information. The loss functions used include classification loss, localization loss, and confidence loss. Calculate the gradients through the backpropagation algorithm and use the optimization algorithm to update the parameters of the model, so that the value of the loss function gradually decreases. During the training process, regularly evaluate the performance of the model on the validation set to prevent overfitting.

[0104] The comparison experiment results between the YOLOv9-Bi FPN-DW-C3 model of the present invention and the traditional YOLOv9 model are as follows in the table:

[0105]

[0106] It can be seen from the comparison of the experimental results that the new model has achieved a large improvement in overall accuracy. The new model reaches 101% performance with only 87.5% of the computational parameters.

[0107] Please refer to Figure 3 , based on the same inventive concept as a steel surface defect detection method in the foregoing embodiment, the embodiment of the present application provides a steel surface defect detection system, including:

[0108] The first receiving unit 201 is used to receive steel surface image data;

[0109] The first analysis unit 202 is used to input the received steel surface image data into the pre-trained YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects. The YOLOV9s-Bi FPN-DW-C3 model includes a backbone network layer, a neck network layer, and a detection head layer. The steps for the YOLOV9s-Bi FPN-DW-C3 model to detect steel surface defects include:

[0110] Input the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a set of feature maps containing semantic information at different levels;

[0111] Input the set of feature maps into the neck network layer, perform upsampling matching, normalization preprocessing, weighted fusion, and multi-path integration operations on the cross-scale feature maps, and output the optimized fused feature maps;

[0112] Input the fused feature map into the detection head layer, traverse the fused feature map based on the preset anchor boxes, and predict the defect class probabilities and bounding box offsets for each anchor box; generate a list of defect detection results containing defect type labels and bounding box coordinates according to the defect class probabilities and bounding box offsets.

[0113] In some embodiments, the backbone network layer includes an initial convolutional downsampling module, an ELAN1 module, an AConv module, and a first C3 module. Among them, an SE attention mechanism is added to the first C3 module. The first analysis unit 202 is specifically further configured to:

[0114] Perform spatial compression and channel expansion on the steel surface image through the initial convolutional downsampling module to generate a first-scale feature map containing underlying texture features;

[0115] Input the first-scale feature map into the ELAN1 module, fuse the feature information of different receptive fields through multi-branch heterogeneous convolutional operations, and output a middle-layer feature map with enhanced expression ability;

[0116] Input the middle-layer feature map into the AConv module for secondary spatial downsampling and channel expansion to generate a second-scale feature map containing high-level semantic information;

[0117] Input the second-scale feature map into the first C3 module, and output a set of feature maps that fuse multi-scale semantic information through multi-path convolutional splitting, heterogeneous processing, and channel splicing integration.

[0118] In some embodiments, the neck network layer includes a Bi FPN module and a second C3 module. Among them, an SE attention mechanism is added to the second C3 module. The first analysis unit 202 is specifically further configured to:

[0119] Upsample the P5 / 32 high-level feature map output by the backbone network layer to the P4 / 16 hierarchical size, and perform normalization preprocessing on the P4 / 16 middle-layer feature map;

[0120] Input the upsampled P5 / 32 feature map and the preprocessed P4 / 16 feature map into the Bi FPN module, and perform cross-level semantic fusion on the input feature maps through a dynamic weight allocation mechanism. The dynamic weight allocation mechanism includes: allocating learnable weight parameters to each input feature map, performing weighted summation after non-linear activation and weight normalization processing to generate a fused feature map with enhanced semantics;

[0121] Input the fused feature map into the second C3 module, and output a discriminatively optimized fused feature map through multi-path convolutional splitting, heterogeneous branch processing, and channel dimension splicing integration. The optimized fused feature map contains detailed features and context semantic information of the steel surface defects.

[0122] In some embodiments, the detection head layer includes a Dual DDetect module, and the first analysis unit 202 is further specifically configured to:

[0123] Input the optimized fusion feature map output by the neck network layer into the Dual DDetect module, traverse the multi-scale feature map through a preset anchor box, and initially locate potential defect regions;

[0124] Synchronously perform the following operations in the Dual DDetect module:

[0125] Extract feature information through the first branch convolutional network and predict the class probabilities of cracks, holes, or scratches existing in each anchor box;

[0126] Calculate the bounding box offset of each anchor box through the second branch convolutional network to accurately locate the spatial coordinates of the defect;

[0127] Based on the non-maximum suppression algorithm, screen the class probabilities and bounding box offsets to remove redundant detection boxes;

[0128] Map the filtered detection results to the original image coordinate system to generate a detection list containing defect type labels and their corresponding bounding box coordinates.

[0129] In some embodiments, the training set and test set used during the training of the YOLOV9s-BiFPN-DW-C3 model are generated through the following steps:

[0130] Input the original steel surface defect images into a pre-trained diffusion latent model, and the diffusion latent model generates new defect images according to the probability distribution inside the original steel surface defect images;

[0131] Unify the conversion of the original data and the enhanced data into the YOLO training format, and divide the training set and test set according to a preset ratio.

[0132] It can be understood that the various modules described in this steel surface defect detection system correspond to the respective steps in the steel surface defect detection method described in the reference Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the steel surface defect detection system and the modules included therein, and will not be elaborated here.

[0133] Please refer to Figure 4, based on the inventive concept of a steel surface defect detection method in the foregoing embodiments, an embodiment of the present application provides an electronic device. The electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device includes a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the ROM 302 (Read-Only Memory) or a program loaded from the storage device 308 into the RAM 303 (Random Access Memory). In the RAM 303, various programs and data required for the operation of the electronic device are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output interface (i.e., the I / O interface 305) is also connected to the bus 304.

[0134] Generally, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data.

[0135] Specifically, according to some embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the method of some embodiments of the present application are executed.

[0136] It should be noted that the computer-readable medium described in some embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0137] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0138] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device can: implement the method steps of any of the above technical solutions.

[0139] Computer program code for performing the operations of some embodiments of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0141] The modules described in some embodiments of the present application may be implemented in software or in hardware. The described modules may also be provided in a processor. Among them, the names of these modules do not constitute a limitation to the module itself.

[0142] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0143] Some embodiments of the present application also provide a computer program product, including a computer program which, when executed by a processor, implements any of the above steel surface defect detection methods.

[0144] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A method for detecting surface defects of steel, characterized in that: The following steps are involved: Receiving steel surface image data; The received steel surface image data is input into the pre-trained YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects. The YOLOV9s-BiFPN-DW-C3 model includes a backbone network layer, a neck network layer and a detection head layer. The steps of detecting steel surface defects by the YOLOV9s-BiFPN-DW-C3 model include: Inputting the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a feature map set containing semantic information at different levels; Inputting the feature map set into the neck network layer, performing upsampling matching, normalization preprocessing, weighted fusion and multi-path integration operations on the cross-scale feature maps, and outputting an optimized fused feature map; The fused feature map is input into the detection head layer, and the fused feature map is traversed based on the preset anchor frame to predict the defect category probability and bounding box offset of each anchor frame; according to the defect category probability and bounding box offset, a defect detection result list including defect type labels and bounding box coordinates is generated.

2. The method for detecting surface defects of steel according to claim 1, characterized in that: The backbone network layer includes an initial convolution downsampling module, an ELAN1 module, an AConv module and a first C3 module, wherein the first C3 module is added with an SE attention mechanism, and the multi-scale feature extraction step includes: The steel surface image is spatially compressed and channel expanded by the initial convolution downsampling module to generate a first scale feature map containing underlying texture features; The first scale feature map is input into the ELAN1 module, and the feature information of different receptive fields is fused through multi-branch heterogeneous convolution operation to output a middle-level feature map with enhanced expression ability; Input the middle-level feature map into the AConv module for secondary spatial downsampling and channel expansion to generate a second-scale feature map containing high-level semantic information; The second-scale feature map is input into the first C3 module, and a set of feature maps integrating multi-scale semantic information is output through multi-path convolution splitting, heterogeneous processing and channel splicing integration.

3. The method for detecting surface defects of steel according to claim 2, characterized in that: The neck network layer includes a BiFPN module and a second C3 module, wherein the second C3 module is added with an SE attention mechanism, and the cross-scale feature map processing step includes: Upsampling the high-level feature map output by the backbone network layer to match the size of the middle-level feature map, and performing normalization preprocessing on the middle-level feature map; The upsampled high-level feature map and the preprocessed middle-level feature map are input into the BiFPN module, and the input feature map is semantically fused across levels through a dynamic weight allocation mechanism, wherein the dynamic weight allocation mechanism includes: assigning a learnable weight parameter to each input feature map, performing weighted summation after nonlinear activation and weight normalization processing, and generating a semantically enhanced fused feature map; The fused feature map is input into the second C3 module, and a discriminatively optimized fused feature map is output through multi-path convolution splitting, heterogeneous branch processing and channel dimension splicing and integration. The optimized fused feature map contains detailed features and contextual semantic information of steel surface defects.

4. The method for detecting surface defects of steel according to claim 3, characterized in that: The detection head layer includes a DualDDetect module, and the defect detection result generation step includes: The optimized fusion feature map output by the neck network layer is input into the DualDDetect module, and the multi-scale feature map is traversed through the preset anchor frame to preliminarily locate the potential defect area; The following operations are performed synchronously in the DualDDetect module: The first branch convolutional network extracts feature information and predicts the probability of cracks, holes or scratches in each anchor frame. The bounding box offset of each anchor box is calculated through the second branch convolutional network to accurately locate the spatial coordinates of the defect; The category probabilities and bounding box offsets are screened based on a non-maximum suppression algorithm to remove redundant detection boxes; The filtered detection results are mapped to the original image coordinate system to generate a detection list containing defect type labels and their corresponding bounding box coordinates.

5. The method for detecting surface defects of steel according to any one of claims 1 to 4, characterized in that: The training set and test set used in the training of the YOLOV9s-BiFPN-DW-C3 model are generated by the following steps: Inputting the original steel surface defect image into a pre-trained diffusion latent model, wherein the diffusion latent model generates a new defect image according to the probability distribution inside the original steel surface defect image; The original data and enhanced data are uniformly converted into the YOLO training format, and the training set and test set are divided according to the preset ratio.

6. A steel surface defect detection system, characterized in that: include: A first receiving unit, used for receiving steel surface image data; The first analysis unit is used to input the received steel surface image data into a pre-trained YOLOV9s-BiFPN-DW-C3 model to detect steel surface defects. The YOLOV9s-BiFPN-DW-C3 model includes a backbone network layer, a neck network layer and a detection head layer. The steps of detecting steel surface defects by the YOLOV9s-BiFPN-DW-C3 model include: Inputting the steel surface image data into the backbone network layer for multi-scale feature extraction to generate a feature map set containing semantic information at different levels; Inputting the feature map set into the neck network layer, performing upsampling matching, normalization preprocessing, weighted fusion and multi-path integration operations on the cross-scale feature maps, and outputting an optimized fused feature map; The fused feature map is input into the detection head layer, and the fused feature map is traversed based on the preset anchor frame to predict the defect category probability and bounding box offset of each anchor frame; according to the defect category probability and bounding box offset, a defect detection result list including defect type labels and bounding box coordinates is generated.

7. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Robot inspection method and device for coal conveying line of thermal power plant

    CN121094778A