Steel defect detection method and device, storage medium and computer equipment

Through the specially made steel defect detection model, the fusion of the re-parameter feature multiplexing module and multi-scale feature is solved, and the existing detection methods are cumbersome and time-consuming and low accuracy are achieved, and the rapid and accurate steel defect detection is achieved.

CN120279017AInactive Publication Date: 2025-07-08XIAN ORDNANCE IND TECH IND DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510758536.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing steel defect detection methods are cumbersome and time-consuming, with low detection accuracy, and manual image extraction is prone to errors.

Method used

A special steel defect detection model is adopted, including input network, backbone network, head network and prediction head network. The re-argument feature multiplexing module and multi-scale feature fusion are used to train the steel surface defect detection data set to achieve fast and accurate detection.

Benefits of technology

It improves the efficiency and accuracy of steel defect detection, reduces detection costs, and can capture details and global information at the same time to accurately detect defects of different types and sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279017A_ABST
    Figure CN120279017A_ABST
Patent Text Reader

Abstract

The invention discloses a steel defect detection method and device, a storage medium and computer equipment, relates to the technical field of steel detection, and can improve the detection efficiency and detection precision of steel defects. Comprising the steps that a preset steel defect detection model is obtained, a backbone network comprises a plurality of heavy parameter feature multiplexing modules, a head network comprises a plurality of feature fusion modules of different scales, and a prediction head network comprises a plurality of steel defect detection heads of different scales; inputting the to-be-detected steel image into a preset steel defect detection model, preprocessing the to-be-detected steel image through the input network, performing image feature extraction on the preprocessed steel image through the backbone network, and fusing multiple image features through the head network to obtain a steel defect detection result; and performing steel defect detection on the fusion features of each feature fusion module through the prediction head network, and determining a target defect detection result in a plurality of steel defect detection results. The method is suitable for steel defect detection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steel detection, and in particular to a steel defect detection method, device, storage medium and computer equipment. Background Art

[0002] The detection of steel surface defects is crucial in industrial production, and the tasks include the identification and location of defects.

[0003] Currently, when identifying steel defects, image features are first manually extracted, and then machine learning classification methods such as K-Nearest Neighbor (KNN) are used for classification detection. However, this detection method for steel defects is executed in multiple steps, with a cumbersome process, time-consuming and laborious. At the same time, incorrect extraction may occur when manually extracting image features, resulting in a low detection accuracy for steel defects. Summary of the Invention

[0004] The present invention provides a steel defect detection method, device, storage medium and computer equipment, mainly capable of improving the detection efficiency and detection accuracy of steel defects.

[0005] According to a first aspect of the present invention, there is provided a steel defect detection method, including: Obtaining a steel image to be detected of a target steel; Obtaining a preset steel defect detection model, wherein the preset steel defect detection model includes an input network for preprocessing an image, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes a plurality of reparameterized feature reuse modules, the head network includes a plurality of feature fusion modules of different scales, and the prediction head network includes a plurality of steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; Inputting the steel image to be detected into the preset steel defect detection model, preprocessing the steel image to be detected through the input network to obtain a preprocessed steel image, extracting image features of the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fusing a plurality of image features through the head network to obtain fused features corresponding to each feature fusion module, performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determining a target defect detection result of the target steel among a plurality of steel defect detection results.

[0006] Optionally, the backbone network sequentially includes multiple CBS (Convolution + BatchNorm + SiLU) modules, a first reparameterized feature reuse module, and multiple feature extraction modules. Each feature extraction module includes a first MPConv (MaxPooling and Conv) module and a reparameterized feature reuse module; Performing image feature extraction on the preprocessed steel material image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, including: Inputting the preprocessed steel material image into the backbone network, and gradually performing image feature extraction on the preprocessed steel material image through multiple CBS modules to obtain intermediate image features output by the last CBS module. Inputting the intermediate image features into the first reparameterized feature reuse module for feature extraction to obtain first reparameterized image features output by the first reparameterized feature reuse module; Sequentially inputting the first reparameterized image features into multiple feature extraction modules to perform feature extraction through the first MPConv module and generate reparameterized image features through the reparameterized feature reuse module, to obtain second reparameterized image features output by the last feature extraction module.

[0007] Optionally, any one of the reparameterized feature reuse modules in the first reparameterized feature reuse module and the multiple feature extraction modules includes multiple structure reparameterization processing layers. The feature extraction process of any one of the reparameterized feature reuse modules includes: Processing the input features of any one of the reparameterized feature reuse modules through each of the structure reparameterization processing layers to obtain output results of each of the structure reparameterization processing layers, and performing fusion processing on the output results of each of the structure reparameterization processing layers and the input features to obtain the output results of any one of the reparameterized feature reuse modules; The first MPConv module includes a first convolutional layer, a slicing layer, a second convolutional layer, a max pooling layer, and a third convolutional layer. The feature extraction process of the first MPConv module includes: Processing the input features of the first MPConv module through the first convolutional layer and the max pooling layer respectively, performing adjacent downsampling processing on the output results of the first convolutional layer through the slicing layer to obtain a preset number of sliced image features output by the slicing layer, performing fusion processing on each of the sliced image features through the second convolutional layer, and processing the output results of the max pooling layer through the third convolutional layer; Performing fusion processing on the output results of the second convolutional layer and the output results of the third convolutional layer to obtain the output results of the first MPConv module.

[0008] Optionally, the head network includes a first feature fusion module, a second feature fusion module, a third feature fusion module, and a fourth feature fusion module; The first feature fusion module includes a first CBS_1EMA (CBS with an added EMA (Exponential Moving Average Attention) attention mechanism module), a first Concat (concatenation) module, a first ELAN-H (Enhanced Local Attention Network with Hierarchical Structure) module, and a first RepConv (Re-parameterized Convolution) module connected in sequence; The second feature fusion module includes a second CBS_1EMA module, a second Concat module, a second ELAN-H module, a third Concat module, a third ELAN-H module, and a second RepConv module connected in sequence; The third feature fusion module includes a third CBS_1EMA module, a fourth Concat module, a fourth ELAN-H module, a fifth Concat module, a fifth ELAN-H module, and a third RepConv module connected in sequence; The fourth feature fusion module includes an SPPCSPC_EMA (SPPCSPC (Spatial Pyramid Pooling - Cross Stage Partial Connections) with an added EMA attention mechanism) module, a sixth Concat module, a sixth ELAN-H module, and a fourth RepConv module connected in sequence; The head network further includes a plurality of feature upsampling modules and a plurality of second MPConv modules. Each feature upsampling module includes a CBS_1 module and an UP (Upsample) module; The output ends of the SPPCSPC_EMA module, the fourth ELAN-H module, and the second ELAN-H module are respectively connected to the input ends of the corresponding feature upsampling modules, and the input ends of the fourth Concat module, the second Concat module, and the first Concat module are respectively connected to the output ends of the corresponding feature upsampling modules; The output ends of the first ELAN-H module, the third ELAN-H module, and the fifth ELAN-H module are respectively connected to the input ends of the corresponding second MPConv module, and the input ends of the third Concat module, the fifth Concat module, and the sixth Concat module are respectively connected to the output ends of the corresponding second MPConv module.

[0009] Optionally, the EMA in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module all includes a Groups (group convolution) layer, a 1×1 convolution layer, a 3×3 convolution layer, a 5×5 convolution layer, a first Sigmoid (activation function) layer, a second Sigmoid layer, a first Re_weight (feature reweighting) layer, a GroupNorm (group normalization) layer, a first Avg_Pool (average pooling) layer, a first Softmax layer, a first Matmu1 (feature transformation) layer, a second Avg_Pool layer, a second Softmax layer, a second Matmu1 layer, a third Avg_Pool layer, a third Softmax layer, a third Matmu1 layer, a third Sigmoid layer, a second Re_weight layer, and an Output (output) layer; The working process of any target module in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module includes: Feature decomposition is performed on the input features of any target module through the Groups layer to respectively obtain a first decomposed and reparameterized image feature, a second decomposed and reparameterized image feature, a third decomposed and reparameterized image feature, a fourth decomposed and reparameterized image feature, and a fifth decomposed and reparameterized image feature. The second decomposed and reparameterized image feature and the third decomposed and reparameterized image feature are respectively input into the 1×1 convolution layer for processing to obtain a first output result and a second output result corresponding to the output of the 1×1 convolution layer. The fourth decomposed and reparameterized image feature is processed through the 3×3 convolution layer to obtain the output result of the 3×3 convolution layer, and the fifth decomposed and reparameterized image feature is processed through the 5×5 convolution layer to obtain the output result of the 5×5 convolution layer; The first output result and the second output result are respectively processed by the first Sigmoid layer and the second Sigmoid layer to obtain the output result of the first Sigmoid layer and the output result of the second Sigmoid layer. The first decomposition re-parameterized image feature, the output result of the first Sigmoid layer, and the output result of the second Sigmoid layer are processed by the first Re_weight layer to obtain the output result of the first Re_weight layer. The output result of the first Re_weight layer is processed by the GroupNorm layer to obtain the output result of the GroupNorm layer. The output result of the GroupNorm layer is successively processed by the first Avg_Pool layer and the first Softmax layer to obtain the output result of the first Softmax layer. The output result of the 3×3 convolutional layer is successively processed by the second Avg_Pool layer and the second Softmax layer to obtain the output result of the second Softmax layer. The output result of the 5×5 convolutional layer is successively processed by the third Avg_Pool layer and the third Softmax layer to obtain the output result of the third Softmax layer; The output result of the first Softmax layer, the output result of the 3×3 convolutional layer, and the output result of the 5×5 convolutional layer are processed by the first Matmu1 layer to obtain the output result of the first Matmu1 layer. The output result of the second Softmax layer and the output result of the GroupNorm layer are processed by the second Matmu1 layer to obtain the output result of the second Matmu1 layer. The output result of the third Softmax layer and the output result of the GroupNorm layer are processed by the third Matmu1 layer to obtain the output result of the third Matmu1 layer; The output results of the first Matmu1 layer, the second Matmu1 layer, and the third Matmu1 layer are fused to obtain a fused result, and the fused result is processed by the third Sigmoid layer to obtain the output result of the third Sigmoid layer. The output result of the third Sigmoid layer and the first decomposition re-parameterized image feature are processed by the second Re_weight layer to obtain the output result of the second Re_weight layer, and the output result of the second Re_weight layer is output through the Output layer as the processing result of any target module.

[0010] Optionally, the prediction head network includes a first steel defect detection head, a second steel defect detection head, a third steel defect detection head, and a fourth steel defect detection head; Performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain the steel defect detection results corresponding to each steel defect detection head, including: Respectively inputting the fused features of each feature fusion module into the first steel defect detection head, the second steel defect detection head, the third steel defect detection head, and the fourth steel defect detection head for defect detection, and respectively obtaining the steel defect detection results corresponding to each steel defect detection head.

[0011] Optionally, before obtaining the preset steel defect detection model, the method further includes: Constructing a preset initial steel defect detection model and obtaining a steel surface defect detection data set as a sample image set, where the sample image set includes multiple sample images with annotation information, and the annotation information includes an annotation box and the annotation type of the steel defect within the annotation box, and the annotation type includes oxide scale, inclusion, crack, scratch, plaque, pitted surface; Dividing the sample image set into a training set, a validation set, and a test set according to a preset ratio, training the preset initial steel defect detection model using the training set, validating the trained preset initial steel defect detection model using the validation set, adjusting the model parameters of the trained preset initial steel defect detection model according to the validation result, and testing the preset initial steel defect detection model after parameter adjustment using the test set, and using the preset initial steel defect detection model after parameter adjustment that meets the test conditions as the preset steel defect detection model.

[0012] According to a second aspect of the present invention, there is provided a steel defect detection device, including: An image acquisition unit for acquiring a to-be-detected steel image of a target steel; A model acquisition unit for acquiring a preset steel defect detection model, where the preset steel defect detection model includes an input network for preprocessing an image, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection, the backbone network includes multiple reparameterized feature reuse modules, the head network includes multiple feature fusion modules of different scales, and the prediction head network includes multiple steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; The defect detection unit is used to input the steel image to be detected into the preset steel defect detection model, preprocess the steel image to be detected through the input network to obtain a preprocessed steel image, extract image features from the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fuse multiple image features through the head network to obtain fused features corresponding to each feature fusion module, perform steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determine the target defect detection result of the target steel from multiple steel defect detection results.

[0013] According to the third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned steel defect detection method is implemented.

[0014] According to the fourth aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above-mentioned steel defect detection method is implemented.

[0015] According to a steel defect detection method, device, storage medium and computer device provided by the present invention, compared with the current method of first manually extracting image features and then using machine learning classification methods such as K-nearest neighbor (KNN) for classification detection, the present invention adopts a special model architecture, including an input network for image preprocessing, a backbone network containing reparameterized feature reuse modules, a head network for multi-scale feature fusion, and a prediction head network for multi-scale steel defect detection, and completes training using a steel surface defect detection data set, and realizes fast and accurate detection of steel defects through a special model. Further, the backbone network of the preset steel defect detection model of the present invention introduces a reparameterized feature reuse module. Since the reparameterized feature reuse module switches the operations originally distributed in different feature spaces waiting for fusion in the inference stage to a single inference in the weight space, it is hardware-friendly and greatly reduces the detection cost of steel defects. Compared with other model structures, the reparameterized feature reuse module has no redundant branch topologies, and each layer in the network structure only takes the output of its only preceding layer as input and feeds the output to its only succeeding layer, thereby being able to improve the inference speed of the model and thus improve the detection efficiency of steel defects. The multi-scale feature fusion and detection mechanism enables the model to capture both detailed information and global information in the steel image simultaneously, thereby more accurately detecting different types and sizes of steel defects, and the present invention can complete the entire process of steel defect detection through a single model, saving time and improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 A flowchart of a steel defect detection method provided by an embodiment of the present invention is shown; Figure 2 A network architecture diagram of a preset steel defect detection model provided by an embodiment of the present invention is shown; Figure 3 A network architecture diagram of a reparameterized feature reuse module provided by an embodiment of the present invention is shown; Figure 4A A network architecture diagram of an original first MPConv module provided by an embodiment of the present invention is shown; Figure 4B A network architecture diagram of an improved first MPConv module provided by an embodiment of the present invention is shown; Figure 5 A network architecture diagram of an EMA attention mechanism provided by an embodiment of the present invention is shown; Figure 6 A schematic diagram of the slicing operation of a slicing layer provided by an embodiment of the present invention is shown; Figure 7 A flowchart of another steel defect detection method provided by an embodiment of the present invention is shown; Figure 8 A schematic structural diagram of a steel defect detection device provided by an embodiment of the present invention is shown; Figure 9 A schematic structural diagram of another steel defect detection device provided by an embodiment of the present invention is shown; Figure 10 A schematic physical structure diagram of a computer device provided by an embodiment of the present invention is shown. Detailed implementation manners

[0017] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0018] Currently, the method of first manually extracting image features and then using machine learning classification methods such as K-Nearest Neighbor (KNN) for steel defect detection is cumbersome, time-consuming and laborious. At the same time, manual extraction of image features may result in extraction errors, thus leading to low detection accuracy of steel defects.

[0019] To solve the above problems, an embodiment of the present invention provides a steel defect detection method, as Figure 1 shown, the method includes: 101. Obtain the steel image to be detected of the target steel.

[0020] Specifically, the steel image to be detected of the target steel can be collected by sensors such as camera devices, and the entire appearance of the target steel is covered in the steel image to be detected.

[0021] 102. Obtain a preset steel defect detection model, where the preset steel defect detection model includes an input network for preprocessing images, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes multiple reparameterized feature reuse modules, the head network includes multiple feature fusion modules of different scales, and the prediction head network includes multiple steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set.

[0022] For the embodiments of the present invention, the preset steel defect detection model can be constructed based on the improvement of the YOLOv7 model. The preset steel defect detection model consists of four parts: an input network, a backbone network, a head network, and a prediction head network. The backbone network includes multiple reparameterized feature reuse modules. Each layer in these reparameterized feature reuse modules only takes the output of its only previous layer as input and feeds the output to its only subsequent layer, without repeated stacking of unit structures with cross-layer operations such as frequent skip connections, which can improve the extraction efficiency of image features. The head network includes multiple feature fusion modules, which are responsible for fusing image features of different scales to simultaneously capture the detailed information and global information in the steel image, so as to be able to detect different types and sizes of steel defects. The prediction head network includes multiple steel defect detection heads, and each steel defect detection head performs object detection on the image features of a specific scale to ensure accurate identification of steel defects from multiple scales. The data set used for training the model is steel images of different scenarios, different types of steel, corresponding to different defect types and different defect sizes, which helps the preset steel defect detection model to detect different steel defects.

[0023] Specifically, as Figure 2 shown, the network architecture of the preset steel defect detection model in the embodiments of the present invention is presented. The backbone network in this network architecture sequentially includes multiple CBS (Convolution + BatchNorm + SiLU, convolutional layer + batch normalization layer + activation function layer) modules, the first reparameterized feature reuse module ( Figure 2 the first reparameterized feature reuse module in the top-down order), multiple feature extraction modules, and each feature extraction module includes a first MPConv (MaxPooling and Conv, max pooling and convolution) module (Figure 2 the first MPConv module, the second MPConv module, and the third MPConv module in the top-down order) and a reparameterized feature reuse module ( Figure 2The second, third, and fourth reparameterized feature reuse modules in the top-down order); Any one of the first reparameterized feature reuse module and the reparameterized feature reuse modules in the multiple feature extraction modules includes multiple structure reparameterization processing layers; The first MPConv module in each feature extraction module includes a first convolutional layer, a slicing layer, a second convolutional layer, a max pooling layer, and a third convolutional layer. The head network includes a first feature fusion module, a second feature fusion module, a third feature fusion module, and a fourth feature fusion module; The first feature fusion module includes a first CBS_1EMA (CBS with the EMA (Exponential Moving Average Attention) attention mechanism module added) module, a first Concat (concatenation) module, a first ELAN-H (Enhanced Local Attention Network with Hierarchical Structure) module, and a first RepConv (Re-parameterized Convolution) module connected in sequence; The second feature fusion module includes a second CBS_1EMA module, a second Concat module, a second ELAN-H module, a third Concat module, a third ELAN-H module, and a second RepConv module connected in sequence; The third feature fusion module includes a third CBS_1EMA module, a fourth Concat module, a fourth ELAN-H module, a fifth Concat module, a fifth ELAN-H module, and a third RepConv module connected in sequence; The fourth feature fusion module includes an SPPCSPC_EMA (SPPCSPC (Spatial Pyramid Pooling - Cross Stage Partial Connections) with the EMA attention mechanism added) module, a sixth Concat module, a sixth ELAN-H module, and a fourth RepConv module connected in sequence; The head network further includes multiple feature upsampling modules and multiple second MPConv modules, and each feature upsampling module includes a CBS_1 module and an UP (Upsample) module; The output ends of the SPPCSPC_EMA module, the fourth ELAN-H module, and the second ELAN-H module are respectively connected to the input ends of the corresponding feature upsampling modules, and the input ends of the fourth Concat module, the second Concat module, and the first Concat module are respectively connected to the output ends of the corresponding feature upsampling modules;The output ends of the first ELAN-H module, the third ELAN-H module, and the fifth ELAN-H module are respectively connected to the input ends of the corresponding second MPConv modules, and the input ends of the third Concat module, the fifth Concat module, and the sixth Concat module are respectively connected to the output ends of the corresponding second MPConv modules; the EMA in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module all includes an Input layer, a Groups (group convolution) layer, an X Avg_Pool layer, a Y Avg_Pool layer, a 1×1 convolution layer (Contact+Conv(1×1)), a 3×3 convolution layer (Conv(3×3)), a 5×5 convolution layer (Conv(5×5)), a first Sigmoid (activation function) layer, a second Sigmoid layer, a first Re_weight (feature reweighting) layer, a GroupNorm (group normalization) layer, a first Avg_Pool (average pooling) layer, a first Softmax layer, a first Matmu1 (feature transformation) layer, a second Avg_Pool layer, a second Softmax layer, a second Matmu1 layer, a third Avg_Pool layer, a third Softmax layer, a third Matmu1 layer, a third Sigmoid layer, a second Re_weight layer, and an Output layer. The prediction head network includes a first steel defect detection head, a second steel defect detection head, a third steel defect detection head, and a fourth steel defect detection head.;

[0024] It should be noted that the first XX module, the second XX module,..., the Nth XX module, and the first XX layer, the second XX layer,..., the Nth XX layer in the embodiments of the present invention are named by distinguishing the same-name modules or layers in the drawings in the order from top to bottom and from left to right.

[0025] In the above embodiments, the multiple CBS modules in the backbone network include multiple CBS_2 modules and multiple CBS_3 modules. The multiple CBS modules all include a convolutional layer, a batch normalization layer, and an activation function. Through the multiple CBS modules, different convolutional kernels and strides can be used at different resolutions to extract richer image features. The reparameterized feature reuse module introduces a structure reparameterization network structure based on the original ELAN module. The structure reparameterization network structure refers to designing a structure that can merge multiple branches into one. During the training phase, multiple branches are used, which is equivalent to expanding the capacity and complexity of the network to extract more effective features. During the inference phase, by fusing the multi-path structure into a single structure, it is possible to inherit the complex parameters while leveraging the good adaptability of the single-path structure to parallel computing devices for fast inference. The reparameterized network structure fuses the multiple branches during training into a single convolutional layer during inference, switching the operations that were originally distributed in different feature spaces waiting to be fused during the inference phase to a single inference performed by the fusion module in the weight space, which is hardware-friendly and greatly reduces the cost of steel defect detection. Compared with the original model structure, the reparameterized feature reuse module newly introduced in the backbone network of the embodiments of the present invention has no redundant branch topologies. Each layer in the reparameterized feature reuse module only takes the output of its unique previous layer as input and feeds the output to its unique subsequent layer. As Figure 3 shown, the main body of the reparameterized feature reuse module only uses a 3×3 depth convolutional layer and ReLU (Rectified Linear Unit). The repeated stacking of this unit structure without cross-layer operations such as frequent skip connections helps to improve the processing speed. The embodiments of the present invention have improved the first MPConv module based on the original model structure. The network structure of the original first MPConv module is as Figure 4A shown (where Stride is the stride). The network structure of the improved first MPConv module is as Figure 4B shown. Compared with the original first MPConv module, a slicing layer is introduced in the improved first MPConv module. The slicing layer can split the input feature map along the channel or spatial dimension to generate multiple sub-feature maps, and each sub-feature map can be processed independently to extract features at different scales or different semantic levels. This multi-scale feature extraction ability helps the model capture richer context information and improve the detection ability for complex steel defects and small steel defects.

[0026] In the above embodiments, the head network includes a first feature fusion module, a second feature fusion module, a third feature fusion module, and a fourth feature fusion module. The head network in the embodiments of the present invention has added a large defect feature fusion module, that is, the fourth feature fusion module, compared with the head network of the original model, as Figure 2The green box part in it can detect steel defects on a larger scale through the large defect feature fusion module. Compared with the head network of the original model, the first feature fusion module, the second feature fusion module, and the third feature fusion module include improved CBS modules, and the fourth feature fusion module contains an improved SPPCSPC module. The specific improvement method is to add an EMA attention mechanism to the CBS module and the SPPCSPC module respectively, so as to obtain the CBS_1EMA module and the SPPCSPC_EMA module in the head network. As Figure 5 shown, the added EMA attention mechanism not only includes a 1×1 convolutional layer, but also introduces 2 parallel subnets (3×3 convolutional layer and 5×5 convolutional layer) to capture multi-scale feature relationships. Compared with a single 1×1 convolutional layer, the additional 2 parallel subnets can perform convolutional operations on a larger receptive field and are more suitable for capturing larger-scale patterns or texture information, thereby further improving the detection accuracy of steel defects. At the same time, if the EMA attention mechanism is added to the front backbone network, the spatial feature map has a large size but lacks the number of channels and insufficient feature generalization. If the EMA attention mechanism is added to the later prediction head network, it will cause too many channels and lead to overfitting problems, affecting the defect decision of the prediction head network. Therefore, in the embodiments of the present invention, the EMA attention mechanism is added to the CBS and SPPCSPC parts of the head network in the middle part, so that the deeper features and the shallower features are not only fused through downsampling, but directly use the EMA attention mechanism for feature fusion. Further, compared with the head network of the original model, the head network in the embodiments of the present invention also introduces a RepConv module and an ELAN-H module. The RepConv module is a model reparameterization module that can combine multiple computational modules into one during the inference stage, improving the efficiency and performance of the model. The RepConv module reparameterizes the parameters of the branches to the main branch during inference, thereby reducing the amount of computation and memory consumption. The ELAN-H module is a variant of the ELAN module, mainly used for feature extraction and channel number adjustment, and contains multiple parallel branches. Each branch can include convolutional layers, batch normalization layers, activation layers, etc. These branches work together during training to enhance the representation ability of the model.

[0027] In the above embodiment, the prediction head network includes multiple steel defect detection heads. On the basis of the existing detection heads of YOLOv7, a large defect feature fusion module is added, that is, the large target steel defect detection head corresponding to the fourth feature fusion module, to detect defects of larger sizes.

[0028] 103. Input the steel image to be detected into a preset steel defect detection model. Preprocess the steel image to be detected through the input network to obtain a preprocessed steel image. Extract image features from the preprocessed steel image through the backbone network to obtain the image features corresponding to each reparameterized feature reuse module. Fuse multiple image features through the head network to obtain the fused features corresponding to each feature fusion module. Detect steel defects for the fused features of each feature fusion module through the prediction head network to obtain the steel defect detection results corresponding to each steel defect detection head, and determine the target defect detection result of the target steel from multiple steel defect detection results.

[0029] For the embodiments of the present invention, in order to detect steel defects, first, the steel image to be detected needs to be input into a preset steel defect detection model. Based on this, the specific preprocessing method includes: inputting multiple steel images to be detected of the target steel into the input network, processing each steel image to be detected into a preset-size steel image through the image size processing layer respectively, performing a scaling operation on multiple preset-size steel images through the image enhancement layer, and combining the scaled multiple preset-size steel images into a steel image with anchor boxes to be processed. Perform adaptive anchor box processing on the steel image with anchor boxes to be processed through the adaptive anchor box layer to obtain a preprocessed steel image with initial steel defect detection boxes. That is, image size processing, image quality enhancement, and adaptive anchor boxes are adopted for the input image. The image size processing scales images with different lengths and widths to corresponding sizes by adding a small amount of black edges. Data enhancement: operations such as randomly scaling and arranging multiple images are performed to combine them into 1 image. The adaptive anchor box calculation calculates the gap between the predicted box and the true box and then updates it in reverse until the most suitable anchor box value is obtained and the steel image is subjected to adaptive anchor box processing to obtain the preprocessed steel image.

[0030] Furthermore, extract features from the preprocessed steel image through the backbone network in the preset steel defect detection model. Among them, the feature extraction process includes: inputting the preprocessed steel image into the backbone network, gradually extracting image features from the preprocessed steel image through multiple CBS modules to obtain the intermediate image features output by the last CBS module, inputting the intermediate image features into the first reparameterized feature reuse module for feature extraction to obtain the first reparameterized image features output by the first reparameterized feature reuse module; sequentially inputting the first reparameterized image features into multiple feature extraction modules to perform feature extraction through the first MPConv module and generate reparameterized image features through the reparameterized feature reuse module to obtain the second reparameterized image features output by the last feature extraction module.

[0031] Among them, the multiple feature extraction modules include a first feature extraction module, a second feature extraction module, and a third feature extraction module. Specifically, the preprocessed steel material image is sequentially input into multiple CBS modules for step-by-step feature extraction to obtain the intermediate image features output by the last CBS module. The intermediate image features are input into the first parameter-sharing feature reuse module for processing to obtain the first parameter-sharing image features. The first parameter-sharing image features are input into the first feature extraction module, and are successively subjected to feature extraction through the first MPConv module and the parameter-sharing feature reuse module in the first feature extraction module to obtain the second parameter-sharing image features output by the parameter-sharing feature reuse module in this module. The second parameter-sharing image features are input into the second feature extraction module, and are successively subjected to feature extraction through the first MPConv module and the parameter-sharing feature reuse module in the second feature extraction module to obtain the third parameter-sharing image features output by the parameter-sharing feature reuse module in this module. The third parameter-sharing image features are input into the third feature extraction module, and are successively subjected to feature extraction through the first MPConv module and the parameter-sharing feature reuse module in the third feature extraction module to obtain the fourth parameter-sharing image features output by the parameter-sharing feature reuse module in this module. Among them, the working process of any one of the parameter-sharing feature reuse modules in each parameter-sharing feature reuse module includes: processing the input features of any one of the parameter-sharing feature reuse modules through each of the structural reparameterization processing layers to obtain the output results of each of the structural reparameterization processing layers, and fusing the output results of each of the structural reparameterization processing layers with the input features to obtain the output results of any one of the parameter-sharing feature reuse modules. The working process of any one of the first MPConv modules in each of the first MPConv modules includes: processing the input features of the first MPConv module through the first convolutional layer and the maximum pooling layer respectively, performing adjacent downsampling processing on the output result of the first convolutional layer through the slicing layer to obtain a preset number of sliced image features output by the slicing layer, performing fusion processing on each of the sliced image features through the second convolutional layer, and processing the output result of the maximum pooling layer through the third convolutional layer; fusing the output result of the second convolutional layer and the output result of the third convolutional layer to obtain the output result of the first MPConv module.

[0032] Among them, the preset number is set according to actual requirements. Specifically, the parameter-sharing feature reuse module realizes feature extraction through the following function:

[0033] Among them, is the output result of the parameter-sharing feature reuse module, represents the feature fusion operation, is the input feature of the parameter-sharing feature reuse module, Represents the output result of the i-th structure reparameterization processing layer. , is the total number of structure reparameterization processing layers. The feature fusion process of the reparameterized feature reuse module is carried out in the weight space and does not introduce any additional inference cost. Therefore, the finally obtained structure is more efficient than the cascaded structure and is suitable for running on high-parallel hardware devices.

[0034] In the embodiment of the present invention, the processing process of the first MPConv module is as follows: During the downsampling process, in order to avoid feature loss caused to the small target network by strided convolution, the embodiment of the present invention introduces a slicing method for processing. The image features are downsampled through the slicing operation of the slicing layer. The principle of the slicing operation is as Figure 6 shown. The slicing operation is performed on the image features. For example, the slicing operation is to take a value every other pixel in an image feature. By a method similar to nearest neighbor downsampling, four image features that are half the size are obtained. The scales of the image features (feature maps) are the same, and the features complement each other. The original information is still retained after slicing. The dimension information is increased in the channel space by increasing the number of channels, expanding the input channels. According to the number of convolutional kernels in the original network, finally, convolutional kernels with the same number are used to perform channel information fusion on the obtained new feature maps, and finally a two-fold downsampled feature map that retains the original feature information is obtained. For example, for a feature map with an input size of C×C and 32 channels in the network, the original network downsamples the feature map through a 3×3 convolutional layer with 128 channels and a stride of 2. In the embodiment of the present invention, the slicing layer is added to separate the feature map into 128 sub-maps of (C / 2)×(C / 2). After the slicing operation, the features can obtain features of size (C / 2)×(C / 2) through a 1×1 convolution with a stride of 1. The downsampling process improves the comprehensiveness of feature extraction.

[0035] Further, after extracting the image features using the backbone network, it is necessary to perform feature fusion on multiple image features through the head network in the preset steel defect detection model. Among them, the process of feature fusion includes: respectively inputting the first reparameterized image feature, the second reparameterized image feature, the third reparameterized image feature, and the fourth reparameterized image feature into the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module, processing the output result of the SPPCSPC_EMA module through the third CBS_1 module and the third UP module in sequence to obtain the output result of the third UP module, performing splicing processing on the output result of the third UP module and the output result of the third CBS_1EMA module through the fourth Concat module to obtain the output result of the fourth Concat module, processing the output result of the fourth Concat module through the fourth ELAN-H module to obtain the output result of the fourth ELAN-H module, processing the output result of the fourth ELAN-H module through the second CBS_1 module and the second UP module in sequence to obtain the output result of the second UP module, performing splicing processing on the output result of the second CBS_1EMA module and the output result of the second UP module through the second Concat module to obtain the output result of the second Concat module, processing the output result of the second Concat module through the second ELAN-H module to obtain the output result of the second ELAN-H module, processing the output result of the second ELAN-H module through the first CBS_1 module and the first UP module in sequence to obtain the output result corresponding to the first UP module, performing splicing processing on the output result of the first CBS_1EMA module and the output result of the first UP module through the first Concat module to obtain the output result of the first Concat module, processing the output result of the first Concat module through the first ELAN-H module to obtain the output result of the first ELAN-H module, and processing the output result of the first ELAN-H module through the first RepConv module to obtain the first fusion feature output by the first RepConv module;The output result of the first ELAN-H module is processed through the corresponding second MPConv module to obtain the output result of the corresponding second MPConv module. The output result of the corresponding second MPConv module and the output result of the second ELAN-H module are concatenated through the third Concat module to obtain the output result of the third Concat module. The output result of the third Concat module is processed through the third ELAN-H module to obtain the output result of the third ELAN-H module. The output result of the third ELAN-H module is processed through the second RepConv module to obtain the second fused feature output by the second RepConv module; The output result of the third ELAN-H module is processed through the corresponding second MPConv module to obtain the output result of the corresponding second MPConv module. The output result of the fourth ELAN-H module and the output result of the corresponding second MPConv module are concatenated through the fifth Concat module to obtain the output result of the fifth Concat module. The output result of the fifth Concat module is processed through the fifth ELAN-H module to obtain the output result of the fifth ELAN-H module. The output result of the fifth ELAN-H module is processed through the third RepConv module to obtain the third fused feature output by the third RepConv module; The output result of the fifth ELAN-H module is processed through the corresponding second MPConv module to obtain the output result of the corresponding second MPConv module. The output result of the corresponding second MPConv module and the output result of the SPPCSPC_EMA module are concatenated through the sixth Concat module to obtain the output result of the sixth Concat module. The output result of the sixth Concat module is processed through the sixth ELAN-H module to obtain the output result of the sixth ELAN-H module. The output result of the sixth ELAN-H module is processed through the fourth RepConv module to obtain the fourth fused feature output by the fourth RepConv module.;

[0036] In the above embodiments, the working process of any target module (EMA module) in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module includes: inputting features through the Input input layer, performing feature decomposition on the input features through the Groups layer to respectively obtain the first decomposed and reparameterized image feature, the second decomposed and reparameterized image feature, the third decomposed and reparameterized image feature, the fourth decomposed and reparameterized image feature, and the fifth decomposed and reparameterized image feature, respectively inputting the second decomposed and reparameterized image feature and the third decomposed and reparameterized image feature into the X Avg_Pool layer and the Y Avg_Pool layer, and jointly inputting the output results of the X Avg_Pool layer and the Y Avg_Pool layer into the 1×1 convolutional layer for processing to obtain the first output result and the second output result output by the 1×1 convolutional layer, processing the fourth decomposed and reparameterized image feature through the 3×3 convolutional layer to obtain the output result of the 3×3 convolutional layer, processing the fifth decomposed and reparameterized image feature through the 5×5 convolutional layer to obtain the output result of the 5×5 convolutional layer; respectively processing the first output result and the second output result through the first Sigmoid layer and the second Sigmoid layer to obtain the output result of the first Sigmoid layer and the output result of the second Sigmoid layer, processing the first decomposed and reparameterized image feature, the output result of the first Sigmoid layer, and the output result of the second Sigmoid layer through the first Re_weight layer to obtain the output result of the first Re_weight layer, processing the output result of the first Re_weight layer through the GroupNorm layer to obtain the output result of the GroupNorm layer, successively processing the output result of the GroupNorm layer through the first Avg_Pool layer and the first Softmax layer to obtain the output result of the first Softmax layer, successively processing the output result of the 3×3 convolutional layer through the second Avg_Pool layer and the second Softmax layer to obtain the output result of the second Softmax layer, successively processing the output result of the 5×5 convolutional layer through the third Avg_Pool layer and the third Softmax layer to obtain the output result of the third Softmax layer;The output result of the first Softmax layer, the output result of the 3×3 convolutional layer, and the output result of the 5×5 convolutional layer are processed by the first Matmu1 layer to obtain the output result of the first Matmu1 layer. The output result of the second Softmax layer and the output result of the GroupNorm layer are processed by the second Matmu1 layer to obtain the output result of the second Matmu1 layer. The output result of the third Softmax layer and the output result of the GroupNorm layer are processed by the third Matmu1 layer to obtain the output result of the third Matmu1 layer. The output results of the first Matmu1 layer, the second Matmu1 layer, and the third Matmu1 layer are fused to obtain a fused result, and the fused result is processed by the third Sigmoid layer to obtain the output result of the third Sigmoid layer. The output result of the third Sigmoid layer and the first decomposed reparameterized image feature are processed by the second Re_weight layer to obtain the output result of the second Re_weight layer, and the output result of the second Re_weight layer is output through the Output layer as the processing result of any target module.

[0037] Specifically, the shared component of the 1×1 convolution is selected from the CA (Coordinate Attention) module of the original model and is called the 1×1 convolutional layer. To aggregate multi-scale relationships, two parallel subnets, a 3×3 convolutional layer and a 5×5 convolutional layer, are added. Different convolutional layers are responsible for processing different features or information. The 3×3 convolutional layer is responsible for capturing fine-grained local features, while the 5×5 convolutional layer is responsible for capturing context information in a larger area. For the input feature , the target module divides it into G sub-features . To collect multi-scale attention weights, the processed input feature is encoded by two 1×1 convolutional layers, one 3×3 convolutional layer, and one 5×5 convolutional layer. After the output of the 1×1 convolutional layer is decomposed into two vectors, two non-linear Sigmoid functions are used to fit the two-state distribution of the two-dimensional plane of the linear convolution. Then, global pooling in the two-dimensional plane (2D) is used to encode global spatial information, as shown in the following formula:

[0038] where represents the average value of channel c, H and W represent the height and width of the image respectively, represents the position on channel c The pixel values are fitted by the non - linear function Softmax after 2D pooling for efficient calculation. The obtained outputs are respectively subjected to matrix dot - product operations with two parallel convolutional layers of 3×3 and 5×5, and the first two spatial attention matrices are obtained. In addition, the 3×3 and 5×5 convolutional layers respectively perform the same pooling and Softmax operations as the 1×1 convolutional layer to encode global information, and the 1×1 convolutional layer is transformed into the corresponding dimensional shape before the joint activation mechanism of channel features. After the dot - product operation, the last two spatial attention matrices are obtained. The four matrices are added, and finally, an attention weight matrix is formed through the Sigmoid function. The input matrix is multiplied by the weight matrix to obtain the output features of the target module.

[0039] Furthermore, after using the head network for image feature fusion, it is necessary to perform steel defect detection on each fused feature through the prediction head network in the preset steel defect detection model. Among them, the process of steel defect detection includes: respectively inputting the fused features of each feature fusion module into the first steel defect detection head, the second steel defect detection head, the third steel defect detection head, and the fourth steel defect detection head for defect detection, and respectively obtaining the steel defect detection results corresponding to each steel defect detection head.

[0040] Specifically, the fused features of each feature fusion module are respectively input into the first steel defect detection head, the second steel defect detection head, the third steel defect detection head, and the fourth steel defect detection head for defect detection, and respectively obtain the first steel defect detection result corresponding to the first steel defect detection head, the second steel defect detection result corresponding to the second steel defect detection head, the third steel defect detection result corresponding to the third steel defect detection head, and the fourth steel defect detection result corresponding to the fourth steel defect detection head. Each steel defect detection result includes information such as defect location, category, and confidence. Finally, the detection result with the highest confidence is selected as the target defect detection result of the target steel.

[0041] A method for detecting steel defects provided by the present invention, compared with the current method of first manually extracting image features and then using machine learning classification methods such as K-Nearest Neighbor (KNN) for classification detection, the present invention adopts a special model architecture, including an input network for image preprocessing, a backbone network containing a reparameterized feature reuse module, a head network for multi-scale feature fusion, and a prediction head network for multi-scale steel defect detection, and completes training using a steel surface defect detection dataset, achieving fast and accurate detection of steel defects through a special model. Further, the backbone network of the preset steel defect detection model of the present invention introduces a reparameterized feature reuse module. Since the reparameterized feature reuse module switches the operations originally distributed in different feature spaces waiting for fusion during the inference stage to a single inference performed by the fusion module in the weight space, it is friendly to hardware and greatly reduces the detection cost of steel defects. Compared with other model structures, the reparameterized feature reuse module has no redundant branch topologies, and each layer in the network structure only takes the output of its only previous layer as input and feeds the output to its only subsequent layer, thus being able to improve the inference speed of the model and further improve the detection efficiency of steel defects. The multi-scale feature fusion and detection mechanism enables the model to capture both detailed information and global information in the steel image simultaneously, thereby more accurately detecting different types and sizes of steel defects.

[0042] Further, to better illustrate the above process of detecting steel defects, as a refinement and extension of the above embodiments, the embodiments of the present invention provide another method for detecting steel defects, as Figure 7 shown, the method includes: 201. Construct a preset initial steel defect detection model.

[0043] 202. Obtain a steel surface defect detection dataset as a sample image set. Among them, the sample image set includes multiple sample images with annotation information, and the annotation information includes annotation boxes and the annotation types of steel defects within the annotation boxes. The annotation types include oxidized scale, inclusions, cracks, scratches, patches, and pitted surfaces.

[0044] 203. Divide the sample image set into a training set, a validation set, and a test set according to a preset ratio. Use the training set to train the preset initial steel defect detection model, use the validation set to validate the trained preset initial steel defect detection model, adjust the model parameters of the trained preset initial steel defect detection model according to the validation results, and use the test set to test the preset initial steel defect detection model after parameter adjustment. Take the preset initial steel defect detection model after parameter adjustment that meets the test conditions as the preset steel defect detection model.

[0045] For the embodiments of the present invention, in order to improve the defect detection accuracy of the preset steel defect detection model, it is first necessary to train and construct the preset steel defect detection model. During the training process, first construct the preset initial steel defect detection model, and the model structure can be specifically as described above. Secondly, download the steel surface defect detection dataset from the network. Ensure that the dataset contains all necessary files, including steel image files and corresponding defect annotation files. Convert the annotation files into a format that the preset initial steel defect detection model can understand. Finally, train and test the model. Specifically, the dataset can be divided first: use a random or specific strategy (such as stratified sampling) to divide the dataset into a training set, a validation set, and a test set. Then use the training set to train the model, and use the validation set to monitor metrics such as the loss value and mAP in the trained model to evaluate the model performance. Adjust the training parameters as needed, such as the learning rate, optimizer, regularization, etc., to optimize the training effect. Finally, test the model: use the test set to test the model after parameter tuning and evaluate its performance on unseen data. Calculate and record metrics such as mAP, precision, and recall on the test set. If the model performance does not meet the requirements, return to the training stage for more iterations or adjustments. In this way, a preset steel defect detection model that meets the requirements is obtained.

[0046] 204. Obtain the steel image to be detected of the target steel.

[0047] 205. Input the steel image to be detected into the preset steel defect detection model. Perform preprocessing on the steel image to be detected through the input network in the preset steel defect detection model to obtain a preprocessed steel image. Extract image features from the preprocessed steel image through the backbone network to obtain the image features corresponding to each reparameterized feature reuse module. Fuse multiple image features through the head network to obtain the fused features corresponding to each feature fusion module. Perform steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain the steel defect detection results corresponding to each steel defect detection head, and determine the target defect detection result of the target steel among multiple steel defect detection results.

[0048] Specifically, the steel image to be detected is input into a trained preset steel defect detection model for image preprocessing, image feature extraction, feature fusion, and steel defect detection. Image preprocessing: The input network is used to process the size of the steel image to be detected, enhance the image quality, and generate adaptive anchor boxes to ensure the image quality, thereby improving the subsequent steel defect detection efficiency and detection accuracy. Image feature extraction: The reparameterized feature reuse module in the backbone network is used to extract features from the input steel image, generating image features at multiple levels. Feature fusion: The head network fuses these feature maps at different levels to generate a richer feature representation for subsequent defect detection. Steel defect detection: Each steel defect detection head in the prediction head network performs defect detection on the fused features at different scales, outputting the steel defect detection results of each steel defect detection head. Finally, based on each steel defect detection result, the target defect detection result of the target steel is determined, including information such as the location, category, and confidence level of the defect. This method does not rely on manual feature extraction and realizes the non-linear modeling of data through the multi-layer neural network in the preset steel defect detection model. In steel defect detection, steel usually has complex texture and shape features, and the preset steel defect detection model in the embodiments of the present invention can effectively capture these non-linear relationships. Compared with the two-stage defect detection method, the embodiments of the present invention directly complete the steel defect detection task in a single forward propagation without generating candidate regions, and the detection speed is relatively fast. The embodiments of the present invention add an improved EMA attention mechanism module, which can enhance the feature extraction ability of the model and reduce the computational complexity; then the MPConv module in the improved network enables the network to retain fine-grained information during downsampling, avoiding the learning of less efficient feature representations; at the same time, a structure reparameterization network structure is introduced to improve the inference speed of the model; finally, a detection layer is added to reduce the missed detection rate. The detection accuracy of the improved model is improved, meeting the requirements of actual steel defect detection.

[0049] Another steel defect detection method provided by the present invention, compared with the current method of first manually extracting image features and then using machine learning classification methods such as K-Nearest Neighbor (KNN) for classification detection, the present invention adopts a special model architecture, including an input network for image preprocessing, a backbone network containing a reparameterized feature reuse module, a head network for multi-scale feature fusion, and a prediction head network for multi-scale steel defect detection, and is trained using a steel surface defect detection dataset, achieving fast and accurate detection of steel defects through a special model. Further, in the backbone network of the preset steel defect detection model of the present invention, by introducing a reparameterized feature reuse module, since the reparameterized feature reuse module switches the operations originally distributed in different feature spaces waiting for fusion in the inference stage to a single inference performed by the fusion module in the weight space, it is hardware-friendly and greatly reduces the detection cost of steel defects. Compared with other model structures, the reparameterized feature reuse module has no redundant branch topologies, and each layer in the network structure only takes the output of its unique previous layer as input and feeds the output to its unique subsequent layer, thereby being able to improve the inference speed of the model and further improve the detection efficiency of steel defects. The multi-scale feature fusion and detection mechanism enables the model to simultaneously capture the detailed information and global information in the steel image, thereby more accurately detecting different types and sizes of steel defects.

[0050] Further, as Figure 1 a specific implementation of Figure 8 shown, the embodiment of the present invention provides a steel defect detection device, as

[0051] shown, the device includes: an image acquisition unit 31, a model acquisition unit 32, and a defect detection unit 33.

[0052] The image acquisition unit 31 can be used to acquire a steel image to be detected of a target steel.

[0053] The defect detection unit 33 can be used to input the steel image to be detected into the preset steel defect detection model, preprocess the steel image to be detected through the input network to obtain a preprocessed steel image, extract image features from the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fuse multiple image features through the head network to obtain fused features corresponding to each feature fusion module, detect steel defects for the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determine the target defect detection result of the target steel from multiple steel defect detection results.

[0054] In a specific application scenario, the backbone network sequentially includes a plurality of CBS (Convolution + BatchNorm + SiLU, convolution layer + batch normalization layer + activation function layer) modules, a first reparameterized feature reuse module, and a plurality of feature extraction modules. Each feature extraction module includes a first MPConv (MaxPooling and Conv, max pooling and convolution) module and a reparameterized feature reuse module. To extract image features from the preprocessed steel image through the backbone network, the defect detection unit 33 can specifically be used to input the preprocessed steel image into the backbone network, gradually extract image features from the preprocessed steel image through a plurality of CBS modules to obtain intermediate image features output by the last CBS module, input the intermediate image features into the first reparameterized feature reuse module for feature extraction to obtain first reparameterized image features output by the first reparameterized feature reuse module; and sequentially input the first reparameterized image features into a plurality of feature extraction modules to perform feature extraction through the first MPConv module and generate reparameterized image features through the reparameterized feature reuse module, so as to obtain second reparameterized image features output by the last feature extraction module.

[0055] In a specific application scenario, any one of the reparameterized feature reuse modules in the first reparameterized feature reuse module and the plurality of feature extraction modules includes a plurality of structural reparameterization processing layers; the first MPConv module includes a first convolutional layer, a slicing layer, a second convolutional layer, a max pooling layer, and a third convolutional layer. To extract image features through any one of the reparameterized feature reuse modules and the first MPConv module, as Figure 9 shown, the defect detection unit 33 includes a feature fusion module 331 and a feature slicing module 332.

[0056] The feature fusion module 331 can be used to process the input features of any reparameterized feature reuse module through each of the structural reparameterization processing layers respectively, obtain the output results of each of the structural reparameterization processing layers, and fuse the output results of each of the structural reparameterization processing layers with the input features to obtain the output results of any reparameterized feature reuse module.

[0057] The feature slicing module 332 can be used to process the input features of the first MPConv module through the first convolutional layer and the max pooling layer respectively, perform adjacent downsampling processing on the output result of the first convolutional layer through the slicing layer to obtain a preset number of sliced image features output by the slicing layer, perform fusion processing on each of the sliced image features through the second convolutional layer, and process the output result of the max pooling layer through the third convolutional layer.

[0058] The feature fusion module 331 can also be used to fuse the output result of the second convolutional layer and the output result of the third convolutional layer to obtain the output result of the first MPConv module.

[0059] In a specific application scenario, the head network includes a first feature fusion module, a second feature fusion module, a third feature fusion module, and a fourth feature fusion module; the first feature fusion module includes a first CBS_1EMA (CBS with an added EMA (Exponential Moving Average Attention) attention mechanism module), a first Concat (concatenation) module, a first ELAN-H (Enhanced Local Attention Network with Hierarchical Structure) module, and a first RepConv (Re-parameterized Convolution) module connected in sequence; the second feature fusion module includes a second CBS_1EMA module, a second Concat module, a second ELAN-H module, a third Concat module, a third ELAN-H module, and a second RepConv module connected in sequence; the third feature fusion module includes a third CBS_1EMA module, a fourth Concat module, a fourth ELAN-H module, a fifth Concat module, a fifth ELAN-H module, and a third RepConv module connected in sequence; the fourth feature fusion module includes a SPPCSPC_EMA (SPPCSPC (Spatial Pyramid Pooling - Cross Stage Partial Connections) with an added EMA attention mechanism) module, a sixth Concat module, a sixth ELAN-H module, and a fourth RepConv module connected in sequence; the head network further includes a plurality of feature upsampling modules and a plurality of second MPConv modules, and each feature upsampling module includes a CBS_1 module and an UP (Upsample) module; the output ends of the SPPCSPC_EMA module, the fourth ELAN-H module, and the second ELAN-H module are respectively connected to the input ends of the corresponding feature upsampling modules, and the input ends of the fourth Concat module, the second Concat module, and the first Concat module are respectively connected to the output ends of the corresponding feature upsampling modules; the output ends of the first ELAN-H module, the third ELAN-H module, and the fifth ELAN-H module are respectively connected to the input ends of the corresponding second MPConv modules, and the input ends of the third Concat module, the fifth Concat module, and the sixth Concat module are respectively connected to the output ends of the corresponding second MPConv modules.

[0060] In a specific application scenario, the EMA in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module all includes a Groups (group convolution) layer, a 1×1 convolution layer, a 3×3 convolution layer, a 5×5 convolution layer, a first Sigmoid (activation function) layer, a second Sigmoid layer, a first Re_weight (feature reweighting) layer, a GroupNorm (group normalization) layer, a first Avg_Pool (average pooling) layer, a first Softmax layer, a first Matmu1 (feature transformation) layer, a second Avg_Pool layer, a second Softmax layer, a second Matmu1 layer, a third Avg_Pool layer, a third Softmax layer, a third Matmu1 layer, a third Sigmoid layer, a second Re_weight layer, and an Output (output) layer; in order to use any one of the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module for feature extraction, the defect detection unit 33 can specifically be used to decompose the input features of any one of the target modules through the Groups layer to respectively obtain a first decomposed and reparameterized image feature, a second decomposed and reparameterized image feature, a third decomposed and reparameterized image feature, a fourth decomposed and reparameterized image feature, and a fifth decomposed and reparameterized image feature, respectively input the second decomposed and reparameterized image feature and the third decomposed and reparameterized image feature into the 1×1 convolution layer for processing to obtain a first output result and a second output result corresponding to the output of the 1×1 convolution layer, process the fourth decomposed and reparameterized image feature through the 3×3 convolution layer to obtain the output result of the 3×3 convolution layer, and process the fifth decomposed and reparameterized image feature through the 5×5 convolution layer to obtain the output result of the 5×5 convolution layer;The first output result and the second output result are respectively processed through the first Sigmoid layer and the second Sigmoid layer to obtain the output result of the first Sigmoid layer and the output result of the second Sigmoid layer. The first decomposition reparameterized image feature, the output result of the first Sigmoid layer, and the output result of the second Sigmoid layer are processed through the first Re_weight layer to obtain the output result of the first Re_weight layer. The output result of the first Re_weight layer is processed through the GroupNorm layer to obtain the output result of the GroupNorm layer. The output result of the GroupNorm layer is sequentially processed through the first Avg_Pool layer and the first Softmax layer to obtain the output result of the first Softmax layer. The output result of the 3×3 convolutional layer is sequentially processed through the second Avg_Pool layer and the second Softmax layer to obtain the output result of the second Softmax layer. The output result of the 5×5 convolutional layer is sequentially processed through the third Avg_Pool layer and the third Softmax layer to obtain the output result of the third Softmax layer. The output result of the first Softmax layer, the output result of the 3×3 convolutional layer, and the output result of the 5×5 convolutional layer are processed through the first Matmu1 layer to obtain the output result of the first Matmu1 layer. The output result of the second Softmax layer and the output result of the GroupNorm layer are processed through the second Matmu1 layer to obtain the output result of the second Matmu1 layer. The output result of the third Softmax layer and the output result of the GroupNorm layer are processed through the third Matmu1 layer to obtain the output result of the third Matmu1 layer. The output results of the first Matmu1 layer, the second Matmu1 layer, and the third Matmu1 layer are fused to obtain a fused result, and the fused result is processed through the third Sigmoid layer to obtain the output result of the third Sigmoid layer. The output result of the third Sigmoid layer and the first decomposition reparameterized image feature are processed through the second Re_weight layer to obtain the output result of the second Re_weight layer, and the output result of the second Re_weight layer is output through the Output layer as the processing result of any target module.;

[0061] In a specific application scenario, the prediction head network includes a first steel defect detection head, a second steel defect detection head, a third steel defect detection head, and a fourth steel defect detection head; in order to detect steel defects, the defect detection unit 33 further includes a defect detection module 333.

[0062] The defect detection module 333 can be used to respectively input the fusion features of each feature fusion module into the first steel defect detection head, the second steel defect detection head, the third steel defect detection head, and the fourth steel defect detection head for defect detection, and respectively obtain the steel defect detection results corresponding to each steel defect detection head.

[0063] In a specific application scenario, in order to train and construct a preset steel defect detection model, the device further includes a construction unit 34.

[0064] The construction unit 34 can be used to construct a preset initial steel defect detection model, and obtain a steel surface defect detection data set as a sample image set. Among them, the sample image set includes multiple sample images with annotation information. The annotation information includes annotation boxes and the annotation types of steel defects within the annotation boxes. The annotation types include oxide scale, inclusion, crack, scratch, plaque, pitting surface; divide the sample image set into a training set, a validation set, and a test set according to a preset ratio, use the training set to train the preset initial steel defect detection model, use the validation set to validate the trained preset initial steel defect detection model, adjust the model parameters of the trained preset initial steel defect detection model according to the validation results, and use the test set to test the preset initial steel defect detection model after parameter adjustment. The preset initial steel defect detection model that meets the test conditions after parameter adjustment is used as the preset steel defect detection model.

[0065] It should be noted that for other corresponding descriptions of each functional module involved in a steel defect detection device provided in an embodiment of the present invention, reference can be made to Figure 1 the corresponding description of the method shown, which will not be elaborated here.

[0066] Based on the above as Figure 1Accordingly, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the following steps are implemented: obtaining a to-be-detected steel image of a target steel; obtaining a preset steel defect detection model, wherein the preset steel defect detection model includes an input network for preprocessing an image, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes a plurality of reparameterized feature reuse modules, the head network includes a plurality of feature fusion modules of different scales, and the prediction head network includes a plurality of steel defect detection heads of different scales. The preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; inputting the to-be-detected steel image into the preset steel defect detection model, preprocessing the to-be-detected steel image through the input network to obtain a preprocessed steel image, extracting image features of the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fusing a plurality of image features through the head network to obtain fused features corresponding to each feature fusion module, performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determining a target defect detection result of the target steel from a plurality of steel defect detection results.

[0067] Based on the above as Figure 1 shown method and as Figure 8 shown device embodiment, an embodiment of the present invention further provides an entity structure diagram of a computer device, as Figure 10As shown in the figure, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor. Both the memory 42 and the processor 41 are set on a bus 43. When the processor 41 executes the program, the following steps are implemented: obtaining a steel image to be detected of a target steel; obtaining a preset steel defect detection model, wherein the preset steel defect detection model includes an input network for preprocessing the image, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes a plurality of reparameterized feature reuse modules, the head network includes a plurality of feature fusion modules of different scales, and the prediction head network includes a plurality of steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; inputting the steel image to be detected into the preset steel defect detection model, preprocessing the steel image to be detected through the input network to obtain a preprocessed steel image, extracting image features of the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fusing a plurality of image features through the head network to obtain fused features corresponding to each feature fusion module, performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determining a target defect detection result of the target steel from among the plurality of steel defect detection results.

[0068] Through the technical solution of the present invention, the present invention adopts a special model architecture, including an input network for image preprocessing, a backbone network containing reparameterized feature reuse modules, a head network for multi-scale feature fusion, and a prediction head network for multi-scale steel defect detection, and completes training using a steel surface defect detection data set, achieving fast and accurate detection of steel defects through a special model. Further, in the backbone network of the preset steel defect detection model of the present invention, by introducing reparameterized feature reuse modules, since the reparameterized feature reuse modules switch the operations that were originally distributed in different feature spaces waiting for fusion in the inference stage to a single inference performed by the fusion module in the weight space, it is hardware-friendly and greatly reduces the detection cost of steel defects. Compared with other model structures, the reparameterized feature reuse modules have no redundant branch topologies, and each layer in the network structure only takes the output of its unique previous layer as input and feeds the output to its unique subsequent layer, thereby being able to improve the inference speed of the model and thus improve the detection efficiency of steel defects. The multi-scale feature fusion and detection mechanism enables the model to simultaneously capture the detailed information and global information in the steel image, thereby more accurately detecting different types and sizes of steel defects.

[0069] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0070] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting steel defects, characterized in that, Including: Obtain a steel image to be detected of the target steel; Obtain a preset steel defect detection model, wherein the preset steel defect detection model includes an input network for preprocessing images, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes a plurality of reparameterized feature reuse modules, the head network includes a plurality of feature fusion modules of different scales, and the prediction head network includes a plurality of steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; Input the steel image to be detected into the preset steel defect detection model, preprocess the steel image to be detected through the input network to obtain a preprocessed steel image, extract image features from the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fuse a plurality of image features through the head network to obtain fused features corresponding to each feature fusion module, perform steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain steel defect detection results corresponding to each steel defect detection head, and determine the target defect detection result of the target steel among the plurality of steel defect detection results.

2. The steel defect detection method according to claim 1, characterized in that The backbone network sequentially includes a plurality of CBS (Convolution + BatchNorm + SiLU) modules, a first reparameterized feature reuse module, and a plurality of feature extraction modules. Each feature extraction module includes a first MPConv (MaxPooling and Conv) module and a reparameterized feature reuse module; The step of extracting image features from the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module includes: Input the preprocessed steel image into the backbone network, gradually extract image features from the preprocessed steel image through a plurality of CBS modules to obtain intermediate image features output by the last CBS module, input the intermediate image features into the first reparameterized feature reuse module for feature extraction to obtain first reparameterized image features output by the first reparameterized feature reuse module; Sequentially input the first reparameterized image features into a plurality of feature extraction modules to perform feature extraction through the first MPConv module and generate reparameterized image features through the reparameterized feature reuse module, and obtain second reparameterized image features output by the last feature extraction module.

3. The steel defect detection method according to claim 2, characterized in that, Any one of the reparameterized feature reuse modules in the first reparameterized feature reuse module and the plurality of feature extraction modules includes a plurality of structure reparameterization processing layers, and the feature extraction process of any one of the reparameterized feature reuse modules includes: Process the input features of any reparameterized feature reuse module through each of the structure reparameterization processing layers respectively to obtain the output results of each structure reparameterization processing layer, and fuse the output results of each structure reparameterization processing layer with the input features to obtain the output result of any reparameterized feature reuse module; The first MPConv module includes a first convolutional layer, a slicing layer, a second convolutional layer, a max pooling layer, and a third convolutional layer. The feature extraction process of the first MPConv module includes: Process the input features of the first MPConv module through the first convolutional layer and the max pooling layer respectively, perform adjacent downsampling on the output result of the first convolutional layer through the slicing layer to obtain a preset number of sliced image features output by the slicing layer, fuse each sliced image feature through the second convolutional layer, and process the output result of the max pooling layer through the third convolutional layer; Fuse the output result of the second convolutional layer and the output result of the third convolutional layer to obtain the output result of the first MPConv module.

4. The steel defect detection method according to claim 1, characterized in that The head network includes a first feature fusion module, a second feature fusion module, a third feature fusion module, and a fourth feature fusion module; The first feature fusion module includes a first CBS_1EMA (CBS with EMA (Exponential Moving Average Attention) attention mechanism module), a first Concat (concatenation) module, a first ELAN-H (Enhanced Local Attention Network with Hierarchical Structure) module, and a first RepConv (Re-parameterized Convolution) module connected in sequence; The second feature fusion module includes a second CBS_1EMA module, a second Concat module, a second ELAN-H module, a third Concat module, a third ELAN-H module, and a second RepConv module connected in sequence; The third feature fusion module includes a third CBS_1EMA module, a fourth Concat module, a fourth ELAN-H module, a fifth Concat module, a fifth ELAN-H module, and a third RepConv module connected in sequence; The fourth feature fusion module includes an SPPCSPC_EMA (SPPCSPC (Spatial Pyramid Pooling - Cross Stage Partial Connections) with EMA attention mechanism) module, a sixth Concat module, a sixth ELAN-H module, and a fourth RepConv module connected in sequence; The head network further includes a plurality of feature upsampling modules and a plurality of second MPConv modules, and each feature upsampling module includes a CBS_1 module and an UP (Upsample) module; The output ends of the SPPCSPC_EMA module, the fourth ELAN-H module, and the second ELAN-H module are respectively connected to the input ends of the corresponding feature upsampling modules, and the input ends of the fourth Concat module, the second Concat module, and the first Concat module are respectively connected to the output ends of the corresponding feature upsampling modules; The output ends of the first ELAN-H module, the third ELAN-H module, and the fifth ELAN-H module are respectively connected to the input ends of the corresponding second MPConv modules, and the input ends of the third Concat module, the fifth Concat module, and the sixth Concat module are respectively connected to the output ends of the corresponding second MPConv modules.

5. The steel defect detection method according to claim 4, wherein, The EMA in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module all includes a Groups (group convolution) layer, a 1×1 convolution layer, a 3×3 convolution layer, a 5×5 convolution layer, a first Sigmoid (activation function) layer, a second Sigmoid layer, a first Re_weight (feature reweighting) layer, a GroupNorm (group normalization) layer, a first Avg_Pool (average pooling) layer, a first Softmax layer, a first Matmu1 (feature transformation) layer, a second Avg_Pool layer, a second Softmax layer, a second Matmu1 layer, a third Avg_Pool layer, a third Softmax layer, a third Matmu1 layer, a third Sigmoid layer, a second Re_weight layer, and an Output layer; The working process of any target module in the first CBS_1EMA module, the second CBS_1EMA module, the third CBS_1EMA module, and the SPPCSPC_EMA module includes: Performing feature decomposition on the input features of the any target module through the Groups layer to respectively obtain a first decomposed reparameterized image feature, a second decomposed reparameterized image feature, a third decomposed reparameterized image feature, a fourth decomposed reparameterized image feature, and a fifth decomposed reparameterized image feature, respectively inputting the second decomposed reparameterized image feature and the third decomposed reparameterized image feature into the 1×1 convolution layer for processing to obtain a first output result and a second output result corresponding to the output of the 1×1 convolution layer, processing the fourth decomposed reparameterized image feature through the 3×3 convolution layer to obtain the output result of the 3×3 convolution layer, and processing the fifth decomposed reparameterized image feature through the 5×5 convolution layer to obtain the output result of the 5×5 convolution layer; The first output result and the second output result are respectively processed through the first Sigmoid layer and the second Sigmoid layer to obtain the output result of the first Sigmoid layer and the output result of the second Sigmoid layer. The first decomposed reparameterized image feature, the output result of the first Sigmoid layer, and the output result of the second Sigmoid layer are processed through the first Re_weight layer to obtain the output result of the first Re_weight layer. The output result of the first Re_weight layer is processed through the GroupNorm layer to obtain the output result of the GroupNorm layer. The output result of the GroupNorm layer is sequentially processed through the first Avg_Pool layer and the first Softmax layer to obtain the output result of the first Softmax layer. The output result of the 3×3 convolutional layer is sequentially processed through the second Avg_Pool layer and the second Softmax layer to obtain the output result of the second Softmax layer. The output result of the 5×5 convolutional layer is sequentially processed through the third Avg_Pool layer and the third Softmax layer to obtain the output result of the third Softmax layer; The output result of the first Softmax layer, the output result of the 3×3 convolutional layer, and the output result of the 5×5 convolutional layer are processed through the first Matmu1 layer to obtain the output result of the first Matmu1 layer. The output result of the second Softmax layer and the output result of the GroupNorm layer are processed through the second Matmu1 layer to obtain the output result of the second Matmu1 layer. The output result of the third Softmax layer and the output result of the GroupNorm layer are processed through the third Matmu1 layer to obtain the output result of the third Matmu1 layer; The output results of the first Matmu1 layer, the second Matmu1 layer, and the third Matmu1 layer are fused to obtain a fused result, and the fused result is processed through the third Sigmoid layer to obtain the output result of the third Sigmoid layer. The output result of the third Sigmoid layer and the first decomposed reparameterized image feature are processed through the second Re_weight layer to obtain the output result of the second Re_weight layer, and the output result of the second Re_weight layer is output through the Output layer as the processing result of any target module.

6. The steel defect detection method according to claim 1, characterized in that, The prediction head network includes a first steel defect detection head, a second steel defect detection head, a third steel defect detection head, and a fourth steel defect detection head; Performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain the steel defect detection results corresponding to each steel defect detection head, including: Respectively inputting the fused features of each feature fusion module into the first steel defect detection head, the second steel defect detection head, the third steel defect detection head, and the fourth steel defect detection head for defect detection, and respectively obtaining the steel defect detection results corresponding to each steel defect detection head.

7. The steel defect detection method according to claim 1, characterized in that, Before obtaining the preset steel defect detection model, the method further includes: Constructing a preset initial steel defect detection model, and obtaining a steel surface defect detection data set as a sample image set. Among them, the sample image set includes multiple sample images with annotation information, and the annotation information includes annotation boxes and annotation types of steel defects within the annotation boxes. The annotation types include oxidized scale, inclusions, cracks, scratches, patches, and pitted surfaces; Dividing the sample image set into a training set, a validation set, and a test set according to a preset ratio, training the preset initial steel defect detection model using the training set, validating the trained preset initial steel defect detection model using the validation set, adjusting the model parameters of the trained preset initial steel defect detection model according to the validation results, and testing the preset initial steel defect detection model after parameter adjustment using the test set. Taking the preset initial steel defect detection model after parameter adjustment that meets the test conditions as the preset steel defect detection model.

8. A steel defect detection device, characterized in that, Including: An image acquisition unit for acquiring a steel image to be detected of the target steel; A model acquisition unit for acquiring a preset steel defect detection model. Among them, the preset steel defect detection model includes an input network for preprocessing images, a backbone network for extracting image features, a head network for fusing image features, and a prediction head network for defect detection. The backbone network includes multiple reparameterized feature reuse modules, the head network includes multiple feature fusion modules of different scales, and the prediction head network includes multiple steel defect detection heads of different scales; the preset steel defect detection model is pre-trained and constructed based on a steel surface defect detection data set; A defect detection unit for inputting the steel image to be detected into the preset steel defect detection model, preprocessing the steel image to be detected through the input network to obtain a preprocessed steel image, extracting image features of the preprocessed steel image through the backbone network to obtain image features corresponding to each reparameterized feature reuse module, fusing multiple image features through the head network to obtain fused features corresponding to each feature fusion module, performing steel defect detection on the fused features of each feature fusion module through the prediction head network to obtain the steel defect detection results corresponding to each steel defect detection head, and determining the target defect detection result of the target steel among multiple steel defect detection results.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the steel defect detection method according to any one of claims 1 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the steel defect detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tank armored vehicle flow detection method based on target detection model and DeepSort

    CN116129312A

  • Strip steel surface defect detection method based on improved YOLO model

    CN117611571A

  • Steel surface defect detection method and system based on computer vision

    CN118037692A

  • Microelectronic device surface defect detection method based on improved YOLOv9

    CN119380099A

  • Surface defect detection method, system, equipment, and terminal thereof

    US12307653B1