Strip steel surface defect detection method based on multi-attention strategy fusion and lightweight

By introducing the MALDA-YOLO network with multi-attention strategy fusion into the YOLOv8 model, the problems of high-precision and lightweight strip defect detection are solved, and real-time and accurate detection of strip surface defects is achieved, thereby improving detection speed and accuracy.

CN120598874APending Publication Date: 2025-09-05ANQING NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510672583.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing strip defect detection methods are difficult to achieve high-precision and real-time detection when faced with complex and changeable defect types, and they also have problems such as lightweight models and an imbalance in the number of individual defects.

Method used

The MALDA-YOLO network based on multi-attention strategy fusion is adopted. By replacing the backbone network and neck network modules of the YOLOv8 model and combining it with the improved loss function HVC Loss, feature extraction and fusion are realized, which reduces the amount of calculation and parameters while improving detection accuracy and speed.

Benefits of technology

It achieves high-precision, real-time detection of strip surface defects, reduces false detection and missed detection rates, improves the model's detection capabilities, and reduces the computational burden, with a detection speed of 170FPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598874A_ABST
    Figure CN120598874A_ABST
Patent Text Reader

Abstract

The invention discloses a strip steel surface defect detection method based on multi-attention strategy fusion and lightweight, and belongs to the technical field of steel surface defect detection, and the method comprises the following steps: S1, data acquisition; s2, constructing a model; s3, model training; and S4, defect detection. On the basis of a YOLOv8 model framework, a multi-feature inverted bottleneck module MAIB based on a multi-scale attention mechanism is designed, and a C2f module of a backbone part is replaced with the MAIB for feature extraction; in the neck part, a feature fusion module DAFG for deeply separating attention based on grouping extraction is provided, and by combining grouping convolution and reduced convolution operation, the quality and the detection precision of feature fusion are improved while the calculated amount and the parameter amount are reduced; in a detection head part, an improved loss function HVC Loss is provided, so that the problem of sample imbalance can be better solved, the model convergence speed is accelerated, the detection capability of a small target is enhanced, and meanwhile, the false detection rate and the omission ratio are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steel surface defect detection, and in particular to a strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting. Background Art

[0002] Monitoring and ensuring product quality is crucial during strip production. Surface defect detection for strip steel not only impacts product safety and reliability but also directly impacts production costs and a company's market competitiveness. Traditional manual inspection methods are inefficient and susceptible to human error, making them inadequate for the high-precision, high-efficiency inspections required by modern production lines. Therefore, the research and application of automated strip steel surface defect detection technology is crucial.

[0003] Methods for detecting surface defects in strip steel are primarily categorized as traditional image processing, machine learning, and deep learning. Traditional image processing methods rely on techniques such as edge detection, morphological processing, and texture analysis, identifying defects through image preprocessing and feature extraction. While effective in certain simple scenarios, they are limited in their ability to detect complex images and diverse defects. Machine learning methods manually design features and utilize machine learning algorithms for classification. While these methods improve detection accuracy to a certain extent, their reliance on feature engineering makes them difficult to handle with complex and diverse defect types. With the development of deep learning technology, deep learning-based strip steel defect detection methods are gaining widespread application, offering advantages such as automatic feature extraction, high accuracy, and strong robustness.

[0004] Existing leading-edge algorithms have improved strip defect detection to varying degrees, but there are still certain challenges in actual application scenarios. Strip defect detection mostly relies on open source data sets, but steel mills themselves are particular and unique. The defects generated by each steel mill are different, and the number of defects varies. Secondly, practical applications pay more attention to the size of the model itself, including the number of parameters and the amount of calculation, which will determine the application efficiency. Therefore, in response to the imbalance in the number of individual defects and the problem of model lightweighting faced by strip defect detection, this paper proposes a strip surface defect detection method based on multi-attention strategy fusion and lightweighting. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to detect strip defects in real time and accurately in actual production, and provides a strip surface defect detection method based on multi-attention strategy fusion and lightweighting.

[0006] The present invention solves the above technical problems through the following technical solutions, which include the following steps:

[0007] S1: Data Collection

[0008] Collect data samples from open source databases and steel plate surface defect data collected on industrial sites, and preprocess the samples;

[0009] S2: Model construction

[0010] Taking the YOLOv8 object detection network framework as the baseline model, all C2f modules in its backbone network are replaced with MAIB modules, and all C2f modules in the neck network are replaced with DAFG modules. The HVC Loss function is designed to replace the CIoU Loss function in the detection head network, thereby obtaining the MALDA-YOLO network. The MAIB module is a multi-feature inverted bottleneck module based on a multi-scale attention mechanism, and the DAFG module is a feature fusion module based on deep separation attention of group extraction.

[0011] S3: Model training

[0012] The MALDA-YOLO network is trained using the training set to obtain a strip surface defect detection model;

[0013] S4: Defect Detection

[0014] Test the strip surface defect detection model on the test set and output the detection results and detection speed.

[0015] Furthermore, in step S1, the specific processing process is as follows:

[0016] S11: Collect data samples from the open source database NEU-DET and steel plate surface defect data collected on industrial sites, and divide the acquired data samples into training and test sets according to the set ratio;

[0017] S12: Perform data enhancement processing on the steel plate surface defect data image samples collected from the industrial field in the training set, and fill the size of the processed image samples to the set size.

[0018] Furthermore, in step S2, the Bottleneck module of the C2f module in the backbone network is replaced by the MUIB module to form the MAIB module. In the MUIB module, the input features are processed in sequence by a 3×3 convolution layer, a 1×1 convolution layer, a 3×3 RepVGGDW convolution layer, a GMSA module, a 1×1 convolution layer, and a 3×3 convolution layer to obtain output features.

[0019] Furthermore, in the GMSA module, the specific processing process is as follows:

[0020] S201: Divide the input features into two groups of features, perform channel attention calculation and spatial attention calculation on the first group of features, fuse the calculated weights, and use the GN layer to normalize them within each group to obtain feature F1;

[0021] S202: Perform 1×1 convolution, 3×3 convolution, and 1×1 convolution on the second set of features to obtain feature F2;

[0022] S203: For feature F1, use 2D global average pooling to encode global spatial information, and use 2D Gaussian mapping to fit the linear transformation through the natural nonlinear function Softmax, and perform feature attention feature fusion with feature F2 to obtain feature F′1;

[0023] S204: For feature F2, use 2D global average pooling to encode global spatial information, and use 2D Gaussian mapping to fit the linear transformation through the natural nonlinear function Softmax, and perform feature attention feature fusion with feature F1 to obtain feature F′2;

[0024] S205: Aggregate the attention weights of features F′1 and F′2, and then process them using the Sigmoid function to obtain new attention weights, which are finally multiplied with the input features and output.

[0025] Furthermore, in step S2, in the DAFG module, the specific processing process is as follows:

[0026] S211: The input features are first processed by 3x3 convolution, and then the features obtained after convolution are divided into two branches, ch1 and ch2, according to the number of channels;

[0027] S212: For branch ch1, a 1x1 convolution is used for dimensionality reduction, a 3x3 depth-analyzable convolution layer based on hard_swish is used, and then the attention weight matrix Q is obtained through the HCA module. ch1 ; For branch ch2, 1x1 convolution is used for dimensionality reduction, and then the attention weight matrix Q is obtained through the HCA module ch2 ;

[0028] S213: The attention weight matrix Q ch1 With Q ch2 Perform feature fusion and compare attention weight matrix Q ch1 With Q ch2 For values ​​of the same position, the maximum value is retained, and the fused features are dimensionally increased using 1x1 convolution to obtain the output features.

[0029] Furthermore, in the HCA module, the specific processing process is as follows:

[0030] The input features are grouped, fused after being processed by the maximum pooling layer and the average pooling layer respectively, and the fused results are passed through a 1x1 convolution and a hard-swish activation function, and then grouped again by a 1x1 convolution, and fused after being processed by the maximum pooling layer and the average pooling layer respectively. The fused results are weighted after being passed through the hard-sigmoid activation function, and multiplied with the input features to obtain the output features.

[0031] Furthermore, in step S2, the calculation formula of the HVC Loss function is as follows:

[0032]

[0033] Among them, α i is the weight coefficient of sample i; p i is the predicted probability of sample i; y i is the true label of sample i, the positive sample is 1 and the negative sample is 0; γ is the focus parameter; λ is the weight parameter used to balance Varifocal Loss and CIoU Loss; f iou is the IoU value between the predicted box and the real box; ρ is the Euclidean distance between the center points of the predicted box and the real box; c is the length of the diagonal of the minimum circumscribed rectangular box; b and b g are the center points of the predicted box and the real box respectively; v is used to measure the aspect ratio consistency of the predicted box and the target box; μ is the parameter used to balance v.

[0034] Furthermore, in the HVC Loss loss function, a weight parameter λ is introduced to control the loss's attention to positive and negative samples and aspect ratios. The calculation formula is as follows:

[0035]

[0036] Among them, δ(x) represents the Sigmoid function, N represents the total number of samples, N p Represents the number of positive samples, N n Represents the number of negative samples, and k is a tuning parameter used to control the sensitivity of weight adjustment.

[0037] Furthermore, in step S3, the specific processing process is as follows:

[0038] S31: Set model training parameters, including learning rate and batch size;

[0039] S32: Input the training set into the MALDA-YOLO network for training. After the training is completed, save the weight parameters to obtain a strip surface defect target detection model.

[0040] Furthermore, in step S4, the specific processing process is as follows:

[0041] S41: Input the test set into the strip surface defect target detection model, and perform non-maximum suppression processing on the original output of the model;

[0042] S42: Output the defect detection results and detection speed of the model.

[0043] Compared with the prior art, the present invention has the following advantages:

[0044] Based on the YOLOv8 model framework, this paper designs a multi-feature inverted bottleneck module MAIB based on the multi-scale attention mechanism, and replaces the C2f module in the backbone part for feature extraction; in the neck part, a feature fusion module DAFG based on deep separation attention of group extraction is proposed, which combines group convolution and reduced convolution operations to reduce the amount of calculation and parameters while improving the quality of feature fusion and detection accuracy; in the detection head part, an improved loss function HVCLoss is proposed, which can better handle the sample imbalance problem, accelerate the model convergence speed, enhance the detection ability of small targets while reducing the false detection and missed detection rates. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 1 is a flow chart of a strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting in an embodiment of the present invention;

[0046] Figure 2 These are image samples of steel plate surface defect data collected at an industrial site in an embodiment of the present invention, where (a) is a stain defect, (b) is a scratch defect, and (c) is a spot defect.

[0047] Figure 3 : are defect image samples in the NEU-DET database in an embodiment of the present invention, where (a) is a crack defect, (b) is an inclusion defect, (c) is a plaque defect, (d) is a pitting surface defect, (e) is a rolling scale defect, and (f) is a scratch defect;

[0048] Figure 4 1 is a diagram of the MALDA-YOLO network structure in an embodiment of the present invention;

[0049] Figure 5 Schematic diagram of the MAIB module and its component structure in an embodiment of the present invention, wherein (a) is the GMSA module, (b) is the MUIB module, and (c) is the MAIB module;

[0050] Figure 6Schematic diagram of the DAFG module and its component structure in an embodiment of the present invention, wherein (a) is the HCA module and (b) is the DAFG module;

[0051] Figure 7 This is a comparison diagram of different models in the embodiment of the present invention on steel plate surface defect data collected at an industrial site;

[0052] Figure 8 This is a comparison diagram of different models on the NEU-DET dataset in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.

[0054] like Figure 1 As shown, this embodiment provides a technical solution: a strip surface defect detection method based on multi-attention strategy fusion and lightweighting to identify and locate surface defects of hot-rolled strip steel, including the following steps:

[0055] S1: Collect data samples from the NEU-DET dataset and steel plate surface defect data collected on industrial sites, and preprocess the samples;

[0056] S2: Based on the YOLOv8 object detection framework, a multi-feature inverted bottleneck module MAIB based on the multi-scale attention mechanism (GMSA) is designed, and it replaces the C2f module in the backbone part for feature extraction. In the neck part, a feature fusion module DAFG based on deep separation attention (HCA) with grouped extraction is proposed. It combines grouped convolution and reduced convolution operations to reduce the amount of computation and parameters while improving the quality of feature fusion and detection accuracy. In the head part, an improved loss function HVCLoss is proposed to better handle the sample imbalance problem, accelerate model convergence, enhance the detection ability of small targets, and reduce false detection and missed detection rates. After improvement, the MALDA-YOLO network is obtained.

[0057] S3: Use the training set to train the MALDA-YOLO network to obtain a strip surface defect detection model;

[0058] S4: Test the strip surface defect detection model on the test set and output the detection results and detection speed.

[0059] In this embodiment, step S1 includes the following steps:

[0060] S11: The performance of the improved module is verified using two datasets: one is the general open-source NEU-DET dataset, and the other is steel plate surface defect data collected on-site. The acquired data samples are divided into training and test sets according to a set ratio.

[0061] S12: For the steel plate surface defect data images collected from industrial sites in the training set, automatic segmentation and feature vector similarity calculation are performed for data enhancement, and the processed image size is padded to the set size.

[0062] In this embodiment, in step S12, the data enhancement process of automatic segmentation and feature vector similarity calculation is as follows:

[0063] First, the feature pixel positions identified by SAM are extracted from the image, and the two-dimensional coordinates Where N is the number of feature pixels. To further extract contour points, assuming that the original grayscale image is I(x, y), the edge points are initially screened using the set gradient amplitude G and the set edge threshold T, as shown in the following formula:

[0064]

[0065] Among them, M is the number of edge points. In order to remove noise and mismatching, a distance threshold ∈ is set; only when the distance between the edge point and the feature point is less than ∈, the feature point is considered to belong to the contour. The Euclidean distance is used to calculate the distance between the feature point in P and the edge point P 边缘 Distance to edge

[0066]

[0067] Among them, C is the final extracted contour feature point set;

[0068] Construct the position data of the contour feature pixels into a vector set. Assume that the feature pixel position data is R:

[0069]

[0070] Use Euclidean distance for index construction:

[0071]

[0072] For each pixel R i , use the FAISS algorithm to find its k nearest neighbor pixels (immediate neighbor points) τ i , and get its corresponding distance matrix D i and index matrix (index position) I i :

[0073] Di ={d(R i ,R j )|R j ∈τ i}

[0074] I i ={j|R j ∈τ i}

[0075] Set the distance threshold β i is the average value of the searched adjacent distances. When the distance value is greater than the threshold, the search is performed according to I i Remove the adjacent point from the position and get the adjacent point τ that meets the conditions i ′ and its index matrix I i ′, distance threshold β i The definition is as follows:

[0076]

[0077] Compare the pixel brightness similarity S of the retained adjacent points, set the similarity threshold to 15, and the similarity calculation formula is:

[0078] S={L(x i,j )-L(x i′,j′ )‖L(x i,j )-L(x i′,j′ )}

[0079] The pixel whose similarity is less than the threshold is indexed according to the position I i ′ is added to the feature area obtained by SAM segmentation, thus completing the extraction of defect feature area; the defect feature area is subjected to random angle inversion, scaling and other transformation operations and added to the background image without defects to generate a sufficient amount of steel plate surface defect data image samples, thereby achieving data enhancement.

[0080] In this embodiment, step S2 includes the following steps:

[0081] S21: Based on the YOLOv8 object detection framework, a multi-feature inverted bottleneck module MAIB based on the multi-scale attention mechanism (GMSA) is designed, and it replaces the C2f module of the Bottleneck layer;

[0082] S22: Design a feature fusion module DAFG based on deep separation attention (HCA) of group extraction and replace the C2f module of the Neck part with it;

[0083] S23: Design an HVC Loss loss function to replace the CIoU Loss loss function of the Head part.

[0084] In this embodiment, in step S21, in the MAIB module, the implementation process of the multi-feature bottleneck inversion module MAIB based on the multi-scale attention mechanism is as follows:

[0085] The input features are divided into two groups x1 and x2. CBMA channel attention (CA) and spatial attention (SA) calculations are performed on x1 respectively. The weights are fused and normalized in each group using the GN layer to obtain the feature variable F1. In order to obtain the deep information of the features while reducing the amount of calculation and enhance the expression ability and receptive field of the model, 1×1 convolution, 3×3 convolution, and 1×1 convolution are performed on x2 in sequence to obtain the deep feature information F2. The global spatial information is encoded using 2D global average pooling, and the natural nonlinear function Softmax is used. 2D Gaussian mapping is used to fit the linear transformation, and the feature attention feature is fused with F1 to obtain F′2; similarly, for F1, 2D global average pooling is used to encode global spatial information, Softmax uses 2D Gaussian mapping to fit the linear transformation, and the feature attention feature is fused with F2 to obtain F′1. F′1 and F′2 are aggregated for attention weights, and the Sigmoid function is used to capture the pixel-level pairwise relationship and highlight the global context of all pixels to obtain the attention weight. Finally, it is multiplied with the input feature and output to obtain the GMSA module. Figure 5 The processing formula for (a) is as follows:

[0086] x1,x2=spilt(x)

[0087] x C =CA(x1),x S =SA(x1)

[0088] F1=GN(x C +x S )

[0089] F2=Conv 1×1 (Conv 3×3 (Conv 1×1 (x2)))

[0090] F′2=Matmul(Softmax(AvgPool(F1)),F2)

[0091] F′1=Matmul(Softmax(AvgPool(F2)),F1)

[0092] F(x)=Sigmoid(F′1×F′2)×x

[0093] In the MUIB module, the input features are sequentially subjected to 3×3 convolution, 1×1 convolution, 3×3RepVGGDW convolution, GMSA module, 1×1 convolution, and 3×3 convolution, see Figure 5 (b) in the equation is as follows:

[0094] F(x)=Conv 3×3 (Conv 1×1 (GMSA(RepVGGDW 3×3 (Conv 1×1 (Conv 3×3 (x))))))

[0095] The MUIB module replaces the Bottleneck module in C2f to form a new multi-feature inverted bottleneck module MAIB

[0096] In this embodiment, in step S22, the implementation process of the feature fusion module DAFG based on deep separation attention (HCA) of group extraction is as follows:

[0097] The input features are grouped and fused after being processed by the maximum pooling layer and the average pooling layer respectively. The fused results are passed through a 1x1 convolution and a hard-swish activation function, and then through a 1x1 convolution for secondary grouping. They are again processed by the maximum pooling layer and the average pooling layer and fused. The fused results are weighted after being activated by the hard-sigmoid function and multiplied with the input features to obtain the HCA module. Figure 6 The processing formula for (a) is as follows:

[0098] x1=Conv 1×1 (hard swish(Conv 1×1 (MaxPool(x)+AvgPool(x))))

[0099] F(x)=x×(hard sigmoid(MaxPool(x1)+AvgPool(x1)))

[0100] like Figure 6 As shown in (b), in the DAFG module, the input feature x is convolved with 3x3 convolution to obtain deep features, and the features are divided into two branches, ch1 and ch2, with a size of c / 2 according to the number of channels c. For branch ch1, 1x1 convolution is used for dimensionality reduction, and a 3x3 deep analyzable convolution based on hard_swish is used to obtain the attention weight matrix Q through the HCA module. ch1 For branch ch2, 1x1 convolution is used for dimensionality reduction, and the attention weight matrix Q is obtained through the HCA module. ch2 . Qch1 With Q ch2 Perform feature fusion and compare Q ch1 With Q ch2 For values ​​of the same position, retain the maximum value, and use 1x1 convolution to increase the dimension of the fused feature to obtain the output feature F(x):

[0101] ch1,ch2=spilt(Conv 3×3 (x))

[0102] Q ch1 =HCA(DWConv 3×3 (Conv 1×1 (ch1)))

[0103] Q ch2 =HCA(Conv 1×1 (ch2))

[0104]

[0105] F(x)=Conv 1×1 (Q v )

[0106] Among them, i and j represent the matrix Q ch1 The rows and columns of v Belong to Q ch1 , then keep Q ch1 The features at the corresponding position, otherwise, Q is retained ch2 Features at corresponding positions.

[0107] In this embodiment, in step S23, the implementation process of the HVC Loss function is as follows:

[0108] HVC Loss is based on the advantages of Varifocal Loss in processing the imbalance of positive and negative samples and combines CIoU Loss with the optimization of the feature information of bounding box overlap, center point distance and aspect ratio consistency. The loss function expression for positive and negative samples is:

[0109]

[0110] Loss q =α i p i γ log(1-p i )

[0111]

[0112] Among them, α i is the weight coefficient of sample i; pi is the predicted probability of sample i; y i is the true label of sample i, the positive sample is 1 and the negative sample is 0; γ is the focusing parameter, and the modulation factor ((1-p i ) γ and p i γ ) also increases; λ is the weight parameter that balances Varifocal Loss and CIoU Loss, f iou is the IoU value between the predicted box and the ground-truth box. ρ is the Euclidean distance between the center points of the predicted box and the ground-truth box, c is the length of the diagonal of the minimum bounding rectangle, b and bg are the center points of the predicted box and the ground-truth box, respectively. v is used to measure the aspect ratio consistency of the predicted box and the target box, and μ is the parameter used to balance v.

[0113] In this embodiment, the variable λ is introduced to control the loss's attention to positive and negative samples and aspect ratio. The calculation formula is:

[0114]

[0115] Among them, δ(x) represents the Sigmoid function, N represents the total number of samples, N p Represents the number of positive samples, N n Represents the number of negative samples. k is a tuning parameter used to control the sensitivity of weight adjustment. In this invention, k is set to 0.5.

[0116] In this embodiment, step S3 includes the following steps:

[0117] S31: Set the model training parameters, learning rate to 0.001, and batch size to 8.

[0118] S32: Input the training set into the MALDA-YOLO network for training. After the training is completed, save the weight parameters to obtain a strip surface defect target detection model.

[0119] In this embodiment, step S4 includes the following steps:

[0120] S41: The test set is input into the strip surface defect target detection model, and non-maximum suppression processing is performed on the original output of the model.

[0121] S42: Output the defect detection results and detection speed of the model.

[0122] In this embodiment, in step S1, the steel plate surface defect data samples are collected at the industrial site. Figure 2As shown in the figure, it mainly includes three categories: scratches (Sc), spots (Pa) and stains (Di); the other type of database is the hot-rolled steel strip surface defect database, in which images of six typical surface defects on the hot-rolled steel strip surface and the position coordinate information of the defects in each image are obtained, such as Figure 3 As shown in the figure, the six typical surface defects are cracks (Cr), inclusions (In), patches (Pa), pitted surface (Ps), rolling scale (RS), and scratches (Sc).

[0123] The MALDA-YOLO network structure in this embodiment is as follows Figure 4 As shown in FIG, it includes a backbone network, a neck network and a detection head network.

[0124] In this embodiment, all C2f modules in the backbone network are replaced with MAIB modules, and in the neck part, all C2f modules are replaced with DAFG modules.

[0125] The neck network includes a first Upsample module, a first splicing (Contcat) module, a first DAFG module, a second Upsample module, a second splicing module, a second DAFG module, a first CBS module (downsampling module, convolution kernel size is 3×3), a third splicing module, a third DAFG module, a second CBS module (downsampling module, convolution kernel size is 3×3), a fourth splicing module, and a fourth DAFG module. The backbone network's SPPF module is connected to the first Upsample module, feeding the backbone network's output P4 layer features and the obtained features into the first concatenation module, which then passes through the first DAFG module to obtain the upsampling P4 layer features. It is connected to the second Upsample module, feeding the backbone network's output P3 layer features and the obtained features into the second concatenation module, which then passes through the second DAFG module to obtain the P3 layer features. After passing through the first CBS module (downsampling module with a convolution kernel size of 3×3), the P4 layer features from the upsampling process are fed into the third concatenation module, which then passes through the third DAFG module to obtain the P4 layer features. The P4 layer features are fed through the second CBS module (downsampling module with a convolution kernel size of 3×3) and the P5 layer features from the backbone SPPF into the fourth concatenation module, which then passes through the fourth DAFG module to obtain the P5 layer features. The P3, P4, and P5 layer features are then input into the detection head network.

[0126] Table 1 Comparison results of different models on the NEU-DET dataset

[0127]

[0128]

[0129] As shown in Table 1 above, MALDA-YOLO (the steel strip surface defect target detection model of the present invention) was trained and tested with MSC-DNet, DNN, WSS-YOLO, SRN-YOLO, DCC-CenterNet, RDD-YOLO, EFD-YOLOv4, DsP-YOLO, YOLOV8, and YOLOV9 in a unified hardware environment and dataset. Comparing the results, we can conclude that MALDA-YOLO achieved the best AP for the three categories of cracks (Cr), rolling scale (RS), and scratches (Sc) on the NEU-DET dataset. It achieved a mAP score of 82.9, 0.6 percentage points higher than the best WSS-YOLO algorithm among other algorithms. The model parameters were reduced by 49% compared to the best YOLOV8, and the model computational complexity was reduced by 27% compared to the best WSS-YOLO, achieving both accuracy and speed. In order to intuitively compare the visual detection effects of the proposed method with the existing methods, the detection results of YOLOV8 and YOLOV9 are compared as follows: Figure 8 shown.

[0130] Table 2 Model detection performance under different combinations

[0131] Data augmentation MAIB DAFG HVC mAP / % Param / M FLOPs / G FPS (frames / s) × × × × 63.7 3.75 9.6 154 × √ × × 65.8 2.43 6.4 159 × × √ × 66.4 1.82 4.6 168 × × × √ 64.3 3.86 9.8 150 × √ √ × 69.7 0.85 3.7 174 × √ × √ 66.9 2.61 6.9 156 × × √ √ 67.3 2.04 5.1 165 × √ √ √ 71.5 0.89 3.9 172 √ × × × 91.2 3.87 9.8 151 √ √ √ √ 95.3 0.94 4.2 170

[0132] As shown in Table 2 above, based on the training results of steel plate surface defect data collected at industrial sites, the present invention introduces a multi-feature inverted bottleneck module MAIB in the backbone part of YOLOv8, captures finer-grained information by dynamically adjusting the weights of features and optimizes the network architecture using MUIB. Compared with the original YOLOv8 model, the mAP is increased by 2.1 percentage points, and the model parameters and computational complexity are reduced by 35% and 33% respectively; the DAFG module is added to the neck unit, and the computational complexity and parameter quantity are greatly reduced by grouping convolution and reducing convolution operations. Combined with the attention mechanism HCA to suppress irrelevant or redundant information, the model parameters are reduced by 51.5% and the computational complexity is reduced by 52.1%; the loss function is changed separately, and HVC Loss is used, and the mAP is increased by 0.6 percentage points. This is because the imbalance of positive and negative samples based on Varifocal Loss and CIoU are considered. Loss adds a certain computational burden to bounding box overlap, center point distance, and aspect ratio consistency. When MAIB and DAFG modules are used simultaneously, network model computation is optimized in the backbone and neck based on their respective feature learning characteristics. A targeted attention mechanism is employed, using the GMSA module to enhance feature expression while simplifying network computation, and the HCA module to suppress redundant information, improving the quality of feature fusion. The mAP is increased by 6, 3.9, and 3.3 percentage points compared to the original YOLOv8 model, the MAIB model alone, and the DAFG model alone, respectively. When the loss function is replaced with HVCLoss, mAP is further improved by 1.8 percentage points. The model parameters and computational complexity are reduced by 76.3% and 59.4%, respectively, compared to the original YOLOv8 model, and the FPS is increased by 18 frames per second. After data enhancement was used on the original YOLOv8 model, the mAP increased by 27.5 percentage points, while the number of model parameters and model calculation increased by 3.2% and 2.1% respectively. Due to the expansion of the effective data set and the increase in feature information, the mAP increased, but the increase in the data set also led to a decrease in the number of model parameters and model calculation. Therefore, on the basis of data enhancement, the MAIB and DAFG modules were added and the HVC Loss was replaced. Compared with the YOLOv8 model using data enhancement, the mAP increased by 4.1 percentage points, the number of model parameters and model calculation decreased by 77% and 57.1% respectively, and the FPS increased by 19 frames / s, realizing real-time and accurate detection of surface defects of strip steel.

[0133] In this embodiment, Figure 7 The figure shows the comparison of the detection results of the original yolov8, the yolov8 after data enhancement, and the improved model of the present invention on three types of defects.

[0134] In summary, the MALDA-YOLO algorithm for strip surface defect detection based on multi-attention strategy fusion and lightweighting in the above embodiment belongs to the technical field of steel surface defect detection. Based on the YOLOv8 model framework, the present invention designs a multi-feature inverted bottleneck module MAIB based on the multi-scale attention mechanism GMSA, which replaces the C2f module in the backbone part for feature extraction. In the neck part, a feature fusion module DAFG based on deep separation attention (HCA) with grouped extraction is proposed. This module combines grouped convolution and reduced convolution operations to reduce the amount of computation and parameters while improving the quality of feature fusion and detection accuracy. In the head part, an improved HVC Loss is proposed to better handle sample imbalance, accelerate model convergence, enhance the detection capability of small targets, and reduce false detection and missed detection rates. The strip surface defect detection model designed by the present invention achieves a mAP of 82.9% on the NEU-DET dataset and a detection speed of 170FPS on an RTX3090, which is superior to all current strip surface defect detection models.

[0135] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting, characterized in that: The following steps are involved: S1: Data Collection Collect data samples from open source databases and steel plate surface defect data collected on industrial sites, and preprocess the samples; S2: Model construction Taking the YOLOv8 object detection network framework as the baseline model, all C2f modules in its backbone network are replaced with MAIB modules, and all C2f modules in the neck network are replaced with DAFG modules. The HVC Loss function is designed to replace the CIoU Loss function in the detection head network, thereby obtaining the MALDA-YOLO network. The MAIB module is a multi-feature inverted bottleneck module based on a multi-scale attention mechanism, and the DAFG module is a feature fusion module based on deep separation attention of group extraction. S3: Model training The MALDA-YOLO network is trained using the training set to obtain a strip surface defect detection model; S4: Defect Detection Test the strip surface defect detection model on the test set and output the detection results and detection speed.

2. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 1 is characterized in that: In step S1, the specific processing process is as follows: S11: Collect data samples from the open source database NEU-DET and steel plate surface defect data collected on industrial sites, and divide the acquired data samples into training and test sets according to the set ratio; S12: Perform data enhancement processing on the steel plate surface defect data image samples collected from the industrial field in the training set, and fill the size of the processed image samples to the set size.

3. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 1 is characterized in that: In step S2, the Bottleneck module of the C2f module in the backbone network is replaced by the MUIB module to form the MAIB module. In the MUIB module, the input features are processed in sequence by a 3×3 convolutional layer, a 1×1 convolutional layer, a 3×3 RepVGGDW convolutional layer, a GMSA module, a 1×1 convolutional layer, and a 3×3 convolutional layer to obtain output features.

4. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 3 is characterized in that: In the GMSA module, the specific processing process is as follows: S201: Divide the input features into two groups of features, perform channel attention calculation and spatial attention calculation on the first group of features, fuse the calculated weights, and use the GN layer to normalize them within each group to obtain feature F1; S202: Perform 1×1 convolution, 3×3 convolution, and 1×1 convolution on the second set of features to obtain feature F2; S203: For feature F1, use 2D global average pooling to encode global spatial information, and use 2D Gaussian mapping to fit linear transformation through natural nonlinear function Softmax, and perform feature attention feature fusion with feature F2 to obtain feature F ′ 1; S204: For feature F2, use 2D global average pooling to encode global spatial information, and use 2D Gaussian mapping to fit linear transformation through natural nonlinear function Softmax, and perform feature attention feature fusion with feature F1 to obtain feature F ′ 2; S205: Feature F ′ 1 and F ′ 2. Aggregate the attention weights, then process them with the Sigmoid function to obtain new attention weights, which are then multiplied with the input features and output.

5. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 4 is characterized in that: In step S2, in the DAFG module, the specific processing process is as follows: S211: The input features are first processed by 3x3 convolution, and then the features obtained after convolution are divided into two branches, ch1 and ch2, according to the number of channels; S212: For branch ch1, a 1x1 convolution is used for dimensionality reduction, a 3x3 depth-analyzable convolution layer based on hard_swish is used, and then the attention weight matrix Q is obtained through the HCA module. ch1 ; For branch ch2, 1x1 convolution is used for dimensionality reduction, and then the attention weight matrix Q is obtained through the HCA module ch2 ; S213: The attention weight matrix Q ch1 With Q ch2 Perform feature fusion and compare attention weight matrix Q ch1 With Q ch2 For values ​​of the same position, the maximum value is retained, and the fused features are dimensionally increased using 1x1 convolution to obtain the output features.

6. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 5 is characterized in that: In the HCA module, the specific processing process is as follows: The input features are grouped, fused after being processed by the maximum pooling layer and the average pooling layer respectively, and the fused results are passed through a 1x1 convolution and a hard-swish activation function, and then grouped again by a 1x1 convolution, and fused after being processed by the maximum pooling layer and the average pooling layer respectively. The fused results are weighted after being passed through the hard-sigmoid activation function, and multiplied with the input features to obtain the output features.

7. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 6 is characterized in that: In step S2, the calculation formula of the HVC Loss function is as follows: Among them, α i is the weight coefficient of sample i; p i is the predicted probability of sample i; y i is the true label of sample i, the positive sample is 1 and the negative sample is 0; γ is the focus parameter; λ is the weight parameter used to balance Varifocal Loss and CIoU Loss; f iou is the IoU value between the predicted box and the real box; ρ is the Euclidean distance between the center points of the predicted box and the real box; c is the length of the diagonal of the minimum circumscribed rectangular box; b and b g are the center points of the predicted box and the real box respectively; v is used to measure the aspect ratio consistency of the predicted box and the target box; μ is the parameter used to balance v.

8. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 7 is characterized in that: In the HVC Loss function, the weight parameter λ is introduced to control the loss's attention to positive and negative samples and aspect ratio. The calculation formula is as follows: Among them, δ(x) represents the Sigmoid function, N represents the total number of samples, N p Represents the number of positive samples, N n Represents the number of negative samples, and k is a tuning parameter used to control the sensitivity of weight adjustment.

9. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 1 is characterized in that: In step S3, the specific processing process is as follows: S31: Set model training parameters, including learning rate and batch size; S32: Input the training set into the MALDA-YOLO network for training. After the training is completed, save the weight parameters to obtain a strip surface defect target detection model.

10. The strip steel surface defect detection method based on multi-attention strategy fusion and lightweighting according to claim 1 is characterized in that: In step S4, the specific processing process is as follows: S41: Input the test set into the strip surface defect target detection model, and perform non-maximum suppression processing on the original output of the model; S42: Output the defect detection results and detection speed of the model.