Steel surface defect detection method based on ACW-YOLO algorithm
By introducing adaptive resolution attention module, C2f_Biformer module and WIoU loss function in the YOLOv5 detection model, the problem of low accuracy of existing steel surface defect detection methods is solved, and more efficient and accurate defect detection is achieved.
Patent Information
- Application Number
- CN202510135922.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The existing steel surface defect detection methods have low accuracy, are prone to missed and missed, and are susceptible to environmental interference.
The steel surface defect detection method based on the ACW-YOLO algorithm is adopted, and the YOLOv5 detection model is improved by introducing an adaptive resolution attention module, replacing the C3 module with a C2f_Biformer module, and introducing the WIoU loss function.
It improves the accuracy and efficiency of steel surface defect detection, enhances the performance of the model in complex scenarios, reduces calculation costs, and achieves a good balance between recognition accuracy and detection speed.
Smart Images

Figure CN120070361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and particularly relates to a method for detecting steel surface defects based on the ACW-YOLO algorithm. Background Art
[0002] Steel materials play an important role in many fields, such as construction, transportation, healthcare, or precision manufacturing. However, surface defects such as cracks, pores, and protrusions that may occur during the production process pose a threat to structural safety and affect economic efficiency. Traditional manual and infrared detection methods are difficult to meet the requirements of modern manufacturing due to low efficiency and high cost. Therefore, developing efficient steel surface defect detection technology is of great significance for ensuring industrial safety and economic efficiency.
[0003] Researchers are exploring new detection technologies to improve the accuracy, efficiency, and cost-effectiveness of defect detection in steel production. For example, Wu Xiuyong et al. optimized feature extraction using the Gaber wavelet kernel local preserving projection algorithm. Then, Zhao Jiuliang et al. developed a multi-scale edge detection algorithm based on wavelet transform. At the same time, Yang Yongmin et al. proposed an accurate image segmentation method using hyperentropy and fuzzy set theory. Although progress has been made, there is still a need to further improve the efficiency and accuracy of detection.
[0004] With the rise of deep learning technology, the importance of object detection algorithms in steel surface defect detection has become increasingly prominent, which can be divided into two-stage and single-stage algorithms. Two-stage algorithms based on R-CNN (Region-based Convolutional Neural Networks) and Faster R-CNN (Faster Region-based Convolutional Neural Networks) improve accuracy through the region proposal network, but the computational cost is very high and not suitable for real-time applications. Single-stage algorithms such as SSD (Single Shot MultiBox Detector) and YOLO (You Only Look Once) provide fast detection but low accuracy. Therefore, the present invention realizes a good balance between recognition accuracy and detection speed by improving the steel surface defect detection method of YOLOv5. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting steel surface defects based on the ACW-YOLO algorithm, so as to solve the technical problems of low accuracy, easy omission and misdetection, and susceptibility to environmental interference of the existing detection methods.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A steel surface defect detection method based on the ACW-YOLO algorithm provided by the present invention includes the following steps:
[0008] Step 1: Obtain image data of steel surface defects, preprocess the image data, and obtain a dataset of steel surface defect images;
[0009] Step 2: Annotate the images in the dataset and divide the dataset according to a ratio to obtain training data;
[0010] Step 3: Introduce an adaptive resolution attention module into the YOLOv5 backbone network, replace the C3 module in the YOLOv5 detection model with a C2f_Biformer module, and introduce a WIoU loss function to obtain an improved YOLOv5 detection model;
[0011] Step 4: Train the improved YOLOv5 detection model based on the training data to obtain a steel surface defect detection model;
[0012] Step 5: Test the existing defect detection model and compare it with the steel surface defect detection model obtained in Step 4, compare different attention mechanisms and loss functions, and then conduct ablation experiments to measure the improvement effect of the three improvements on improving the model accuracy. Record the test results after the test is completed;
[0013] Step 6: Evaluate the defect detection model according to the test results, and identify steel surface defects through the defect detection model that passes the evaluation; among them, the test results include accuracy, recall rate, and mAP value to verify whether the performance of the model needs to be improved.
[0014] Further, in Step 1, the defect types of the dataset include inclusions, scratches, pressed-in scale, cracks, pitting, and patches.
[0015] Further, in Step 2, the dataset is divided into a training set, a test set, and a validation set in a ratio of 8:1:1 for model training.
[0016] Further, the adaptive resolution attention module improves the computational efficiency and performance of the model by dynamically adjusting the resolution of the input feature map, including:
[0017] Use a 1x1 convolution for channel reduction to reduce computational complexity, and the mathematical expression is:
[0018]
[0019] where w kis the convolution kernel weight, x k is the channel of the input feature map, and K is the total number of channels of the input feature map;
[0020] Max pooling is used for downsampling to reduce the spatial dimension of the feature map. The mathematical formula for max pooling is:
[0021]
[0022] where, Ω represents the index set of the pooling window, and i and j respectively represent the row and column indices of the pooling window on the feature map;
[0023] The channel dimension is restored through 1x1 convolution, and the feature map is upsampled to the original size using bilinear interpolation. The mathematical expression for bilinear interpolation is:
[0024]
[0025] where, w ij is the interpolation weight, and I(x i , y j ) is the corresponding pixel value;
[0026] The Sigmoid function is used to generate attention weights. The Sigmoid function maps the input to the interval (0, 1), assigns different weights to each position of the feature map, and dynamically adjusts the importance of different regions;
[0027] The generated attention weights are applied to the original feature map through element-wise multiplication. The mathematical expression is:
[0028] Output = x ⊙ Attention.Weights
[0029] where, ⊙ represents element-wise multiplication.
[0030] Furthermore, the C2f_Biformer module combines convolution operations with the Transformer-based BiFormerBlock to achieve efficient feature extraction and fusion, including:
[0031] The channels of the input feature map are expanded through 1x1 convolution, and the number of channels of the expanded feature map increases to 2·c. The mathematical expression for convolution is:
[0032] y = W * x + b
[0033] where, W is the convolution kernel, x is the input feature map, and b is the bias term;
[0034] The expanded feature map is split into multiple parts in the channel dimension through the chunk operation, and each part enters the BiFormerBlock for processing in sequence. The BiFormerBlock is one of the core components of this module, which captures global and local dependencies through the self-attention mechanism. The mathematical expression of the self-attention mechanism is as follows:
[0035]
[0036] where Q, K, and V represent the query, key, and value matrices respectively, and d k is the scaling factor;
[0037] The processed multiple feature blocks are concatenated in the channel dimension, combined with the features output by different BiFormerBlocks, and form a feature representation with richer information;
[0038] The concatenated feature map compresses the number of channels back to the output channel number c2 through the second 1x1 convolution cv2 to obtain the final output feature. The mathematical expression of the convolution is:
[0039] y out = W 2 *concat(y 1 , y 2 ,..., y n )
[0040] where, W 2 is the weight matrix, and y n represents the output or feature map of the nth layer in the neural network.
[0041] Furthermore, the WIoU loss function optimizes the bounding box regression of the object detection model through a dynamic non-monotonic focusing mechanism; in WIoU, a minimum bounding box is defined. The minimum bounding box is a rectangle that simultaneously contains the anchor box B and the target box B gt . The anchor box B is defined by its center point (x, y) and width and height (W, H). The target box B gt is defined by its center point (x gt , y gt ) and width and height (w gt , h gt ). Denote the anchor box as the target box as The area S u of the minimum bounding box is obtained by subtracting the intersection area of the anchor box and the target box from the sum of their areas, that is:
[0042] S u = w × h + x gt × y gt - W i × Hi
[0043] Among them, W i and H i are respectively the width and height of the intersection area between the anchor box and the target box.
[0044] Furthermore, the WIoU loss function is designed as:
[0045]
[0046] Among them, is the loss term based on the outlier degree, and R i is a penalty term used to further optimize the localization performance of the model; The definition of
[0047]
[0048] When W i and H i are greater than 0, it means there is an overlap between the anchor box and the target box. At this time, will be adjusted according to the size of the overlapping area. If there is no overlap, that is, W i = 0 or H i = 0, then is 1, indicating the maximum loss;
[0049] In the WIoU loss function, there is also a key component used to measure the center point distance between the anchor box and the target box, and dynamically adjust the gradient gain of the loss function according to this distance. The value of quantifies the relative distance between the center point of the anchor box and the center point of the target box. Relative to the size of the target box, the smaller this value is, the closer the center point of the anchor box is to the center point of the target box, which means the higher the quality of the anchor box. On the contrary, if
[0050]
[0051] represents a loss function that combines the localization quality of the anchor box with the traditional intersection over union loss to evaluate the localization quality between the predicted box and the target box, so that the loss function can more precisely reflect the accuracy of the predicted box:
[0052]
[0053] Based on the above technical solutions, the embodiments of the present invention can at least produce the following technical effects:
[0054] (1) The steel surface defect detection method based on the ACW-YOLO algorithm provided by the present invention effectively applies the attention mechanism to the feature map through the adaptive resolution attention module, focuses on more important regions, improves the response of tiny targets of steel surface defects in the feature map, and enhances the performance of the model in complex scenarios. This design not only reduces the computational cost, but also enables the model to process the input features more efficiently by dynamically adjusting the resolution and focusing on the features of tiny targets of steel surface defects, thus improving the overall performance and computational efficiency.
[0055] (2) In the steel surface defect detection, the design of the C2f_Biformer module in the steel surface defect detection method based on the ACW-YOLO algorithm provided by the present invention is very advantageous. Through the local self-attention mechanism of the BiFormerBlock, the fine-grained features of tiny defects on the steel surface can be effectively captured. At the same time, the module can also handle the global dependencies of the targets. This is crucial in the detection of tiny defects on the steel surface because the saliency of small targets often depends on their context relationship with the surrounding environment. Through the splitting and splicing of features, the network can obtain more comprehensive feature information from multi-scale fusion, enhancing the detection effect on small targets. In addition, the combination of convolutional operations and self-attention enables the module to balance local and global features while ensuring computational efficiency, ensuring that the model has higher detection accuracy for tiny defects on the steel surface in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0057] Figure 1 is a schematic diagram of the network structure of the ACW-YOLO algorithm of the present invention;
[0058] Figure 2 is a schematic diagram of the principle structure of the adaptive resolution attention module of the present invention;
[0059] Figure 3 is a structural diagram of the C2f_Biformer module of the present invention;
[0060] Figure 4 is a concept diagram of IoU of the present invention;
[0061] Figure 5 is a concept diagram of WIoU of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0063] YOLOv5 is the fifth-generation object detection framework of the YOLO series, adopting a single-stage detection mechanism. It has an input layer for preprocessing, a backbone for feature extraction, a neck for multi-scale feature integration, and a prediction head for bounding boxes and probabilities. It provides multiple versions (YOLOv5s, YOLOv5m, YOLOv5l, YOLOv5x) to meet different computing requirements and application scenarios. The present invention introduces a method for detecting steel surface defects based on the ACW-YOLO algorithm, and its network framework diagram is as Figure 1 shown, and specifically includes the following steps:
[0064] Step 1: Obtain the image data of steel surface defects, preprocess the image data, and obtain a data set of steel surface defect images;
[0065] The defect types of this data set altogether include 6 types of defects: inclusion, scratch, pressed-in scale, crack, pitting, and plaque.
[0066] Step 2: Annotate the images in the data set and divide the data set according to a ratio to obtain training data;
[0067] Specifically, the data set is divided into a training set, a test set, and a validation set according to a ratio of 8:1:1 for model training.
[0068] Step 3: Introduce an adaptive resolution attention module into the YOLOv5 backbone network, replace the C3 module in the YOLOv5 detection model with a C2f Biformer module, and introduce a WIoU loss function to obtain an improved YOLOv5 detection model.
[0069] The adaptive resolution attention module improves the computational efficiency and performance of the model by dynamically adjusting the resolution of the input feature map, and its principle framework diagram is as Figure 2 shown. In the forward propagation process, first, 1x1 convolution is used for channel reduction to reduce the computational complexity, which can be expressed by the mathematical formula:
[0070]
[0071] Among them, w k is the convolution kernel weight, x k is the channel of the input feature map, and K is the total number of channels of the input feature map; this operation reduces the subsequent computational amount by reducing the number of channels;
[0072] Max pooling is used for downsampling to reduce the spatial dimension of the feature map. The mathematical formula of max pooling is:
[0073]
[0074] Among them, Ω represents the index set of the pooling window, and i and j respectively represent the row and column indices of the pooling window on the feature map; the pooling operation reduces the height and width of the feature map by half, providing a more compact feature representation for subsequent processing;
[0075] The channel dimension is restored through 1x1 convolution, and the feature map is upsampled to the original size using bilinear interpolation. The mathematical expression of bilinear interpolation is:
[0076]
[0077] Among them, w ij is the interpolation weight, and I(x i , y j ) is the corresponding pixel value; through upsampling, the adaptive resolution attention module restores the feature map to the same spatial resolution as the original input;
[0078] The Sigmoid function is used to generate attention weights. The formula of Sigmoid is:
[0079]
[0080] The Sigmoid function maps the input to the interval (0, 1), assigns different weights to each position of the feature map, and dynamically adjusts the importance of different regions;
[0081] The generated attention weights are applied to the original feature map through element-wise multiplication. The mathematical expression is:
[0082] Output = x ⊙ Attention_Weights,
[0083] Among them, ⊙ represents element-wise multiplication. The module can effectively apply the attention mechanism to the feature map, focus on more important regions, enhance the response of small targets of steel surface defects in the feature map, and improve the performance of the model in complex scenarios. This design not only reduces the computational cost, but also enables the model to process input features more efficiently by dynamically adjusting the resolution and focusing on the features of small targets of steel surface defects, improving the overall performance and computational efficiency.
[0084] Meanwhile, the model improves the original C3 module to the C2f_Biformer module. The C3 module of YOLOv5 adopts an efficient feature processing structure, similar to the idea of residual connection, and consists of three convolutional layers (CBS, convolutional-batchnormal-SiLu) and several bottleneck modules. The input features of the C3 module are divided into two parts. One part is directly processed by CBS, and the other part is processed by the combination of CBS and BottleNeck (bottleneck structure module). There are two forms of bottleneck modules: bottleneck 1 and bottleneck 2. BottleNeck1 reduces the number of channels in the feature map through 1x1 convolution and extracts features through 3x3 convolution, while BottleNeck2 undergoes one 1x1 convolution and one 3x3 convolution. This design not only reduces the computational complexity, but also realizes feature fusion through shortcut connections, enhancing the model's ability to capture features of different scales.
[0085] The C2f_Biformer module combines convolutional operations with the Transformer-based BiFormerBlock to achieve efficient feature extraction and fusion, showing significant advantages especially in the detection of small targets of steel surface defects. Its principle framework is as Figure 3 shown. The module first expands the channels of the input feature map through 1x1 convolution. The mathematical expression of the convolution is:
[0086] y = W * x + b,
[0087] where W is the convolution kernel, x is the input feature map, and b is the bias term; the number of channels of the expanded feature map increases to 2·c, which provides richer feature dimensions for subsequent feature processing;
[0088] The expanded feature map is divided into multiple parts in the channel dimension through the chunk operation, and each part enters the BiFormerBlock for processing in turn. The BiFormerBlock is one of the core components of this module, which captures global and local dependencies through the self-attention mechanism. The mathematical expression of the self-attention mechanism is:
[0089]
[0090] where Q, K, and V represent the query, key, and value matrices respectively, and d k is the scaling factor; the introduction of BiFormerBlock enables the network to capture fine-grained features within local windows and enhance the representation ability of target features through global dependence modeling;
[0091] The multiple processed feature blocks are concatenated in the channel dimension, combining the features output by different BiFormerBlocks to form a feature representation with richer information;
[0092] The concatenated feature map compresses the number of channels back to the output channel number c2 through the second 1x1 convolution cv2 to obtain the final output feature. The mathematical expression of the convolution is:
[0093] y out = W 2 *concat(y 1 , y 2 ,..., y n ),
[0094] where, W 2 is the weight matrix, and y n represents the output or feature map of the nth layer in the neural network.
[0095] IoU (Intersection over Union, intersection ratio) is an important metric in the field of object detection, used to evaluate the overlap degree between the predicted bounding box and the actual bounding box. It is a value between 0 and 1, where 1 indicates complete overlap and 0 indicates no overlap. Based on two rectangular regions: the prediction (A) and the ground truth (B). As Figure 4 shown, IoU calculates the ratio of the intersection and union of the anchor box B and the target box A. The calculation formula is as follows:
[0096]
[0097] WIoU is an innovative IoU-based loss function. As Figure 5 shown, it optimizes the bounding box regression of the object detection model through a dynamic non-monotonic focusing mechanism (FM). WIoU includes three different loss calculation strategies, namely WIoU-v1, WIoU-v2, and WIoU-v3. Each strategy aims to adjust the gradient allocation in different ways to adapt to training samples of different qualities. Experiments on the NEU-DET dataset show that the WIoU-v1 strategy exhibits the highest efficiency in loss calculation.
[0098] In WIoU, the smallest enclosing box is first defined, which is a rectangle that can contain both the anchor box (B) and the target box (B gt ) simultaneously. The anchor box B is defined by its center point (X, y) and width-height (W, H), while the target box B gt is defined by its center point (x gt , y gt ) and width-height (w gt : h gt ).
[0099] Denote the anchor box as and the target box as The area (S u ) of the smallest enclosing box is obtained by subtracting the intersection area of the anchor box and the target box from the sum of their areas, that is:
[0100] S u = w × h + x gt × y gt W i × H i
[0101] where, W i and H i are the width and height of the intersection region of the anchor box and the target box respectively.
[0102] Traditional IoU measurement methods face some problems during calculation. For example, when the anchor box and the target box have no overlap (W i = 0 or H i = 0), IoU is 0, resulting in the disappearance of the backpropagation gradient, thus unable to update the parameters W i ,
[0103]
[0104] To solve this problem, WIoU adopts a measurement based on outliers instead of directly using IoU. Outliers are measurements of the differences between the anchor box and the target box, which can be used to evaluate the quality of the anchor framework. Low outliers indicate high quality of the anchor box, and the gradient gain strategy optimizes the gradient distribution for different outlier anchor boxes through differential adjustment. Its advantage is that it strengthens the gradient retention of high-quality anchor frameworks while suppressing the gradients of low-quality anchor frameworks. This strategy particularly focuses on applying gradient gain to medium-quality anchor frameworks to improve the overall detection effect.
[0105] The loss function of WIoU is designed as:
[0106]
[0107] where, is a loss term based on outlier degree, while R i is a penalty term used to further optimize the localization performance of the model. is defined as follows:
[0108]
[0109] When W i and H i are greater than 0, it means there is an overlap between the anchor box and the target box. At this time, will be adjusted according to the size of the overlapping area. If there is no overlap, that is, W i = 0 or H i = 0, then is 1, indicating the maximum loss.
[0110] In addition, in the WIoU loss function, is a quantity used to measure the distance between the center point of the anchor box and the center point of the target box, and the gradient gain of the loss function is dynamically adjusted according to this distance. The value of quantifies the relative distance between the center point of the anchor box and the center point of the target box relative to the size of the target box. The smaller this value is, the closer the center point of the anchor box is to the center point of the target box, which means the higher the quality of the anchor box. On the contrary, if
[0111]
[0112] represents a loss function that combines the localization quality of the anchor box (reflected by ) with the traditional intersection over union loss to evaluate the localization quality between the predicted box and the target box, enabling the loss function to more precisely reflect the accuracy of the predicted box. In this way, during the training process, the model not only considers the overlapping degree between the predicted box and the ground truth box but also the localization accuracy of the center point of the predicted box. Through this method, helps to improve the localization performance of the model for the steel surface defect detection task. Especially when dealing with samples of different qualities, it can more effectively allocate gradients and optimize the learning process.
[0113]
[0114] WIoU reduces the gradient competition of high-quality anchor boxes by dynamically adjusting the gradient gain, and at the same time reduces the harmful gradients caused by low-quality samples, enabling the model to focus more on medium-quality anchor boxes. This strategy not only improves the robustness of the model to low-quality samples but also improves the overall performance of the detector.
[0115] Step 4: Based on the training data, train the improved YOLO v 5 detection model to obtain a defect detection model for the steel surface;
[0116] The operating system used in the experiment is the Windows operating system, NVIDIA GeForce RTX 3050 GPU, 16 GB of memory (4800 MHz), the deep learning framework is PyTorch 1.8 + cuda 11.1, the programming language is Python 3.8.10, the initial learning rate is set to 0.01, the number of worker threads for data loading is 4, the training batch size is set to 16500 iterations, and during the training process, it is executed (when the perfect match is achieved, the model stops the training process);
[0117] Step 5: Test the existing defect detection model and compare it with the defect detection model for the steel surface obtained in Step 4, compare different attention mechanisms and loss functions, and then conduct ablation experiments to measure the improvement effect of the three improvements on improving the model accuracy. Record the test results after the test is completed;
[0118] Step 6: Evaluate the defect detection model according to the test results, and use the defect detection model that passes the evaluation to identify defects on the steel surface; among them, the test results include accuracy, recall rate, and mAP value to verify whether the performance of the model needs to be improved.
[0119] Accuracy calculation formula: TP represents the number of true positives, while FP represents the number of false positives, and FN represents the number of samples that are mispredicted as negative classes, that is, the positive classes that the model fails to identify; Recall rate calculation formula: mAP calculation formula: C represents the set of categories, and AP is the average precision for category c.
[0120] In this embodiment, in order to verify the effectiveness of the method described in this solution, a comparative experiment is conducted between this method and the existing method, and the results are shown in Table 1 below:
[0121] Table 1 Comparative experiments of different methods on the NEU-DET dataset
[0122]
[0123] As can be seen from the table, the present invention significantly improves indicators such as the detection accuracy, recall rate, and mAP@0.5. The Gflops value of the algorithm of the present invention indicates that it is computationally relatively efficient and does not require a large number of floating-point operations to process images. This makes the algorithm more suitable for deployment in resource-constrained environments, such as mobile devices or edge computing devices.
[0124] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will also have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A steel surface defect detection method based on ACW-YOLO algorithm, characterized in that: The following steps are involved: Step 1: Acquire image data of steel surface defects, preprocess the image data, and obtain a data set of steel surface defect images; Step 2: Obtain training data by annotating the images in the dataset and dividing the dataset according to proportion; Step 3: Introduce the adaptive resolution attention module into the YOLOv5 backbone network, replace the C3 module in the YOLOv5 detection model with the C2f_Biformer module, and introduce the WIoU loss function to obtain the improved YOLOv5 detection model; Step 4: Train the improved YOLOv5 detection model based on the training data to obtain a defect detection model for the steel surface; Step 5: Test the existing defect detection model and compare it with the steel surface defect detection model obtained in step 4. Compare different attention mechanisms and loss functions. Then conduct an ablation experiment to measure the improvement effect of the three improvements on improving the model accuracy. After the test is completed, record the test results. Step 6: Evaluate the defect detection model based on the test results, and identify steel surface defects through the defect detection model that has passed the evaluation. The test results include accuracy, recall rate, and mAP value to verify whether the performance of the model needs to be improved.
2. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 1, characterized in that: In step 1, the defect types of the data set include inclusions, scratches, pressed oxide scales, cracks, pits, and plaques.
3. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 1, characterized in that: In step 2, the dataset is divided into training set, test set, and validation set in a ratio of 8:1:1 for model training.
4. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 1, characterized in that: The adaptive resolution attention module improves the computational efficiency and performance of the model by dynamically adjusting the resolution of the input feature map, including: Use 1x1 convolution to reduce the channel size to reduce the computational complexity. The mathematical expression is: Among them, w k is the convolution kernel weight, x k is the channel of the input feature map, K is the total number of channels of the input feature map; Use maximum pooling for downsampling to reduce the spatial dimension of the feature map. The mathematical formula for maximum pooling is: Among them, Ω represents the index set of the pooling window, i and j represent the row and column indexes of the pooling window on the feature map respectively; The channel dimension is restored through 1x1 convolution, and the feature map is upsampled to the original size using bilinear interpolation. The mathematical expression of bilinear interpolation is: Among them, w ij is the interpolation weight, I(x i ,y j ) is the corresponding pixel value; The attention weight is generated by the Sigmoid function. The Sigmoid function maps the input to the (0,1) interval, assigns different weights to each feature map position, and dynamically adjusts the importance of different areas. The generated attention weights are applied to the original feature map through element-wise multiplication, mathematically expressed as: Output=x☉Attention_Weights, where ⊙ represents element-wise multiplication.
5. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 1, characterized in that: The C2f_Biformer module combines convolution operations with the Transformer-based BiFormerBlock to achieve efficient feature extraction and fusion, including: The channels of the input feature map are expanded by 1x1 convolution. The number of channels of the expanded feature map increases to 2·c. The mathematical expression of convolution is: y=W*x+b, Among them, W is the convolution kernel, x is the input feature map, and b is the bias term; The expanded feature map is divided into multiple parts in the channel dimension through the chunk operation, and each part enters the BiFormerBlock for processing in turn. BiFormerBlock is one of the core components of this module, which captures global and local dependencies through the self-attention mechanism. The mathematical expression of the self-attention mechanism is: Where Q, K, and V represent query, key, and value matrices, respectively. k is the scaling factor; The processed multiple feature blocks are spliced in the channel dimension, combining the features output by different BiFormerBlocks to form a feature expression with richer information; The concatenated feature map is compressed back to the output channel number c2 through the second 1x1 convolution cv2 to obtain the final output feature. The mathematical expression of the convolution is: and out =W2*concat(y1,y2,...,y n ), Among them, W2 is the weight matrix, y n is the output or feature map of the nth layer in the neural network.
6. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 1, characterized in that: The WIoU loss function optimizes the bounding box regression of the target detection model through a dynamic non-monotonic focusing mechanism; in WIoU, a minimum bounding box is defined, which contains both the anchor box B and the target box B gt The anchor box B is defined by its center point (x, y) and width and height (W, H), and the target box B gt From its center point (x gt ,y gt ) and width and height (w gt ,h gt ) definition, the anchor box is The target frame is The area of the minimum bounding box S u It is obtained by subtracting the intersection area of the anchor box and the target box from the sum of their areas, that is: S u =w×h+x gt ×y gt -W i ×H i , Among them, W i and H i are the width and height of the intersection area of the anchor box and the target box, respectively.
7. The steel surface defect detection method based on ACW-YOLO algorithm according to claim 6, characterized in that: The WIoU loss function is designed as: in, is the loss term based on outlier degree, and R i is a penalty term used to further optimize the positioning performance of the model; is defined as follows: When W i and H i When it is greater than 0, it means that the anchor box and the target box overlap. It will be adjusted according to the size of the overlapping area. If there is no overlap, that is, W i =0 or H i =0, then 1 means the loss is the largest; In the WIoU loss function, a key component is also included It is used to measure the distance between the center point of the anchor box and the target box, and dynamically adjust the gradient gain of the loss function according to this distance. The value of quantifies the relative distance between the center point of the anchor box and the center point of the target box. Relative to the size of the target box, the smaller this value is, the closer the center point of the anchor box is to the center point of the target box, which means the quality of the anchor box is higher. On the contrary, if If the value of is large, it means that the positioning of the anchor box is not accurate enough. The expression is: It represents a loss function that compares the positioning quality of the anchor box with the traditional intersection loss Combined with the prediction box, the positioning quality between the prediction box and the target box is evaluated, so that the loss function can more finely reflect the accuracy of the prediction box. The expression is:
Citation Information
Patent Citations
Multi-scale attention-fused traffic helmet small target detection system and method
CN116665156A
Multi-scale nixie tube detection method based on improved YOLO adaptive attention-feature enhancement network
CN117095155A
Power transmission line icing detection method of improved YOLOv8 network
CN117911837A
Steel surface defect detection method and device based on YOLOv5 improvement
CN118134850A
Defective solder ball detection method based on improved YOLO v8 network
CN119048516A