Strip steel surface defect detection method integrating global local sensing and layered feature fusion
By integrating the method of integrating global local perception and layered feature fusion, the accuracy and adaptability of strip surface defect detection are improved, the shortcomings of existing methods in background noise and unstructured defect detection are solved, and more efficient feature extraction and detection effects are achieved.
Patent Information
- Application Number
- CN202510136687.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
Existing strip surface defect detection methods are susceptible to interference from background noise and irrelevant information, it is difficult to effectively extract defect characteristics, and the detection ability of unstructured defects is insufficient.
The strip surface defect detection method that integrates global local perception and hierarchical features is adopted. Through feature enhancement modules, hierarchical fusion networks and global local perception modules, the model's ability to extract defect features and detect unstructured defects is improved.
The model's detection accuracy of strip surface defects is improved, especially in the detection of complex backgrounds and unstructured defects, and the model's consumption of computing resources is reduced, making it more suitable for actual industrial scenarios.
Smart Images

Figure CN120070365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to a strip surface defect detection method integrating global-local perception and hierarchical feature fusion. Background Art
[0002] During the strip production process, defects such as scratches, cracks, and patches are inevitable on the surface. If these defects are not detected and processed in time, they may cause fracture or corrosion problems during the use of the finished strip, thereby posing a potential threat to the safety and service life of the project. Therefore, designing an efficient and accurate strip surface defect detection method has important practical significance for improving the strip production quality and ensuring project safety.
[0003] Currently, strip surface defect detection algorithms are mainly divided into two-stage algorithms and one-stage algorithms. Two-stage object detection algorithms represented by Faster-RCNN and Mask-RCNN have attracted much attention due to their excellent detection effects. Such algorithms divide the object detection task into two stages: in the first stage, the original image is decomposed into multiple candidate regions that may generate objects through a region proposal network (RPN); in the second stage, a regression loss function is used to accurately locate the object, and a classification loss function is used to identify the object category. Two-stage algorithms have the main advantages of accurate positioning and high detection accuracy and perform excellently in detection tasks. However, their multi-stage inference process and high computational complexity limit the detection speed and it is difficult to meet the requirements of real-time detection of strip surface defects.
[0004] In contrast, one-stage object detection algorithms represented by SSD and YOLO have more advantages in terms of detection speed and model complexity. By simultaneously completing object detection and feature extraction through a single network, one-stage algorithms significantly improve the inference efficiency. Their regression mechanism directly predicts the object position and category from the input image, reducing the computational overhead of multiple inferences, so they perform excellently in high-real-time applications. Although the detection accuracy of one-stage algorithms is slightly lower than that of two-stage algorithms, their significant advantage in detection speed makes this accuracy loss negligible in practical applications.
[0005] With the development of single-stage target detection algorithms, the YOLO series of algorithms have achieved a good balance between speed and accuracy with their innovative design, and have gradually become the mainstream method in the field of target detection. In strip surface defect detection, the YOLO series of algorithms also show excellent performance. The algorithm achieves efficient reasoning and low resource consumption through a lightweight network structure, fully meeting the dual requirements of industrial detection for real-time performance and accuracy. Therefore, many scholars have optimized and improved it based on this, and successfully applied it to strip surface defect detection, promoting the further development of industrial detection technology.
[0006] In the related technologies, Wang Yanshu et al. proposed an adaptive tree-based candidate box extraction network based on tree-based Parzen estimation in the YOLOv3 model, which realized the dynamic optimization of anchor box parameters. In addition, they designed a global positioning regression algorithm, which broke through the limitations of the traditional anchor box mechanism by accurately calculating the offset between the feature map unit and the true calibration box. Li et al. developed an innovative structure based on YOLOv4, which introduced an enhancement path composed of convolution encoder-decoder modules in the residual block, further enhancing the feature representation capability. In addition, they designed a multi-scale module based on hierarchical residual connection to expand the receptive field, thereby effectively improving the detection accuracy. Chen et al. introduced deformable convolution and adaptive histogram equalization data enhancement methods in the YOLOX model, which effectively dealt with the problems of irregular shape of strip surface defects and low image contrast. Li et al. adopted a multi-scale feature extraction module, realized multi-scale feature extraction through a three-branch structure of convolution kernels of different sizes, and used a feature fusion module to integrate the backbone network and neck network features, thereby improving the model's detection capability for complex scale defects. Dong Yongfeng et al. proposed the MADD-Net model, which optimizes the segmentation and classification networks simultaneously by jointly optimizing the objective function. Based on the spatial attention module with multiple receptive fields, the model's receptive field and ability to perceive tiny defects are improved through dilated convolutions with different expansion rates.
[0007] The above research has improved the accuracy of model detection to a certain extent, but there are still many problems to be solved. First, the existing defect detection algorithms are often interfered by background noise and irrelevant information, making it difficult to fully extract defect features. Secondly, the current mainstream models usually introduce feature pyramid structures to solve the problem of scale complexity of strip surface defects. However, when fusing multi-scale features, such methods do not fully consider the problem of insufficient information transmission between non-adjacent features, thereby limiting the effect of feature fusion. In addition, the unstructured characteristics of strip surface defects further increase the difficulty of detection, especially when the defect features show a fragmented appearance or irregular distribution, the adaptability of the existing models is obviously insufficient. Summary of the invention
[0008] In view of this, the present invention provides a strip surface defect detection method integrating global-local perception and hierarchical feature fusion, which can solve the problems that the strip surface defect detection of the existing method is easily interfered by background noise and irrelevant information, it is difficult to effectively extract defect features, and the detection ability for unstructured defects is insufficient.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] An embodiment of the present invention provides a strip surface defect detection method integrating global-local perception and hierarchical feature fusion, including the following steps:
[0011] S1. Obtain the strip surface image to be detected;
[0012] S2. Input the strip surface image to be detected into the trained strip surface defect detection model; the strip surface defect detection model is YOLOv8n, including: multiple feature enhancement modules, a hierarchical fusion network, and multiple global-local perception modules; the feature enhancement module replaces the original C2f module in its backbone network;
[0013] S3. Output the corresponding detection result.
[0014] Further, the data processing process of the strip surface defect detection model includes:
[0015] S21. Use the feature enhancement module in the backbone network to extract features from the input image; the feature enhancement module includes two convolutional layers, a splitting operation, multiple Bottleneck modules, and a splicing operation, and realizes the enhancement of target features and the suppression of background noise by dynamically adjusting feature weights;
[0016] S22. Use the hierarchical fusion network in the neck network to perform multi-scale feature fusion on the features output by the feature enhancement module in the backbone network, gradually fuse adjacent layer features layer by layer, and support direct interaction of non-adjacent features; then input the features into the feature enhancement module in the neck network through upsampling and splicing operations;
[0017] S23. Use the global-local perception module to perform global and local feature fusion on the features output by the feature enhancement module in the neck network; input the features output by the global-local perception module into the detection head for classification and regression to obtain the category and bounding box position of the defect.
[0018] Further, the step S21 includes:
[0019] The feature enhancement module in the backbone network performs preliminary extraction and channel transformation on the input feature map through convolution operations to generate an intermediate feature representation X;
[0020] Divide X into two sub - feature maps along the channel dimension through a splitting operation, where one sub - feature map X 0 After being processed by the Bottleneck module, it serves as the initial input for the subsequent Bottleneck; the features generated by each operation are described by the recurrence formula in Equation (1):
[0021] X i =CBS 2 (X i-1 ) + SE(X i-1 ) + X i-1 , i = 1, 2,..., n(1)
[0022] Among them, CBS 2 (·) represents the feature mapping function of two convolution operations, and SE(·) is the feature mapping function of the channel attention module;
[0023] The feature enhancement module in the backbone network concatenates all deep features (X 1 , X 2 , …, X n ) with the original feature map X to generate the final feature representation Y, and its calculation formula is as follows:
[0024] Y = CBS(Concat(X, X 1 , X 2 ,..., X n ))(2)
[0025] In the formula, Concat represents the concatenation operation; CBS represents the convolution operation.
[0026] Furthermore, the processing process in the feature mapping function of the channel attention module includes:
[0027] Compress the spatial information of each channel into global semantic features through global average pooling operation;
[0028] Successively perform two - layer fully - connected network operations to generate channel weights, and normalize them using the Sigmoid activation function.
[0029] Furthermore, the step S22 includes:
[0030] 1) Feature processing:
[0031] The backbone network extracts multi - scale feature representations of the input image as C3, C4, C5, corresponding to low - dimensional, medium - dimensional, and high - dimensional feature maps respectively;
[0032] 2) Hierarchical fusion:
[0033] Use formulas (3) and (4) to perform feature fusion of C3 and C4:
[0034] C3=Concat(conv 1 (C3),↑ 2 (C4))(3)
[0035] C4=Concat(↓ 2 (C3),conv 1 (C4))(4)
[0036] ↑ 2 Indicates that the features are upsampled using 1×1 convolution and bilinear interpolation methods, and 2 is the upsampling ratio; conv 1 Indicates using 1×1 convolution to adjust the number of channels;↓ 2 It means that downsampling is achieved by using 3×3 convolution with stride lengths of 2 and 4 respectively, and 2 is the downsampling ratio; Concat means the concatenation operation to obtain the fused features;
[0037] Use formula (5)-formula (7) to fuse C3, C4 and C5 features:
[0038] P3 = Concat(conv 1 (C3),↑ 2 (C4),↑ 4 (C5))(5)
[0039] P4=Concat(↓ 2 (C3),conv 1 (C4),↑ 2 (C5))(6)
[0040] P5=Concat(↓ 4 (C3),↓ 2 (C4),conv 1 (C5))(7)
[0041] ↑ 4 Indicates that the features are upsampled using 1×1 convolution and bilinear interpolation methods, and 4 is the upsampling ratio; conv 1 Indicates using 1×1 convolution to adjust the number of channels;↓ 4 It means that downsampling is achieved by using 3×3 convolution with stride of 2 and 4 respectively, and 4 is the downsampling ratio;
[0042] 3) Fusion feature output:
[0043] After the above layer-by-layer fusion process, a new set of multi-scale fusion features P3, P4, and P5 are obtained; each feature contains rich information of different scales, and the semantic gap between levels is effectively narrowed.
[0044] Further, the step S23 includes:
[0045] 1) Local feature extraction:
[0046] For the feature X output by the feature enhancement module in the neck network, global average pooling is performed on each channel to compress the spatial dimension features into a scalar, forming a channel description vector;
[0047] Perform a one-dimensional convolution operation on the channel description vector to construct a local dependence model; and dynamically adjust the convolution kernel size according to the number of channels;
[0048] The convolution result is transformed through the Sigmoid activation function to generate the weight w of each channel L ; All weights act on the input feature X in a channel-by-channel weighted manner to obtain the local feature representation F L ;
[0049] 2) Global feature extraction:
[0050] The feature X generates query Q, key K, and value V feature maps through three independent convolutional layers respectively;
[0051] Reshape the Q, K, and V feature maps into matrices with the shape of HW×C; H and W are the spatial dimensions of the feature map, and C is the number of channels;
[0052] Divide the Q, K, and V feature maps along the channel dimension into h subspaces, each subspace containing C i channels; within each subspace, perform a dot product operation on Q and K to generate a self-attention relationship matrix; multiply the self-attention relationship matrix by the corresponding V to obtain the self-attention feature F of this subspace i G ;
[0053] Concatenate the self-attention features of all h subspaces along the channel dimension to integrate into the final global feature representation F G ;
[0054] 3) Feature fusion:
[0055] Concatenate the local feature F L and the global feature F G along the channel dimension, use 1x1 convolution to adjust the number of channels and perform weight allocation to generate the final output feature.
[0056] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following advantages:
[0057] This method improves the detection accuracy of the model for strip surface defects, especially in the detection of complex backgrounds and unstructured defects, reduces the consumption of computing resources by the model, makes it more suitable for actual industrial scenarios, and further improves the detection ability of the model for defects of different scales. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0059] Figure 1 It is a flowchart of the strip surface defect detection method integrating global-local perception and hierarchical feature fusion provided by the present invention.
[0060] Figure 2 It is an overall architecture diagram of the strip surface defect detection model provided by the present invention.
[0061] Figure 3 It is a data processing process diagram of the strip surface defect detection model provided by the present invention.
[0062] Figure 4 It is a structural diagram of the hierarchical fusion network provided by the present invention.
[0063] Figure 5 It is a structural diagram of the global-local perception module provided by the present invention.
[0064] Figure 6 It is a heat map of the model for detecting strip surface defects provided by the present invention.
[0065] Figure 7 It is a comparison diagram of the detection results of multiple models provided by the present invention. Detailed Embodiments
[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0067] Referring to Figure 1 As shown, the embodiments of the present invention disclose a strip surface defect detection method integrating global-local perception and hierarchical feature fusion, including the following steps:
[0068] S1. Obtain the strip surface image to be detected;
[0069] S2. Input the strip surface image to be detected into the trained strip surface defect detection model; the strip surface defect detection model is YOLOv8n, including: multiple feature enhancement modules, a hierarchical fusion network, and multiple global-local awareness modules; the feature enhancement module replaces the original C2f module in its backbone network;
[0070] S3. Output the corresponding detection result.
[0071] In step S2 of the present invention, first, aiming at the problem that the strip defect detection result is vulnerable to background interference, a feature enhancement module (Feature Enhanced Module, FEM) is proposed. Its core lies in dynamically adjusting the feature weights to strengthen the target features and suppress the background noise, improving the model's attention to defect information, thereby enhancing the model's feature extraction effect. Then, to alleviate the problem of insufficient information transmission between non-adjacent features during the feature fusion process, a hierarchical fusion network (Hierarchical Fusion Network, HFN) is designed. By hierarchically and gradually fusing adjacent layer features and supporting direct interaction between non-adjacent features, more efficient feature fusion is achieved. Finally, aiming at the unstructured defects on the strip surface and their irregular shape distributions, a global-local awareness module (Global-Local Awareness Module, GLAM) is proposed. This module enhances the model's understanding ability of the macroscopic structure and local features of defects by integrating global context information and local features, thereby improving the model's adaptability to complex and unstructured defects.
[0072] Refer to Figure 2 As shown, it is the overall architecture of the strip surface defect detection model; first, in the Backbone part, this embodiment uses the FEM module to replace the C2f module in the original network, reducing background noise information and enhancing the model's attention to defect features. Then, the feature layer processed by the Backbone is input into the HFN module. By hierarchically and gradually fusing and supporting direct interaction between non-adjacent feature layers, the model's performance in object detection at different scales is further improved. Subsequently, the GLAM module is introduced to capture global and local semantic information, thereby enhancing the model's processing ability for unstructured defects. Finally, the feature map processed by the above modules is input into the decoupled detection head to perform classification and regression tasks to determine the category and bounding box position of the defect.
[0073] For the specific data processing process, refer to Figure 3 As shown, it includes:
[0074] S21. Use the feature enhancement module in the backbone network to extract features from the input image; the feature enhancement module includes two convolutional layers, a split operation, multiple Bottleneck modules, and a concatenation operation, and realizes the enhancement of target features and the suppression of background noise by dynamically adjusting feature weights;
[0075] S22. Use the hierarchical fusion network in the neck network to perform multi-scale feature fusion on the features output by the feature enhancement module in the backbone network, gradually fuse adjacent layer features layer by layer, and support direct interaction of non-adjacent features; then input the features into the feature enhancement module in the neck network through upsampling and concatenation operations;
[0076] S23. Use the global-local perception module to perform global and local feature fusion on the features output by the feature enhancement module in the neck network; input the features output by the global-local perception module into the detection head for classification and regression to obtain the category and bounding box position of the defect.
[0077] The above steps S21 to S23 are described in detail below:
[0078] 1. Refer to Figure 3 As shown, it is the feature enhancement module in step S21:
[0079] In strip surface defect detection, background interference is an important factor affecting detection performance. Strip surface defects usually have significant local features, and the importance of these features in different channels varies. To improve detection performance, the response of channels related to defect features should be highlighted, while suppressing the response of channels dominated by background noise. However, the feature expression ability of traditional convolution operations in different channels is equal, and it is unable to dynamically adjust the response to defect features according to the characteristics of the input data, thus limiting the detection performance to a certain extent.
[0080] To alleviate the above problems, the present invention designs a feature enhancement module FEM. Its detailed module structure is as shown in Figure 3 shown. This module consists of two convolutional layers (CBS), a split operation (Split), multiple Bottleneck modules, and a concatenation operation (Concat), aiming to enhance the model's response ability to defect features through an efficient branch feature processing mechanism and a lightweight SE attention module.
[0081] Specifically, the FEM module first performs preliminary extraction and channel transformation on the input feature map through convolution operations to generate an intermediate feature representation X. Subsequently, X is evenly divided into two sub-feature maps along the channel dimension through a split operation, where one sub-feature map X 0After being processed by the Bottleneck module, it serves as the initial input for the subsequent Bottleneck. Specifically, the features generated by each operation can be described by the recurrence formula in Equation (1):
[0082] X i = CBS 2 (X i-1 ) + SE(X i-1 ) + X i-1 , i = 1, 2,..., n (1)
[0083] Among them, CBS 2 (·) represents the feature mapping function of two convolutional operations, and SE(·) is the feature mapping function of the channel attention module. For the SE module, first, through the global average pooling operation, the spatial information of each channel is compressed into global semantic features. This process extracts the overall structural information of the image and avoids the interference of local feature redundancy on the expression of global features. Subsequently, two fully connected network operations are sequentially performed to generate channel weights, and the Sigmoid activation function is used to normalize them. This weight allocation mechanism essentially realizes the dynamic adjustment of features, enabling the model to optimize the response to defect features in real-time according to the characteristics of the input data. Finally, the FEM module cascades all deep features (X 1 , X 2 , …, X n ) with the original feature map X to generate the final feature representation Y, and its calculation formula is as follows:
[0084] Y = CBS(Concat(X, X 1 , X 2 ,..., X n )) (2)
[0085] Compared with the traditional feature stacking method, this splitting and splicing operation can capture the diversity of defect features more comprehensively, while retaining the local details of the defect area and avoiding information loss during the fusion process. Through this design, the feature identification ability of the model is significantly improved, enabling it to more accurately distinguish target features from interference information in complex backgrounds. After integrating the FEM module into YOLOv8, the feature extraction ability and detection accuracy of the model are further improved. Compared with the C2f module in the original network, the FEM module only introduces a small number of additional training parameters, showing better performance and having less impact on the detection speed. Therefore, it has certain application potential under complex background conditions.
[0086] 2. The hierarchical fusion network in step S22:
[0087] The Feature Pyramid Network (FPN) is a classic feature fusion method that propagates semantic information step by step through a top-down feature transfer mechanism and uses lateral connections to achieve the fusion of features at different scales. However, during the fusion process, FPN mainly focuses on the interaction of features in adjacent layers and lacks consideration for the interaction between non-adjacent features. High-level semantic information must be transmitted to the low level through intermediate layers step by step, and during this multi-level transmission process, high-level features may experience information loss or degradation due to layer-by-layer processing. Similarly, when the Path Aggregation Network (PANet) transfers low-level feature details to the high level through a bottom-up path, it also faces similar problems, thus weakening the effect of non-adjacent layer feature fusion.
[0088] To alleviate the above problems, the present invention proposes a Hierarchical Fusion Network (HFN). Inspired by the dynamic fusion method of the Asymptotic Feature Pyramid Network (AFPN) for object detection, HFN fuses adjacent-dimensional features in a hierarchical and step-by-step manner and supports direct interaction between non-adjacent feature layers. The detailed module structure is as Figure 4 shown. The backbone network extracts multi-scale feature representations of the input image as C3, C4, and C5, corresponding to low-dimensional, medium-dimensional, and high-dimensional feature maps respectively. To align the feature scales and dimensions and prepare for feature fusion, the following operations are adopted in this embodiment to process the features:
[0089] Upsample the features using a 1×1 convolution and bilinear interpolation, denoted as ↑ s , where s is the upsampling ratio; use a 1×1 convolution conv 1 to adjust the number of channels; use 3×3 convolutions with strides of 2 and 4 respectively to perform downsampling, denoted as ↓ s , where s is the downsampling ratio; adopt the Concat operation to splice the features with aligned sizes in the channel dimension. HFN first fully fuses adjacent low-dimensional features as shown in Equations (3) and (4):
[0090] C3 = Concat(conv 1 (C3), ↑ 2 (C4)) (3)
[0091] C4 = Concat(↓ 2 (C3), conv 1 (C4)) (4)
[0092] After the fusion of C3 and C4 features is completed, the high-dimensional feature C5, i.e., the most abstract feature layer, is gradually introduced, as shown in Equations (5), (6), and (7). This will make the semantic information of features at different levels closer during the fusion process.
[0093] P3 = Concat(conv 1 (C3), ↑ 2 (C4), ↑ 4 (C5))(5)
[0094] P4 = Concat(↓ 2 (C3), conv 1 (C4), ↑ 2 (C5))(6)
[0095] P5 = Concat(↓ 4 (C3), ↓ 2 (C4), conv 1 (C5))(7)
[0096] Through the above layer-by-layer fusion process, a new set of multi-scale fusion features P3, P4, and P5 is obtained. Each feature contains rich information at different scales, and at the same time, the semantic gap between levels is effectively reduced. Compared with the traditional FPN and PANet structures, HFN avoids the information loss problem caused by the fusion of non-adjacent level features through a layer-by-layer progressive feature fusion strategy, further improving the fusion effect of multi-scale features. In addition, HFN performs excellently in terms of computational efficiency. The generated multi-scale features not only retain the detailed information of the underlying features but also possess the semantic abstraction expression of high-level features, thereby enhancing the model's detection ability for defects at different scales.
[0097] 3. The global-local perception module in step S23:
[0098] Deep learning detection algorithms based on convolutional neural networks usually rely on convolutional kernels to extract local features of images. However, for some unstructured defects on the steel surface, it is difficult to achieve accurate detection relying only on local information. Such defects often span a large range and lack fixed shape and distribution characteristics, making it difficult for traditional methods to capture their global characteristics. In contrast, the multi-head self-attention mechanism (MHA) in Transformer can effectively capture global information but tends to ignore local details. Therefore, to make full use of their complementary advantages, this embodiment proposes a global-local awareness module (GLAM), which combines the advantages of MHA and CNN. GLAM gives full play to the global perception ability of MHA in capturing long-range dependencies and global context information, and at the same time utilizes the sharp local detail capture ability of CNN to enhance the model's detection ability for complex defects. The structure of the detailed module is as Figure 5 shown.
[0099] First, global average pooling is performed on the feature map X to compress the spatial dimension features of each channel into a scalar, thereby generating a channel description vector. Then, one-dimensional convolution is used to model local dependencies in the channel dimension, and the convolution kernel size is dynamically adjusted to adapt to different numbers of channels. The convolution result is converted into the weight w of each channel through the Sigmoid activation function L , and these weights act on the input feature map in a channel-wise weighted manner, where σ is the activation function and ⊙ is broadcast multiplication, so as to obtain the local feature representation F L .
[0100] w L = σ(Conv(AvgPool(X)))(8)
[0101] F L = w L ⊙ X(9)
[0102] Next, the feature map X generates query Q, key K, and value V feature maps through three independent convolutional layers respectively. Subsequently, these feature maps are reshaped into matrices with the shape of HW×C, where H and W are the spatial dimensions of the feature map and C represents the number of channels. This operation aims to provide a convenient representation form for subsequent matrix operations. Then, Q, K, and V are evenly divided into h subspaces along the channel dimension, denoted as Q i , K i , V i , each subspace contains C i channels, where i = 1, 2,..., h. Figure 5Part (b) details the specific implementation under a subspace. Through this multi-head partitioning method, the model can parallelly calculate the relationships between features within different subspaces, thereby capturing semantic information at different levels from multiple perspectives. For each subspace, Q i and K i perform a dot product operation to generate a self-attention relationship matrix, which reflects the global dependencies between different positions in the input feature map. Subsequently, the generated relationship matrix is multiplied by the corresponding V i to obtain the self-attention features of this subspace In this way, each head independently generates a feature containing global context information, reflecting the feature interactions and semantic expressions within a specific subspace. Finally, the self-attention features of all h subspaces are concatenated along the channel dimension to be integrated into the final global feature representation F G .
[0103]
[0104] Finally, the features processed locally and globally are concatenated along the channel dimension, and the number of channels is adjusted and weight distribution is performed through 1×1 convolution to generate the final output features. The GLAM module effectively integrates the advantages of CNN and Transformer in processing local and global information, enabling the model to fully focus on context information and local details, thereby improving the accuracy of defect detection.
[0105] Next, the advantages of the technical solution of the present invention will be illustrated through specific experiments:
[0106] 4. Experimental Results and Analysis:
[0107] In this embodiment, based on the Python-based Pytorch experimental platform, the PC is configured with an Intel i9-12900KF CPU, an NVIDIA GeForce RTX 3090Ti GPU, a CUDA version of 11.6, and a cuDNN version of 8.9. During the training process, the image input size is 640×640, the batch size is 32, the initial learning rate is 0.01, the Stochastic Gradient Descent (SGD) optimizer is used to optimize the network parameters, the momentum is set to 0.937, and the number of training epochs is set to 200.
[0108] 4.1 Introduction to the Dataset
[0109] This embodiment conducts research using the NEU-DET strip surface defect dataset open-sourced by Northeastern University and the GC10-DET steel surface defect dataset collected in real industry. The NEU-DET dataset contains 6 types of strip surface defects, namely Crazing (Cr), Inclusion (In), Patches (Pa), Pitted Surface (Ps), Rolled-in Scale (Rs), and Scratches (Sc). There are 300 defect images for each type, totaling 1800 images.
[0110] The GC10-DET dataset covers 10 types of steel surface defects, namely Punching (Pu), Welding Line (Wl), Crescent Gap (Cg), Water Spot (Ws), Oil Spot (Os), Silk Spot (Ss), Inclusion (In), Rolled Pit (Rp), Crease (Cr), and Waist Folding (Wf), with a total of 2294 images. Both datasets are randomly generated into training sets, validation sets, and test sets at a ratio of 8:1:1.
[0111] 4.2 Evaluation Metrics
[0112] In object detection, when evaluating the network performance, both Precision (P) and Recall (R) need to be considered simultaneously. Therefore, this embodiment uses average precision (AP) as the evaluation metric for each defect category, and mean average precision (mAP) to evaluate the performance of the entire network model. AP represents the detection accuracy of a certain type of defect, and its calculation formula is shown in Equation (12). mAP is the mean of the detection accuracies of all categories, and its calculation formula is shown in Equation (13):
[0113]
[0114] Among them, i represents a certain type of defect, AP(i) is the detection accuracy of a certain type, and n is the total number of categories. P represents precision, R represents recall, and the calculation formulas for P and R are respectively:
[0115]
[0116] Wherein, TP is the correctly predicted positive example, FN is the positive example misjudged as a negative example, and FP is the negative example misjudged as a positive example. TP + FP represents the total number of positive samples detected, and TP + FN represents the total number of all positive samples. By calculating IoU and setting a threshold, it is determined whether the classification result is correct, thereby determining the quantities of TP, FP, and FN.
[0117] 4.3 Ablation Experiment
[0118] To verify the effectiveness of the method in this embodiment, the present invention uses YOLOv8n as the baseline model and conducts progressive performance tests on each improvement point on the NEU-DET dataset, including the FEM, HFN, and GLAM modules. Table 1 shows the ablation experiment results after adding each improvement point. The first row is the performance result of the baseline model, and the symbol represents the network model after adding this improvement method.
[0119] Table 1 Ablation Experiment Results on NEU-DET Dataset
[0120]
[0121] From the experimental results in Table 1, it can be seen that the feature enhancement module (FEM), hierarchical fusion network (HFN), and global-local awareness module (GLAM) proposed by the method of the present invention have significantly improved the detection performance of the network. Compared with the original YOLOv8n, the method of the present invention has increased the mean average precision (mAP) by 3.8%.
[0122] First, in terms of feature extraction, the FEM module avoids the interference of local feature redundancy on the global feature expression. At the same time, by quantifying the importance of each channel feature map and performing weighted processing, it effectively filters background and noise information, enhances the model's attention to defect features, thereby improving the network's feature extraction ability and enhancing the robustness of the model in different background and noise environments. The experimental results show that after adding the FEM module, the mAP of the model has increased by 1.2%, while the number of parameters has only increased from 3.157M to
[0123] 3.161M, with almost no increase in parameter overhead, which indicates the effectiveness of this module in optimizing feature extraction.
[0124] Secondly, the present invention optimizes the multi-scale feature fusion strategy, and improves the multi-scale detection ability by hierarchically fusing low-level and high-level features step by step. Initially, low-level features are fused, and higher-level features are gradually introduced, and finally fused with the top-level features of the backbone network. This design effectively alleviates the problem of insufficient information transmission between non-adjacent layers caused by indirect interaction in traditional methods. During the fusion process, low-level features provide rich detailed information, while high-level features provide powerful semantic information, ensuring that the detected objects have sufficient details and semantic understanding. Especially in the detection of strip surface defects, this strategy enables the model to handle defect features of different scales, improves the detection accuracy and robustness, and increases the mAP value of the improved model to 79.1%.
[0125] Finally, after adding the GLAM module to the model, the mAP further increases from 79.1% to 80.7%, with an increase of 1.6%. The introduction of the GLAM module increases the model's parameter count from 3.397M to 4.487M. Although the parameter overhead of the GLAM module is large, its performance gain is also significant. This improvement stems from the fact that the GLAM module combines the advantages of CNN and Transformer, fully leveraging their complementary advantages and making up for the deficiency of traditional convolutional networks in only being able to extract local features, thereby enhancing the model's performance in complex defect detection tasks. Figure 5 The heatmaps of the model's detection of strip surface defects before and after adding the GLAM module are shown. In the heatmap, the darker the color, the higher the degree of attention of the model. Compared with the baseline model, the method of the present invention pays more precise attention to the defect area. This benefits from the GLAM module's ability to simultaneously capture global information with long-range dependencies and fine-grained local features, enhancing the model's understanding of the macroscopic structure and microscopic features of defects, thus effectively highlighting the key areas, preventing the model from being distracted by irrelevant areas, and making the model's attention range more concentrated and more consistent with the actual defect area.
[0126] To more intuitively evaluate the performance of the method of the present invention, the present invention conducts a visualization analysis of the baseline model YOLOv8n and the method of the present invention. The two methods respectively detect the same target defect, and the results are as Figure 7 shown. The labeled image, the detection result of the baseline model YOLOv8n, and the detection result of the method of the present invention are respectively shown. The comparison results show that the method of the present invention is superior to the baseline model in the detection of multiple defect categories, including Cr, In, Ps, Rs, and Sc defects. The prediction boxes generated by the method of the present invention have a higher degree of matching with the actual defect area, especially in the case of complex backgrounds and irregular defect morphologies, demonstrating stronger robustness.
[0127] Specifically, for Ps and Rs defects, the detection boxes of the method of the present invention fit the labeled images better, showing higher detection accuracy. When detecting Ps defects, the detection boxes of the baseline model failed to fully cover the defect area, reflecting its deficiency in detecting large-scale defect targets. However, with the introduction of the hierarchical fusion network (HFN) in the method of the present invention, more efficient fusion of multi-scale features is achieved, enhancing the detection ability for large-scale defects. In addition, when detecting Cr defects, the baseline model failed to accurately identify the edge area of Cr defects. Its limitation lies in that the baseline model only relies on the local feature extraction mechanism of the convolutional neural network and is difficult to capture the global structural information of the defects. In contrast, the method of the present invention shows advantages in identifying the unstructured features of Cr defects, and the generated detection boxes match the actual defect areas more closely.
[0128] 4.4 Comparative Experiments
[0129] Table 2 Average Precision Results of Defect Detection of Each Model on the NEU-DET Dataset
[0130]
[0131]
[0132] To verify the effectiveness of the method of the present invention, under the same dataset division conditions, the method of the present invention is compared with five object detection models, namely CenterNet, YOLOv5, YOLOX, YOLOv7, and RT-DETR series. The results are shown in Table 2. Among them, the number of parameters (Params) represents the total number of parameters that the model needs to train, and the mean average precision (mAP) is the mean of the detection accuracies of all categories, which is used to comprehensively evaluate the overall performance of the network model.
[0133] The experimental results show that the method of the present invention achieves the highest mean average precision (mAP) among all the compared models, reaching 80.7%, and is superior to other models in terms of detection accuracy. From the specific data in Table 1, it can be seen that the method of the present invention performs the best in the detection tasks of Cr and Sc defects, with detection accuracies reaching 65.3% and 92.8% respectively, achieving the optimal results. At the same time, in the detection tasks of In, Ps, and Rs defects, the method achieves sub-optimal performance. In addition, for the detection of other defect categories, the method of the present invention also maintains a high precision, achieving the overall best performance of the mAP index.
[0134] From the detection results, it can be seen that CenterNet, YOLOv5, YOLOX, YOLOv7, and RT-DETR perform well in detecting defect types such as In, Pa, Ps, and Sc. However, these methods all have deficiencies in the detection of Cr defects and are difficult to effectively capture the characteristics of this type of defect. Specifically, the method proposed in the present invention achieves a detection accuracy of 65.3% for Cr defects, which is 9.3% higher than that of the sub-optimal model. The main reason for this improvement is that such defects often have a large distribution range, lack fixed shapes and distribution characteristics, which brings difficulties in accurate positioning to the detection network. To address this challenge, the present invention proposes a global-local perception module that can simultaneously capture the global information of long-distance dependencies and fine-grained local features. This module enhances the model's ability to understand the macroscopic structure and microscopic characteristics of defects, thus improving the detection accuracy of Cr defects.
[0135] In addition to the improvement in detection accuracy, the method of the present invention also performs well in terms of the number of parameters. Compared with YOLOv5, the mAP of the method of the present invention is increased by 6.5%, and the detection accuracy for multiple types of defects is higher than that of YOLOv5. At the same time, the number of parameters is only 0.64 times that of YOLOv5. Further comparing the method of the present invention with the RT-DETR model based on Transformer, the mAP of the method of the present invention is increased by 4.6%, while the number of parameters is only 0.22 times that of RT-DETR. This result shows that the method of the present invention achieves the best detection performance while reducing the number of parameters, fully reflecting the balance between computational efficiency and accuracy, and demonstrating high practical application value.
[0136] 4.5 Model Generalization
[0137] Table 3 Average Precision Results of Defect Detection of Each Model on the GC10-DET Dataset
[0138]
[0139] To verify the generalization performance of the model, the present invention conducted experiments on the real industrial scenario dataset GC10-DET, and the results are shown in Table 3. The experiments show that the proposed model achieved 66.9% in the mAP metric, showing a significant improvement compared to 64.9% of the baseline model YOLOv8. As can be seen from Table 3, there are differences in the performance of different models in the detection of various types of defect targets, but the overall detection effect of the model of the present invention is better than that of other models. In terms of specific categories, the method of the present invention achieved the optimal average precision in two types of defects, Ws and In, which are 91.2% and 34.6% respectively, and achieved the sub-optimal average precision in two types of defects, Rp and Wf, which are 31.1% and 88.2% respectively. At the same time, it also maintained a high precision in the detection of other defect categories, reflecting its adaptability and generalization ability in multi-category defect detection. Compared with the RT-DETR model, the model parameter quantity of the method of the present invention is 4.5M, which is lower than 19.9M of RT-DETR, but it is 0.4% higher in the mAP metric. In addition, compared with models such as YOLOX and YOLOv7, the method of the present invention maintains a lightweight design while achieving high detection accuracy. In summary, the experimental results prove the good generalization ability of the method of the present invention on different datasets.
[0140] The present invention deeply analyzes problems such as background noise, insufficient feature fusion, and defect scale variation in strip surface defect detection. In response to these challenges, a strip surface defect detection method integrating global-local perception and hierarchical feature fusion is proposed to improve the model's capabilities in feature extraction, multi-scale feature expression, and unstructured defect detection. First, to address the problem that existing detection algorithms are vulnerable to background noise and irrelevant information, the present invention designs a feature enhancement module, which effectively filters background noise during the feature extraction process, focuses on defect features, and thus improves the detection accuracy. Second, a hierarchical fusion network is introduced into the backbone network to fully fuse semantic and detail information at different levels, effectively dealing with the scale complexity of strip surface defects. Finally, the global-local perception module is added to the model. By integrating global context information and local features, the model macroscopically understands the defect structure while paying attention to local details, improving the detection ability for unstructured defects and effectively enhancing the detection accuracy of strip surface defects.
[0141] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0142] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A strip surface defect detection method integrating global local perception and hierarchical feature fusion, characterized in that: The following steps are involved: S1, obtaining a surface image of a steel strip to be inspected; S2. Inputting the strip surface image to be detected into the trained strip surface defect detection model; the strip surface defect detection model is YOLOv8n, including: multiple feature enhancement modules, a hierarchical fusion network and multiple global-local perception modules; The feature enhancement module replaces the original C2f module in its backbone network; S3. Output the corresponding detection result.
2. The method for detecting surface defects of a steel strip according to claim 1, characterized in that: The data processing process of the strip surface defect detection model includes: S21, using the feature enhancement module in the backbone network to extract features from the input image; the feature enhancement module includes two convolutional layers, a split operation, multiple Bottleneck modules and a splicing operation, and dynamically adjusts the feature weights to achieve the enhancement of target features and the suppression of background noise; S22, using the hierarchical fusion network in the neck network to perform multi-scale feature fusion on the features output by the feature enhancement module in the backbone network, gradually fusing adjacent layer features by layering, and supporting direct interaction of non-adjacent features; and then inputting the features into the feature enhancement module in the neck network through upsampling and splicing operations; S23, using the global-local perception module to perform global and local feature fusion on the features output by the feature enhancement module in the neck network; inputting the features output by the global-local perception module into the detection head for classification and regression to obtain the category and bounding box position of the defect.
3. The method for detecting surface defects of a strip steel according to claim 2, characterized in that: The step S21 comprises: The feature enhancement module in the backbone network performs preliminary extraction and channel transformation on the input feature map through convolution operations to generate an intermediate feature representation X; Through the split operation, X is divided into two sub-feature maps along the channel dimension, one of which is a sub-feature map X 0 After being processed by the Bottleneck module, it serves as the initial input of the subsequent Bottleneck. The features generated by each operation are described by the recursive formula of formula (1): X i =CBS 2 (X i-1 )+SE(X i-1 )+X i-1 , i=1,2,...,n (1) Among them, CBS 2 (·) represents the feature mapping function of two convolution operations, SE(·) is the feature mapping function of the channel attention module; The feature enhancement module in the backbone network converts all deep features (X 1 ,X 2 ,…,X n ) is concatenated with the original feature map X to generate the final feature representation Y, which is calculated as follows: Y=CBS(Concat(X,X 1 ,X 2 ,...,X n )) (2) In the formula, Concat represents concatenation operation; CBS represents convolution operation.
4. The method for detecting surface defects of a steel strip according to claim 3, characterized in that: The processing process in the feature mapping function of the channel attention module includes: Through the global average pooling operation, the spatial information of each channel is compressed into global semantic features; Two layers of fully connected network operations are performed in sequence to generate channel weights, which are normalized using the Sigmoid activation function.
5. The method for detecting surface defects of a steel strip according to claim 2, characterized in that: The step S22 comprises: 1) Feature processing: The backbone network extracts multi-scale features of the input image and represents them as C3, C4, and C5, which correspond to low-dimensional, medium-dimensional, and high-dimensional feature maps respectively; 2) Layered Fusion: Formula (3) and formula (4) are used to fuse the features of C3 and C4: C3=Concat(conv1(C3),↑2(C4)) (3) C4=Concat(↓2(C3),conv1(C4)) (4) ↑2 means using 1×1 convolution and bilinear interpolation method to upsample the features, 2 is the upsampling ratio; conv1 means using 1×1 convolution to adjust the number of channels; ↓2 means using 3×3 convolution with step sizes of 2 and 4 to downsample, 2 is the downsampling ratio; Concat means concatenation operation to obtain the fused features; Use formula (5)-formula (7) to fuse C3, C4 and C5 features: P3=Concat(conv1(C3),↑2(C4),↑4(C5)) (5) P4=Concat(↓2(C3),conv1(C4),↑2(C5)) (6) P5=Concat(↓4(C3),↓2(C4),conv1(C5)) (7) ↑4 means using 1×1 convolution and bilinear interpolation method to upsample the features, 4 is the upsampling ratio; conv1 means using 1×1 convolution to adjust the number of channels; ↓4 means using 3×3 convolution with step size of 2 and 4 respectively to achieve downsampling, 4 is the downsampling ratio; 3) Fusion feature output: After the above layer-by-layer fusion process, a new set of multi-scale fusion features P3, P4, and P5 are obtained; each feature contains rich information of different scales, and the semantic gap between levels is effectively narrowed.
6. The method for detecting surface defects of a steel strip according to claim 2, characterized in that: The step S23 comprises: 1) Local feature extraction: For the feature X output by the feature enhancement module in the neck network, each channel is globally averaged pooled to compress the spatial dimension feature into a scalar to form a channel description vector; Perform one-dimensional convolution operation on the channel description vector to build a local dependency model; and dynamically adjust the convolution kernel size according to the number of channels; The convolution result is transformed by the Sigmoid activation function to generate the weight w of each channel. L ; All weights are applied to the input feature X in a channel-by-channel weighted manner to obtain the local feature representation F L ; 2) Global feature extraction: Feature X is passed through three separate convolutional layers to generate query Q, key K, and value V feature maps respectively; Reshape the Q, K, and V feature maps into matrices of shape HW×C; H and W are the spatial dimensions of the feature maps, and C is the number of channels; Divide the Q, K and V feature maps into h subspaces along the channel dimension, each subspace contains C i channels; in each subspace, Q and K perform a dot product operation to generate a self-attention relationship matrix; multiply the self-attention relationship matrix with the corresponding V to obtain the self-attention feature F of the subspace i G ; The self-attention features of all h subspaces are concatenated in the channel dimension and integrated into the final global feature representation F G ; 3) Feature Fusion: The local feature F L and the global feature F G Splicing is done on the channel dimension, and 1x1 convolution is used to adjust the number of channels and perform weight assignment to generate the final output features.
Citation Information
Cited By
Steel surface defect detection method based on multi-scale edge enhancement
CN120852428A
Steel surface defect detection method based on multi-scale edge enhancement
CN120852428B
Induction cooker glass thickness measuring method
CN121112920A
Photovoltaic module defect detection method based on multi-scale feature and global information interaction
CN121147138A
Road surface defect detection method based on improved YOLOv8 model
CN121527082A