Hot rolled steel surface defect detection method and device based on double-strip window pyramid network model
By adopting a double-band window pyramid network model in hot-rolled steel surface defect detection, combined with CSPDarkNet53 and DSwinTransformer modules, the problem of insufficient detection accuracy and real-time in complex industrial noise environments in the prior art is solved, and high-precision and fast defect detection are achieved.
Patent Information
- Application Number
- CN202510305552.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology has problems such as high false alarm rate, poor generalization ability, insufficient fusion of multi-scale defects, small target miss detection and limited real-time limitations in hot-rolled steel surface defects in complex industrial noise environments.
The detection method based on the double-striped windowed pyramid network (DSWPN) model is adopted, combined with the CSPDarkNet53 and DSwinTransformer module, replace the Neck part with BiFPN, and add the multi-head context integration (MHCI) attention mechanism module before the detection head. Through multi-level feature extraction and bidirectional cross-scale feature fusion, detection accuracy and real-time performance are improved.
Average detection accuracy (mAP) of 91.9% and 84.0% was achieved on the NEU-DET and GC10-DET datasets, and achieved a real-time inference speed of 45FPS, which was significantly better than the existing technology.
Smart Images

Figure CN120219339A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and specifically, relates to a method and device for detecting surface defects of hot-rolled steel based on a double-strip window pyramid network model. Background Art
[0002] As a key material in industrial production, steel is widely used in fields such as shipbuilding, automobiles, and bridges. Its surface defects (such as cracks, scratches, and inclusions) directly affect the safety and reliability of products. Traditional manual inspection is inefficient and relies on subjective experience, and there is an urgent need for automated inspection technology. However, the production environment of hot-rolled steel is complex, and process defects are prone to cause missed inspections, leading to potential safety hazards. Therefore, a defect detection method with high precision and high efficiency has become the core requirement for industrial intelligent upgrading.
[0003] Traditional methods rely on manual feature design. For example, statistical methods (such as Tsai et al. based on the weighted variance matrix), spectral methods (such as Jasper et al. extracting frequency domain features through wavelet transform), and model methods (such as Cohen et al. using Markov random field modeling) are effective in stable scenarios, but have a high false alarm rate and poor generalization ability when facing complex industrial noise. Deep learning methods (such as Faster R-CNN, YOLO series) have improved the detection accuracy by data-driven feature extraction, but still have problems such as insufficient multi-scale defect fusion, missed detection of small targets, limited real-time performance, and rely on large-scale labeled data, making it difficult to adapt to the challenge of unbalanced industrial samples.
[0004] The current core difficulties in steel defect detection include: (1) The types of defects are diverse and the features are easily interfered by light and rust; (2) Small targets (such as micron-scale scratches) are prone to missed detection due to feature loss, and large-scale defects have blurred positioning due to insufficient resolution. (3) The performance of existing models drops significantly in new environments or when the data distribution shifts, and they cannot adapt to dynamic industrial demands. Therefore, there is an urgent need for a detection method that integrates multi-scale feature fusion, global context modeling, and efficient data augmentation to improve the robustness, accuracy, and real-time performance in complex industrial scenarios and meet dynamic production requirements. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a method for detecting surface defects of hot-rolled steel based on a double-strip window pyramid network model, which improves the inference speed while enhancing the detection accuracy of surface defects of hot-rolled strip steel.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a method for detecting surface defects of hot-rolled steel based on a double-strip window pyramid network (DSWPN) model, specifically including the following steps:
[0008] Step 1: Collect the surface defect images of hot-rolled steel, construct a defect image set, and perform data augmentation on the images in the defect image set to obtain an enhanced defect data set;
[0009] Step 2: Use CSPDarkNet53 as the backbone network, add a DSwinTransformer module at its backend, replace the original Neck part with BiFPN, and add a multi-head context integration (MHCI) attention mechanism module before the detection head to construct a double-strip window pyramid network;
[0010] Step 3: Input the enhanced defect data set into the trained DSWPN network, and train it based on the CIOU loss function until the loss function converges to complete the training of the defect detection model;
[0011] Step 4: Collect the surface images of hot-rolled steel in real time, input them into the trained DSWPN network, and obtain the surface defect detection results of hot-rolled steel.
[0012] Furthermore, input the original industrial image, perform normalization processing on the image, scale it to 200×200 pixels, and perform standardization; then apply Mosaic augmentation, Mixup augmentation and graffiti noise data augmentation strategies.
[0013] Furthermore, the construction process of the double-strip window pyramid network is specifically as follows: Based on the improvement of the YOLOv8 structure, use CSPDarkNet53 as the backbone feature extraction network; add a DSwinTransformer module after CSPDarkNet53 to improve the feature modeling ability; replace the original Neck structure and use the BiFPN module for multi-scale feature fusion; add a multi-head context integration attention mechanism module before the detection head to enhance the feature expression ability.
[0014] Furthermore, the network structure of the DSwinTransformer module includes layer normalization (LN), double-strip window self-attention (DSwinAttention), LN for re-normalization, and a multi-layer perceptron (MLP) connected in sequence. Among them, the input features are first normalized by LN to stabilize the data distribution. The normalized features enter DSwinAttention, which adopts a strip window division strategy to calculate self-attention in the horizontal and vertical directions respectively, and fuses the attention results in different directions by splicing to enhance the feature expression ability. The calculated features are connected with the input features before normalization by residual connection, and then normalized by LN again and input into MLP. MLP performs feature transformation through full connection and activation function, and the output result is connected with the normalized features before entering MLP by residual connection, and finally forms the output of the DSwinTransformer Block.
[0015] Furthermore, the operating principle of BiFPN is as follows: adopt bidirectional feature flow (top-down and bottom-up); introduce a learnable weight assignment mechanism to optimize the fusion of features at different scales; adopt a normalization fusion strategy to improve the feature integration effect.
[0016] Furthermore, the process of the MHCI module is specifically as follows:
[0017] Y = ((α · ADynamic + β · AStatic) · V) · W2 + X
[0018] Wherein, ADynamic is an affinity matrix dynamically generated from the input X, AStatic is a static affinity matrix, α and β are learnable parameters, V is a linear transformation of X, W2 is a learnable parameter, and X is the input feature map.
[0019]
[0020] Where R IoU (M, N) represents the IoU value between the target prediction box and the image label box; ρ 2 (M ctr , N ctr ) represents the Euclidean distance between the center points of the target prediction box and the image label box; m represents the diagonal distance of the smallest circumscribed rectangle containing the target prediction box and the image label box; α is a weight coefficient; υ represents a parameter for measuring the aspect ratio consistency between the target prediction box and the image label box.
[0021] The second aspect of the present invention relates to a hot-rolled steel surface defect detection device based on a double-strip window pyramid network (DSWPN), including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the hot-rolled steel surface defect detection method based on the double-strip window pyramid network of the present invention.
[0022] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the hot-rolled steel surface defect detection method based on the double-strip window pyramid network of the present invention.
[0023] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes an innovative method in the field of industrial defect detection. By combining CSPNet and Swin Transformer for multi-level feature extraction and using the Bidirectional Feature Pyramid Network (BiFPN) to achieve bidirectional cross-scale feature fusion. The above strategy improves the disadvantages of the YOLO model, such as the easy omission of detection of small targets (such as micron-scale scratches) due to feature loss and the blurred positioning of large-scale defects due to insufficient resolution. Experimental results show that the average detection accuracy (mAP) of this method on the NEU-DET and GC10-DET datasets reaches 91.9% and 84.0% respectively, and the real-time inference speed reaches 45 FPS, which is better than the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the flowchart of the method of the present invention.
[0025] Figure 2 is the framework of the double-striped window transformation pyramid network model of the present invention.
[0026] Figure 3 is the structural diagram of the double-striped window transformation module of the present invention.
[0027] Figure 4 is the structural diagram of the double-striped window transformation attention of the present invention.
[0028] Figure 5 is the schematic diagram of the BiFPN of the present invention.
[0029] Figure 6 is the heat map of the comparison of different defects between the DSWPN and YOLOv8 of the present invention. Among them, column (a) is cracks, column (b) is inclusions, column (c) is patches, column (d) is unevenness, column (e) is oxides, and column (f) is scratches.
[0030] Figure 7 (a)- Figure 7 (f) is the visualization result diagram of different defects on the NEU-DET dataset of the present invention. Among them, Figure 7 (a) is cracks, Figure 7 (b) is inclusions, Figure 7 (c) is patches, Figure 7 (d) is unevenness, Figure 7 (e) is oxides, Figure 7 (f) is scratches.
[0031] Figure 8 is the schematic diagram of the device of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] The technical solution of the present invention will be further explained below with reference to the accompanying drawings.
[0033] Example 1
[0034] As Figure 1 FIG. is the overall flowchart of the hot-rolled steel surface defect detection method based on the double-strip window pyramid network model of the present invention. The hot-rolled strip surface defect detection method specifically includes the following steps:
[0035] Step 1: Collect hot-rolled steel surface defect images, construct a defect image set, and perform data enhancement processing on the images in the defect image set to obtain an enhanced defect data set; the data enhancement strategy in the present invention includes normalizing the images, scaling them to 200×200 pixels, and standardizing them; subsequently, apply Mosaic enhancement, Mixup enhancement, and graffiti noise data enhancement strategies.
[0036] Step 2: As Figure 2As shown in the overall network framework, in order to reduce costs and model complexity and improve the detection speed, CSPDarkNet53 is used as the backbone network, integrating the DSwinTransformer attention mechanism. The original Neck part is replaced with BiFPN, and a multi-head context integration (MHCI) attention mechanism module is added before the detection head to construct a dual-strip window pyramid network (DSWPN) model. Specifically, the CSPDarknet53 backbone network is responsible for extracting features from the surface defect images of hot-rolled strip steel, thereby capturing the semantic information and structural information in the surface defect images of hot-rolled strip steel. It consists of convolutional layers, pooling layers, and activation layer components, and can effectively convert the input surface defect images of hot-rolled strip steel into high-level feature representations. The DSwin Transformer is inserted behind the CSPDarknet53 feature extraction network to integrate the local feature extraction ability of CSPDarknet53 and the global attention mechanism of DSwin Transformer. The local receptive field is extended through horizontal and vertical strip self-attention grouping calculations (DSwin Attention) to enhance the representation ability of multi-scale defect features. The neck network is located between the backbone network and the detection head and is responsible for further processing and fusing the features extracted by the backbone network to improve the feature expression ability and semantic information. It consists of convolutional layers, upsampling operations, and cross-layer connections, and is used to achieve feature refinement and fusion. As the network deepens and undergoes multiple convolutional operations, the feature information of small-sized defect parts in the surface defect images of hot-rolled strip steel will be lost in the high-level feature maps. The bidirectional feature pyramid network (BiFPN) is used to replace the original Yolov8 neck network, adopting a bidirectional cross-scale connection structure and a fast normalization weighted fusion strategy to dynamically allocate the weights of feature maps with different resolutions, which can optimize the detection accuracy of multi-scale defects. Finally, a context attention module (CA) is introduced before the detection head, combining the dynamic affinity matrix and the static prior weight to suppress industrial background noise and improve the sensitivity to tiny defects (such as microcracks).
[0037] In the present invention, the DSwinTransformer module inserted behind the CSPDarknet53 feature extraction network, as Figure 3As shown in the figure, it includes Layer Normalization (LN), Double-strip Window Self-attention (DSwinAttention), LN for re-normalization, and a Multi-Layer Perceptron (MLP) connected in sequence. Among them, the input features are first normalized by LN to stabilize the data distribution. The normalized features enter DSwinAttention, where self-attention is calculated in the horizontal and vertical directions respectively, and the attention results in different directions are fused by concatenation to enhance the feature expression ability. The calculated features are subjected to residual connection with the input features before normalization, and then are normalized by LN again and input into the MLP. The MLP performs feature transformation through fully connected layers and activation functions, and the output result is subjected to residual connection with the normalized features before entering the MLP, finally forming the output of the DSwinTransformerBlock.
[0038] As Figure 3 , the specific process of DSwin Attention used in the DSwinTransformer module in the present invention is as follows: self-attention operations are performed on the attention heads on the horizontal and vertical strips, and the results are Concat together to obtain the final output, solving the problem of the excessively high complexity of the original window attention mechanism. Assume the input feature X ∈ R H×W×C , and it is first projected linearly onto K heads. Each head performs local self-attention within the horizontal strip or vertical strip, as Figure 4 shown. For horizontal strip self-attention, X is evenly divided into M non-overlapping horizontal strips [X1, …, X M , the width of each strip is s w , and each strip contains s w ×W tokens. Assume the dimensions of the query, key, and value of the k-th head are all d k , then the output of the horizontal strip self-attention of the k-th head can be expressed as:
[0039]
[0040] where and represent the projection matrices of the query, key, and value of the k-th head respectively, and The vertical strip self-attention can be derived similarly, and the output of its k-th head is denoted as V-Attention k (X). The K heads are evenly divided into two groups (K / 2 heads in each group). The first group of heads performs horizontal strip self-attention, and the second group of heads performs vertical strip self-attention. Finally, the outputs of the two parallel groups will be concatenated back together to obtain the final self-attention output:
[0041] As Figure 4, the specific process of implementing feature refinement and fusion using the Bidirectional Feature Pyramid Network (BiFPN) in the present invention is as follows: Based on the traditional pyramid network PANet, the bidirectional feature streams (top-down and bottom-up) are retained, enabling features of different granularities to interact in the neck network; in PANet, there are nodes with only one in-degree, and the information of these nodes is highly similar to that of the input nodes. To reduce data redundancy and streamline the network, the nodes with only one in-degree are deleted; to compensate for the resulting data fusion degree, skip connections from the input nodes to the output nodes are added, and each bidirectional path is regarded as a feature network layer to optimize cross-scale connections. BiFPN optimizes the feature fusion process by adding weights to each input feature, enabling the network to pay more attention to features with greater information content, and adding additional edges between the original input and output nodes at the same level. This weighting strategy can dynamically adjust the contribution of each input feature map according to its importance, thereby more effectively combining feature information at different scales, improving the quality of feature expression and the performance of the detection task. The calculation formula is as follows:
[0042]
[0043] Among them, I i represents the i-th input feature; w i represents the weight of the i-th input feature, which is a learnable parameter; ∈ represents a very small positive number used for numerical stability to prevent division-by-zero errors; ∑ j w j : The sum of the weights of all input features, ensuring the normalization of the weights.
[0044] The specific process of adding the MHCI module in front of the detection head in the present invention is as follows:
[0045] Y = ((α · ADynamic + β · AStatic) · V) · W2 + X
[0046] Among them, ADynamic is an affinity matrix dynamically generated from the input X, AStatic is a static affinity matrix, α and β are learnable parameters, V is a linear transformation of X, W2 is a learnable parameter, and X is the input feature map.
[0047] Step 3: Input the enhanced defect dataset into the trained DSWPN network and train it based on the CIOU loss function until the loss function converges to complete the training of the defect detection model;
[0048] The specific calculation of the CIOU loss function in the present invention is as follows:
[0049]
[0050] Among them, R IoU(M, N) represents the IoU value between the target prediction box and the image label box; ρ 2 (M ctr , N ctr ) represents the Euclidean distance between the center points of the target prediction box and the image label box; m represents the diagonal distance of the smallest circumscribed rectangle containing the target prediction box and the image label box; α is the weight coefficient; υ represents the parameter for measuring the aspect ratio consistency between the target prediction box and the image label box.
[0051] Step 4: Real-time collect the surface images of hot-rolled steel, input them into the trained DSWPN network, obtain the detection results of the surface defects of hot-rolled steel, and mark the defect-free hot-rolled strip as a qualified product.
[0052] The present invention evaluates the improved model on the NEU-DET and GC10-DET datasets, and verifies its superior defect detection performance. The comparison experiment results on the NEU-DET dataset are shown in Table 1. The improved model performs excellently in terms of the average detection accuracy (mAP), reaching 91.90%, significantly higher than other mainstream models, verifying the improvement effect of the model in terms of accuracy. The comparison experiment results on the GC10-DET dataset are shown in Table 2. The improved model also performs outstandingly, achieving an mAP of 84.0%, further proving its powerful detection ability in different application scenarios.
[0053] Table 1 Comparison experiments on NEU-DET
[0054]
[0055]
[0056] Table 2 Comparison experiments on GC10-DET
[0057]
[0058] To further verify the actual effect of the improved model, the present invention uses the heatmap visualization technology (CAM technology) to display the attention areas during the model detection process. Figure 6 Shows the comparison results of the heatmaps of DSWPN and YOLOv8. The model's attention in the defect area is significantly more concentrated. Compared with the Baseline, the heatmap shows the effective focus of the method in this paper on the target area, indicating that the improved model has stronger feature extraction and area attention capabilities in complex backgrounds. This result further proves the feature extraction ability and area focusing effect of the present invention in complex backgrounds. In addition, a visual analysis of the prediction results of the NEU-DET dataset was also carried out. Figure 7 (a)- Figure 7(f) shows the prediction results of the improved model on this dataset. These visualization results demonstrate the accurate annotation and recognition of different types of defects in the images by the model, further verifying the superior performance of the present invention in industrial defect detection tasks. By intuitively presenting the location and category information of the defects, the improved model can not only improve the detection accuracy but also effectively enhance the accuracy and reliability of the detection.
[0059] Embodiment 2
[0060] Referring to Figure 8 , this embodiment relates to a hot-rolled steel surface defect detection device based on a double-strip window pyramid network (DSWPN), including a memory and one or more processors. An executable code is stored in the memory, and when the one or more processors execute the executable code, it is used to implement the hot-rolled steel surface defect detection method based on the double-strip window pyramid network in Embodiment 1.
[0061] Embodiment 3
[0062] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the hot-rolled steel surface defect detection method based on the double-strip window pyramid network in Embodiment 1.
[0063] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept.
Claims
1. A hot-rolled steel surface defect detection method based on a double-strip window pyramid network model, characterized in that: The specific steps include: Step 1: collect the surface defect images of hot-rolled steel, construct a defect image set, and perform data enhancement processing on the images in the defect image set to obtain an enhanced defect data set; Step 2: Use CSPDarkNet53 as the backbone network, add the DSwinTransformer module to its back end, replace the original Neck part with BiFPN, and add the multi-head context integration MHCI attention mechanism module before the detection head to construct a dual-strip window pyramid network DSWPN; Step 3: Input the enhanced defect data set into the trained DSWPN network and perform training based on the CIOU loss function until the loss function converges, thus completing the defect detection model training; Step 4: Collect the hot-rolled steel surface image in real time, input it into the trained DSWPN network, and obtain the hot-rolled steel surface defect detection result.
2. A hot rolled steel surface defect detection method based on a double strip window pyramid network as claimed in claim 1, characterized in that: Step 1 specifically includes: inputting the original industrial image I a , the image is normalized, scaled to 200×200 pixels, and standardized; then Mosaic enhancement, Mixup enhancement, and graffiti noise data enhancement strategies are applied.
3. The method for detecting surface defects of hot rolled steel based on a double-strip window pyramid network model according to claim 1, characterized in that: The construction process of the dual-strip window pyramid network in step 2 is as follows: based on the improvement of the YOLOv8 structure, CSPDarkNet53 is used as the backbone feature extraction network; the DSwinTransformer module is added after CSPDarkNet53 to improve the feature modeling capability; The original Neck structure is replaced and the BiFPN module is used for multi-scale feature fusion. A multi-head context integration MHCI attention mechanism module is added before the detection head to enhance the feature expression capability.
4. According to claim 3, a method for detecting surface defects of hot rolled steel based on a double-strip window pyramid network is characterized in that: The network structure of the DSwinTransformer module includes a layer normalization LN, a dual-strip window self-attention DSwinAttention, a normalized LN and a multi-layer perceptron MLP connected in sequence, wherein the input features are first normalized by LN to stabilize the data distribution; the normalized features enter DSwinAttention, which adopts a strip window division strategy to calculate self-attention in the horizontal and vertical directions respectively, and merge the attention results in different directions by splicing to enhance the feature expression ability. The calculated features are residually connected with the input features before normalization, and then normalized by LN again and input into MLP; MLP performs feature transformation through full connection and activation function, and the output results are residually connected with the normalized features before entering MLP, finally forming the output of DSwinTransformer Block.
5. The method for detecting surface defects of hot rolled steel based on a double-strip window pyramid network according to claim 3, characterized in that: The operating principle of BiFPN is as follows: adopting bidirectional feature flow; introducing a learnable weight allocation mechanism to optimize the fusion of features of different scales; adopting a normalized fusion strategy to improve the feature integration effect.
6. The method for detecting surface defects of hot rolled steel based on a double-strip window pyramid network according to claim 3, characterized in that: The MHCI module process is as follows: Y=((α·ADynamic+β·AStatic)·V)·W2+X Among them, ADynamic is the affinity matrix dynamically generated from the input X, AStatic is the static affinity matrix, α and β are learnable parameters, V is the linear transformation of X, W2 is a learnable parameter, and X is the input feature map.
7. The method for detecting surface defects of hot rolled steel based on a double-strip window pyramid network according to claim 1, characterized in that: The CIOU loss function in step 3 is calculated as follows: Where R IoU (M, N) represents the IoU value between the target prediction box and the image label box; ρ 2 (M ctr ,N ctr ) represents the Euclidean distance between the center points of the target prediction box and the image label box; m represents the diagonal distance of the minimum enclosing rectangle containing the target prediction box and the image label box; α is the weight coefficient; υ represents the parameter used to measure the consistency of the length and width ratio of the target prediction box and the image label box.
8. A hot rolled steel surface defect detection device based on a dual strip window pyramid network (DSWPN), characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the hot-rolled steel surface defect detection method based on a double-strip window pyramid network as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the hot-rolled steel surface defect detection method based on a double-strip window pyramid network as described in any one of claims 1-2 is implemented.
Citation Information
Cited By
Steel surface defect detection method based on edge enhancement and double-flow fusion
CN121304654A