Method and system for detecting surface defects of steel material
Patent Information
- Application Number
- CN202410514388.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-26
AI Technical Summary
[0004]发明人发现,虽然上述模型已经在缺陷检测领域发挥了很好的性能,但是应用其对钢材表面的缺陷进行检测时,仍然存在很高的漏检和误检率
[0020]本发明提供了一种钢材表面缺陷检测方法及系统,对检测器的模型体积、算法速度和检测精度进行研究,在模型中加入了稀疏自注意力SA模型结构、卷积和Transformer相互协同的CTR模型结构以及GDC瓶颈卷积结构,提升了缺陷检测速度和检测精度,解决了现有模型在钢材数据集上存在的检测精度低、漏检率高、模型参数量大,不利用工业部署等问题。
Smart Images

Figure CN118314114B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of defect detection technology, and in particular relates to a method and system for detecting defects on the surface of steel. Background Technology
[0002] Steel, as a crucial building and manufacturing material, is susceptible to surface defects during production due to various factors, such as scratches, oxidation, and stains. These defects not only affect the product's appearance but, more importantly, its performance and durability. Defect detection technology is a key and challenging task in computer-aided industrial production, enhancing the safety and reliability of steel products and ensuring the safety of life and property. Previously, steel surface defects were detected manually using visual inspection. However, manual visual inspection is subject to subjective differences and prone to operator fatigue, significantly reducing efficiency and posing safety hazards. Furthermore, the large-scale production of steel requires substantial time and human resources for inspection. Existing detection models also have limitations in terms of efficiency, applicability, and automation.
[0003] In recent years, the requirements for storage and computing power in defect detection equipment have been increasing, and deep learning technology has been widely applied in steel surface defect detection. Numerous researchers have proposed different deep learning models. For example, the ResNet model proposed by Kaiming He et al. has made a significant contribution to the development of convolutional neural networks. The Google team proposed the MobileNet model, which uses depthwise separable convolutions, reducing the number of model parameters and improving model performance. In 2020, the team proposed the ViT model, applying the Transformer to the vision domain and achieving remarkable results. LeViT does not use patch stems but reintroduces convolutional stems at the beginning of the network to learn low-resolution features, improving model performance. EdgeViT introduces Local-Global-Local blocks to integrate self-attention modules and convolutional neural networks, allowing the model to capture spatial tokens with different ranges and exchange information between them. EfficientFormer combines convolutional layer modules and self-attention layer structures in a hybrid approach to achieve a balance between accuracy and efficiency.
[0004] The inventors discovered that although the aforementioned model has demonstrated excellent performance in the field of defect detection, it still suffers from high false negative and false negative rates when applied to detect defects on steel surfaces. In particular, it performs poorly for minute defects such as dents and stains, and it fails to capture global features effectively from the steel dataset. Furthermore, the model suffers from large training size, slow recognition response speed, and is not well-suited for industrial deployment. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, this invention provides a method and system for detecting defects on the surface of steel. The model volume, algorithm speed and detection accuracy of the detector are studied. A sparse self-attention (SA) model structure, a CTR model structure in which convolution and Transformer work together, and a GDC bottleneck convolution structure are added to the model to improve the defect detection speed and detection accuracy.
[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of this invention provides a method for detecting defects on the surface of steel.
[0008] The method for detecting defects on the surface of steel includes the following steps:
[0009] Acquire images of the steel surface;
[0010] The steel surface image is input into multiple cascaded CTR modules of the CSTRNet model. In each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules, respectively, and feature fusion is performed.
[0011] The features extracted by each intermediate CTR module are sequentially input into the cascaded two-layer GDC module. Convolutional neural networks are used to extract features, and the shallow and deep features extracted by the CTR module are fused bidirectionally using the two-layer GDC module. The predicted box position, defect confidence, and defect classification of steel surface defects are obtained through the two-layer GDC module.
[0012] A second aspect of the present invention provides a steel surface defect detection system.
[0013] A steel surface defect detection system, including:
[0014] The image acquisition module is configured to acquire images of the steel surface.
[0015] The CTR module feature fusion module is configured to: input the steel surface image into multiple cascaded CTR modules of the CSTRNet model; in each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules respectively, and feature fusion is performed.
[0016] The GDC module feature fusion module is configured to: input the features extracted by each intermediate CTR module into the serially connected two-layer GDC module in sequence, extract features using a convolutional neural network, and use the two-layer GDC module to perform bidirectional fusion of the shallow and deep features extracted by the CTR module. The two-layer GDC module is used to obtain the predicted box position, defect confidence and defect classification of steel surface defects.
[0017] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps of the steel surface defect detection method as described in the first aspect of the present invention.
[0018] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the steel surface defect detection method as described in the first aspect of the present invention.
[0019] The above one or more technical solutions have the following beneficial effects:
[0020] This invention provides a method and system for detecting defects on steel surfaces. It studies the model size, algorithm speed, and detection accuracy of the detector. The model incorporates a sparse self-attention (SA) model structure, a CTR model structure that combines convolution and Transformer, and a GDC bottleneck convolution structure, which improves the defect detection speed and accuracy. This solves the problems of low detection accuracy, high false negative rate, large number of model parameters, and unsuitability for industrial deployment in existing models on steel datasets.
[0021] The SA module provided by this invention reduces the spatial size of K and V by applying 1×1 convolution and 3×3 depthwise separable convolution to shrink the feature map before the self-attention operation, thereby emphasizing the local spatial context and focusing self-attention on the shrunken feature map. This helps the model capture low-frequency global information, further reduces the number of parameters, and improves the detection speed.
[0022] The CTR model provided by this invention adopts a dual-branch design structure, which enables the CTR model to efficiently capture high-frequency and low-frequency information, achieving a balance between optimizing accuracy and efficiency, and comprehensively improving the detection performance of the model.
[0023] The bottleneck convolutional structure GDC module provided by this invention adopts a bottleneck structure design, which improves detection accuracy while reducing the number of parameters. In the CSTRNet model, the GDC module fuses different features and merges the shallow and deep features extracted by the CTR module, which can contain more information that is helpful for detection.
[0024] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0026] Figure 1 This is a flowchart of the method in the first embodiment.
[0027] Figure 2 This is a data flow diagram for the SA module.
[0028] Figure 3 This is a data flow diagram for the CTR model.
[0029] Figure 4 This is a data flow diagram for the GDC module.
[0030] Figure 5 The graph shows the evaluation results of the optimal model.
[0031] Figure 6 The figure shows the validation results of using the optimal model on the steel test set. Detailed Implementation
[0032] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0033] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0034] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0035] The overall concept proposed in this invention is as follows:
[0036] This invention proposes a method for detecting surface defects in steel using a collaborative approach of convolutional neural networks and a visual Transformer (ViT). The first step involves selecting, processing, and analyzing a steel dataset. The second step is establishing a deep learning detection model. Within this model, firstly, based on the visual Transformer, this invention proposes a sparse self-attention method, which improves detection accuracy while reducing the number of parameters. Secondly, this invention proposes a model structure using convolutional modules and ViT in parallel, working synergistically to significantly improve the model's detection performance on the steel dataset. Finally, this invention proposes a convolutional combination module structure to further enhance the model's detection accuracy on the steel dataset. The third step is training and evaluation.
[0037] This invention solves the problems of low detection accuracy, high false negative rate, large number of model parameters, and unsuitability for industrial deployment in existing models on steel datasets.
[0038] Example 1
[0039] This embodiment discloses a method for detecting defects on the surface of steel.
[0040] like Figure 1 As shown, the method for detecting defects on the surface of steel includes the following steps:
[0041] Acquire images of the steel surface;
[0042] The steel surface image is input into multiple cascaded CTR modules of the CSTRNet model. In each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules, respectively, and feature fusion is performed.
[0043] The features extracted by each intermediate CTR module are sequentially input into the cascaded two-layer GDC module. Convolutional neural networks are used to extract features, and the shallow and deep features extracted by the CTR module are fused bidirectionally using the two-layer GDC module. The predicted box position, defect confidence, and defect classification of steel surface defects are obtained through the two-layer GDC module.
[0044] The CSTRNet model proposed in this embodiment has the following specific structure: Figure 1 As shown, the CSTRNet model incorporates a designed sparse self-attention (SA) model structure, a CTR model structure that combines convolution and Transformer, and a GDC bottleneck convolution structure. This enables the CSTRNet model to achieve excellent performance in extracting local minor defects and global features, while balancing accuracy and speed, surpassing existing model detectors and making it more suitable for industrial deployment.
[0045] I. Sparse Self-Attention Model Structure (SA)
[0046] While ViT is highly effective at capturing interactions between distant pixels, the computational complexity of the ViT self-attention mechanism increases quadratically with the number of tokens, thus still requiring significant computational resources. Furthermore, in large datasets of steel samples captured by industrial cameras, the defect features are quite similar, leading to a large number of duplicate tokens. Therefore, we propose an improved self-attention mechanism to reduce the number of tokens, improving detection accuracy while minimizing computational costs. The SA module structure is as follows: Figure 2 As shown.
[0047] Before the self-attention operation, we reduce the spatial size of K and V by applying 1×1 convolutions and 3×3 depthwise separable convolutions to shrink the feature map. This emphasizes the local spatial context while focusing self-attention on the shrunk feature map. This helps the model capture low-frequency global information and further reduces the number of parameters. The MLP layer then maps the original features to a new high-dimensional feature space through two fully connected layers, thereby better capturing the nonlinear relationships between features. The formula is shown below:
[0048] Q = Linear(x) = xW Q #(1)
[0049] K = Linear(dF) 3x3 (F 1x1 (x)))=dF 3x3 (F 1x1 (x))W K #(2)
[0050] V = Linear(dF) 3x3 (F 1x1 (x)))=dF 3x3 (F 1x1 (x))W V #(3)
[0051]
[0052] ω=(x+At)+MLP(x+At)#(5)
[0053] x is the feature map representing the input. Q is the query vector, K is the key vector, and V is the value vector. W Q W K W V This is the learned weight matrix; `Linear` performs the linear transformation operation. dF 3x3 This represents a depthwise separable convolution with a 3x3 kernel, F 1x1 This represents a standard convolution with a 1x1 kernel. This means dividing the dot product of the query vector and the key vector by the scaling factor. The attention score obtained, where d k is the dimension of the key vector, the softmax function represents the conversion of attention scores to 0s and 1s, At represents the output of the calculation result, and ω represents the output of the SA module.
[0054] II. Model Structure of Convolution and Transformer Working Together (CTR)
[0055] While the SA module demonstrates excellent performance in extracting global information, reducing the feature map size to such a small scale and decreasing the number of tokens may ultimately lead to the loss of local information. Furthermore, the Transformer may, to some extent, degrade high-frequency local information, such as local texture and collision details. Moreover, because the captured steel dataset contains a large number of minute local and global features, these shortcomings make it difficult for ordinary detection models to adequately address them, resulting in a sharp decline in model performance.
[0056] To address the above issues, this embodiment optimizes self-attention calculation while preserving local information. It also draws inspiration from multi-branch topology to design a CTR model structure that follows the general MetaFormer architecture, such as... Figure 3 As shown. It is a block of two parallel streams:
[0057] One is the convolutional module used to extract local information from feature maps. The reason why CNNs are good at extracting local features is mainly due to their use of convolutional operations and parameter sharing mechanisms. Convolutional operations identify local patterns by sliding the convolutional kernel on the input data, while parameter sharing allows the network to share weights at different locations, improving its ability to learn local invariance.
[0058] The other input stream is the SA module, which is used to extract global features.
[0059] Finally, a simple method was used to fuse the outputs of local and global branches, and then pass them through an MLP layer to capture the non-linear relationships between features, extracting more important and obvious features.
[0060] This dual-branch design structure enables the CTR model to efficiently capture both high-frequency and low-frequency information, achieving a balance between accuracy and efficiency, and comprehensively improving the model's detection performance. The formula is shown below:
[0061] ψ(x)=BnSi(F 1x1 (BnSi(dF 3x3 (x))))#(6)
[0062] ρ(x)=x+ω+ψ(x)#(7)
[0063] x is the feature map representing the input, dF 3x3 This represents a depthwise separable convolution with a 3x3 kernel, F 1x1 represents a standard convolution with a 1x1 kernel. BnSi represents normalization and the SiLu activation function. ψ(x) represents the output of the convolutional module, and ρ(x) represents the output of the CTR module.
[0064] III. Bottleneck Convolutional Structure (GDC)
[0065] The convolutional structure GDC proposed in this embodiment mainly consists of grouped convolutions, depthwise separable convolutions, and ordinary convolutions. The advantages of grouped convolutions are primarily reflected in reducing the number of parameters, enabling parallel processing, enhancing local feature learning, and improving model adaptability. Depthwise separable convolutions are an advantageous operation in convolutional neural networks, not only reducing storage requirements and improving computational efficiency, but also being suitable for resource-constrained mobile devices and embedded systems.
[0066] The Bottleneck Convolutional Structure (GDC) first uses grouped convolutions and regular convolutions for spatial downsampling, enabling cross-channel information interaction and extracting richer features. Second, it employs 3x3 depthwise separable convolutions to capture spatial feature information, allowing the model to better capture image features both spatially and in terms of channels. Finally, it uses convolutions again for upsampling to restore the feature map dimensions.
[0067] This bottleneck structure design reduces the number of parameters while improving detection accuracy. In the CSTRNet model, the GDC module can also fuse different features, combining shallow and deep features extracted by the CTR module—features that contain more information helpful for detection. The formula is shown below:
[0068]
[0069] x is the input feature map, gF 3x3 This represents a grouped convolution with a 3x3 kernel, F 1x1 This represents a standard convolution with a 1x1 kernel, dF 3x3 This represents a depthwise separable convolution with a 3x3 kernel. This indicates the output of the GDC module.
[0070] The main function of the GDC module is to extract feature information, while the function of the "Detect" layer is to convert these feature maps into detection results, that is, to determine the objects in the image and their locations.
[0071] 1. Decoding bounding boxes: Multiple bounding boxes are generated at each location in the feature map, and each bounding box represents a region that may contain an object.
[0072] 2. Calculate confidence score: For each bounding box, the model calculates the confidence score that it contains an object. This confidence score reflects the model's level of confidence in whether the bounding box contains an object.
[0073] 3. Category prediction: The model also predicts the category of the object for each bounding box.
[0074] 4. Non-maximum suppression (NMS): Non-maximum suppression is performed on the detection results to eliminate overlapping bounding boxes and select the bounding box most likely to contain the object. This ensures that each object is detected only once and selects the most accurate bounding box.
[0075] 5. Output Results: Finally, the "Detect" layer will output results containing the bounding boxes of detected objects, confidence scores, and object categories. These results are typically composed of the bounding box coordinates, confidence scores, and category labels.
[0076] IV. Experimental Verification
[0077] Step 1: Dataset selection.
[0078] This embodiment uses the publicly available steel dataset provided by Northeastern University (NEU-DET). This dataset contains 1800 grayscale images, covering six different types of typical surface defects: rolled oxide scale, patches, cracks, pitting, inclusions, and scratches. Each type of defect contains 300 samples, providing abundant data for training and evaluating the model, reducing the risk of overfitting, and helping the model learn the characteristics and patterns of different defect types, thus improving the model's generalization ability.
[0079] Step 2: Image preprocessing.
[0080] When preprocessing images from a steel image dataset, a series of operations are required to increase data diversity and robustness. Specifically, the following preprocessing steps can be taken:
[0081] 1. By randomly changing the brightness, contrast, and color of an image, it is possible to simulate image changes under different lighting conditions and environments. This helps the model to better adapt to changes in lighting and color shifts.
[0082] 2. Random rotation and flipping: Randomly rotating and flipping images can increase the diversity of data and improve the robustness of the model.
[0083] 3. Random cropping and stitching: Random cropping and stitching of images can simulate the missing or covered areas in an image.
[0084] 4. Noise addition: Adding random noise, such as Gaussian noise or salt and pepper noise, to the image can simulate noise interference in different environments, thereby improving the model's ability to resist noise interference.
[0085] The above image preprocessing steps can generate more diverse and complex training data, effectively expanding the scale and diversity of the dataset, thereby improving the model's generalization ability and robustness.
[0086] Step 3: Observe and analyze the dataset.
[0087] The steel dataset contains numerous minute defects and slender flaws such as scratches. Because the dataset consists of close-up images taken with industrial cameras, many similarities exist. Analyzing the characteristics and problems of the dataset lays the foundation for model design and innovation.
[0088] Step 4: Model building.
[0089] First, to address the issues of model complexity and large parameter count, an SA self-attention module was designed. Second, a CTR model structure was designed to solve the problem of difficulty in extracting localized minor defects and global defects such as scratches in the steel dataset. Finally, a GDC module was designed, using convolutional neural networks to extract features and fuse different features to establish the overall CSTRNet model structure, allowing the model to achieve optimal performance.
[0090] Step 5: Model training.
[0091] The preprocessed steel data is fed into the CSTRNet model for training. By adjusting the training parameters and training multiple times, the training results are generated.
[0092] Step six, model evaluation.
[0093] Precision, recall, and mean AP (mAP) are used as evaluation metrics. Precision refers to the proportion of samples correctly classified as positive out of all samples predicted as positive. It measures how many samples the model predicts to be positive are actually positive. The formula is shown below.
[0094]
[0095] In this context, TP indicates that the true class of the sample is positive (P), and the model predicts it to be positive (P) as well, in which case the prediction is correct. FP indicates that the true class of the sample is negative (N), but the model predicts it to be positive (P), in which case the prediction is incorrect.
[0096] Recall rate refers to the proportion of samples that are correctly classified as positive out of all true positive samples.
[0097] It measures how many positive class samples the model can correctly identify. The formula is shown below.
[0098]
[0099] Here, FN indicates that the true class of the sample is positive sample P, but the model predicts it as negative sample N, which is a prediction error.
[0100] mAP stands for Mean Precision, which is the average AP across all classes and is used to express the performance of multi-class label prediction. A higher mAP indicates better performance. mAP@.5 refers to the mAP value when the Interchange of Units (IOU) is 0.5. mAP@.5:.95 refers to the average mAP when the IOU is in the range (0.5:0.95:0.05). The formulas for AP and mAP are shown below.
[0101] AP=∫0 t P(R)dR#(11)
[0102]
[0103] Based on the evaluation scheme, the model with the highest accuracy is selected from all training results.
[0104] The evaluation results of the optimal model are as follows Figure 5 As shown. In Figure 5 In the diagram, the x-axis of each graph represents the number of training epochs (500 epochs in total). On the training dataset:
[0105] box_loss is the box regression loss, which measures the difference between the predicted box position and the ground truth box position in the object detection model. When box_loss is low, it means that the model can accurately predict the position of the object.
[0106] obj_loss represents the target confidence loss, which measures the accuracy of the target detection model in judging whether a target exists in an image. When obj_loss is low, it means that the model can correctly identify the target in the image.
[0107] cls_loss represents the category classification loss, which measures the accuracy of the object detection model in classifying different categories. When cls_loss is low, it means that the model has high accuracy in distinguishing different categories.
[0108] Meanwhile, the box_loss, obj_loss, and cls_loss values on the validation dataset are all low, indicating good model performance.
[0109] The optimal model was validated on a steel test set, and the results are as follows: Figure 6 As shown.
[0110] Example 2
[0111] This embodiment discloses a steel surface defect detection system.
[0112] A steel surface defect detection system, including:
[0113] The image acquisition module is configured to acquire images of the steel surface.
[0114] The CTR module feature fusion module is configured to: input the steel surface image into multiple cascaded CTR modules of the CSTRNet model; in each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules respectively, and feature fusion is performed.
[0115] The GDC module feature fusion module is configured to: input the features extracted by each intermediate CTR module into the serially connected two-layer GDC module in sequence, extract features using a convolutional neural network, and use the two-layer GDC module to perform bidirectional fusion of the shallow and deep features extracted by the CTR module. The two-layer GDC module is used to obtain the predicted box position, defect confidence and defect classification of steel surface defects.
[0116] Example 3
[0117] The purpose of this embodiment is to provide a computer-readable storage medium.
[0118] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the steel surface defect detection method as described in Embodiment 1 of this disclosure.
[0119] Example 4
[0120] The purpose of this embodiment is to provide an electronic device.
[0121] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the steel surface defect detection method as described in Embodiment 1 of this disclosure.
[0122] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0123] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0124] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for detecting surface defects in steel, characterized in that, Includes the following steps: Acquire images of the steel surface; The CSTRNet model is constructed, which includes multiple cascaded CTR modules and a GDC module that merges the shallow and deep features extracted by the CTR modules, wherein: The CTR module consists of two parallel streams: a convolutional module for extracting local information from the feature map and a sparse self-attention module for extracting global features. Finally, the outputs of the local and global branches are fused and passed through an MLP layer to capture the non-linear relationships between features. The GDC module includes grouped convolution, depthwise separable convolution, and ordinary convolution: First, spatial downsampling is performed using grouped convolution and ordinary convolution to interact information across channels and extract richer features; second, depthwise separable convolution is used to capture spatial feature information, enabling the model to better capture image features in both space and channels; finally, convolution is used again for upsampling to restore the feature map dimension. The steel surface image is input into multiple cascaded CTR modules of the CSTRNet model. In each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules, respectively, and feature fusion is performed. Specifically, in each CTR module: By utilizing a sparse self-attention module, the feature map is reduced by concatenated standard convolution and depthwise separable convolution before the self-attention operation, and the output of the sparse self-attention module is obtained. In the convolutional module, features are extracted using concatenated depthwise separable convolutions and standard convolutions to obtain the output of the convolutional module; The output of the sparse self-attention module, the output of the convolution module, and the input feature map are fused to obtain the output of the CTR module; The features extracted by each intermediate CTR module are sequentially input into the cascaded two-layer GDC module. Convolutional neural networks are used to extract features, and the shallow and deep features extracted by the CTR module are fused bidirectionally using the two-layer GDC module. The predicted box position, defect confidence, and defect classification of steel surface defects are obtained through the two-layer GDC module. Specifically, a two-layer GDC module is used to bidirectionally fuse the shallow and deep features extracted by the CTR module, as follows: The features extracted by the top-level CTR module are scaled by the SPPF module and then input into the first-level GDC module; The first-layer GDC module implements feature propagation from back to front, while the second-layer GDC module implements feature propagation from front to back. The first-layer GDC module is connected to the second-layer GDC module to achieve bidirectional feature transfer; In each GDC module: First, spatial downsampling is performed using grouped convolution and regular convolution. Then, depthwise separable convolution is used to capture feature information in the feature map space. Finally, regular convolution and grouped convolution are used for upsampling to restore the feature map dimension.
2. The method for detecting surface defects in steel as described in claim 1, characterized in that, The output of the CTR module is expressed as follows: It represents the feature map of the input. This represents a depthwise separable convolution with a 3x3 kernel. This represents a standard convolution with a 1x1 kernel; BnSi represents normalization, and SiLu represents the activation function. Represents the output of the convolutional module. This represents the output of the CTR module. This represents the output of the sparse self-attention module.
3. The method for detecting surface defects in steel as described in claim 2, characterized in that, The output of the sparse self-attention module is represented as: It is a query vector. It is a key vector. It is a value vector; , , It is the learned weight matrix. It performs a linear transformation operation; This means dividing the dot product of the query vector and the key vector by the scaling factor. The attention scores obtained, where It is the dimension of the key vector. The function represents converting the attention score to a value between 0 and 1, and At represents the output of the calculation result.
4. A steel surface defect detection system, characterized in that, include: The image acquisition module is configured to acquire images of the steel surface. The CSTRNet model is constructed, which includes multiple cascaded CTR modules and a GDC module that merges the shallow and deep features extracted by the CTR modules, wherein: The CTR module consists of two parallel streams: a convolutional module for extracting local information from the feature map and a sparse self-attention module for extracting global features. Finally, the outputs of the local and global branches are fused and passed through an MLP layer to capture the non-linear relationships between features. The GDC module includes grouped convolution, depthwise separable convolution, and ordinary convolution: First, spatial downsampling is performed using grouped convolution and ordinary convolution to interact information across channels and extract richer features; second, depthwise separable convolution is used to capture spatial feature information, enabling the model to better capture image features in both space and channels; finally, convolution is used again for upsampling to restore the feature map dimension. The CTR module feature fusion module is configured to: input the steel surface image into multiple cascaded CTR modules of the CSTRNet model; in each CTR module, the global and local features of the steel surface image are extracted by parallel sparse self-attention modules and convolution modules respectively, and feature fusion is performed. Specifically, in each CTR module: By utilizing a sparse self-attention module, the feature map is reduced by concatenated standard convolution and depthwise separable convolution before the self-attention operation, and the output of the sparse self-attention module is obtained. In the convolutional module, features are extracted using concatenated depthwise separable convolutions and standard convolutions to obtain the output of the convolutional module; The output of the sparse self-attention module, the output of the convolution module, and the input feature map are fused to obtain the output of the CTR module; The GDC module feature fusion module is configured to: input the features extracted by each intermediate CTR module into the serial double-layer GDC module in sequence, extract features using a convolutional neural network, and use the double-layer GDC module to perform bidirectional fusion of the shallow and deep features extracted by the CTR module. The double-layer GDC module is used to obtain the predicted box position, defect confidence and defect classification of steel surface defects. Specifically, a two-layer GDC module is used to bidirectionally fuse the shallow and deep features extracted by the CTR module, as follows: The features extracted by the top-level CTR module are scaled by the SPPF module and then input into the first-level GDC module; The first-layer GDC module implements feature propagation from back to front, while the second-layer GDC module implements feature propagation from front to back. The first-layer GDC module is connected to the second-layer GDC module to achieve bidirectional feature transfer; In each GDC module: First, spatial downsampling is performed using grouped convolution and regular convolution. Then, depthwise separable convolution is used to capture feature information in the feature map space. Finally, regular convolution and grouped convolution are used for upsampling to restore the feature map dimension.
5. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the steel surface defect detection method as described in any one of claims 1-3.
6. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the steel surface defect detection method as described in any one of claims 1-3.
Citation Information
Patent Citations
Aerial photography insulator spontaneous explosion defect detection model, method and equipment and storage medium
CN114677357A
Defect detection method based on joint optimization and mixed attention feature fusion
CN115294038A