Lightweight steel surface multi-class defect detection method based on improved RT-DETR
By improving the RT-DETR network, combined with the CSPDarknet backbone network, CSP-CDMSA module, AIFI-DPB module and DAMSFPN feature fusion module, the missed detection and missed detection in steel surface defect detection is solved, and efficient multi-category and multi-scale defect detection is achieved.
Patent Information
- Application Number
- CN202510462888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art has missed and missed detection in steel surface defect detection, making it difficult to take into account both lightweight and detection accuracy in complex contexts, especially in the detection of multiple categories and multi-scale defects.
Using the improved RT-DETR network, the CSPDarknet backbone network, CSP-CDMSA module, AIFI-DPB module and DAMSFPN feature fusion module are introduced to optimize multi-scale feature extraction and spatial relationship modeling to enhance defect detection capabilities.
It realizes high-precision detection of multi-category and multi-scale defects in complex contexts, reduces computing overhead, and improves the balance between detection accuracy and lightweight design.
Smart Images

Figure CN120374557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and steel surface defect detection, and more specifically, to a lightweight multi-class defect detection method for steel surfaces based on improved RT-DETR. Background Technique
[0002] As a core material for infrastructure and industrial manufacturing, the detection of surface defects in steel is directly related to the reliability and safety of products. During the production process of steel, various types of surface defects may be formed, such as cracks, inclusions, and rolling marks. If these defects are not detected and processed in a timely manner, the material properties will be significantly reduced, and even failures may occur. The multi-class targets of steel surface defects have complex characteristics such as diverse types, significant differences in morphology and size, fuzzy features, and multi-scale features, requiring the detection system to have efficient classification and recognition capabilities to cope with the diversity of various defects and background interference. Therefore, the research and development of accurate and efficient steel surface defect detection technology is crucial for improving product quality and optimizing production efficiency.
[0003] In the field of industrial defect detection, with the rapid development of computer vision and deep learning technologies, detection methods based on convolutional neural networks (CNNs) have become a research hotspot, mainly divided into single-stage and two-stage algorithms. Two-stage algorithms (such as Faster R-CNN) generate candidate regions through a region proposal network (RPN), and then perform precise classification and boundary regression through classification and regression networks. However, they rely on high-quality training samples, have a slow inference speed, and require a large amount of computing resources, making it difficult to meet the industrial requirements of real-time detection. Single-stage algorithms (such as YOLO, RetinaNet, etc.) can directly predict the target category and bounding box, with a fast inference speed. However, due to the lack of a candidate region generation step, they are easily interfered by the background, resulting in lower detection accuracy. When dealing with multi-class or multi-scale defects, the detection ability still needs to be enhanced.
[0004] In recent years, Transformer-based object detection algorithms represented by RT-DETR have shown significant advantages in the field of defect detection. Compared with two-stage algorithms, it omits the region proposal step, avoids error accumulation, and uses a global attention mechanism to more accurately capture long-range dependencies. Compared with one-stage algorithms, it optimizes multi-scale features and combines object queries to avoid the bias brought by anchors, improving the accuracy. Its hybrid encoder optimizes multi-scale feature processing, while the Transformer decoder captures global dependencies through the multi-head self-attention mechanism, simplifying the process and enhancing the ability to adapt to complex scenarios.
[0005] The detection of steel surface defects faces significant challenges, mainly due to the diversity of defect categories, including characteristics such as a large number of categories, high similarity between some defects, diverse scales, and low contrast. The traditional RT-DETR detection framework is vulnerable to background interference and dilution of global information, resulting in missed detections and false detections; at the same time, enhancing the ability to extract detailed features requires more computing resources, further exacerbating the difficulty of balancing lightweight and accuracy. Summary of the Invention
[0006] To overcome the above-mentioned defects of missed detections and false detections in the existing technology, as well as the difficulty of balancing lightweight and detection accuracy, the present invention provides a lightweight multi-category defect detection method for steel surfaces based on improved RT-DETR. The designed LMCD-RTDETR network effectively improves the detection ability for large defect scale differences and irregular shapes through multi-module collaborative optimization, achieves an effective balance between detection accuracy and lightweight design, and can meet the challenges of category diversity and scale differences in steel surface defect detection under complex backgrounds.
[0007] To solve the above technical problems, the technical solution of the present invention is as follows:
[0008] A lightweight multi-category defect detection method for steel surfaces based on improved RT-DETR, comprising the following steps:
[0009] S1: Obtain a steel surface defect dataset and perform preprocessing, and divide the preprocessed steel surface defect dataset into a training set and a test set; the steel surface defect dataset includes several steel images with different surface defects respectively;
[0010] S2: Select the RT-DETR R18 model as the basic model, and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR;
[0011] The LMCD-RTDETR includes a backbone network, an encoder, a decoder, and a detection head connected in sequence; among them, the CSPDarknet network is used to replace the ResNet-18 network in the RT-DETR R18 model as the backbone network, and the CSP-CDMSA module is introduced into the backbone network; the encoder includes an AIFI-DPB module and a DAMSFPN feature fusion module connected in sequence;
[0012] S3: Iteratively train LMCD-RTDETR using the training set to obtain a trained LMCD-RTDETR;
[0013] S4: Input the test set into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set.
[0014] Preferably, in the step S1, the preprocessing includes: cropping all steel images in the steel surface defect dataset to a unified size and respectively converting them into grayscale images;
[0015] Manually annotating the positions and types of surface defects in each grayscale image to complete the preprocessing.
[0016] Preferably, in the step S2, the backbone network includes, connected in sequence: a first convolutional layer, a second convolutional layer, a first CSP-CDMSA module, a third convolutional layer, a second CSP-CDMSA module, a fourth convolutional layer, a third CSP-CDMSA module, a fifth convolutional layer, and a fourth CSP-CDMSA module;
[0017] The first CSP-CDMSA module, the second CSP-CDMSA module, and the third CSP-CDMSA module respectively output feature maps P2, P3, and P4 to the DAMSFPN feature fusion module;
[0018] The output of the fourth CSP-CDMSA module is connected to the AIFI-DPB module, and the AIFI-DPB module outputs a feature map P5 to the DAMSFPN feature fusion module;
[0019] The DAMSFPN feature fusion module fuses the feature maps P2 to P5 and then inputs them into the decoder.
[0020] Preferably, each CSP-CDMSA module has the same structure and includes, connected in sequence: a sixth convolutional layer, a splitting layer, several CDMSA secondary modules, a first splicing layer, and a sixth convolutional layer; the outputs of the splitting layer and each CDMSA secondary module are also respectively connected to the first splicing layer;
[0021] The structure of each CDMSA secondary module includes, connected in sequence: a 3×3 convolutional layer, a 5×5 depthwise separable convolutional layer, a 7×7 depthwise separable convolutional layer, a second splicing layer, a 1×1 convolutional layer, and an SCSA spatial and channel collaborative attention tertiary module; the outputs of the 3×3 convolutional layer and the 5×5 convolutional layer are also respectively connected to the second splicing layer; the input of the 3×3 convolutional layer is also connected to the output of the SCSA spatial and channel collaborative attention tertiary module to form a residual addition connection.
[0022] Preferably, in the step S2, the structure of the AIFI-DPB module includes, connected in sequence: a seventh convolutional layer, a DPB attention secondary module, a first normalization layer, an FFN layer, a second normalization layer, and an eighth convolutional layer; the input of the DPB attention secondary module and the input of the first normalization layer also form a residual addition connection; the input of the FFN layer also forms a residual addition connection with the input of the second normalization layer;
[0023] The DPB attention secondary module adjusts the attention score by dynamically calculating the relative position bias B ij The attention mechanism of the DPB attention secondary module is expressed as:
[0024]
[0025] B ij = DPB(Δx ij , Δy ij ) = MLP(LayerNorm(Δx ij , Δy ij ))
[0026] where, Attention(Q i , K j , V j ) represents the attention score of the i-th and j-th and features; Q i represents the query matrix of the i-th feature; K j represents the key matrix of the j-th feature; V j represents the value matrix of the j-th feature; Softmax represents the Softmax activation function; d k represents the dimension of the key matrix K j ; (Δx ij , Δy ij ) represents the relative offsets in the horizontal and vertical directions; MLP represents a multi-layer perceptron; LayerNorm represents layer normalization;
[0027] The output of the fourth CSP-CDMSA module is connected to the seventh convolutional layer; the eighth convolutional layer outputs the feature map P5 to the DAMSFPN feature fusion module.
[0028] Preferably, in the step S2, the structure of the DAMSFPN feature fusion module includes, arranged in parallel: a first branch, a second branch, and a third branch; each branch includes, connected in sequence: a ninth convolutional layer, a first fusion layer, a first CSP-DMSC secondary module, a second fusion layer, and a second CSP-DMSC secondary module;
[0029] Input the feature maps P3, P4, and P5 into the ninth convolutional layer of the first branch, the second branch, and the third branch respectively; downsample the feature map P2 and input it into the first fusion layer of the first branch;
[0030] Downsample the output of the ninth convolutional layer of the first branch and input it into the first fusion layer of the second branch; downsample the output of the ninth convolutional layer of the second branch and input it into the first fusion layer of the third branch;
[0031] Downsample the output of each CSP-DMSC secondary module of the first branch and input it into the second fusion layer of the second branch; downsample the output of each CSP-DMSC secondary module of the second branch and input it into the second fusion layer of the third branch;
[0032] Perform EUCB efficient upsampling on the output of the first CSP-DMSC secondary module of the second branch and input it into each fusion layer of the first branch; perform EUCB efficient upsampling on the output of the first CSP-DMSC secondary module of the third branch and input it into each fusion layer of the second branch;
[0033] Input the outputs of the second CSP-DMSC secondary modules of each branch into the decoder together.
[0034] Preferably, the calculation formula of each said fusion layer is expressed as:
[0035]
[0036] where, F fusion is the output of the fusion layer; F i is the i-th feature map; ω i is the weight of the i-th feature map; N is the number of feature maps input into the fusion layer.
[0037] Preferably, the formula of the EUCB efficient upsampling is:
[0038]
[0039] where, is the output of the EUCB efficient upsampling; X is the input of the EUCB efficient upsampling; represents the composite operation of depth convolution after upsampling, C is the number of channels, and H and W are the height and width of the input feature map respectively; represents the channel rearrangement operation, g represents the number of groups of rearrangement, represents the pointwise convolution operation, and C′ represents the number of channels after pointwise convolution.
[0040] Preferably, the calculation formula of each said CSP-DMSC secondary module is:
[0041]
[0042] Among them, is the output of the CSP-DMSC secondary module; X is the input of the CSP-DMSC secondary module; and respectively represent the first and second pointwise convolution operations. C and C' are the original number of channels and the number of channels after the first pointwise convolution respectively, and H and W are the height and width of the input feature map respectively; represents a depth convolution operation using a kernel size of k i where k is the number of preset convolution kernel sizes; represents an effective channel attention operation, represents a channel rearrangement operation.
[0043] Preferably, in step S4, it further includes:
[0044] using a preset evaluation metric to evaluate the performance of the trained LMCD-RTDETR; the evaluation metric includes at least any one or more of precision P, recall R, mean average precision mAP, number of parameters Params, and GFLOPs.
[0045] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0046] The present invention provides a lightweight multi-class defect detection method for steel surface based on improved RT-DETR. First, a steel surface defect dataset is obtained and preprocessed, and the preprocessed steel surface defect dataset is divided into a training set and a test set; then the RT-DETR R18 model is selected as the base model, and the RT-DETR R18 model is improved to construct a lightweight steel surface defect detection model LMCD-RTDETR; then the training set is used to iteratively train LMCD-RTDETR to obtain the trained LMCD-RTDETR; finally, the test set is input into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set;
[0047] The lightweight steel surface defect detection model LMCD-RTDETR provided by the present invention has the following advantages:
[0048] 1) CSPDarknet is adopted as the backbone architecture to effectively extract deep semantic features while minimizing redundant calculations. Meanwhile, a cascaded dynamic multi-scale convolution module (CSP-CDMSA) is designed. This module dynamically calibrates the importance of features at different scales, integrates multi-scale convolutions to enhance the target perception ability, and reduces the calculation of irrelevant features. It operates in conjunction with the spatial-channel collaborative attention mechanism to precisely highlight key defect regions and enhance the recognition of low-contrast and edge-blurred defects.
[0049] 2) The present invention proposes an AIFI-DPB module based on the dynamic position bias attention mechanism (DPB). This AIFI variant dynamically generates relative position biases to adjust self-attention calculations, thereby enhancing the spatial relationship modeling ability, improving the discrimination ability between categories, and reducing the false detection rate and missed detection rate.
[0050] 3) The present invention constructs a dynamic adaptive multi-scale feature pyramid network (DAMSFPN). Through multi-scale weighted fusion, efficient upsampling techniques, multi-scale convolution modules, and a global heterogeneous kernel selection mechanism, it coordinates deep semantic information and shallow detail features, dynamically adjusts the importance of features at different scales, expands the receptive field, can effectively handle large differences in scale changes, and enhances the expression ability of multi-scale and irregular features. Description of the Drawings
[0051] Figure 1 It is a flowchart of a lightweight steel surface multi-class defect detection method based on the improved RT-DETR provided in Embodiment 1.
[0052] Figure 2 It is a structural diagram of the LMCD-RTDETR model provided in Embodiment 2.
[0053] Figure 3 It is a structural diagram of the CSP-CDMSA module provided in Embodiment 2.
[0054] Figure 4 It is a structural diagram of the AIFI-DPB module provided in Embodiment 2.
[0055] Figure 5 It is a structural diagram of the CSP-DMSC secondary module provided in Embodiment 2.
[0056] Figure 6 It is a comparison diagram of the visualization experiment results provided in Embodiment 2. Detailed Implementation Manner
[0057] The drawings are only for illustrative purposes and should not be construed as a limitation to this application.
[0058] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;
[0059] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0060] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.
[0061] Embodiment 1
[0062] As Figure 1 shown, this embodiment provides a lightweight multi-class defect detection method for steel surface based on improved RT-DETR, including the following steps:
[0063] S1: Obtain the steel surface defect data set and perform preprocessing, and divide the preprocessed steel surface defect data set into a training set and a test set; the steel surface defect data set includes several steel images with different surface defects respectively;
[0064] S2: Select the RT-DETR R18 model as the basic model, and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR;
[0065] The LMCD-RTDETR includes a backbone network, an encoder, a decoder and a detection head connected in sequence; among them, the CSPDarknet network is used to replace the ResNet-18 network in the RT-DETR R18 model as the backbone network, and the CSP-CDMSA module is introduced into the backbone network; the encoder includes an AIFI-DPB module and a DAMSFPN feature fusion module connected in sequence;
[0066] S3: Iteratively train the LMCD-RTDETR using the training set to obtain the trained LMCD-RTDETR;
[0067] S4: Input the test set into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set.
[0068] In the specific implementation process, first obtain the steel surface defect data set and perform preprocessing, and divide the preprocessed steel surface defect data set into a training set and a test set;
[0069] Then select the RT-DETR R18 model as the basic model, and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR;
[0070] RT-DETR R18 is an object detection model based on the Transformer architecture. It uses ResNet-18 as the backbone network, and effectively alleviates the problem of gradient disappearance in deep networks through residual connections. The model processes multi-scale feature maps through a hybrid encoder, and enhances feature expression by using the intra-scale interaction and cross-scale fusion mechanisms. The decoder part adopts the Transformer structure and uses the multi-head self-attention mechanism. Through a fixed number of object queries, it directly predicts the category and location of the object, omitting the steps of candidate box generation and post-processing in traditional object detection methods, and realizing an end-to-end detection process. The model uses a multi-task loss function and performs object matching through the Hungarian algorithm to ensure the precise alignment between the decoder output and the real object. Although RT-DETR R18 shows good performance in defect detection tasks, how to optimize the model structure under limited computing resources to reduce the computational burden, while improving the detection ability of multi-scale and small targets and the category discrimination ability, remains a key problem in current research;
[0071] Based on RT-DETR R18, this embodiment proposes a lightweight LMCD-RTDETR network, aiming to improve the detection performance of small targets, multi-categories and multi-scale defects in complex backgrounds, and balance model lightweight and detection accuracy. The network uses CSPDarknet to replace ResNet-18 as the backbone network, which can extract deep features more efficiently, reduce redundant calculations, and improve the detection ability of defects in complex scenarios. A lightweight cascaded dynamic multi-scale attention module (CSP-CDMSA) is designed. By dynamically fusing convolutional features and spatial channel collaborative attention mechanisms, it enhances the recognition and discrimination ability of low-contrast and edge-blurred defects, and reduces the computational overhead. The AIFI-DPB module is proposed to optimize the scale interaction mechanism. Based on the dynamic relative position bias, it adjusts the self-attention mechanism, strengthens the interaction between local and global features, and improves the classification ability and feature discrimination. The DAMSFPN feature fusion module is proposed to optimize CCFM, which fuses multi-scale efficient convolution and global heterogeneous kernel selection mechanisms, strengthens the multi-scale object detection ability, and enhances the detail recovery through an efficient upsampling strategy. Using the feature weighted fusion method of BIFPN further improves the detection accuracy and reduces the computational complexity. Through the collaborative optimization of multiple modules, LMCD-RTDETR effectively improves the detection ability for large defect scale differences and irregular shapes, and realizes the balance between lightweight and high precision;
[0072] After that, the training set is used to iteratively train LMCD-RTDETR to obtain the trained LMCD-RTDETR;
[0073] Finally, the test set is input into the trained LMCD-RTDETR for defect detection, and the defect detection results of each steel image in the test set are obtained;
[0074] This embodiment proposes a lightweight defect detection algorithm based on RT-DETR R18: LMCD-RTDETR. By optimizing the model structure and integrating a lightweight cascaded dynamic multi-scale convolution module, a dynamic position perception module, and an improved feature fusion network, it balances lightweight and high precision while reducing the computational overhead and improving the detection accuracy. This method can enhance the model's ability to detect and distinguish diverse defects in terms of category and size in complex backgrounds.
[0075] Embodiment 2
[0076] This embodiment provides a lightweight multi-category defect detection method for steel surfaces based on improved RT-DETR, including the following steps:
[0077] S1: Obtain a steel surface defect dataset and perform preprocessing. Divide the preprocessed steel surface defect dataset into a training set and a test set. The steel surface defect dataset includes several steel images with different surface defects respectively.
[0078] S2: Select the RT-DETR R18 model as the basic model and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR.
[0079] The LMCD-RTDETR includes a backbone network, an encoder, a decoder, and a detection head connected in sequence. Among them, the CSPDarknet network is used to replace the ResNet-18 network in the RT-DETR R18 model as the backbone network, and the CSP-CDMSA module is introduced into the backbone network. The encoder includes an AIFI-DPB module and a DAMSFPN feature fusion module connected in sequence.
[0080] S3: Iteratively train LMCD-RTDETR using the training set to obtain the trained LMCD-RTDETR.
[0081] S4: Input the test set into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set.
[0082] In the step S1, the preprocessing includes: cropping all steel images in the steel surface defect dataset to a unified size and converting them into grayscale images respectively.
[0083] Manually annotate the positions and types of surface defects in each grayscale image to complete the preprocessing.
[0084] In step S2, the backbone network includes, in sequence: a first convolutional layer, a second convolutional layer, a first CSP-CDMSA module, a third convolutional layer, a second CSP-CDMSA module, a fourth convolutional layer, a third CSP-CDMSA module, a fifth convolutional layer, and a fourth CSP-CDMSA module;
[0085] The first CSP-CDMSA module, the second CSP-CDMSA module, and the third CSP-CDMSA module respectively output feature maps P2, P3, and P4 to the DAMSFPN feature fusion module;
[0086] The output of the fourth CSP-CDMSA module is connected to the AIFI-DPB module, and the AIFI-DPB module outputs a feature map P5 to the DAMSFPN feature fusion module;
[0087] The DAMSFPN feature fusion module fuses the feature maps P2 to P5 and then inputs them into the decoder;
[0088] Each CSP-CDMSA module has the same structure and includes, in sequence: a sixth convolutional layer, a splitting layer, several CDMSA secondary modules, a first splicing layer, and a sixth convolutional layer; the outputs of the splitting layer and each CDMSA secondary module are also respectively connected to the first splicing layer;
[0089] The structure of each CDMSA secondary module includes, in sequence: a 3×3 convolutional layer, a 5×5 depthwise separable convolutional layer, a 7×7 depthwise separable convolutional layer, a second splicing layer, a 1×1 convolutional layer, and an SCSA spatial and channel collaborative attention tertiary module; the outputs of the 3×3 convolutional layer and the 5×5 convolutional layer are also respectively connected to the second splicing layer; the input of the 3×3 convolutional layer is also connected to the output of the SCSA spatial and channel collaborative attention tertiary module to form a residual addition connection;
[0090] In step S2, the structure of the AIFI-DPB module includes, in sequence: a seventh convolutional layer, a DPB attention secondary module, a first normalization layer, an FFN layer, a second normalization layer, and an eighth convolutional layer; the input of the DPB attention secondary module and the input of the first normalization layer also form a residual addition connection; the input of the FFN layer and the input of the second normalization layer also form a residual addition connection;
[0091] The DPB attention secondary module dynamically calculates the relative position bias B ij to adjust the attention score, and the attention mechanism of the DPB attention secondary module is expressed as:
[0092]
[0093] Bij = DPB(Δx ij , Δy ij ) = MLP(LayerNorm(Δx ij , Δy ij ))
[0094] Wherein, Attention(Q i , K j , V j ) represents the attention score of the i-th and j-th features; Q i represents the query matrix of the i-th feature; K j represents the key matrix of the j-th feature; V j represents the value matrix of the j-th feature; Softmax represents the Softmax activation function; d k represents the dimension of the key matrix K j ; (Δx ij , Δy ij ) represents the relative offsets in the horizontal and vertical directions; MLP represents a multi-layer perceptron; LayerNorm represents layer normalization;
[0095] The output of the fourth CSP-CDMSA module is connected to the seventh convolutional layer; The eighth convolutional layer outputs the feature map P5 to the DAMSFPN feature fusion module;
[0096] In the step S2, the structure of the DAMSFPN feature fusion module includes: a first branch, a second branch, and a third branch arranged in parallel; Each branch includes, connected in sequence: a ninth convolutional layer, a first fusion layer, a first CSP-DMSC secondary module, a second fusion layer, and a second CSP-DMSC secondary module;
[0097] The feature maps P3, P4, and P5 are respectively input into the ninth convolutional layers of the first branch, the second branch, and the third branch; The feature map P2 is downsampled and then input into the first fusion layer of the first branch;
[0098] The output of the ninth convolutional layer of the first branch is downsampled and then input into the first fusion layer of the second branch; The output of the ninth convolutional layer of the second branch is downsampled and then input into the first fusion layer of the third branch;
[0099] The outputs of each CSP-DMSC secondary module of the first branch are respectively downsampled and then input into the second fusion layer of the second branch; The outputs of each CSP-DMSC secondary module of the second branch are respectively downsampled and then input into the second fusion layer of the third branch;
[0100] The output of the first CSP-DMSC secondary module of the second branch is input into each fusion layer of the first branch after EUCB efficient upsampling; the output of the first CSP-DMSC secondary module of the third branch is input into each fusion layer of the second branch after EUCB efficient upsampling;
[0101] The outputs of the second CSP-DMSC secondary modules of each branch are jointly input into the decoder;
[0102] The calculation formula of each said fusion layer is expressed as:
[0103]
[0104] where, F fusion is the output of the fusion layer; F i is the i-th feature map; ω i is the weight of the i-th feature map; N is the number of feature maps input into the fusion layer;
[0105] The formula of the said EUCB efficient upsampling is:
[0106]
[0107] where, is the output of the EUCB efficient upsampling; X is the input of the EUCB efficient upsampling; represents the composite operation of depth convolution after upsampling, C is the number of channels, and H and W are the height and width of the input feature map respectively; represents the channel rearrangement operation, g represents the number of groups for rearrangement, represents the pointwise convolution operation, C′ represents the number of channels after pointwise convolution;
[0108] The calculation formula of each said CSP-DMSC secondary module is:
[0109]
[0110] where, is the output of the CSP-DMSC secondary module; X is the input of the CSP-DMSC secondary module; and respectively represent the first and second pointwise convolution operations, C and C′ are the original number of channels and the number of channels after the first pointwise convolution respectively, and H and W are the height and width of the input feature map respectively; represents the depth convolution operation with a kernel size of k i where k is the number of preset convolution kernel sizes; represents the effective channel attention operation, represents the channel rearrangement operation;
[0111] In step S4, it further includes:
[0112] Use a preset evaluation metric to evaluate the performance of the trained LMCD-RTDETR; the evaluation metric includes at least any one or more of precision P, recall rate R, mean average precision mAP, number of parameters Params, and GFLOPs.
[0113] In the specific implementation process, first obtain the steel surface defect dataset and perform preprocessing, and divide the preprocessed steel surface defect dataset into a training set and a test set; in this embodiment, all steel images in the steel surface defect dataset are cropped to a unified size and converted into grayscale images respectively; the positions and types of surface defects in each grayscale image are manually labeled to complete the preprocessing.
[0114] Then select the RT-DETR R18 model as the basic model and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR.
[0115] As Figure 2 shown, it is the structure diagram of the LMCD-RTDETR network proposed in this embodiment; its specific structure is introduced as follows:
[0116] 1) Aiming at the problems of fuzzy defect features and low contrast in multi-class defect detection on the steel surface, use CSPDarknet to replace the ResNet-18 network as the feature extraction backbone network; some defects on the steel surface are highly similar to the surface texture, or the defect boundaries are unclear, which makes the detection and classification tasks complex; due to the limited network depth and receptive field of ResNet-18, it is difficult to effectively distinguish such subtle differences; during multiple downsampling processes, key detail features are easily lost, resulting in indistinguishable fuzzy defect features and background textures; CSPDarknet enhances the extraction ability of fuzzy defect features through a cross-stage partial connection mechanism, and its structural design realizes the effective fusion of shallow texture features and deep semantic features, improving the network's perception ability of subtle texture differences; at the same time, the residual connection structure effectively alleviates the problem of gradient disappearance, ensuring the stability of the network when extracting fuzzy defect features; CSPDarknet enhances the model's expression ability while avoiding redundant calculations.
[0117] 2) To solve the ambiguity problems such as similar defect features and surface textures and unclear defect boundaries, a cascaded dynamic multi-scale feature extraction module (CSP-CDMSA) is designed, and the structure is as Figure 3As shown; this module integrates multi-scale feature extraction, dynamic weight fusion, and spatial-channel collaborative attention mechanism (SCSA) to specifically address the challenge of feature ambiguity in multi-category defects on the steel surface; the core of CSP-CDMSA is the CDMSA unit. Inspired by the design concepts of PartialConv and GhostNet, it adopts a cascaded partial convolution structure to extract multi-scale spatial information under different receptive fields. The formula is:
[0118] X split ={X1,X2}=Split(Conv 3×3 (X))
[0119] Y1={Y 1,1 ,Y 1,2}=Split(DWConv 5×5 (X1))
[0120] Y2=DWConv 7×7 (Y 1,1 )
[0121] Among them, DWConv represents depthwise separable convolution. By gradually expanding the convolution kernel size (3×3→5×5→7×7), it expands the receptive field and enhances the ability to capture fuzzy defect features at different scales; the hierarchical receptive field design enables the model to simultaneously focus on local texture details and global structural information, improving the ability to distinguish low-contrast targets; at the same time, the module adopts grouped convolution and channel splitting strategies to effectively avoid the processing overhead of redundant regions; the cascaded structure ensures that part of the original features are retained after each level of convolution operation, preventing fuzzy features from being overly abstracted and lost in the deep network;
[0122] The dynamic fusion weight mechanism of the CDMSA unit assigns adaptive weights to different-scale features through learnable parameters, enhancing the response intensity to fuzzy defect features; according to the feature differences of different types of fuzzy defects, it adaptively adjusts the importance of each scale feature, improving the expression ability for defects with unclear boundaries and low contrast; after the dynamically weighted multi-scale features are concatenated, 1×1 convolution is used for channel fusion and feature compression to enhance the interaction between different-scale information and effectively capture the discriminative features of fuzzy defects; after fusing the features, it is processed through the spatial-channel collaborative attention mechanism. The formula is:
[0123] F fused =Conv 1×1 (F concat )
[0124] A spatial (F)=σ(Norm(f spatial (F h ,F w )))
[0125] A channel (F) = σ(Norm(f channel (A spatial (F))))
[0126] F enhanced = A channel (F fused )·F fused
[0127] Among them, F h and F w respectively represent the feature statistics in the horizontal and vertical directions, f spatial represents the multi-scale spatial feature extraction function, f channel represents the channel feature extraction function based on multi-head self-attention, σ represents the activation function, and Norm represents the normalization operation;
[0128] The SCSA mechanism uses one-dimensional convolutions of different scales to model the feature map in the horizontal and vertical directions, and performs adaptive weighting according to the importance of spatial distribution, enhancing the model's perception ability of defect edges and shape changes; in the channel dimension, SCSA highlights discriminative feature channels through dimensionality reduction and multi-head attention mechanisms, thereby enhancing the response intensity to defects and suppressing background texture interference; finally, the features processed by SCSA are added to the original input features through residual connections, ensuring the retention of the original feature information and preventing the further loss of blurred features in the deep network;
[0129] The CSP-CDMSA module significantly improves the model's recognition ability for feature-blurred and low-contrast defects; compared with traditional feature extraction methods, this module can more accurately capture defect features similar to background textures and enhance the localization accuracy of defects with unclear boundaries; through hierarchical partial convolution and grouping strategies, it effectively reduces the calculation of irrelevant features and scales, maintaining a lightweight design while improving detection accuracy;
[0130] 3) The traditional AIFI module is based on the standard self-attention mechanism and mainly relies on the semantic similarity between features to achieve information interaction; however, this mechanism does not explicitly model the relative spatial relationship between features. In the multi-class object dispersion scenario, it is easy to ignore the spatial distribution structure of features, resulting in insufficient spatial localization ability, thereby increasing the detection difficulty; in the steel surface defect detection task, different types of defects have similar appearance textures and are relatively dispersed. The insufficient ability of the AIFI module to model the spatial dependence between classes exacerbates the risk of misjudgment and missed judgment;
[0131] To address these limitations, this embodiment proposes an AIFI-DPB module based on dynamic position bias (DPB); as Figure 4As shown, this module encodes the relative position information between feature points through a multi-layer perceptron, dynamically generates a spatial bias for adjusting attention calculation, and realizes the collaborative modeling of semantic similarity and spatial position information; DPB adjusts the attention score through the dynamically calculated relative position bias B ij The adjustment formula of the attention score is as follows:
[0132]
[0133] where Q i and K j represent the query and key vectors of the i-th and j-th features respectively, and B ij is the dynamically generated relative position bias; enabling the model to adjust the attention distribution weight according to the relative position between features, thus significantly improving the modeling ability of spatial relationships; by giving the relative position offset (Δx ij , Δy ij ) between feature points i and j, DPB calculates the relative position bias through a multi-layer perceptron structure:
[0134] B ij = DPB(Δx ij , Δy ij ) = MLP(LayerNorm(Δx ij , Δy ij ))
[0135] where (Δx ij , Δy ij ) represents the relative offsets in the horizontal and vertical directions. Through this dynamic calculation process, DPB can flexibly adjust the relative position bias to ensure that the model can effectively model spatial relationships under different input sizes and different target scales;
[0136] The AIFI-DPB module explicitly models the relative spatial relationship between features through position bias, enhancing the model's perception ability of targets with different position distributions; the position bias is dynamically generated by MLP, which can adapt to the shapes and scales of different targets and has good dynamic adaptability; by introducing spatial position information into self-attention calculation, while retaining semantic information, it enhances the modeling of spatial dependence relationships between features, enabling the model to consider both the semantic content and spatial arrangement of features; it can utilize the spatial distribution differences between defects to reduce the confusion of visually similar defects and improve the category discrimination ability; it promotes the effective integration of features at different scales through spatial relationship encoding and achieves good cross-scale fusion;
[0137] Compared with traditional attention mechanisms, AIFI-DPB significantly enhances the model's sensitivity to spatial structures while capturing local position dependencies and global semantic contexts; in multi-class steel defect recognition, it strengthens the discrimination ability between classes through spatially aware position biases; for example, creases and welds are similar in texture but different in spatial distribution, and AIFI-DPB can learn and utilize these differences to reduce inter-class confusion, accurately distinguish multiple types of defects with similar features, and avoid missed and false detections; this module not only enhances the spatial relationship modeling ability but also provides strong feature support for high-precision localization and multi-class recognition.
[0138] 4) The multi-class features of steel surface defects have the characteristics of large size variations and irregular shapes, which require the network to effectively detect defects of different scales; the existing cross-scale feature fusion module (CCFM) in RT-DETR uses top-down and bottom-up paths for feature fusion, relying on channel similarity calculation and direct concatenation (Concat) operations, resulting in an expansion of the feature channel dimension, significantly increasing the model parameters and computational cost; it is difficult to effectively coordinate the fusion of shallow details and deep semantic information, and when dealing with defects with large scale differences, the features of small targets are often suppressed or lost, thus affecting the detection accuracy; therefore, this method proposes a dynamic adaptive multi-scale feature pyramid network (DAMSFPN) structure, which combines multi-scale weighted fusion, efficient upsampling, multi-scale convolution modules, and a global heterogeneous kernel selection mechanism to construct a cross-scale feature fusion structure suitable for defect detection tasks of different sizes and irregular shapes in complex backgrounds.
[0139] 4.1) Multi-scale weighted fusion: Traditional multi-scale feature fusion strategies use the method of directly concatenating (Concat) feature maps of different scales, significantly increasing the dimension of the feature channels and causing a sharp rise in the number of parameters and computational overhead of the subsequent convolutional layers; to solve this problem, this method uses the design idea of BiFPN and adopts a weighted feature fusion mechanism to replace the traditional concatenation operation with a weighted summation operation, reducing the computational load, adaptively weighting and fusing features of different scales using learnable weight coefficients, dynamically adjusting the fusion ratio of each scale feature, and balancing the importance and information integrity of the features. The specific formula is as follows:
[0140]
[0141] where, F fusion is the output of the fusion layer; F i is the i-th feature map; ω iis the weight of the i-th feature map; N is the number of feature maps in the input fusion layer; this adaptive weighting mechanism dynamically adjusts its fusion ratio according to the information contribution degree of features at different scales, suppresses redundant features while ensuring information diversity, and provides a more refined and semantically rich feature input for subsequent detection;
[0142] 4.2) Efficient upsampling: In RT-DETR, the main purpose of upsampling is to increase the resolution of low-resolution feature maps to high resolution through upsampling operations. However, during the interpolation process, details will be lost, and the detailed structure in the image cannot be accurately restored. Blurring or distortion will be introduced during the magnification of the feature map, thus reducing the detection accuracy for small targets and complex features. DAMSFPN adopts EUCB efficient upsampling, which combines upsampling operations with depthwise separable convolutions to improve the feature expression ability while maintaining computational efficiency. This module doubles the resolution of the feature map through interpolation operations, and then uses 3×3 depthwise separable convolutions for feature refinement, effectively reducing the computational complexity and the number of parameters. The channel rearrangement operation realizes the information recombination and interaction of features between different channels, enhancing the expression of irregular textures and boundaries. Finally, a 1×1 pointwise convolution is used to adjust the channel dimension to ensure that it is consistent with the dimension of the feature map of the skip connection. The formula is as follows:
[0143]
[0144] Among them, is the output of EUCB efficient upsampling; X is the input of EUCB efficient upsampling; represents the composite operation of upsampling followed by depth convolution, C is the number of channels, and H and W are the height and width of the input feature map respectively; represents the channel rearrangement operation, g represents the number of groups for rearrangement, represents the pointwise convolution operation, and C′ represents the number of channels after pointwise convolution;
[0145] EUCB significantly improves the ability to restore detailed information and the fusion efficiency of multi-scale features, while effectively reducing the computational burden;
[0146] 4.3) Global Heterogeneous Kernel Selection Mechanism: There are significant differences in the sizes of steel surface defects, ranging from tiny dot-like defects to large-area regional damages. This requires the detection network to have good multi-scale perception and feature adaptation capabilities. Inspired by the heterogeneous convolution kernel design in YOLO-MS, the global heterogeneous kernel selection mechanism is systematically integrated into the DAMSFPN architecture, significantly enhancing the model's adaptability to diverse defects. By dynamically allocating the optimal convolution kernel set for feature layers with different resolutions, it adaptively matches the receptive field size required by the current feature layer to enhance the parsing ability of context information at different scales. Large-size convolution kernels can significantly expand the receptive field and enhance the modeling ability of long-distance context and structural information. Small-size convolution kernels, on the other hand, effectively maintain the response to edges and details, avoiding feature degradation caused by over-smoothing, and thus showing stronger sensitivity and resolution ability when capturing local features. It achieves a dynamic balance between global perception and local details, overcoming the deficiencies of traditional models in limited receptive field coverage, single information interaction, and insufficient feature expression. Under the synergistic effect of the multi-scale convolution module and feature fusion, this mechanism improves the model's semantic modeling and discrimination capabilities for defects in small targets, large targets, and complex backgrounds, enhancing the overall detection performance and generalization ability.
[0147] 4.4) CSP-DMSC: Traditional methods expand the receptive field by increasing the convolution kernel size or stacking convolution layers, but this significantly increases the computational cost and the number of parameters, making it difficult to meet the efficiency requirements in resource-constrained scenarios. To achieve multi-scale perception ability at low computational cost, this method integrates the core design ideas of MSCB and ESE to design an efficient and highly expressive multi-scale convolution module: CSP-DMSC (see Figure 5 ); This module aims to strengthen the perception ability and feature expression ability for defect targets at different scales, and the formula is as follows:
[0148]
[0149] Among them, is the output of the second-level module of CSP-DMSC; X is the input of the second-level module of CSP-DMSC; and respectively represent the first and second pointwise convolution operations. C and C′ are the original number of channels and the number of channels after the first pointwise convolution respectively, and H and W are the height and width of the input feature map; represents the depth convolution operation using a kernel size of k i , where k is the preset number of convolution kernel sizes; represents the effective channel attention operation, represents the channel rearrangement operation;
[0150] The CSP-DMSC module first expands the channels of the input features through 1×1 pointwise convolution, and parallelly introduces depthwise separable convolutions of multiple different sizes to model local details and global context respectively, achieving feature extraction under multi-scale receptive fields and constructing the multi-angle perception ability for irregular morphological features. Subsequently, it fuses feature representations of different scales to fully handle the size and morphological variations of defects. The ESE module generates channel attention maps through global average pooling and fully connected layers, dynamically adjusts the weight distribution of each channel, strengthens the key feature channels and suppresses redundant information, significantly improving the perception ability for low-contrast defects such as fine wire marks and targets with large scale variations. It rearranges the information flow among channels through channel shuffle operation, breaks the information islands between convolutional channels, and enhances the effective fusion of different morphological features in the channel dimension. It uses 1×1 convolution to map the features back to the original channel dimension, realizing the compact compression and expression consistency of semantic information. To prevent gradient vanishing and enhance the semantic consistency of features, the CSP-DMSC introduces a residual path to fuse the input features with the enhanced features, effectively retaining the original context information and strengthening the training stability of the network. Compared with traditional residual blocks, the CSP-DMSC module significantly improves the expression ability for multi-scales, effectively balances the deep semantic information and shallow detail information, ensures the accurate expression and localization of defects, and adapts to the variation differences of target scales and morphologies.
[0151] DAMSFPN significantly improves the performance of multi-scale and irregular defect detection through the collaborative optimization of multiple modules and mechanisms. While reducing the computational consumption, this network enhances the sensitivity of the network to irregularly shaped defects, overcomes the problem that morphological information is averaged or weakened in traditional methods, and demonstrates good detection accuracy and generalization ability.
[0152] After that, the LMCD-RTDETR is iteratively trained using the training set to obtain the trained LMCD-RTDETR.
[0153] Finally, the test set is input into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set.
[0154] To verify the effectiveness of LMCD-RTDETR for steel surface defect detection, this embodiment also conducts ablation and comparison experiments on the GC10-DET dataset, and generalization experiments on the PCB dataset and the NEU-DET dataset.
[0155] Specifically, the GC10-DET dataset was processed to remove mislabeled data. A total of 2,293 steel surface defect images were included in the experimental data, which were divided into 8:1:1, covering 10 types of defects: punching (Pu), welding line (Wl), crescent gap (Cg), water stain (Ws), oil stain (Os), wire stain (Ss), inclusion (In), rolling pit (Rp), crease (Cr), and waist fold (Wf).
[0156] The NEU-DET dataset contains 1,800 images, and the defect types include cracks, inclusions, spots, pitting, rolled-in scale, and scratches, also divided into 8:1:1.
[0157] The DeepPCB dataset contains 1,500 PCB defect images, with 6 types of defects: open circuit, short circuit, mouse bite, thorn, false copper, and pinhole. The dataset is also divided into 8:1:1.
[0158] The evaluation metrics used in the experiment include Precision (P), Recall (R), mean Average Precision (mAP), number of parameters (Params), and GFLOPs, etc., to comprehensively evaluate the algorithm. The relevant formulas are as follows:
[0159]
[0160] Among them, TP represents the number of true positives, FP represents the number of false positives, and FN represents the number of false negatives.
[0161] The experimental environment is shown in Table 1:
[0162] Table 1 Hardware and software configurations of the experimental environment
[0163] Environmental Component Parameter Configuration GPU RTX 3090 (24GB) CPU 15vCPU Intel(R)Xeon(R)Platinum 8358P Programming Language Python 3.8 Deep Learning Framework PyTorch 1.11.0 IDE Pycharm
[0164] Result analysis:
[0165] 1) Ablation experiments on the GC10-DET dataset are shown in Table 2. The precision of the baseline model RT-DETR R18 is 69.3%, the recall is 67%, the mAP is 65.3%, the number of parameters is 19.88M, and the computational cost is 57 GFLOPs. After adding CSP-CDMSA, the mAP is increased to 65.9%. At the same time, the number of parameters is reduced to 14.12M, and the computational cost is decreased to 47.7 GFLOPs, which can effectively capture detailed features and reduce the number of parameters and complexity. Adding the AIFI-DPB module increases the precision to 72.2% and the mAP reaches 66.3%. Although the recall decreases slightly, the detection ability is improved. After adding DAMSFPN, the precision reaches 72.8%, the recall is 67.2%, the mAP is increased to 66.5%, the number of parameters is 18.24M, and the computational cost is 48.7 GFLOPs, demonstrating the effectiveness of this module for feature fusion. The combination of AIFI-DPB and DAMSFPN further improves the performance, with the precision jumping to 76.2%, the recall increasing to 68.9%, and the mAP reaching 67.5%, while maintaining a relatively low number of parameters (18.24M) and computational cost (48.9 GFLOPs). Finally, the proposed LMCD-RTDETR integrates three key modules to achieve the optimal performance: precision of 77.2%, recall of 67.9%, and mAP of 68.1%, which is 2.8 percentage points higher than the baseline model. At the same time, the number of parameters is only 12.52M, and the computational cost is 40.3 GFLOPs, demonstrating a perfect balance between performance and efficiency;
[0166] Table 2 Results of ablation experiments
[0167]
[0168] 2) Comparative experiments on the DAMSFPN structure; To verify the effectiveness of the DAMSFPN module, in this embodiment, it is compared with other feature fusion networks (MAFPN, BIMAFPN (BIFPN + MAFPN), BIFPN, and Slimneck); As shown in Table 3, the accuracy of DAMSFPN is 72.8%, higher than that of other feature fusion networks, indicating that it can more accurately and effectively distinguish targets from the background in complex backgrounds; The recall rate is 67.2%, the highest, significantly better than BIMAFPN (64.1%) and BIFPN (65.1%), and far exceeding SlimNeck (61.4%), indicating that it reduces the phenomenon of missed detections; The mAP of DAMSFPN is 66.5%, performing best among all the comparison models, further verifying its advantages in detection performance; In terms of model complexity, the number of parameters of DAMSFPN is 18.24M, the smallest among all the comparison models, and the computational complexity is 48.7 GFLOPs, significantly lower than other networks; This makes DAMSFPN more advantageous in resource-constrained scenarios; Generally speaking, DAMSFPN is a better choice for feature fusion networks, enhancing the detection ability for multi-scale and irregular defects in complex backgrounds;
[0169] Table 3 Comparative experiments on the DAMSFPN structure
[0170] Model P / % R / % Params / M GFLOPs / G mAP / % MAFPN 72.6 63 22.94 56.4 64.8 BIFPN 72.4 65.1 20.31 64.3 66.1 BIMAFPN 72.2 64.1 20.12 57.5 65 SlimNeck 72.7 61.4 19.31 53.3 64.3 DAMSFPN 72.8 67.2 18.24 48.7 66.5
[0171] 3) Comparative experiments on the AIFI-DPB module; To verify the effectiveness of AIFI-DPB, ablation experiments are conducted with other modules (AIFI-HiLO, AIFI-LPE, AIFI-EAA), as shown in Table 4; The results show that AIFI-DPB performs better in terms of accuracy, recall rate, and mAP; Specifically, the accuracy of AIFI-DPB is 72.2%, higher than that of AIFI-HiLO (71.9%), AIFI-EAA (72%), and AIFI-LPE (71.4%), showing a more reliable feature recognition ability; In terms of the recall rate, AIFI-DPB leads other modules with a performance of 65.5%, especially showing a significant improvement compared to AIFI-EAA (59.9%), enhancing the ability to capture targets in complex scenarios and reducing missed detections; The mAP of AIFI-DPB reaches 66.3%, better than other modules; In terms of computing resources, the number of parameters of AIFI-DPB is 19.89M, and the computational complexity is 57.2G, comparable to other modules, indicating that its performance improvement does not rely on increasing complexity. By dynamically adjusting the relative position bias to optimize feature interaction, it improves the classification ability and discrimination, enhances the detection accuracy and stability, and reduces the false detection rate and missed detection rate;
[0172] Table 5 Comparative experiments on the AIFI-DPB module
[0173] Model P / % R / % Params / M GFLOPs / G mAP / % AIFI-HiLO 71.9 65.1 19.85 57.1 65.9 AIFI-LPE 71.4 63.3 19.99 57 64.1 AIFI-EAA 72.0 59.9 19.88 57.2 63.5 AIFI-DPB 72.2 65.5 19.89 57.2 66.3
[0174] 4) To verify the effectiveness of the proposed algorithm, in this embodiment, LMCD-RTDETR is also compared with other models on the GC10-DET dataset; the experimental results are shown in Table 5, and LMCD-RTDETR is significantly superior to other models in terms of comprehensive performance; first of all, the mAP of LMCD-RTDETR reaches 68.1%, which is the highest among all models, and is 2.8% higher than that of RT-DETR R18; in terms of computational efficiency, the GFLOPs of LMCD-RTDETR is 40.3G, which is the lowest among all models, significantly reducing the computational amount and maintaining high precision; compared with RT-DETR R18 (57 GFLOPs), it is reduced by 29%, and the precision-close YOLOV8m is 78.7G, and RTDETR-r50 is 129.6G, indicating that LMCD-RTDETR has an obvious advantage in computational efficiency; the Precision and Recall of LMCD-RTDETR are 77.2% and 67.9% respectively, which are the highest among all models, indicating that it can significantly reduce false detections and missed detections and provide more accurate object recognition; in summary, LMCD-RTDETR has a better balance between precision and computational efficiency;
[0175] Table 5 Comparative experiments with other models
[0176]
[0177] 5) To verify the effectiveness of the proposed model, this embodiment conducted generalization experiments on two datasets, NEU-DET and DeepPCB, verifying the effectiveness and generalization of the model design, as shown in Table 6; on the NEU-DET dataset, the precision of LMCD-RTDETR was 71.7%, and the recall rate was 71.9%, both higher than those of RT-DETR R18 (71.5%, 70.5%); the mAP reached 74.6%, an increase of 1.1% compared to RT-DETR R18; the performance improvement indicates that LMCD-RTDETR has a stronger ability to capture features in complex backgrounds; on the DeepPCB dataset, the precision of LMCD-RTDETR was 98.9%, and the recall rate was 96.9%, slightly higher than 98.6% and 96.5% of RT-DETR R18 respectively; its mAP was 98.3%, slightly higher than that of RT-DETR R18; although the improvement was small, based on high-precision detection, the number of parameters decreased significantly, maintaining light weight and high performance; in summary, LMCD-RTDETR performed excellently on these two datasets, achieving an overall performance improvement compared to RT-DETR R18 and excellent performance on two different types of datasets, verifying its good generalization ability;
[0178] Table 6 Generalization Experiment
[0179]
[0180]
[0181] 6) Visual comparison: As Figure 6 shown, Figure 6 in A is the original image, B is the detection result using the RT-DETR R18 algorithm, and C is the detection result using the LMCD-RTDETR algorithm; as Figure 6 shown, LMCD-RTDETR is superior to the RT-DETR R18 algorithm in both confidence distribution and heatmap performance, showing higher confidence scores and a more concentrated and accurate heat distribution; compared with the baseline model, LMCD-RTDETR shows stronger detection ability in the detection tasks of low contrast, strong target similarity, and small target defects; it can be observed from the figure that the creases are similar to the welds, and this algorithm reduces its false detection rate and the missed detection rate of defects such as oil spots and rolling pits, and at the same time has a significant improvement in the recognition effect of low-contrast and irregular defects such as wire spots, foreign objects, and waist creases; this performance improvement benefits from the multi-module collaborative optimization mechanism of LMCD-RTDETR, effectively enhancing the model's multi-scale and irregular feature recognition ability and similarity category discrimination ability, and showing stronger detection efficiency under multi-category target detection and complex background interference conditions;
[0182] This embodiment proposes a lightweight defect detection algorithm based on RT-DETR R18 - LMCD-RTDETR, aiming to address challenges such as multi-category, shape diversity, and size differences of defects in complex backgrounds. Among them, the CSP-CDMSA module enhances the extraction and recognition ability of multi-category fuzzy defect features through a lightweight cascaded multi-scale receptive field structure. The AIFI-DPB module uses dynamic position offsets to improve the discrimination ability and spatial positioning accuracy of visually similar defects. A DAMSFPN feature fusion mechanism is constructed to achieve adaptive weight assignment and optimized integration of multi-scale features, effectively balancing deep semantic and shallow detail information, and effectively improving the detection ability of the model under different defect morphologies and scales. Experimental results show that LMCD-RTDETR significantly outperforms the baseline model and other object detection methods in terms of detection accuracy on multiple datasets, and also shows obvious advantages in computational complexity and the number of parameters, demonstrating good generalization ability. This method not only makes a breakthrough in detection accuracy, but also realizes lightweight design, effectively balancing computational efficiency and detection accuracy, and ensuring the accurate recognition and positioning of diverse defects.
[0183] The same or similar reference numerals correspond to the same or similar components.
[0184] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this application.
[0185] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR, characterized in that, It includes the following steps: S1: Obtain the steel surface defect dataset and perform preprocessing. Divide the preprocessed steel surface defect dataset into a training set and a test set; the steel surface defect dataset includes several steel images with different surface defects respectively; S2: Select the RT-DETR R18 model as the basic model and improve the RT-DETR R18 model to construct a lightweight steel surface defect detection model LMCD-RTDETR; The LMCD-RTDETR includes a backbone network, an encoder, a decoder, and a detection head connected in sequence; among them, the CSPDarknet network is used to replace the ResNet-18 network in the RT-DETR R18 model as the backbone network, and the CSP-CDMSA module is introduced into the backbone network; the encoder includes an AIFI-DPB module and a DAMSFPN feature fusion module connected in sequence; S3: Use the training set to iteratively train LMCD-RTDETR to obtain the trained LMCD-RTDETR; S4: Input the test set into the trained LMCD-RTDETR for defect detection to obtain the defect detection results of each steel image in the test set.
2. A lightweight steel surface multi-class defect detection method based on improved RT-DETR according to claim 1, characterized in that In step S1, the preprocessing includes: cropping all the steel images in the steel surface defect dataset to a unified size and converting them into grayscale images respectively; Manually annotate the positions and types of surface defects in each grayscale image to complete the preprocessing.
3. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 1, characterized in that In step S2, the backbone network includes, connected in sequence: a first convolutional layer, a second convolutional layer, a first CSP-CDMSA module, a third convolutional layer, a second CSP-CDMSA module, a fourth convolutional layer, a third CSP-CDMSA module, a fifth convolutional layer, and a fourth CSP-CDMSA module; The first CSP-CDMSA module, the second CSP-CDMSA module, and the third CSP-CDMSA module respectively output feature maps P2, P3, and P4 to the DAMSFPN feature fusion module; The output of the fourth CSP-CDMSA module is connected to the AIFI-DPB module, and the AIFI-DPB module outputs feature map P5 to the DAMSFPN feature fusion module; The DAMSFPN feature fusion module fuses the feature maps P2 to P5 and then inputs them into the decoder.
4. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 3, characterized in that Each CSP-CDMSA module has the same structure and includes, connected in sequence: a sixth convolutional layer, a segmentation layer, several CDMSA secondary modules, a first splicing layer, and a sixth convolutional layer; the outputs of the segmentation layer and each CDMSA secondary module are also respectively connected to the first splicing layer; The structure of each CDMSA secondary module includes, connected in sequence: a 3×3 convolutional layer, a 5×5 depthwise separable convolutional layer, a 7×7 depthwise separable convolutional layer, a second splicing layer, a 1×1 convolutional layer, and an SCSA spatial and channel collaborative attention tertiary module; The outputs of the 3×3 convolutional layer and the 5×5 convolutional layer are also respectively connected to the second splicing layer; the input of the 3×3 convolutional layer is also connected to the output of the SCSA spatio-channel co-attention three-level module in a residual addition connection.
5. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 4, characterized in that In the step S2, the structure of the AIFI-DPB module includes, connected in sequence: a seventh convolutional layer, a DPB attention secondary module, a first normalization layer, an FFN layer, a second normalization layer, and an eighth convolutional layer; the input of the DPB attention secondary module and the input of the first normalization layer also form a residual addition connection; the input of the FFN layer and the input of the second normalization layer also form a residual addition connection; The DPB attention secondary module adjusts the attention score by dynamically calculating the relative position bias B. ij The attention mechanism of the DPB attention secondary module is expressed as: Among them, Attention(Q i , K j , V j ) represents the attention score of the i-th and j-th features; Q i represents the query matrix of the i-th feature; K j represents the key matrix of the j-th feature; V j represents the value matrix of the j-th feature; Softmax represents the Softmax activation function; d k represents the dimension of the key matrix K j ; (Δx ij , Δy ij ) represents the relative offsets in the horizontal and vertical directions; MLP represents a multi-layer perceptron; LayerNorm represents layer normalization; The output of the fourth CSP-CDMSA module is connected to the seventh convolutional layer; the eighth convolutional layer outputs the feature map P5 to the DAMSFPN feature fusion module.
6. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 5, characterized in that, In the step S2, the structure of the DAMSFPN feature fusion module includes, arranged in parallel: a first branch, a second branch, and a third branch; each branch includes, connected in sequence: a ninth convolutional layer, a first fusion layer, a first CSP-DMSC secondary module, a second fusion layer, and a second CSP-DMSC secondary module; The feature maps P3, P4, and P5 are respectively input into the ninth convolutional layers of the first branch, the second branch, and the third branch; The downsampled feature map P2 is input into the first fusion layer of the first branch; The output of the ninth convolutional layer of the first branch is downsampled and then input into the first fusion layer of the second branch; The output of the ninth convolutional layer of the second branch is downsampled and then input into the first fusion layer of the third branch; The outputs of each CSP-DMSC secondary module of the first branch are respectively downsampled and then input into the second fusion layer of the second branch; The outputs of each CSP-DMSC secondary module of the second branch are respectively downsampled and then input into the second fusion layer of the third branch; The output of the first CSP-DMSC secondary module of the second branch is subjected to EUCB efficient upsampling and then input into each fusion layer of the first branch; The output of the first CSP-DMSC secondary module of the third branch is subjected to EUCB efficient upsampling and then input into each fusion layer of the second branch; The outputs of the second CSP-DMSC secondary modules of each branch are jointly input into the decoder.
7. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 6, characterized in that, The calculation formula of each fusion layer is expressed as: Among them, F fusion is the output of the fusion layer; F i is the feature map of the i-th one; ω i is the weight of the i-th feature map; N is the number of feature maps input to the fusion layer.
8. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 6, characterized in that, The formula of the EUCB efficient upsampling is: Among them, is the output of the EUCB efficient upsampling; X is the input of the EUCB efficient upsampling; represents a composite operation of depth convolution following upsampling, where C is the number of channels, and H and W are the height and width of the input feature map respectively; represents a channel rearrangement operation, and g represents the number of groups for rearrangement, represents a pointwise convolution operation, and C′ represents the number of channels after pointwise convolution.
9. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to claim 6, characterized in that, The calculation formula of each CSP-DMSC secondary module is: Among them, is the output of the CSP-DMSC secondary module; X is the input of the CSP-DMSC secondary module; and respectively represent the first and second pointwise convolution operations, C and C′ are the original number of channels and the number of channels after the first pointwise convolution respectively, and H and W are the height and width of the input feature map respectively; represents a depth convolution operation with a kernel size of k i , where k is the number of preset convolution kernel sizes; represents an effective channel attention operation, represents a channel rearrangement operation.
10. A lightweight multi-class defect detection method for steel surface based on improved RT-DETR according to any one of claims 1 to 9, characterized in that, In the step S4, it further includes: Using a preset evaluation metric to evaluate the performance of the trained LMCD-RTDETR; the evaluation metric includes at least any one or more of: precision P, recall rate R, mean average precision mAP, number of parameters Params, and GFLOPs.
Citation Information
Cited By
X-ray weld defect detection method based on structure degradation correction and parallel response recombination
CN121981993A