Automobile paint surface damage detection method based on RT-DETR improvement
By using the improved end-to-end Transformer architecture and multi-scale feature enhancement strategy of RT-DETR, the problems of missed detection and loss of details in the existing technology are solved, and high-precision and robust automotive paint damage detection is achieved.
Patent Information
- Application Number
- CN202510910174.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-21
AI Technical Summary
Existing automotive paint damage detection methods rely on non-end-to-end architectures, leading to missed detections/false detections. Traditional CNN structures are insufficient in modeling global semantic relationships, resulting in the loss of detailed information and making it difficult to detect complex and subtle damage.
An end-to-end Transformer architecture based on RT-DETR is adopted, which combines a global attention mechanism and a multi-scale feature enhancement strategy. Through the dynamic convolutional hybrid module DCMB and the full-scale frequency attention structure CSPO, long-range contextual association and preservation of fine damage features in the damaged region are achieved.
It significantly improves the accuracy and robustness of complex paint surface damage detection, enhances the ability to detect fine-grained damage areas, and balances model lightweighting and efficient inference performance.
Smart Images

Figure CN120823167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of paint surface detection, and in particular to an improved automobile paint surface damage detection method based on RT-DETR. Background Art
[0002] Automotive paint damage detection is a critical component of vehicle maintenance and insurance claims. Paint damage types primarily include scratches, dents, peeling, and rust. Rapidly and accurately identifying damage types is not only fundamental for assessing repair costs but also a key requirement for improving automated inspection efficiency.
[0003] Computer vision-based inspection technology is increasingly being used in automotive surface inspections due to its non-contact, high-efficiency, and low-cost advantages. Current mainstream methods primarily employ CNN-based object detection algorithms (such as the YOLO series) to achieve damage location and classification. While some progress has been made, significant drawbacks remain:
[0004] First, non-end-to-end architectures rely on post-processing. Detectors such as YOLO require the introduction of post-processing operations such as non-maximum suppression (NMS). Manually tuning the threshold can easily lead to missed detections or false detections, especially for subtle damage to the paint surface (such as hairline scratches). Secondly, the dominance of local features will lead to loss of details. The pure CNN structure focuses too much on local textures and has difficulty in modeling the global semantic association of paint damage (such as the continuity characteristics of large-area scratches). In addition, the multi-level downsampling process will weaken the feature expression of minor damage. Summary of the Invention
[0005] Purpose of the invention: To address the problems of missed detection / false detection caused by existing automobile paint damage detection methods due to their reliance on non-end-to-end architecture, as well as the defects of traditional CNN structures in insufficient modeling of global semantic associations and loss of detail information, the present invention proposes an improved paint damage detection method based on RT-DETR. By using an end-to-end Transformer architecture to eliminate the dependency on post-processing parameters, the global attention mechanism is used to model the long-range contextual association of the damaged area, and the multi-scale feature enhancement strategy is combined to retain the structural information of fine damage (such as hairline scratches and pitting rust), ultimately achieving high-precision and robust detection of complex paint damage.
[0006] Technical solution: The present invention proposes an improved automobile paint damage detection method based on RT-DETR, comprising the following steps:
[0007] Step 1: Construct a car paint damage dataset and add negative samples and interference samples, then perform image enhancement and preprocessing;
[0008] Step 2: Based on the multi-layer convolution backbone network of ResNet18, a dynamic convolutional mixing module (DCMB) is introduced to extract multi-scale features. The DCMB includes a parallel dynamic Inception mixer and a convolutional gating unit (GLU). The dynamic Inception mixer implements dynamic multi-scale feature mixing, and the convolutional gating unit (GLU) enhances channel modeling capabilities.
[0009] Step 3: Add a full-scale frequency attention structure CSPO to the neck network Neck to extract cross-domain structural information. The full-scale frequency attention structure CSPO includes an SPDConv module and a CSP-OmniKernel feature integration module, which receives shallow features from the backbone network Backbone and extracts multi-scale information.
[0010] Step 4: The fused feature map is input to the Transformer decoder, and the query vector output by the decoder is input to the prediction head, which generates the bounding box coordinates and category probabilities.
[0011] Furthermore, the specific method of step 1 is:
[0012] Images of damaged car paint surfaces were collected and manually classified and labeled with bounding boxes. Then, three other types of interference data were added: negative samples of intact vehicles, samples of interference from other object surface textures, and images of paint reflection artifacts. In the data augmentation stage, a combination of multiple strategies was used, including basic spatial transformation, random rotation, scaling, cropping, brightness and contrast adjustment, and local noise injection. Finally, the training, validation, and test sets were divided proportionally.
[0013] Furthermore, the backbone network Backbone includes a 9-layer structure, namely 5 layers of Conv modules and 4 layers of dynamic convolution mixing modules DCMB. The input features pass through the first layer of Conv module to obtain P1 layer features, and then are sequentially input to the second layer of Conv module and the first layer of dynamic convolution mixing module DCMB to obtain P2 layer features, and then sequentially input to the third layer of Conv module and the second layer of dynamic convolution mixing module DCMB to obtain P3 layer features, and sequentially input to the fourth layer of Conv module and the third layer of dynamic convolution mixing module DCMB to obtain P4 layer features, and finally input to the fifth layer of Conv module and the fourth layer of dynamic convolution mixing module DCMB to obtain P5 layer features.
[0014] Furthermore, the dynamic convolutional mixing module DCMB is a combination module of the residual structure, including a dynamic Inception mixer and a convolutional gating unit GLU; the input features are first batch normalized BatchNorm and then passed through the dynamic Inception mixer and convolutional gating unit GLU respectively to achieve dynamic multi-scale feature mixing and channel modeling capability enhancement. The outputs of the two branches are respectively subjected to LayerScale scaling and DropPath random depth processing, and then returned to the trunk part to be additively fused with the original input through residual connection, and finally output a multi-scale feature tensor.
[0015] Furthermore, the dynamic Inception mixer uses two dynamic Inception depthwise separable convolutions to form paths of different scales, achieving multi-scale modeling in the channel dimension. First, the input feature map is divided into two groups according to the number of channels. Each group is fed into a dynamic Inception depthwise separable convolution with a different convolution kernel size to extract features and capture spatial information at different scales. Subsequently, all processing results are spliced in the channel dimension and fused through 1×1 convolution.
[0016] Furthermore, the dynamic Inception deep separable convolution extracts channel-level global context by performing global average pooling on the input features, and then uses 1×1 convolution to generate weight coefficients corresponding to the three branches, which are used as dynamic attention guidance after softmax normalization; the input features are processed by square kernel convolution, horizontal strip convolution, and vertical strip convolution, and the outputs of the three convolution branches are weightedly fused according to the weight coefficients corresponding to the three branches of dynamic attention guidance, so as to adaptively adjust the expression strength of the receptive field and directional features for different inputs; the fusion result is then processed by batch normalization and SiLU activation function, and the output is a dynamic enhanced feature that contains both local texture and rich structural semantics.
[0017] Furthermore, the specific structure of the full-scale frequency attention structure CSPO in step 3 is:
[0018] The P2 layer features of the backbone network are processed by SPDConv, which preserves pixel-level details through space-to-depth transformation. The features rich in small object information are then fused with the P3 layer features. The fused features are then processed by the CSP-OmniKernel feature integration module.
[0019] The output feature F1 of the CSP-OmniKernel feature integration module is processed by the first RepC3 module and the Conv module to obtain the feature F2;
[0020] The backbone network P4 layer features are processed by the Conv module and the P5 layer features are fused after being processed by AIFI and upsample. After fusion, they are processed by the second RepC3 module and the Conv module to obtain feature F3.
[0021] After the features F2 and F3 are fused, they are processed by the third RepC3 module and the Conv module and then fused with the P5 layer features. After fusion, they are passed through the fourth RepC3 module.
[0022] The features output by the first RepC3 module, the third RepC3 module, and the fourth RepC3 module are input to the decoder for decoding processing.
[0023] Furthermore, the specific architecture of the CSP-OmniKernel feature integration module is as follows:
[0024] After channel mapping through Conv 1×1, the feature channel is split into two sub-branches in a ratio of 1:3: one is the main OmniKernel processing branch, which is 1 / 4 channel, and the other is the residual path;
[0025] The main OmniKernel processing branch first performs a 1×1 convolution transformation on the input and activates it. It then simultaneously inputs three parallel branches: 1) The local branch Local uses 1×1 depthwise separable convolution to focus on fine-grained textures; 2) The large-scale branch Large extracts long and short domain texture information through depthwise convolution in three directions: horizontal, vertical, square, and point convolution; 3) The global branch Global achieves a global receptive field through dual-domain channel attention (DCAM) and frequency-gated mechanism FSAM.
[0026] After fusion, the three parallel branches are processed by Conv 1×1 and then fused with the residual path of 3 / 4 channels and processed by Conv 1×1 to output the final features.
[0027] Furthermore, the specific method of step 4 is:
[0028] The fused multi-scale feature map is first input into the Transformer decoder, which performs cross-attention interaction with the feature map through a learnable ObjectQuerie. During the decoding process, each query vector dynamically focuses on a specific target area in the feature map, outputting a set of attention-enhanced instantiated feature vectors. The feature vectors are input in parallel into the prediction head to generate bounding box coordinates and category probabilities.
[0029] Beneficial effects:
[0030] The above-mentioned improved automobile paint damage detection method based on RT-DETR can significantly improve the detection accuracy of complex and fine-grained damage areas while maintaining the lightweight model and efficient inference performance.
[0031] 1. This paper introduces a variety of interferences and negative samples to construct a diversified training set, thereby enhancing the robustness of the model in actual scenarios.
[0032] 2. The dynamic convolutional mixing module designed in this invention uses two branches to achieve adaptive modeling of features in different directions and scales, effectively integrating local texture and long-range dependency information, thereby improving feature expression capabilities; the dynamic Inception mixer performs grouping and multi-scale processing in the channel dimension, further enhancing the model's ability to perceive small damaged areas in complex backgrounds; the dynamic Inception deep separable convolution introduces a dynamic attention mechanism based on global context, improving the model's response flexibility and discriminability to different spatial structures.
[0033] 3. The full-scale frequency attention structure proposed in this invention effectively integrates high-resolution small target features through SPDConv and CSP-OmniKernel modules while keeping the amount of computation under control, significantly enhancing the detection ability of edge details and minor damage.
[0034] The overall technical solution takes into account detection accuracy, computational efficiency and actual deployment requirements, providing a high-performance, highly robust intelligent solution for automotive paint quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is the overall flow chart of the improved automobile paint damage detection method based on RT-DETR;
[0036] Figure 2 This is the improved backbone network structure diagram;
[0037] Figure 3 This is the structure diagram of the dynamic convolution hybrid module (DCMB);
[0038] Figure 4 This is the structure diagram of the dynamic Inception mixer;
[0039] Figure 5 This is a diagram of the dynamic Inception depth separable convolution structure;
[0040] Figure 6 This is the structure diagram of the full-scale frequency attention (CSPO);
[0041] Figure 7 This is the CSP-OmniKernel structure diagram;
[0042] Figure 8 This is the DCAM structure diagram in CSP-OmniKernel;
[0043] Figure 9 This is the FSAM structure diagram in CSP-OmniKernel. DETAILED DESCRIPTION
[0044] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0045] The present invention discloses a paint surface damage detection method based on RT-DETR improvement, see Figure 1 The process specifically includes the following steps:
[0046] Step 1: Images of eight common types of automotive paint damage, including scratches, dents, flaking, sagging, orange peel, rust, cracking, and fading, were collected. Three types of interfering data were then introduced: intact vehicle images as negative samples; images of other surface textures, such as walls and rocks; and vehicle body images with strong reflective or mirror artifacts as artifact interfering samples. The data augmentation stage employed a combination of strategies, including random horizontal flipping, affine transformation, random rotation (±15°), scaling (0.8–1.2x), center and random cropping, brightness / contrast adjustment (range ±30%), and Gaussian and salt-and-pepper noise injection. Two images were augmented for each image, and the data was ultimately split into training, validation, and test sets in an 8:1:1 ratio for model training and evaluation.
[0047] Step 2: Introduce the dynamic convolutional mixing module DCMB into the backbone network based on ResNet18 to extract features; the dynamic convolutional mixing module DCMB includes a parallel dynamic Inception mixer and a convolutional gating unit GLU. The dynamic Inception mixer implements dynamic multi-scale feature mixing, and the convolutional gating unit GLU enhances channel modeling capabilities. Figure 2 .
[0048] After the image is input into the backbone network, it passes through the first convolutional layer, where a 3×3 convolution with a stride of 2 downsamples the image size from 640×640 to 320×320. At the same time, the number of channels is increased to 64, which is used to extract low-level features such as edges and textures. It then passes through another convolutional layer with the same settings, further compressing the spatial size to 160×160 and increasing the number of channels to 128. It then enters the first dynamic convolutional hybrid module (DCMB). Figure 3In this module, features are first fed into a normalization layer (BatchNorm) and then split into two subpaths. The main path passes through a dynamic Inception mixer, which uses three depthwise separable convolutions of different sizes to model the features at multiple scales. It also uses a dynamic weighting mechanism to adaptively assign importance to the convolution outputs based on the input, simulating an attention selection mechanism. The output of the main path is then multiplied by a learnable LayerScale parameter and, after passing through a DropPath, residual summed with the input features. The bypass path further normalizes the main path output before feeding it into a convolutional gated unit (GLU), introducing nonlinearity and inter-channel dependencies to enhance feature representation. It is then also multiplied by an independent LayerScale weight and, after passing through a DropPath, residual summed again to form the final output. After the first dynamic convolutional mixer (DCMB) completes, it is fed into a convolutional layer for further downsampling (to 80×80). The above structure is then repeated before entering the next dynamic convolutional mixer (DCMB). The process is almost identical, but with different channels. The entire network continuously downsamples and extracts higher-level features by alternating between Conv modules and DCMB modules until entering the last dynamic convolutional hybrid module DCMB, ultimately obtaining a deep semantic feature map with a spatial resolution of 20×20 and 384 channels as the output of the entire backbone network.
[0049] In the overall network, the Dynamic Convolutional Mixing Module (DCMB) is repeatedly stacked across multiple resolution layers, from P2 to P5, downsampling and increasing the number of channels layer by layer. In particular, the Dynamic Convolutional Mixing Module (DCMB) is stacked three times at the P5 layer to fully extract deep semantic features and provide the detection head with rich context and multi-scale representations. By combining dynamic convolutional fusion, a multi-branch structure, gated activation, and a stable residual mechanism, this architecture effectively improves the model's ability to perceive fine-grained defects and complex texture interference in tasks such as paint scratch detection.
[0050] See also Figure 4 and Figure 5The Dynamic Inception Mixer uses two Dynamic Inception Depthwise Separable Convolutions (DSCs) to form paths of different scales, achieving multi-scale modeling in the channel dimension. Specifically, the input feature map is first divided into two groups based on the number of channels. Each group is fed into a Dynamic Inception Depthwise Separable Convolution with a different kernel size to extract features, capturing spatial information at different scales. All processed results are then concatenated in the channel dimension and fused via a 1×1 convolution. The Dynamic Inception Depthwise Separable Convolution extracts channel-level global context by performing global average pooling on the input features. A 1×1 convolution is then used to generate weight coefficients for the three branches, which are then normalized using softmax to guide dynamic attention. The input features are processed through square kernel convolution, horizontal strip convolution, and vertical strip convolution. The outputs of the three convolution branches are weighted and fused according to the weight coefficients corresponding to the three branches of the dynamic attention guidance. This allows for adaptive adjustment of the receptive field and directional feature representation for different inputs. The fused results are then batch normalized and processed with the SiLU activation function, resulting in dynamic enhanced features that contain both local texture and rich structural semantics.
[0051] Step 3: Add a full-scale frequency attention structure CSPO to the neck network Neck to extract cross-domain structural information. The full-scale frequency attention structure CSPO includes an SPDConv module and a CSP-OmniKernel feature integration module, which receives shallow features of the backbone network Backbone and extracts multi-scale information.
[0052] For detailed structure of the neck network, see Figure 6 The P2 layer features of the backbone network Backbone are processed by SPDConv, and the pixel-level details are retained through space-to-depth transformation. Then the features rich in small target information are fused with the P3 layer to obtain the fused multi-scale low-level features, which are then sent to the CSP0mniKernel module to further extract cross-domain structural information.
[0053] The output feature F1 of the CSP-OmniKernel feature integration module is processed by the first RepC3 module and the Conv module to obtain the feature F2.
[0054] The P4 layer features of the backbone network Backbone are processed by the Conv module and the P5 layer features are fused after upsample processing. After fusion, they are processed by the second RepC3 module and the Conv module to obtain feature F3.
[0055] After the features F2 and F3 are fused, they are processed by the third RepC3 module and the Conv module and then fused with the P5 layer features. After fusion, they are passed through the fourth RepC3 module.
[0056] The features output by the first RepC3 module, the third RepC3 module, and the fourth RepC3 module are input to the decoder for decoding processing.
[0057] CSP-OmniKernel feature integration module see Figure 7 , combining frequency-domain attention (FCA), spatial attention (SCA), and multi-modal direction-aware convolution. CSP-OmniKernel first splits the input into two sub-branches through Conv 1×1: one is the main OmniKernel processing branch (covering 1 / 4 channels), and the other is the residual path (covering 3 / 4 channels). The main OmniKernel processing branch first performs a 1×1 convolution on the input and activates it. This is then fed into three parallel branches: 1) the local branch (Local) uses 1×1 depthwise separable convolution to focus on fine-grained textures; 2) the large-scale branch (Large) extracts long- and short-domain texture information through depthwise convolution in three directions: horizontal, vertical, square, and point convolution; and 3) the global branch (Global) achieves a global receptive field through dual-domain channel attention (DCAM) and frequency gated mechanism (FSAM). After fusion, the three parallel branches are processed through Conv 1×1, fused with the residual path of 3 / 4 channels, and then processed through Conv 1×1 to output the final features.
[0058] For the global branch, see Figure 8 and Figure 9 , DCAM (Dual-Domain Channel Attention Module) and FSAM (Frequency-Domain-Based Spatial Attention Module). DCAM first extracts frequency-domain channel attention information through the FCA module. This involves performing global average pooling on the input feature map and generating channel attention weights through 1×1 convolution. Subsequently, the original feature map is converted to the frequency domain via FFT (Fast Fourier Transform), multiplied channel-by-channel with the weights, and then returned to the spatial domain via IFFT (Inverse Fast Fourier Transform), resulting in a frequency-domain enhanced feature representation. The SCA module then further extracts channel attention in the spatial domain. Global average pooling and 1×1 convolution are performed on the FCA output again to generate new channel attention weights. These weights are then multiplied channel-by-channel with the FCA output, achieving fine-grained modulation at the spatial level. The FSAM module then feeds the DCAM output into two parallel 1×1 convolutions. The output of one branch is converted to the frequency domain via FFT, while the output of the other branch remains in the spatial domain. The two branches are element-wise multiplied and returned to the spatial domain via IFFT, forming a frequency-space fusion spatially enhanced feature.
[0059] Step 4. The fused multi-scale feature maps (feature maps output by the first RepC3 module, the third RepC3 module, and the fourth RepC3 module) are first input into the Transformer decoder, which performs cross-attention interaction with the feature maps through learnable Object Queries (usually 300 256-dimensional vectors). During the decoding process, each query vector dynamically focuses on a specific target area in the feature map and outputs a set of instantiated feature vectors that have been enhanced with attention. These feature vectors are then input in parallel into the dual-branch structure of the prediction head: the bounding box branch directly regresses the normalized bounding box coordinates through a 4-layer fully connected network to achieve pixel-level positioning; the category branch outputs the category probability distribution through a single-layer linear transformation and a Softmax activation function.
[0060] The experimental training data comes from web crawler images, with a total of 8,300 images. Two augmentations were performed on each image, and the data was ultimately divided into training, validation, and test sets in an 8:1:1 ratio.
[0061] The specific parameters of the experimental environment of this study are shown in Table 1.
[0062]
[0063] In this example, we conducted an ablation experiment to verify the performance of the proposed improved RT-DETR algorithm and the rationality of the design of each module. In the experiment, "√" indicates that the corresponding module is enabled, and "×" indicates that it is not used. The experimental results are shown in Table 2. All data are based on the verification results of the test set.
[0064]
[0065]
[0066] To further verify the effectiveness of our improved algorithm, this section compares it with other different benchmark models. The experiments were conducted under the same parameters and environment configuration. The results are shown in Table 3. All data are based on the test set verification results.
[0067]
[0068] Experimental results show that the present invention significantly improves the performance of the RT-DETR model by proposing a dynamic convolution hybrid module DCMB and a full-scale frequency attention structure CSPO. On a self-built dataset, the mAP50 is increased by 2.33%, while the number of parameters is reduced by 23% to 15.3M. It comprehensively surpasses the YOLO series and Faster R-CNN, verifying the effectiveness of the improved scheme in multi-scale feature expression and small target detection, and achieving high-precision and robust detection of automobile paint damage.
[0069] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. Equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. An improved automobile paint damage detection method based on RT-DETR, characterized in that: The steps include: Step 1: Construct a car paint damage dataset and add negative samples and interference samples, then perform image enhancement and preprocessing; Step 2: Based on the multi-layer convolution backbone network of ResNet18, a dynamic convolutional mixing module (DCMB) is introduced to extract multi-scale features. The DCMB includes a parallel dynamic Inception mixer and a convolutional gating unit (GLU). The dynamic Inception mixer implements dynamic multi-scale feature mixing, and the convolutional gating unit (GLU) enhances channel modeling capabilities. Step 3: Add a full-scale frequency attention structure CSPO to the neck network Neck to extract cross-domain structural information. The full-scale frequency attention structure CSPO includes an SPDConv module and a CSP-OmniKernel feature integration module, which receives shallow features from the backbone network Backbone and extracts multi-scale information. Step 4: The fused feature map is input to the Transformer decoder, and the query vector output by the decoder is input to the prediction head, which generates the bounding box coordinates and category probabilities.
2. The improved automobile paint surface damage detection method based on RT-DETR according to claim 1, characterized in that: The specific method of step 1 is: Collect images of car paint damage, manually classify them, and annotate their bounding boxes. Then add three other types of interference data: negative samples of intact vehicles, samples of interference from other object surface textures, and images of paint reflection artifacts. The data enhancement stage adopts a combination of multiple strategies, including basic spatial transformation, random rotation, scaling, cropping, brightness and contrast adjustment, and local noise injection, and finally divides the training, validation and test sets proportionally.
3. The improved automobile paint surface damage detection method based on RT-DETR according to claim 1 is characterized in that, The backbone network Backbone includes a 9-layer structure, namely 5 layers of Conv modules and 4 layers of dynamic convolution mixing modules DCMB. The input features pass through the first layer of Conv module to obtain P1 layer features, and then are sequentially input to the second layer of Conv module and the first layer of dynamic convolution mixing module DCMB to obtain P2 layer features, and then sequentially input to the third layer of Conv module and the second layer of dynamic convolution mixing module DCMB to obtain P3 layer features, and sequentially input to the fourth layer of Conv module and the third layer of dynamic convolution mixing module DCMB to obtain P4 layer features, and finally input to the fifth layer of Conv module and the fourth layer of dynamic convolution mixing module DCMB to obtain P5 layer features.
4. The improved automobile paint surface damage detection method based on RT-DETR according to claim 1 or 3, characterized in that, The dynamic convolutional mixing module (DCMB) is a combination module of a residual structure, including a dynamic Inception mixer and a convolutional gating unit (GLU). The input features are first batch-normalized (BatchNorm) and then passed through the dynamic Inception mixer and the convolutional gating unit (GLU) to achieve dynamic multi-scale feature mixing and channel modeling capability enhancement. The outputs of the two branches are respectively subjected to LayerScale scaling and DropPath random depth processing, and then returned to the trunk part for additive fusion with the original input through residual connections, ultimately outputting a multi-scale feature tensor.
5. The improved automobile paint surface damage detection method based on RT-DETR according to claim 4 is characterized in that, The dynamic Inception mixer uses two dynamic Inception depthwise separable convolutions to form paths of different scales, achieving multi-scale modeling in the channel dimension. First, the input feature map is divided into two groups according to the number of channels. Each group is fed into a dynamic Inception depthwise separable convolution with a different convolution kernel size to extract features and capture spatial information at different scales. Then, all processed results are spliced in the channel dimension and fused through 1×1 convolution.
6. The improved automobile paint surface damage detection method based on RT-DETR according to claim 5 is characterized in that: The dynamic Inception depth-wise separable convolution extracts channel-level global context by performing global average pooling on the input features, and then uses 1×1 convolution to generate weight coefficients corresponding to the three branches, which are used as dynamic attention guides after softmax normalization; The input features are processed by square kernel convolution, horizontal strip convolution, and vertical strip convolution, and the outputs of the three convolution branches are weightedly fused according to the weight coefficients corresponding to the three branches guided by dynamic attention, so as to adaptively adjust the expression strength of the receptive field and directional features for different inputs; the fusion results are then processed by batch normalization and SiLU activation function, and the output is a dynamic enhanced feature that contains both local texture and rich structural semantics.
7. The improved automobile paint surface damage detection method based on RT-DETR according to claim 1 is characterized in that, The specific structure of the full-scale frequency attention structure CSPO in step 3 is: The P2 layer features of the backbone network are processed by SPDConv, which preserves pixel-level details through space-to-depth transformation. The features rich in small object information are then fused with the P3 layer features. The fused features are then processed by the CSP-OmniKernel feature integration module. The output feature F1 of the CSP-OmniKernel feature integration module is processed by the first RepC3 module and the Conv module to obtain the feature F2; The backbone network P4 layer features are processed by the Conv module and the P5 layer features are fused after being processed by AIFI and upsample. After fusion, they are processed by the second RepC3 module and the Conv module to obtain feature F3. After the features F2 and F3 are fused, they are processed by the third RepC3 module and the Conv module and then fused with the P5 layer features. After fusion, they are passed through the fourth RepC3 module. The features output by the first RepC3 module, the third RepC3 module, and the fourth RepC3 module are input to the decoder for decoding processing.
8. The improved automobile paint surface damage detection method based on RT-DETR according to claim 4 is characterized in that, The specific architecture of the CSP-OmniKernel feature integration module is as follows: After channel mapping through Conv 1×1, the feature channel is split into two sub-branches in a ratio of 1:3: one is the main OmniKernel processing branch, which is 1 / 4 channel, and the other is the residual path; The main OmniKernel processing branch first performs a 1×1 convolution transformation on the input and activates it, and then simultaneously inputs three parallel branches: 1) The local branch uses a 1×1 depth-separable convolution to focus on fine-grained textures; 2) The large-scale branch Large extracts long- and short-domain texture information through three directions of depth convolution: horizontal, vertical, square, and point convolution; 3) The global branch Global achieves a global receptive field through dual-domain channel attention (DCAM) and frequency-gated mechanism FSAM; After fusion, the three parallel branches are processed by Conv 1×1 and then fused with the residual path of 3 / 4 channels and processed by Conv1×1 to output the final features.
9. The improved automobile paint surface damage detection method based on RT-DETR according to claim 1, characterized in that: The specific method of step 4 is: The fused multi-scale feature map is first input into the Transformer decoder, which performs cross-attention interaction with the feature map through a learnable object query. During the decoding process, each query vector dynamically focuses on a specific target area in the feature map, outputting a set of attention-enhanced instantiated feature vectors. The feature vectors are input in parallel into the prediction head to generate bounding box coordinates and category probabilities.
Citation Information
Cited By
Under-vehicle bolt corrosion detection method and device based on improved RT-DETR and medium
CN121564697A
Tunnel lining crack detection method based on improved RT-DETR
CN121685461A
Intelligent detection method for damage of PDC drill bit
CN122244030A