Lightweight Road Damage Detection Method Improved Based on RT-DETR
By improving the RT-DETR model and adopting a multi-scale feature extraction and fusion network, the problems of low efficiency and insufficient accuracy of road damage detection in the prior art are solved, and more efficient and accurate road damage detection is achieved.
Patent Information
- Application Number
- CN202510279394.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing road damage detection methods have problems such as low detection efficiency, high missed detection rate, insufficient ability to accurately evaluate road conditions, and degraded detection performance in complex environments.
A lightweight road damage detection method based on RT-DETR is proposed, using an end-to-end multi-scale feature extraction and fusion network, and through a multi-scale edge information enhancement module and a multi-branch hollow convolutional pyramid network module, the accuracy and efficiency of detection are improved.
It significantly improves the accuracy and detection speed of road damage detection, enhances the real-time and accuracy of detection, and can identify road damage faster and more accurately, meeting the needs of drones to detect road damage in real time.
Smart Images

Figure CN119784759B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of road monitoring, and particularly relates to a lightweight road damage detection method improved based on RT-DETR. Background Art
[0002] With the acceleration of the urbanization process, the scale and complexity of the road network have experienced explosive growth. However, with the rapid expansion of the road scale, the problem of road damage has become increasingly severe. Damages such as cracks and potholes not only affect the aesthetics and service life of the road, but also pose a serious threat to the safety of vehicles and passengers.
[0003] In the technical field of road monitoring, road damage detection is crucial for ensuring road safety and service performance. Traditional road damage detection methods mainly include (1) manual visual inspection, (2) detection by a detection vehicle equipped with high-precision sensors, and (3) detection based on image processing technology.
[0004] The manual visual inspection method highly depends on the professional knowledge and practical experience of the inspectors, is significantly affected by environmental factors, has low detection efficiency and a high missed detection rate, and cannot meet the requirements of large-scale and high-efficiency detection in the field of highway detection, severely limiting its wide application in the field of highway detection.
[0005] The high-precision sensors carried by special detection vehicles usually include ground-penetrating radar, laser scanners, etc. The laser scanning technology has strict requirements for environmental conditions, and factors such as changes in light intensity and the presence of obstacles will have a significant impact on the scanning results. It can only roughly reflect the deformation trend of the road surface, is difficult to provide accurate deformation data, and cannot meet the requirements for accurate assessment of the road surface conditions. Using ground-penetrating radar technology to detect road damages such as cracks and potholes has the advantages of non-destructive detection, high detection efficiency, and relatively high accuracy. However, this technology has many limitations in practical applications, such as being unable to automatically classify road surface cracks, being severely interfered by surrounding environmental clutter, and being greatly affected when detecting in inhomogeneous media. At the same time, the cost of using and maintaining this technology is high, and manual assistance is still required, facing great difficulties in popularizing and applying it in daily road surface detection work.
[0006] Traditional image - processing - based detection methods mainly use algorithms such as wavelet transform and threshold segmentation to extract road damage features. These methods mainly rely on artificially constructed features for detection. When road damage features change due to various factors (such as differences in road surface materials, diversity of damage types, etc.), the detection performance of these algorithms will decline sharply. In addition, this method is easily interfered by complex background factors such as water stains, shadows, and color differences, and has poor anti - interference ability during the detection process, severely restricting the further improvement of the accuracy of road damage recognition. Due to its insufficient robustness and generalization ability, it is difficult to effectively cope with the complex challenges of road damage detection in different actual environments and cannot meet diverse detection requirements.
[0007] To solve the above problems, deep - learning technology has been widely used in the field of road damage detection, significantly improving the accuracy and efficiency of detection. Such as two - stage object detection algorithms based on R - CNN, Faster R - CNN, Mask R - CNN, etc., single - stage object detection algorithms based on SSD, YOLO, and real - time end - to - end object detection models based on the Transformer architecture, such as RT - DETR, etc. Deep learning has shown great advantages in the field of road damage detection and is expected to completely replace traditional road damage detection methods. However, there are many types of road damage, including cracks, potholes, bumps, etc., which show complex and variable characteristics in terms of shape, size, and color. Current deep - learning models still have problems such as difficult recognition of multi - scale road damage and large numbers of parameters in high - precision models, making it difficult to apply them to resource - constrained devices. Summary of the Invention
[0008] The present invention proposes a lightweight road damage detection method improved based on RT - DETR, using an end - to - end multi - scale feature extraction and fusion network for road damage detection, which can greatly improve the detection accuracy and speed.
[0009] To achieve the above object, the technical solution of the present invention is realized as follows:
[0010] A lightweight road damage detection method improved based on RT - DETR, comprising:
[0011] S1. Obtain road damage image data and divide the data set;
[0012] S2. Modify the backbone network of the RT-DETR model to be composed of multiple multi-scale edge information enhancement modules for extracting multi-scale features of the input image data; modify the hybrid encoder of the RT-DETR model to be composed of an attention-based intra-scale feature interaction module and a multi-branch dilated convolutional pyramid network module. The attention-based intra-scale feature interaction module AIFI processes the multi-scale features, and the multi-branch dilated convolutional pyramid network module fuses the multi-scale features to convert features at different levels into a sequence of image features. The decoder adopts an IoU-based query selection mechanism to select a set of image features from the sequence of image features output by the modified hybrid encoder as the initial object query, and iteratively optimizes to generate prediction boxes and confidence scores through an auxiliary prediction head. After modification, the RDD-DETR model is obtained.
[0013] S3. Input the training set of the road damage image data obtained by division into the RDD-DETR model for training, adjust the model through the validation set, and evaluate the model performance through the test set. The trained model is used for road damage detection in road image data.
[0014] Further, in step S2, the method for constructing the backbone network includes: alternating arrangement of multiple multi-scale edge information enhancement modules and multiple convolutional layers.
[0015] Furthermore, the multi-scale edge information enhancement module realizes multi-scale edge information enhancement through three parts: multi-scale pooling, edge enhancement, and feature fusion. The multi-scale pooling extracts multi-level features by performing adaptive average pooling on the input feature map at different scales through an adaptive average pooling operation. The edge enhancement adaptively enhances the edge information of the input feature map by strengthening the edge details of each level of feature through an independent EdgeInfoEnhancer module to obtain multi-scale edge features. The feature fusion concatenates the original local features and the multi-scale edge features in the channel dimension to form unified features.
[0016] Still further, the multi-scale edge information enhancement module also introduces the HiLo Attention attention mechanism. The high-frequency attention Hi-Fi enhances the ability to capture edge details, and the low-frequency attention Lo-Fi optimizes the ability to perceive the global structure.
[0017] Preferably, the EdgeInfoEnhancer module first uses average pooling to locally smooth the input feature map to extract low-frequency information, then subtracts the smoothed feature map from the input feature map to obtain high-frequency information, adjusts the weights of the high-frequency information of different channels through a convolutional kernel, and finally superimposes the enhanced edges on the input feature map in a residual form.
[0018] Further, in step S2, the backbone network outputs a shallow feature map S3, a middle feature map S4, and a deep feature map S5; the attention-based intra-scale feature interaction module AIFI processes the deep feature map S5 to perform intra-scale interaction and applies the self-attention mechanism to the deep features rich in semantic information; the multi-branch dilated convolutional pyramid network module fuses the shallow feature map S3, the middle feature map S4, and the features processed by AIFI.
[0019] Furthermore, the multi-branch dilated convolutional pyramid network module is jointly composed of a parallel upsampling module, a parallel downsampling module, and a local multi-scale dilated convolutional module PMAC. The parallel upsampling module and the parallel downsampling module provide multiple feature extraction paths, and the local multi-scale dilated convolutional module PMAC realizes multi-scale feature extraction through multi-path dilated convolutions.
[0020] Preferably, the parallel upsampling module sets two upsampling paths. One path realizes upsampling through transposed convolution, and the other path completes upsampling through linear interpolation. The outputs of the two paths are concatenated; the concatenated feature map is multiplied by the attention weight, and a convolutional layer is used to adjust the number of channels of the fused feature map to obtain the final upsampled feature map; the attention weight is generated through global pooling operation and the HardSigmoid activation function.
[0021] Preferably, the parallel downsampling module sets two downsampling paths. One path realizes downsampling through convolution, and the other path completes downsampling through max pooling. The outputs of the two paths are concatenated; the concatenated feature map is multiplied by the attention weight, and a convolutional layer is used to adjust the number of channels of the fused feature map to obtain the final downsampled feature map; the attention weight is generated through global pooling operation and the HardSigmoid activation function.
[0022] Preferably, the local multi-scale dilated convolutional module PMAC includes a multi-scale dilated convolutional module MAC, and also introduces a convolutional layer to independently extract features and fuse them with the output of MAC; the multi-scale dilated convolutional module MAC captures features of different scales by parallelly processing dilated convolutions with different dilation rates, and then concatenates the features of different scales.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. The present invention proposes a multi-scale edge information enhancement structure in the backbone network part to improve the quality of feature extraction and the ability to represent edge information. Through three core functions of multi-scale feature extraction, edge enhancement, and feature fusion, the model significantly enhances the perception ability of image features and can effectively handle object detection tasks that require precise edge information.
[0025] 2. The present invention proposes a multi-branch dilated convolutional pyramid network structure in the design of the neck encoding network. By combining multi-branch sampling, feature selection, and dilated convolution, the diversity and effectiveness of feature expression are improved.
[0026] 3. The present invention realizes model lightweighting, has practicality and efficiency for real-scene applications when used for road damage detection, and enhances the real-time performance and accuracy of road damage detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic flowchart of an embodiment of the present invention;
[0028] Figure 2 is a schematic structural diagram of the improved RDD-DETR model according to an embodiment of the present invention;
[0029] Figure 3 is a schematic structural diagram of the multi-scale edge information enhancement module (MEIE) according to an embodiment of the present invention;
[0030] Figure 4 is a schematic structural diagram of the EdgeInfoEnhancer module according to an embodiment of the present invention;
[0031] Figure 5 is a schematic structural diagram of the parallel upsampling module according to an embodiment of the present invention;
[0032] Figure 6 is a schematic structural diagram of the parallel downsampling module according to an embodiment of the present invention;
[0033] Figure 7 is a schematic structural diagram of the PMAC module according to an embodiment of the present invention;
[0034] Figure 8 is a schematic diagram of the detection effect of an embodiment of the present invention Figure 1 ;
[0035] Figure 9 is a schematic diagram of the detection effect of an embodiment of the present invention Figure 2 ;
[0036] Figure 10 is a schematic diagram of the detection effect of an embodiment of the present invention Figure 3 ;
[0037] Figure 11 is a schematic diagram of the detection effect of an embodiment of the present invention Figure 4 。 DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0039] To make the objectives and features of this invention patent more obvious and understandable, the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0040] In this embodiment, the present invention is applied to the scenario of road damage detection by drones, aiming to solve the problems of low detection accuracy and insufficient real-time performance in the prior art. The specific method is as Figure 1 shown and includes:
[0041] S1. Obtain the road damage image data collected by the drone and divide the data set;
[0042] S2. Modify the backbone network of the RT-DETR model to be composed of multiple multi-scale edge information enhancement modules, and modify the hybrid encoder of the RT-DETR model to be composed of an attention-based intra-scale feature interaction module and a multi-branch dilated convolutional pyramid network module. After modification, the RDD-DETR model is obtained and deployed on the drone;
[0043] S3. Input the training set of the divided road damage image data into the RDD-DETR model for training, adjust the model through the validation set, and evaluate the model performance through the test set. The trained model is used for the drone to collect road image data for road damage detection.
[0044] In this embodiment, the division of the data set in step S1 means that after obtaining the road damage image data collected by the drone, after preprocessing and annotating the image data, the image data is divided into a training set, a test set, and a validation set according to a ratio of 7:2:1 or 6:2:2.
[0045] In step S2, the RT-DETR model is modified to the RDD-DETR model. The structure of the modified RDD-DETR model is as Figure 2 shown, and the specific modification process is as follows:
[0046] Modify the backbone network to be composed of multiple Multi-scale Edge Information Enhancers (MEIEs). The multiple MEIEs and multiple convolutional layers (Conv) are arranged alternately. Through this alternating arrangement structure, the image data can be gradually penetrated to capture feature information at different levels. The backbone network extracts multi-scale feature representations from the original image and outputs feature maps at three stages (shallow feature map s3, middle feature map s4, and deep feature map s5), and inputs these features into the hybrid encoder of the neck encoding network. The hybrid encoder is composed of an Attention-based Intrascale Feature Interaction (AIFI) module and a Multiple Atrous Convolution Pyramid Network (MAPN) module. AIFI processes the deep feature map s5 from the backbone network, performs intrascale interaction only on the feature layer of S5, and applies the self-attention mechanism to the deep features rich in semantic information, which can significantly reduce the computational cost. MAPN fuses features at all scales, including the shallow feature map S3, the middle feature map S4, and the features processed by AIFI. Convert features at different levels into a sequence of image features to enhance the accuracy of object recognition. The decoder of the decoding prediction network adopts an IoU-based query selection mechanism to select a set of image features from the sequence output by the encoder as the initial object query, which improves the quality of the decoder's initial object query and iteratively optimizes the generation of prediction boxes and confidence scores through the auxiliary prediction head, enhancing the detection performance of the decoder.
[0047] The MEIE proposed in this embodiment is used to improve the problem of insufficient feature extraction ability of the original backbone network. The MEIE module realizes multi-scale edge information enhancement through three parts: multi-scale pooling, edge enhancement, and feature fusion.
[0048] In this embodiment, the specific structure of the MEIE module is as Figure 3As shown in the figure, first, the information of different scales corresponding to different branches is subjected to an Adaptive Average Pooling operation on the input feature map to perform adaptive average pooling of different scales to extract multi-level features. This pooling method can automatically adjust the size of the pooling kernel according to the preset target size, so as to extract local information at different resolutions, capture multi-level features in the image, and take into account both global and local information. Then, the feature map is reduced in dimension through two convolutional layers Conv to reduce the computational amount while separating feature channels of different scales and extracting local features. Subsequently, the feature map of each branch is restored to the original input size through linear interpolation Upsample to ensure that multi-scale features can be stitched together, and then passed through an independent EdgeInfoEnhancer module to enhance edge details, adaptively enhancing the edge information of the input feature map, which helps the model capture more detailed features. Finally, the original local features after passing through one convolutional layer are stitched together with the multi-scale edge features enhanced by EdgeInfoEnhancer in the channel dimension to form a unified feature representation. In addition, the MEIE module also introduces HiLo Attention, enabling the multi-scale edge information enhancement module to process multi-scale features and edge information more efficiently. High-frequency attention (Hi-Fi) enhances the ability to capture edge details, and low-frequency attention (Lo-Fi) optimizes the perception ability of the global structure. The combination of the two makes the fusion of multi-scale features more efficient, not only improving the sensitivity of the model to edge information but also significantly enhancing the computational efficiency.
[0049] In the MEIE module, an EdgeInfoEnhancer module (edge information enhancement module) is set because edge information plays a key role in the road damage detection task. Clear edge information helps to more accurately locate the target boundary. The EdgeInfoEnhancer module is used to enhance edge features. Edge information usually includes important details such as the contours, textures, and shapes of objects in the image. By enhancing this information, the network can more sensitively capture the edge features in the image. The structure of the EdgeInfoEnhancer module is as Figure 4 shown. First, average pooling AvgPool is used to perform local smoothing on the input feature map to extract low-frequency information. Subsequently, the original feature map is subtracted from the smoothed feature map to obtain high-frequency information, that is, details such as edges and textures. Then, the weights of the high-frequency information of different channels are adjusted through the convolutional layer Conv. Finally, the enhanced edges are superimposed on the original input in the form of residuals, which not only retains the original features but also strengthens the edge details.
[0050] The multi-branch atrous convolution pyramid network MAPN in the hybrid encoder of the neck coding network also belongs to one of the improvements to the RT-DETR model. The multi-branch atrous convolution pyramid network module MAPN is jointly composed of a parallel upsampling module, a parallel downsampling module, and a local multi-scale atrous convolution module PMAC.
[0051] The parallel upsampling module is as Figure 5 shown, using two different upsampling paths to extract features: one is the upsampling implemented by transposed convolution ConvTranspose, and the other is the upsampling completed by linear interpolation Upsample followed by convolution. These two paths work in parallel, aiming to capture important information in the input feature map from different perspectives. At the same time, the module generates attention weights through global pooling operation GlobalPool and HardSigmoid activation function, and these weights reflect the importance of each channel. In the feature fusion stage, the outputs of the two upsampling paths are first concatenated and fused in the channel dimension to form a richer feature representation. The concatenated feature map is multiplied by the attention weights calculated by the previous parallel upsampling module. This multiplication operation realizes feature weighted fusion based on attention weights, thereby strengthening the features that are more important for the task and suppressing irrelevant or redundant features. Finally, the number of channels of the fused feature map is adjusted through a convolutional layer to obtain the final upsampled feature map.
[0052] Similar to the upsampling module, the parallel downsampling module is as Figure 6 shown, and also adopts two parallel downsampling paths to extract features: one is the downsampling implemented by convolution with a stride of 2, and the other is the downsampling completed by max pooling MaxPool followed by convolution. These two paths also execute in parallel, capturing key information in the input feature map from different perspectives. At the same time, the parallel downsampling module also generates attention weights through global pooling operation GlobalPool and HardSigmoid activation function. In the feature fusion stage, the outputs of the two downsampling paths are first concatenated (Concat) in the channel dimension to form a richer feature representation. Then, these concatenated feature maps are multiplied by the attention weights calculated by the previous parallel downsampling module. This multiplication operation also realizes feature weighted fusion based on attention weights, strengthening important features and suppressing irrelevant features. Finally, the number of channels of the fused feature map is adjusted through a convolutional layer to obtain the final downsampled feature map.
[0053] The multi-branch dilated convolutional pyramid network module MAPN provides multiple feature extraction paths for the network through parallel upsampling modules and parallel downsampling modules. This multi-path design enriches the diversity of feature representation, enabling the network to capture target information from different scales and levels and extract diverse features through different sampling strategies such as transposed convolution, interpolation, pooling, etc. In addition, it also introduces a gating mechanism to selectively enhance the sampled features. By learning the importance weights of the features, the gating mechanism can strengthen the features related to road damage while suppressing redundant features, thus improving the effectiveness of feature representation. The parallel upsampling module and the parallel downsampling module correspond to Figure 2 U and D in
[0054] The local multi-scale dilated convolutional module PMAC (Partial Multiple Atrous Convolution) module is one of the core structures of the multi-branch dilated convolutional pyramid network module. As Figure 7 shown, it realizes multi-scale feature extraction through multi-path dilated convolution. In PMAC, in addition to multi-path dilated convolution, an additional convolutional layer Conv is introduced. This convolutional layer can independently extract features and then perform Concat splicing fusion with the output of the multi-scale dilated convolutional module MAC, further enhancing the feature expression ability and being able to reduce the computational amount while improving the performance, which has significant advantages for resource-constrained devices or application scenarios that require fast inference. Dilated convolution expands the receptive field by introducing a dilation rate in the convolutional kernel, enabling the network to capture a larger range of context information without increasing the computational amount. Different dilation rates of dilated convolution are used in the multi-scale dilated convolutional module MAC (Multiple Atrous Convolution) in PMAC, so that it can extract local details and global context information simultaneously.
[0055] The structure of the MAC module is as Figure 7As shown, different dilation rate AtrousConvs are processed in parallel to capture features at different scales. The AtrousConv branches with different dilation rates can capture features at different scales. Larger dilation rates can capture more extensive context information, while smaller dilation rates can retain more detailed information. This multi-scale feature extraction ability enables the MAC module to better understand local and global information in the input feature map. The outputs of the AtrousConv branches with different dilation rates and the output of the convolutional layer are concatenated in the channel dimension. This concatenation operation fuses features at different scales to form a richer feature representation. This feature fusion helps retain feature information at different scales and improves the model's ability to recognize complex patterns. Finally, a convolutional layer Conv processes the concatenated feature map to generate the final output feature map. This convolutional layer can be used to adjust the number of channels to match the number of channels in the subsequent layer. In the road damage detection task, small target damages usually require more refined local features, while large target damages require more extensive context information. The MAC module can meet both of these requirements simultaneously through parallel multi-scale feature extraction, significantly improving the network's ability to represent complex features.
[0056] The improved RDD-DETR model is deployed on the drone and trained using the partitioned dataset. The trained model is used for the drone to collect road image data for road damage detection, which can significantly improve the efficiency of road damage detection.
[0057] Figure 8 、 Figure 9 、 Figure 10 、 Figure 11 Shows the performance of the improved RDD-DETR model in the partial image detection results in the public dataset UAPD (Unmanned Aerial Vehicle Image Dataset for Detecting Road Cracks). Figures 8 - 11 Longitudinal crack, Transverse crack, Oblique crack, Alligator crack, Repair in represent damage types. Longitudinal crack is a longitudinal crack, Transverse crack is a transverse crack, Oblique crack is an oblique crack, Alligator crack is a network crack, and Repair represents repair. The number after the damage type represents the detection accuracy. Through Figures 8 - 11 's display, it can be seen that the present invention optimizes the accuracy and speed of road damage image recognition, can recognize road damage faster and more accurately, thereby improving the accuracy of the detection results, and meeting the needs of real-time road damage detection by drones.
[0058] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight road damage detection method based on RT-DETR improvement, characterized in that: include: S1. Obtain road damage image data and divide the data set; S2, modifying the backbone network of the RT-DETR model to consist of multiple multi-scale edge information enhancement modules to extract multi-scale features of the input image data; The hybrid encoder of the RT-DETR model is modified to consist of an attention-based intra-scale feature interaction module and a multi-branch dilated convolutional pyramid network module. The attention-based intra-scale feature interaction module AIFI processes the multi-scale features, and the multi-branch dilated convolutional pyramid network module fuses the multi-scale features to convert features at different levels into a sequence of image features. The decoder adopts an IoU-based query selection mechanism to select a set of image features from a sequence of image features output by the modified hybrid encoder as the initial object query, and generates prediction boxes and confidence scores through iterative optimization of the auxiliary prediction head; After modification, the RDD-DETR model is obtained; The multi-scale edge information enhancement module realizes multi-scale edge information enhancement through multi-scale pooling, edge enhancement and feature fusion. The edge enhancement is realized by an independent EdgeInfoEnhancer module. The EdgeInfoEnhancer module first uses average pooling to locally smooth the input feature map to extract low-frequency information, then subtracts the smoothed feature map from the input feature map to obtain high-frequency information, and then adjusts the high-frequency information weights of different channels through convolution kernels. Finally, the enhanced edge is superimposed on the input feature map in the form of a residual. The multi-branch dilated convolutional pyramid network module is composed of a parallel upsampling module, a parallel downsampling module, and a local multi-scale dilated convolution module PMAC. The parallel upsampling module and the parallel downsampling module provide multiple feature extraction paths. The local multi-scale dilated convolution module PMAC realizes multi-scale feature extraction through multi-path dilated convolution. S3. Input the training set of the divided road damage image data into the RDD-DETR model for training, adjust the model through the validation set, and evaluate the model performance through the test set; the trained model is used for road damage detection in road image data.
2. The improved lightweight road damage detection method based on RT-DETR according to claim 1 is characterized in that: In step S2, the method for constructing the backbone network includes: alternatingly arranging a plurality of multi-scale edge information enhancement modules and a plurality of convolutional layers.
3. The improved lightweight road damage detection method based on RT-DETR according to claim 2 is characterized in that: The multi-scale edge information enhancement module realizes multi-scale edge information enhancement through three parts: multi-scale pooling, edge enhancement and feature fusion; the multi-scale pooling is to extract multi-level features by performing adaptive average pooling of different scales on the input feature map through adaptive average pooling operation; The edge enhancement is to strengthen the edge details of each level of features through an independent EdgeInfoEnhancer module, adaptively enhance the edge information of the input feature map, and obtain multi-scale edge features; the feature fusion is to splice the original local features with the multi-scale edge features in the channel dimension to form a unified feature.
4. The improved lightweight road damage detection method based on RT-DETR according to claim 3 is characterized in that: The multi-scale edge information enhancement module also introduces the HiLo Attention mechanism. The high-frequency attention Hi-Fi enhances the ability to capture edge details, and the low-frequency attention Lo-Fi optimizes the perception of the global structure.
5. The improved lightweight road damage detection method based on RT-DETR according to claim 1 is characterized in that: In step S2, the backbone network outputs a shallow feature map S3, a middle feature map S4, and a deep feature map S5; the attention-based intra-scale feature interaction module AIFI processes the deep feature map S5, performs intra-scale interaction, and applies a self-attention mechanism to the deep features rich in semantic information; the multi-branch atrous convolutional pyramid network module fuses the shallow feature map S3, the middle feature map S4, and the features processed by AIFI.
6. The improved lightweight road damage detection method based on RT-DETR according to claim 1 is characterized in that: The parallel upsampling module sets two upsampling paths, one of which implements upsampling through transposed convolution and the other completes upsampling through linear interpolation, and the outputs of the two paths are spliced; the spliced feature map is multiplied by the attention weight, and the number of channels of the fused feature map is adjusted through a convolution layer to obtain the final upsampling feature map; the attention weight is generated by a global pooling operation and a HardSigmoid activation function.
7. The improved lightweight road damage detection method based on RT-DETR according to claim 1 is characterized in that: The parallel downsampling module sets two downsampling paths, one for downsampling through convolution and the other for downsampling through maximum pooling, and the outputs of the two paths are spliced; the spliced feature map is multiplied by the attention weight, and the number of channels of the fused feature map is adjusted through a convolution layer to obtain the final downsampled feature map; the attention weight is generated by a global pooling operation and a HardSigmoid activation function.
8. The improved lightweight road damage detection method based on RT-DETR according to claim 1 is characterized in that: The local multi-scale dilated convolution module PMAC includes a multi-scale dilated convolution module MAC, and also introduces a convolution layer to independently extract features and fuse them with the output of MAC; the multi-scale dilated convolution module MAC captures features of different scales by processing dilated convolutions with different expansion rates in parallel, and then splices the features of different scales.
Citation Information
Patent Citations
Road inspection robot obstacle detection method and system based on RT-DETR-Sat
CN118864424A
Real-time traffic target detection method for low-illuminance scene
CN119007149A