Construction Method and Application of Road Damage Detection Model Based on Improved RT-DETR Model
By adopting a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network in the RT-DETR model, the problem of insufficient multi-scale semantic information aggregation ability in road damage detection is solved, and a higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202510443723.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the road damage detection, traditional RT-DETR models have problems such as insufficient multi-scale semantic information aggregation ability, difficulty in retaining small texture differences in shallow features and being easily disturbed by background noise, resulting in low detection accuracy.
The backbone network and neck network of the RT-DETR model are replaced by the cross-stage part-multi-scale edge feature fusion backbone network (CSP-MEFFN) and the rectangular context fusion pyramid network (RCFPN) to improve feature extraction and context perception capabilities.
The model's sensitivity to road damage edge features and its ability to identify targets in complex contexts is significantly improved, thereby improving the accuracy of road damage detection.
Smart Images

Figure CN119962404B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the target detection technology in the field of computer technology, and particularly relates to a method for constructing and applying a road damage detection model based on an improved RT-DETR model. Background Art
[0002] As a core technology in the fields of intelligent transportation and infrastructure maintenance, road damage detection plays a crucial role in aspects such as road safety assessment, preventive maintenance decision-making, and full-life cycle management.
[0003] Road damage detection technologies can be divided into two major categories: manual detection methods and image recognition methods. Conventional manual detection methods generally rely on manual inspections and semi-automatic detection equipment, but they have inherent defects such as low detection efficiency and large subjective evaluation biases. With the rapid development of computer operations, machine learning, and deep learning technologies, image recognition methods have received increasing attention. Patent No. 202311220593.8 proposes a rapid road damage detection method based on a multi-head attention mechanism. This method introduces multi-head attention into the Resnet-18 backbone network, fuses multi-layer sampling results, and achieves a high accuracy rate in road damage detection; Patent No. 202410706333.X proposes a road damage detection method based on improved YOLOv8. This method improves the accuracy rate of road damage detection by introducing LSKA attention and designing a focused diffusion pyramid network FDPN. The above methods can all effectively implement the function of pavement state classification, but there are challenges such as global context modeling (understanding the overall scene) and local detail encoding (accurately identifying boundaries and subtle features) in the actual application process.
[0004] RT-DETR (Real-Time Detection Transformer) is an end-to-end real-time object detection model based on the Transformer architecture. By combining the local feature extraction advantages of CNN and the global relationship modeling ability of Transformer, it significantly enhances the object detection performance. Its dynamic label assignment strategy improves the training efficiency through an adaptive matching mechanism. Combined with the real-time optimized decoder design, it reduces the computational complexity while maintaining the end-to-end detection characteristics, and is considered to have great development potential in road damage detection. However, since the Resnet18 backbone network adopted by the traditional RT-DETR model only relies on simple cross-layer connections for hierarchical feature fusion, it is difficult to effectively aggregate multi-scale semantic information, resulting in weak representation ability of road damage edge features. At the same time, the feature fusion network CCFM (Cross-Scale Feature Fusion Module) of the traditional RT-DETR uses simple splicing or summation operations, resulting in its inability to retain the tiny texture differences of the shallow features of road damage targets, and is easily interfered by the noise of adjacent background features during mid- and high-level feature fusion, leading to semantic degradation of target features. In addition, the complexity of damage feature representation involved in the road damage detection scenario is high, and the detection accuracy is also affected by uncertain factors such as weather and similar backgrounds, which poses great technical difficulties. At present, there is no relevant research on implementing road damage detection based on the RT-DETR model. The reason is that the Resnet18 feature extraction backbone network and the CCFM feature fusion neck network of the traditional RT-DETR model are not sufficient to provide high efficiency and accuracy.
[0005] Therefore, by making modifications such as constructing a new feature extraction backbone and feature fusion neck network, it is of great research significance to design a road damage detection model based on an improved RT-DETR model. Summary of the Invention
[0006] The embodiment of this application provides a construction method and application of a road damage detection model based on an improved RT-DETR model. It improves the basic framework of the RT-DETR model, adopts the Cross-Stage Partial-Multi-Scale Edge Feature Fusion Network (CSP-MEFFN) as the backbone network to significantly enhance the sensitivity of the model to road damage edges; and designs the Rectangular Context Fusion Pyramid Network (RCFPN) as the neck network to achieve dynamic adaptive fusion of shallow high-frequency details, middle-level edge structures and deep semantic features, improving the model's context awareness and the recognition ability of targets under complex backgrounds, thereby improving the accuracy of the traditional RT-DETR model for road damage detection.
[0007] In a first aspect, an embodiment of the present application provides a method for constructing a road damage detection model based on an improved RT-DETR model, including the following steps:
[0008] S1: Obtain original road damage images and label the damage type of each original road damage image to form a training set;
[0009] S2: Construct a road damage detection model based on the improved RT-DETR model:
[0010] The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network;
[0011] The cross-stage partial-multi-scale edge feature fusion backbone network includes a plurality of cascaded convolutional modules and cross-stage partial-multi-scale edge feature fusion modules, and each cross-stage partial-multi-scale edge feature fusion module includes a plurality of multi-scale edge feature fusion modules for cross-stage feature interaction; the rectangular context fusion pyramid network adopts a three-level recursive architecture built in sequence by a primary fusion stage, an intermediate interaction stage, and an ultimate optimization stage;
[0012] The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module are output. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module, after passing through the AIFI module of the RT-DETR basic framework, are input into the rectangular context fusion pyramid network together with other hierarchical features; the features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage. The cross-layer features of different levels are strengthened in relevance in the intermediate interaction stage by using a dynamic interpolation fusion unit and a multi-scale fusion module to generate corresponding strengthened features. The strengthened features are subjected to feature fusion using a reparameterized convolution module in the ultimate optimization stage to obtain output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results;
[0013] S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train the road damage detection model based on the improved RT-DETR model.
[0014] Second aspect, the embodiments of the present application provide an application method of a road damage detection model based on an improved RT-DETR model, including the following steps: Input a road image into the road damage detection model based on the improved RT-DETR model constructed by the construction method of the road damage detection model based on the improved RT-DETR model, and output the damage type, where the damage type is selected from any one of the six damage types of longitudinal cracks, transverse cracks, crocodile cracks, diagonal cracks, repairs, and potholes.
[0015] The main contributions and innovations of the present invention are as follows:
[0016] 1. The cross-stage partial-multi-scale edge feature fusion network designed in the present invention replaces the original backbone network of the RT-DETR model. According to the collaborative optimization mechanism of multi-scale context dynamic aggregation and high-frequency edge residual enhancement, it effectively enhances the feature expression ability and effectively solves the local coding problem.
[0017] 2. The rectangular context fusion pyramid network designed in the present invention replaces the neck network of the RT-DETR model, realizes the dynamic adaptive fusion of the neck network for high-frequency details, middle-layer edge structures, and deep semantic features, improves the global context modeling ability, and at the same time improves the performance of the model in focusing on important information and suppressing lightweight features (noise introduced by external factors such as lighting and weather), thereby improving the accuracy of road damage detection.
[0018] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0020] Figure 1 is a flowchart of the construction method of the road damage detection model based on the improved RT-DETR model according to the embodiments of the present application;
[0021] Figure 2 is a basic architecture diagram of the road damage detection model based on the improved RT-DETR model according to the present application;
[0022] Figure 3 is a basic architecture diagram of the cross-stage partial-multi-scale edge feature fusion backbone network according to the present application;
[0023] Figure 4 is a basic architecture diagram of the rectangular context fusion pyramid network according to the present application;
[0024] Figure 5 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0025] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0026] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0027] Embodiment 1
[0028] As Figure 1 shown, this solution provides a method for constructing a road damage detection model based on an improved RT-DETR model, including the following steps:
[0029] S1: Obtain the original road damage images and label the damage type of each original road damage image to form a training set;
[0030] S2: Construct a road damage detection model based on the improved RT-DETR model:
[0031] The road damage detection model replaces the backbone network and the neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network;
[0032] The cross-stage partial-multi-scale edge feature fusion backbone network includes a plurality of cascaded convolutional modules and a cross-stage partial-multi-scale edge feature fusion module, and each cross-stage partial-multi-scale edge feature fusion module includes a plurality of multi-scale edge feature fusion modules for cross-stage feature interaction; the rectangular context fusion pyramid network adopts a three-level recursive architecture built in sequence by a primary fusion stage, an intermediate interaction stage, and an ultimate optimization stage;
[0033] The feature map input passes through the cross-stage partial multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial multi-scale edge feature fusion module pass through the AIFI module of the RT-DETR basic framework and are input into the rectangular context fusion pyramid network together with other hierarchical features; the features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage. The cross-layer features of different levels are enhanced in relevance in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are subjected to feature fusion using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results;
[0034] S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train the road damage detection model based on the improved RT-DETR model.
[0035] Step S1 further includes the steps:
[0036] S11: Obtain the original road damage images;
[0037] S12: Mark the damage type of each original road damage image according to the road damage features in each original road damage image to obtain a data set;
[0038] S13: Perform image processing on the original road damage images in the data set and divide them into a training set, a validation set, and a test set.
[0039] Furthermore, the original road damage images of this solution are obtained by a vehicle-mounted camera shooting the actual road damage features and being augmented by a public data set. The original road damage images include at least one image area corresponding to at least one road damage feature.
[0040] Specifically, a USB camera is set at the front of the vehicle hood to shoot and obtain the road with road damage. The original road damage images are obtained by extracting frames from the captured video, and then a public data set is introduced for augmentation, including Crowdsensing-based Road Damage Detection Challenge, UAV-PDD 2023, CNRDD, IRRDD.
[0041] Furthermore, the damage types of this solution are any one of the six damage types of longitudinal cracks, transverse cracks, crocodile cracks, oblique cracks, repairs, and potholes.
[0042] Furthermore, the method for image processing of the original road damage images in the dataset of this solution includes image screening and image enhancement. The means of image screening include, but are not limited to, removing blurred and duplicate photos. The means of image enhancement include, but are not limited to, rotation, scaling, brightness adjustment, etc.
[0043] Furthermore, a labeling software (such as LabelImg software) is used to label the six damage types of the original road damage images.
[0044] In some embodiments, the dataset of more than 50,000 original road damage images in this solution is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0045] Regarding the road damage detection model based on the improved RT-DETR model constructed in step S2 of this solution, the structural framework schematic diagram of the road damage detection model based on the improved RT-DETR model is as Figure 2 shown.
[0046] The road damage detection model based on the improved RT-DETR model replaces the original backbone network with a Cross Stage Partial-Multi-Scale Edge Feature Fusion backbone network CSP-MEFFN on the basis of the RT-DETR basic framework, and replaces the neck network with a Rectangular Context Fusion Pyramid Network RCFPN, and retains the AIFI module and the decoder of the RT-DETR basic framework. Specifically, the road damage detection model based on the improved RT-DETR model includes a Cross Stage Partial-Multi-Scale Edge Feature Fusion backbone network CSP-MEFFN, an AIFI module, a Rectangular Context Fusion Pyramid Network, and a decoder connected in sequence. Regarding the AIFI module and the decoder of the road damage detection model, the original RT-DETR basic framework is followed, so no redundant introduction will be made.
[0047] The following will focus on introducing the Cross Stage Partial-Multi-Scale Edge Feature Fusion backbone network CSP-MEFFN of this solution, as Figure 3 shown is the structural schematic diagram of the Cross Stage Partial-Multi-Scale Edge Feature Fusion backbone network CSP-MEFFN of this solution.
[0048] The Cross Stage Partial-Multi-Scale Edge Feature Fusion backbone network CSP-MEFFN of this solution includes a plurality of cascaded convolutional modules and a Cross Stage Partial-Multi-Scale Edge Feature Fusion module CSP-MEFF. The convolutional modules and the Cross Stage Partial-Multi-Scale Edge Feature Fusion module are alternately cascaded, that is, the features are input into the Cross Stage Partial-Multi-Scale Edge Feature Fusion module CSP-MEFF for feature extraction after being convolved by the convolutional module, and then input into the next convolutional module for convolution processing.
[0049] As Figure 3 shown, the cross-stage multi-scale edge feature fusion backbone network CSP-MEFFN of this solution includes 5 convolutional modules and 4 cross-stage multi-scale edge feature fusion modules CSP-MEFF. Each cross-stage multi-scale edge feature fusion module CSP-MEFF outputs the hierarchical features of the corresponding level. In some specific embodiments, a road damage image with three-channel RGB and a resolution of 640×640 is input into the cross-stage multi-scale edge feature fusion backbone network CSP-MEFFN. The road damage image passes through the first convolutional module to output features of 320×320×64, then passes through the second convolutional module to output features of 160×160×127, passes through the first cross-stage multi-scale edge feature fusion module CSP-MEFF to output features of 160×160×127, then passes through the third convolutional module to output features of 80×80×256, passes through the second cross-stage multi-scale edge feature fusion module CSP-MEFF to output features of 80×80×256, then passes through the fourth convolutional module to output features of 40×40×512, passes through the third cross-stage multi-scale edge feature fusion module CSP-MEFF to output features of 40×40×512, then passes through the fifth convolutional module to output features of 20×20×512, and passes through the fourth cross-stage multi-scale edge feature fusion module CSP-MEFF to output features of 20×20×512.
[0050] Specifically, for each cross-stage multi-scale edge feature fusion module CSP-MEFF: The cross-stage multi-scale edge feature fusion module CSP-MEFF includes multiple multi-scale edge feature fusion modules MEFF for cross-stage feature interaction. The features input into the cross-stage multi-scale edge feature fusion module CSP-MEFF are sequentially processed by a convolutional layer and a separation layer to obtain multiple separated features. The multiple separated features are input into the multi-scale edge feature fusion modules MEFF connected in a multi-level cascade. And after the output feature of each multi-scale edge feature fusion module MEFF is added to the separated feature, convolution processing is performed to obtain the output feature of the current cross-stage multi-scale edge feature fusion module CSP-MEFF.
[0051] Specifically, the features input into the cross-stage multi-scale edge feature fusion module CSP-MEFF are optimized for the gradient flow through the cross-stage feature interaction mechanism. The formula is as follows:
[0052] ;
[0053] where Concat() is the concatenation operation in the channel dimension, X is the feature map input into the current cross-stage multi-scale edge feature fusion module CSP-MEFF, and Conv is the convolutional layer. Denotes the addition operation for k = 1 to n, Split k Is the k-th separated feature, MEFF k Denotes the processing of the k-th branch within the multi-scale edge feature fusion module MEFF module.
[0054] Each multi-scale edge feature fusion module MEFF is composed of three parts: multi-scale feature encoding, edge enhancer, and cross-scale feature fusion. Through the processing flow of the three parts, the damaged edges of the input features are strengthened. By jointly optimizing local context awareness and high-frequency edge feature extraction, the model's ability to represent details in complex scenes is improved.
[0055] Specifically, each multi-scale edge feature fusion module MEFF includes multiple parallel edge feature processing layers and a single convolutional layer. Each edge feature processing layer includes an adaptive average pooling layer, a double convolutional layer, an upsampling layer, and an edge enhancer connected in sequence. The features input to each multi-scale edge feature fusion module MEFF are processed through multi-scale encoding and edge feature processing of multiple edge feature processing layers to obtain multiple edge enhancement features. The multiple edge enhancement features are concatenated with the local features extracted by the single convolutional layer to obtain a concatenated feature. The concatenated feature is further convolved to obtain the output feature of the current multi-scale edge feature fusion module MEFF, and this output feature is a multi-scale fusion feature.
[0056] The cross-scale feature fusion of each multi-scale edge feature fusion module MEFF in this scheme refers to: concatenating the edge enhancement features output by the multiple edge feature processing layers with the local features extracted by the single convolutional layer along the channel dimension, and finally eliminating redundancy and enhancing feature consistency through the convolutional layer to form a multi-scale fusion feature.
[0057] It should be noted that the multi-scale feature encoding of each multi-scale edge feature fusion module MEFF in this scheme is composed of the adaptive average pooling layer and the double convolutional layer of the multiple edge feature processing layers. The feature maps input to each multi-scale edge feature fusion module MEFF are parallelly input into the adaptive average pooling layers of multiple edge feature processing layers and downsampled to a preset scale respectively to capture local context information under different receptive fields. The features after downsampling are input into the double convolutional layer to further extract scale-related semantic features. And in order to restore the spatial resolution, the features of each edge feature processing layer are upsampled to the original size through bilinear interpolation to ensure the spatial alignment of multi-scale features.
[0058] Furthermore, the double convolutional layer in each edge feature processing layer includes two convolutional layers, and each convolutional layer is composed of a 3×3 convolution, batch normalization, and ReLU activation.
[0059] To strengthen the representation of the high-frequency edge information of the features obtained through multi-scale encoding, the features input into the edge enhancer generate low-frequency features after average pooling, and then the features input into the edge enhancer are subtracted based on the low-frequency features through element-wise subtraction to obtain high-frequency features. The high-frequency features are then processed through convolution to obtain the edge feature response values, and the edge feature response values are added to the features input into the edge enhancer to obtain the edge-enhanced features of the current layer's edge feature processing layer.
[0060] Specifically, the low-frequency features after average pooling can suppress low-frequency information and filter out high-frequency noise; the high-frequency features are separated from the original features through element-wise subtraction operations, so that the high-frequency features focus on edges and texture details; then the edge feature response values are obtained through convolution processing to further optimize the response intensity of the edge features. The mathematical expression for this edge enhancer is as follows:
[0061] ;
[0062] where are the features obtained through upsampling in different edge feature processing layers, AvgPool is the average pooling operation, Conv is the convolutional layer, and EdgeEnhancer is the edge enhancer.
[0063] The multi-scale edge feature fusion module of this solution captures multi-granularity context through adaptive pooling and combines high-frequency residual calculation to strengthen edge localization to obtain multi-scale encoded features, realizing complementary enhancement of details and semantics, greatly enhancing the extraction of road damage contour features. At the same time, due to the design of parallel multi-scale edge feature processing layers and shared convolutional parameters, while improving feature diversity, the computational complexity is controlled.
[0064] In addition, it should be noted that the feature map is input into the cross-stage part - multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage part - multi-scale edge feature fusion module. The hierarchical features of the last cross-stage part - multi-scale edge feature fusion module are input into the AIFI module of the RT-DETR basic framework and then input into the rectangular context fusion pyramid network together with other hierarchical features.
[0065] In a specific embodiment, when the cross-stage part - multi-scale edge feature fusion backbone network includes four cross-stage part - multi-scale edge feature fusion modules, the fourth cross-stage part - multi-scale edge feature fusion module is input into the AIFI module of the RT-DETR basic framework for further encoding, and the encoded features, together with the hierarchical features output by the second cross-stage part - multi-scale edge feature fusion module and the third cross-stage part - multi-scale edge feature fusion module, are used as the input features of the rectangular context fusion pyramid network.
[0066] As Figure 4 shown, it is a schematic structural diagram of the rectangular context fusion pyramid network RTFPN of this solution. As Figure 4 shown, the rectangular context fusion pyramid network RTFPN includes a pyramid context extraction module, a rectangular self-correction module, a dynamic interpolation fusion unit, a multi-scale fusion module, and a reparameterized convolution module. The features input into the rectangular context fusion pyramid network RTFPN generate corresponding cross-layer features through the pyramid context extraction module, and each obtains intermediate features through the rectangular self-correction module. The cross-layer features at each level are input into the multi-scale fusion module at the corresponding level to obtain enhanced features, and the intermediate features at the current level and the enhanced features at the previous level are input into the dynamic interpolation fusion unit together. The features output by the dynamic interpolation fusion unit are input into the multi-scale fusion module at the next level. The output value of the multi-scale fusion module at the first level from top to bottom is input into the rectangular self-correction module at the second level; the enhanced features at the current level are concatenated with the features after convolution of the output features at the next level, and the concatenated features are processed by the reparameterized convolution module to obtain the output features at the current level; among them, the enhanced features at the last level from top to bottom are concatenated with the enhanced features at the previous level after convolution.
[0067] Specifically, in the step of "the features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage", the cross-layer features generated by the pyramid context extraction module are used as context descriptors for constraints for subsequent intermediate interaction stages and ultimate optimization stages.
[0068] As Figure 4 shown, the context fusion pyramid network of this solution extracts multi-granularity context features of the features input into the rectangular context fusion pyramid network through parallel asymmetric adaptive average pooling operations, and uses a cascaded rectangular self-correction module to achieve cross-level multi-granularity context feature channel compression and information aggregation, and then performs a separation operation to obtain corresponding cross-layer features, while reducing the computational complexity and retaining fine-grained damage features.
[0069] In some specific embodiments, the number of features input into the rectangular context fusion pyramid network RTFPN is 3. The pyramid context extraction module includes parallel 3x3, 5x5, and 7x5 adaptive average pooling, and each adaptive pooling is cascaded with a rectangular self-correction module. The calculation formula is as follows:
[0070] ;
[0071] Among them, AdaAvgPool is adaptive average pooling, RCA is residual channel attention, γ is a learnable layer scaling parameter, Norm is batch normalization, MLP is a multi-layer perceptron, and Spilt is a splitting operation.
[0072] The rectangular self-correction module includes a residual channel attention layer, a batch normalization layer, and a multi-layer perceptron layer connected in sequence. The features input into the rectangular self-correction module pass through the residual channel attention layer, the batch normalization layer, and the multi-layer perceptron layer in sequence, and then are added to the original features input into the rectangular self-correction module to obtain the output features of the rectangular self-correction module.
[0073] In some embodiments, the dynamic interpolation fusion unit includes a parallel local perception branch and a global association branch. The local perception branch corresponds to high-level features, and the global association branch corresponds to low-level features. The local perception branch includes a bilinear interpolation upsampling layer and a convolutional layer connected in sequence. The high-level features input into the dynamic interpolation fusion unit pass through the bilinear interpolation upsampling process of the bilinear interpolation upsampling layer and the convolutional process of the convolutional layer in sequence to obtain interpolation features, and the interpolation features are added to the low-level features to obtain the output features of the current dynamic interpolation fusion unit.
[0074] Specifically, the dynamic interpolation fusion unit of this solution first upsamples the high-level features to the low-level resolution by bilinear interpolation, and applies convolution to the features to extract local details to obtain interpolation features. The interpolation features are added to the low-level features of the global association branch to obtain the output of the dynamic interpolation fusion unit. The specific formula is as follows:
[0075] ;
[0076] Among them, Input1 corresponds to the low-level features input by the global association branch, Input2 corresponds to the high-level features input by the local perception branch, Interpolate() is bilinear interpolation upsampling, which is used to keep the channel numbers aligned, and Conv is the convolutional layer.
[0077] In the embodiments of this solution, the intermediate features of the current level are used as high-level features, and the enhanced features of the previous level are used as low-level features and are input into the dynamic interpolation fusion unit for processing.
[0078] The multi-scale fusion module of this solution also includes a parallel high-level feature branch and a low-level feature branch. The high-level feature branch includes a convolutional layer, an activation function, and a bilinear interpolation upsampling layer. The high-level features input into the high-level feature branch pass through the convolutional layer, the activation function, and the bilinear interpolation upsampling layer in sequence to obtain weight features, and the weight features are fused with the low-level features according to the weights to obtain the output features of the multi-scale fusion module.
[0079] Specifically, the multi-scale fusion module of this solution performs a convolution operation on the high-level features and activates them with the H-Sigmoid function to generate spatial attention weights. After that, it performs bilinear interpolation upsampling to obtain weighted features. The weighted features are fused with the low-level features according to the weights to strengthen the features input to the current multi-scale fusion module. The specific formula is as follows: ;
[0080] Among them, Input1 corresponds to the low-level features input by the low-level feature branch, Input2 corresponds to the high-level features input by the high-level feature branch, Interpolate() is bilinear interpolation upsampling, which is used to keep the number of channels aligned, Conv is the convolutional layer, and H-Sigmoid is the Hard-sigmoid function.
[0081] In the embodiments of this solution, the cross-layer features of each level are input into the multi-scale fusion module as high-level features, and the features output by the dynamic interpolation fusion unit of the previous level are input into the multi-scale fusion module as low-level features.
[0082] The dynamic interpolation fusion unit and multi-scale fusion module in this rectangular context fusion pyramid network are co-optimized through the differentiable bilinear interpolation algorithm and learning-based weight coefficients to dynamically adjust the scale transformation process of the feature map and improve the alignment accuracy of the crack edge.
[0083] In some embodiments, the features output from this rectangular context fusion pyramid network are sequentially input into the decoder after being selected by IOU-aware query to output the decoding result as the output features.
[0084] In step S3, the training set is input into the road damage detection model based on the improved RT-DETR model for training. TensorBoard is used to draw the loss function and accuracy curves in real time. By observing these curves, identify whether the model has overfitting or underfitting phenomena. According to the changing trend of the curves, adjust the training parameters such as the learning rate and batch size to ensure that the model can steadily converge to the best state and obtain the optimal road damage detection model.
[0085] In some embodiments, after training is completed, the obtained optimal model is tested on the test set. The model is comprehensively evaluated through six evaluation indicators: recall rate, precision rate, mAP50 value, mAP50-95 value, FPS, and inference time to verify its detection ability in unknown road damage scenarios.
[0086] Embodiment 2
[0087] Based on the same concept, this solution provides an application method of a road damage detection model based on the improved RT-DETR model, that is, a road damage detection method based on the improved RT-DETR model, including the following steps:
[0088] The road image is input into the road damage detection model based on the improved RT-DETR model obtained by training in the first embodiment, and the damage type is output, where the damage type is selected from any one of the six damage types of longitudinal crack, transverse crack, crocodile crack, inclined crack, repair, and pothole.
[0089] Embodiment III
[0090] This embodiment also provides an electronic device. Refer to Figure 5 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any of the above embodiments of the application method and construction method of the road damage detection model based on the improved RT-DETR model.
[0091] Specifically, the above processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0092] Among them, the memory 404 may include a mass memory 404 for data or instructions. The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0093] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any of the above application methods and construction methods of the road damage detection model based on the improved RT-DETR model:
[0094] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0095] The transmission device 406 may be used to receive or send data via a network. Specific examples of the above network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0096] The input / output device 408 is used to input or output information. In this embodiment, the input information may be an original road damage image, etc., and the output information may be a road damage detection model based on the improved RT-DETR model, etc.
[0097] Optionally, in this embodiment, the above-mentioned processor 402 may be configured to execute the following steps through a computer program:
[0098] S1: Obtain the original road damage images and label the damage types of each original road damage image to form a training set;
[0099] S2: Construct a road damage detection model based on the improved RT-DETR model:
[0100] The backbone network and neck network of the RT-DETR basic framework of the road damage detection model are replaced with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network;
[0101] The cross-stage partial-multi-scale edge feature fusion backbone network includes a plurality of cascaded convolution modules and a cross-stage partial-multi-scale edge feature fusion module, and each cross-stage partial-multi-scale edge feature fusion module includes a plurality of multi-scale edge feature fusion modules for cross-stage feature interaction; the rectangular context fusion pyramid network adopts a three-level recursive architecture built in sequence by a primary fusion stage, an intermediate interaction stage, and an ultimate optimization stage;
[0102] The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module pass through the AIFI module of the RT-DETR basic framework and are input into the rectangular context fusion pyramid network together with other hierarchical features; the features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage. The cross-layer features of different levels are strengthened in correlation in the intermediate interaction stage by using a dynamic interpolation fusion unit and a multi-scale fusion module to generate corresponding strengthened features. The strengthened features are subjected to feature fusion by using a reparameterized convolution module in the ultimate optimization stage to obtain output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results;
[0103] S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train the road damage detection model based on the improved RT-DETR model.
[0104] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and alternative embodiments, and will not be elaborated herein.
[0105] Generally, various embodiments can be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.
[0106] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components that are configured to execute the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow in the figure can represent a program step, or interconnected logical circuits, boxes, and functions, or a combination of program steps and logical circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.
[0107] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0108] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as a limitation to the scope of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for constructing a road damage detection model based on an improved RT-DETR model, characterized in that: The following steps are involved: S1: Obtain original road damage images and annotate the damage type of each original road damage image to form a training set; S2: Construct a road damage detection model based on the improved RT-DETR model: The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network; The cross-stage partial-multi-scale edge feature fusion backbone network includes a plurality of convolution modules and a cross-stage partial-multi-scale edge feature fusion module arranged in cascade, and each cross-stage partial-multi-scale edge feature fusion module includes a plurality of multi-scale edge feature fusion modules for cross-stage feature interaction; The rectangular context fusion pyramid network adopts a three-level recursive architecture consisting of a primary fusion stage, an intermediate interaction stage, and a final optimization stage. The features input into the cross-stage partial-multi-scale edge feature fusion module are processed by the convolution layer and the separation layer in turn to obtain multiple separation features. The multiple separation features are input into the multi-level cascade-connected multi-scale edge feature fusion module, and the output features of each multi-scale edge feature fusion module are added to the separation features and then convolved to obtain the output features of the current cross-stage partial-multi-scale edge feature fusion module. The formula is as follows: ; Concat() is a concatenation operation on the channel dimension, X is the feature map of the input current cross-stage part-multi-scale edge feature fusion module CSP-MEFF, Conv is a convolutional layer, Indicates the addition operation for k=1~n, Split k is the kth separated feature, MEFF k Indicates that the kth branch is processed in the multi-scale edge feature fusion module MEFF module; The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are input into the rectangular context fusion pyramid network together with other hierarchical features after passing through the AIFI module of the RT-DETR basic framework; the features input into the rectangular context fusion pyramid network are used to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage, and the cross-layer features of different levels are strengthened in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are fused using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results; S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train and obtain the road damage detection model based on the improved RT-DETR model.
2. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: The road damage detection model based on the improved RT-DETR model includes a sequentially connected cross-stage partial-multi-scale edge feature fusion backbone network, an AIFI module, a rectangular context fusion pyramid network and a decoder.
3. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: Each multi-scale edge feature fusion module includes multiple layers of edge feature processing layers and a single convolution layer in parallel, wherein each edge feature processing layer includes an adaptive average pooling layer, a double convolution layer, an upsampling layer and an edge enhancer connected in sequence, and the features input to each multi-scale edge feature fusion module are subjected to multi-scale encoding and edge feature processing by the multiple layers of edge feature processing layers to obtain multi-layer edge enhancement features, and the multi-layer edge enhancement features are spliced with the local features extracted by the single convolution layer to obtain spliced features, and the spliced features are convoluted again to obtain the output features of each current multi-scale edge feature fusion module.
4. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 3, characterized in that: The features input to the edge enhancer are average pooled to generate low-frequency features, and then the features input to the edge enhancer are subtracted by element-by-element subtraction based on the low-frequency features to obtain high-frequency features. The high-frequency features are then convolved to obtain edge feature response values. The edge feature response values and the features input to the edge enhancer are added to obtain the edge enhancement features of the edge feature processing layer of the current layer.
5. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: The rectangular context fusion pyramid network includes a pyramid context extraction module, a rectangular self-correction module, a dynamic interpolation fusion unit, a multi-scale fusion module and a reparameterized convolution module. The features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module, and each of them is passed through the rectangular self-correction module to obtain intermediate features. The cross-layer features of each level are input into the multi-scale fusion module of the corresponding layer to obtain enhanced features, and the intermediate features of the current level are input into the dynamic interpolation fusion unit together with the enhanced features of the previous level, and the features output by the dynamic interpolation fusion unit are input into the multi-scale fusion module of the next level, and the output value of the multi-scale fusion module at the first level from top to bottom is input into the rectangular self-correction module of the second level; the enhanced features of the current level are spliced with the output features of the next level after convolution, and the spliced features are processed by the reparameterized convolution module to obtain the output features of the current level; wherein the enhanced features at the last level from top to bottom are spliced with the enhanced features of the previous level after convolution.
6. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 5, characterized in that: The dynamic interpolation fusion unit includes parallel local perception branches and global association branches, wherein the local perception branches correspond to high-level features, and the global association branches correspond to low-level features. The local perception branches include bilinear interpolation upsampling layers and convolution layers connected in sequence. The high-level features input into the dynamic interpolation fusion unit are sequentially processed by bilinear interpolation upsampling layers and convolution layers to obtain interpolation features. The interpolation features are added to the low-level features to obtain the output features of the current dynamic interpolation fusion unit.
7. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 5, characterized in that: The multi-scale fusion module also includes parallel high-level feature branches and low-level feature branches. The high-level feature branch includes convolution layer, activation function and bilinear interpolation upsampling layer. The high-level features input into the high-level feature branch are successively passed through convolution layer, activation function and bilinear interpolation upsampling layer to obtain weighted features. The weighted features are fused with the low-level features according to the weights to obtain the output features of the multi-scale fusion module.
8. An application method of a road damage detection model based on an improved RT-DETR model, characterized in that: The following steps are involved: The road image is input into the road damage detection model based on the improved RT-DETR model constructed by the method for constructing a road damage detection model based on the improved RT-DETR model described in any one of claims 1 to 7, and the damage type is output, wherein the damage type is selected from any one of the six damage types of longitudinal cracks, transverse cracks, crocodile cracks, oblique cracks, repairs and potholes.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a method for constructing a road damage detection model based on an improved RT-DETR model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Road rapid damage detection method and system based on multi-head attention mechanism
CN117132956A
Road damage detection method based on improved YOLOv8
CN118521869A
Vehicle detection method and system based on unmanned aerial vehicle and improved YOLO algorithm
CN119516412A
Lightweight road damage detection method based on RT-DETR improvement
CN119784759A