Construction method and application of road damage detection model based on improved RT-DETR model
By adopting a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network in the RT-DETR model, the problems of insufficient feature expression ability and weak edge feature characterization in road damage detection are solved, and a higher accuracy of road damage detection is achieved.
Patent Information
- Application Number
- CN202510443723.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the detection of road damage, traditional RT-DETR models have problems such as insufficient feature expression ability, weak edge feature characterization, and insufficient identification of complex backgrounds and uncertainties.
Cross-stage partial-multi-scale edge feature fusion backbone network (CSP-MEFFN) and rectangular context fusion pyramid network (RCFPN) are used to replace the backbone network and neck network of the RT-DETR model, enhancing feature extraction and fusion capabilities.
The model's sensitivity to road damage edges and its ability to identify targets in complex contexts has been significantly improved, thereby improving the accuracy of road damage detection.
Smart Images

Figure CN119962404A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to target detection technology in the field of computer technology, and in particular to a method for constructing and applying a road damage detection model based on an improved RT-DETR model. Background Art
[0002] As a core technology in the field of smart transportation and infrastructure maintenance, road damage detection plays a key role in road safety assessment, preventive maintenance decisions and full life cycle management.
[0003] Road damage detection technology can be divided into two categories: manual detection method and image recognition method. Conventional manual detection methods generally rely on manual inspection and semi-automatic detection equipment, but they have inherent defects such as low detection efficiency and large subjective evaluation deviation. With the rapid development of computer computing, machine learning, and deep learning technology, image recognition methods have received more and more attention. Patent No. 202311220593.8 proposed a highway rapid damage detection method based on a multi-head attention mechanism. This method achieved a high accuracy in road damage detection by introducing multi-head attention in the Resnet-18 backbone network and fusing multi-layer sampling results; Patent No. 202410706333.X proposed a road damage detection method based on improved YOLOv8. This method improves the accuracy of road damage detection by introducing LSKA attention and designing a focused diffusion pyramid network FDPN. The above methods can effectively realize the road state classification function, but in actual application, there are challenges such as global context modeling (understanding the overall scene) and local detail encoding (accurately identifying boundaries and subtle features).
[0004] RT-DETR (Real-Time Detection Transformer) is an end-to-end real-time target detection model based on the Transformer architecture. It significantly enhances the target detection performance by combining the local feature extraction advantages of CNN with the global relationship modeling capabilities of Transformer. Its dynamic label assignment strategy improves training efficiency through an adaptive matching mechanism. Combined with the real-time optimized decoder design, it reduces computational complexity while maintaining end-to-end detection characteristics. It is considered to have great development potential in road damage detection. However, since the Resnet18 backbone network used in the traditional RT-DETR model only relies on simple cross-layer connections for hierarchical feature fusion, it is difficult to achieve effective aggregation of multi-scale semantic information, resulting in weak road damage edge feature representation capabilities. At the same time, the feature fusion network CCFM (Cross-Scale Feature Fusion Module) of the traditional RT-DETR uses simple splicing or addition operations, which makes it unable to retain the slight texture differences of the shallow features of the road damage target. At the same time, it is easy to be disturbed by the noise of adjacent background features when the mid- and high-level features are fused, resulting in semantic degradation of the target features. In addition, the damage feature representation involved in the road damage detection scenario is highly complex, and the detection accuracy is also affected by uncertain factors such as weather and similar backgrounds, which is technically difficult. At present, there is no relevant research on the implementation of road damage detection based on the RT-DETR model. The reason is that the Resnet18 feature extraction backbone network and CCFM feature fusion neck network of the traditional RT-DETR model are not sufficient to provide high efficiency and accuracy.
[0005] Therefore, it is of great research significance to design a road damage detection model based on the improved RT-DETR model by constructing a new feature extraction backbone and feature fusion neck network. Summary of the invention
[0006] The embodiment of the present application provides a method for constructing and applying a road damage detection model based on an improved RT-DETR model. The basic framework of the RT-DETR model is improved, and a cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN is adopted as the backbone network to significantly improve the model's sensitivity to road damage edges; and a rectangular context fusion pyramid network RCFPN is designed as a neck network to achieve dynamic adaptive fusion of shallow high-frequency details, mid-level edge structures, and deep semantic features, thereby improving the model's context perception and ability to recognize targets in complex backgrounds, thereby improving the accuracy of traditional RT-DETR models for road damage detection.
[0007] In a first aspect, an embodiment of the present application provides a method for constructing a road damage detection model based on an improved RT-DETR model, comprising the following steps: S1: Obtain original road damage images and annotate the damage type of each original road damage image to form a training set; S2: Construct a road damage detection model based on the improved RT-DETR model: The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network; The cross-stage partial-multi-scale edge feature fusion backbone network includes multiple convolution modules and cross-stage partial-multi-scale edge feature fusion modules in cascade arrangement, and each cross-stage partial-multi-scale edge feature fusion module includes multiple cross-stage feature interaction multi-scale edge feature fusion modules; the rectangular context fusion pyramid network adopts a three-level recursive architecture constructed in sequence by the primary fusion stage, the intermediate interaction stage and the ultimate optimization stage; The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are input into the rectangular context fusion pyramid network together with other hierarchical features after passing through the AIFI module of the RT-DETR basic framework; the features input into the rectangular context fusion pyramid network are used to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage, and the cross-layer features of different levels are strengthened in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are fused using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results; S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train and obtain the road damage detection model based on the improved RT-DETR model.
[0008] In the second aspect, an embodiment of the present application provides an application method of a road damage detection model based on an improved RT-DETR model, comprising the following steps: inputting a road image into a road damage detection model based on an improved RT-DETR model obtained by a construction method of a road damage detection model based on an improved RT-DETR model, and outputting a damage type, wherein the damage type is selected from any one of six damage types: longitudinal cracks, transverse cracks, crocodile cracks, oblique cracks, repairs and potholes.
[0009] The main contributions and innovations of the present invention are as follows: 1. The cross-stage partial-multiscale edge feature fusion network designed by the present invention replaces the original backbone network of the RT-DETR model, and effectively enhances the feature expression capability based on the collaborative optimization mechanism of multi-scale context dynamic aggregation and high-frequency edge residual enhancement, thereby effectively solving the local coding problem.
[0010] 2. The present invention designs a rectangular context fusion pyramid network to replace the neck network of the RT-DETR model, so as to realize the dynamic adaptive fusion of high-frequency details, mid-level edge structures and deep semantic features of the neck network, improve the global context modeling capability, and at the same time improve the performance of the model in focusing on important information and suppressing lightweight features (noise introduced by external factors such as lighting and weather), thereby improving the accuracy of road damage detection.
[0011] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a flow chart of a method for constructing a road damage detection model based on an improved RT-DETR model according to an embodiment of the present application; Figure 2 It is a basic architecture diagram of a road damage detection model based on an improved RT-DETR model according to the present application; Figure 3 It is a basic architecture diagram of the cross-stage partial-multi-scale edge feature fusion backbone network according to the present application; Figure 4 It is a basic architecture diagram of a rectangular context fusion pyramid network according to the present application; Figure 5 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0014] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0015] Embodiment 1 like Figure 1 As shown, this scheme provides a method for constructing a road damage detection model based on an improved RT-DETR model, including the following steps: S1: Obtain original road damage images and annotate the damage type of each original road damage image to form a training set; S2: Construct a road damage detection model based on the improved RT-DETR model: The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network; The cross-stage partial-multi-scale edge feature fusion backbone network includes multiple convolution modules and cross-stage partial-multi-scale edge feature fusion modules in cascade arrangement, and each cross-stage partial-multi-scale edge feature fusion module includes multiple cross-stage feature interaction multi-scale edge feature fusion modules; the rectangular context fusion pyramid network adopts a three-level recursive architecture constructed in sequence by the primary fusion stage, the intermediate interaction stage and the ultimate optimization stage; The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are input into the rectangular context fusion pyramid network together with other hierarchical features after passing through the AIFI module of the RT-DETR basic framework; the features input into the rectangular context fusion pyramid network are used to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage, and the cross-layer features of different levels are strengthened in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are fused using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results; S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train and obtain the road damage detection model based on the improved RT-DETR model.
[0016] Step S1 further comprises the steps of: S11: Acquire original road damage image; S12: Marking the damage type of each original road damage image according to the road damage feature in each original road damage image to obtain a data set; S13: The original road damage images in the data set are processed and divided into a training set, a validation set and a test set.
[0017] Furthermore, the original road damage image of the present solution is obtained by capturing actual road damage features with a vehicle-mounted camera and expanding a public data set. The original road damage image includes at least one image region corresponding to at least one road damage feature.
[0018] Specifically, a USB camera is set on the front of the vehicle hood to capture the road with road damage, and the original road damage image is obtained by extracting frames from the captured video. Then, public data sets are introduced for expansion, including Crowdsensing-based Road Damage Detection Challenge, UAV-PDD 2023, CNRDD, and IRRDD.
[0019] Furthermore, the damage type of the present solution is any one of the six damage types of longitudinal cracks, transverse cracks, alligator cracks, oblique cracks, repairs and potholes.
[0020] Furthermore, the method of performing image processing on the original road damage images in the data set in this solution includes image screening and image enhancement, wherein the means of image screening include but are not limited to removing blurred and duplicate photos, and the means of image enhancement include but are not limited to rotation, scaling, brightness adjustment, etc.
[0021] Furthermore, labeling software (such as LabelImg software) is used to label the original road damage image with six types of damage.
[0022] In some embodiments, the present solution divides a dataset of more than 50,000 original road damage images into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0023] Regarding the road damage detection model based on the improved RT-DETR model constructed in step S2 of this solution, the structural framework diagram of the road damage detection model based on the improved RT-DETR model is as follows: Figure 2 shown.
[0024] The road damage detection model based on the improved RT-DETR model replaces the original backbone network with the cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN on the basis of the RT-DETR basic framework, replaces the neck network with the rectangular context fusion pyramid network RCFPN, and retains the AIFI module and decoder of the RT-DETR basic framework. Specifically, the road damage detection model based on the improved RT-DETR model includes a cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN, an AIFI module, a rectangular context fusion pyramid network and a decoder connected in sequence. The AIFI module and decoder of the road damage detection model follow the original RT-DETR basic framework, so no cumbersome introduction is given.
[0025] The following will focus on the cross-stage part of this solution - the multi-scale edge feature fusion backbone network CSP-MEFFN, such as Figure 3 Shown is a schematic diagram of the structure of the cross-stage part-multi-scale edge feature fusion backbone network CSP-MEFFN of this scheme.
[0026] The cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN of this scheme includes multiple convolution modules and a cross-stage partial-multi-scale edge feature fusion module CSP-MEFF arranged in cascade, wherein the convolution modules and the cross-stage partial-multi-scale edge feature fusion modules are alternately cascaded, that is, the features are input into the cross-stage partial-multi-scale edge feature fusion module CSP-MEFF for feature extraction after being processed by the convolution modules, and then input into the next convolution module for convolution processing again.
[0027] like Figure 3As shown, the cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN of the present solution includes 5 convolution modules and 4 cross-stage partial-multi-scale edge feature fusion modules CSP-MEFF, and each cross-stage partial-multi-scale edge feature fusion module CSP-MEFF outputs hierarchical features of the corresponding hierarchical level. In some specific embodiments, a three-channel RGB road damage image with a resolution of 640×640 is input to the cross-stage partial-multi-scale edge feature fusion backbone network CSP-MEFFN. The road damage image outputs 320×320×64 features after the first convolution module, and then outputs 160×160×127 features after the second convolution module. After the first cross-stage partial-multi-scale edge feature fusion module CSP-MEFF outputs 160×160×127 features, and then outputs 80×80×2 56 features, after the second cross-stage part-multi-scale edge feature fusion module CSP-MEFF outputs 80×80×256 features, and then after the fourth convolution module outputs 40×40×512 features, after the third cross-stage part-multi-scale edge feature fusion module CSP-MEFF outputs 40×40×512 features, and then after the witch-th convolution module outputs 20×20×512 features, and after the fourth cross-stage part-multi-scale edge feature fusion module CSP-MEFF outputs 20×20×512 features.
[0028] Specifically, with respect to each cross-stage partial-multi-scale edge feature fusion module CSP-MEFF: the cross-stage partial-multi-scale edge feature fusion module CSP-MEFF includes multiple cross-stage feature interactive multi-scale edge feature fusion modules MEFF, the features input into the cross-stage partial-multi-scale edge feature fusion module CSP-MEFF are successively processed by convolution layer and separation layer to obtain multiple separation features, the multiple separation features are input into the multi-scale edge feature fusion module MEFF with multiple layers of cascade connection, and the output features of each multi-scale edge feature fusion module MEFF and the separation features are added and then convolution is performed to obtain the output features of the current cross-stage partial-multi-scale edge feature fusion module CSP-MEFF.
[0029] Specifically, the features of the input cross-stage partial-multi-scale edge feature fusion module CSP-MEFF are optimized through the cross-stage feature interaction mechanism, and the formula is as follows: ; Concat() is a concatenation operation on the channel dimension, X is the feature map of the input current cross-stage part-multi-scale edge feature fusion module CSP-MEFF, Conv is a convolutional layer, Indicates the addition operation for k=1~n, Splitk is the kth separated feature, MEFF k Indicates that the kth branch is processed in the multi-scale edge feature fusion module MEFF module.
[0030] Each multi-scale edge feature fusion module MEFF consists of three parts: multi-scale feature encoding, edge enhancer and cross-scale feature fusion. The three-part processing flow realizes the enhancement of the damaged edges of the input features, and improves the model's ability to represent the details of complex scenes by collaboratively optimizing local context perception and high-frequency edge feature extraction.
[0031] Specifically, each multi-scale edge feature fusion module MEFF includes multiple layers of edge feature processing layers and a single convolution layer in parallel, wherein each edge feature processing layer includes an adaptive average pooling layer, a double convolution layer, an upsampling layer and an edge enhancer connected in sequence, and the features input to each multi-scale edge feature fusion module MEFF are subjected to multi-scale encoding and edge feature processing by the multiple layers of edge feature processing layers to obtain multi-layer edge enhancement features, and the multi-layer edge enhancement features are spliced with the local features extracted by the single convolution layer to obtain spliced features, and the spliced features are convoluted again to obtain the output features of each current multi-scale edge feature fusion module MEFF, and the output features are multi-scale fusion features.
[0032] The cross-scale feature fusion of each multi-scale edge feature fusion module MEFF in this scheme refers to: splicing the edge enhancement features output by the multi-layer edge feature processing layer with the local features extracted by the single convolution layer along the channel dimension, and finally eliminating redundancy and enhancing feature consistency through the convolution layer to form a multi-scale fusion feature.
[0033] It should be noted that the multi-scale feature encoding of each multi-scale edge feature fusion module MEFF of the present scheme is composed of an adaptive average pooling layer of multiple edge feature processing layers and a double-layer convolution layer. The feature maps input into each multi-scale edge feature fusion module MEFF are input into the adaptive average pooling layers of multiple edge feature processing layers in parallel and are downsampled to preset scales respectively to capture local context information under different receptive fields. The downsampled features are input into the double-layer convolution layer to further extract scale-related semantic features. In order to restore the spatial resolution, the features of each edge feature processing layer are upsampled to the original size by bilinear interpolation to ensure the spatial alignment of multi-scale features.
[0034] Furthermore, the double-layer convolutional layer in each edge feature processing layer includes two convolutional layers, and each convolutional layer consists of 3×3 convolution, batch normalization and ReLU activation.
[0035] In order to strengthen the representation of high-frequency edge information of the features obtained by multi-scale encoding, the features input to the edge enhancer are averaged and pooled to generate low-frequency features, and then the features input to the edge enhancer are subtracted by element-by-element subtraction based on the low-frequency features to obtain high-frequency features. The high-frequency features are then convolved to obtain edge feature response values, and the edge feature response values and the features input to the edge enhancer are added to obtain the edge enhancement features of the edge feature processing layer of the current layer.
[0036] Specifically, the low-frequency features after average pooling can suppress low-frequency information and filter out high-frequency noise; the high-frequency features are separated from the original features through element-by-element subtraction operations so that the high-frequency features focus on edge and texture details; then the edge feature response value is obtained through convolution processing to further optimize the response strength of the edge feature. The mathematical expression of the edge enhancer is as follows: ; in It is the feature obtained by upsampling in different edge feature processing layers, AvgPool is the average pooling operation, Conv is the convolution layer, and EdgeEnhancer is the edge enhancer.
[0037] The multi-scale edge feature fusion module of this scheme captures multi-granularity contexts through adaptive pooling and combines high-frequency residual calculation to strengthen edge positioning to obtain multi-scale encoded features, thereby achieving complementary enhancement of details and semantics, greatly enhancing the extraction of road damage contour features. At the same time, due to the parallelization of multi-scale edge feature processing layers and the design of shared convolution parameters, the computational complexity is controlled while improving feature diversity.
[0038] In addition, it should be noted that the feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are passed through the AIFI module of the RT-DETR basic framework and then input into the rectangular context fusion pyramid network together with other hierarchical features.
[0039] In a specific embodiment, when the cross-stage partial-multi-scale edge feature fusion backbone network includes four cross-stage partial-multi-scale edge feature fusion modules, the fourth cross-stage partial-multi-scale edge feature fusion module is input into the AIFI module of the RT-DETR basic framework for further encoding, and the encoded features are used together with the hierarchical features output by the second cross-stage partial-multi-scale edge feature fusion module and the third cross-stage partial-multi-scale edge feature fusion module as input features of the rectangular context fusion pyramid network.
[0040] like Figure 4 The following is a schematic diagram of the structure of the rectangular context fusion pyramid network RTFPN of this solution. Figure 4 As shown, the rectangular context fusion pyramid network RTFPN includes a pyramid context extraction module, a rectangular self-correction module, a dynamic interpolation fusion unit, a multi-scale fusion module and a reparameterized convolution module. The features input into the rectangular context fusion pyramid network RTFPN generate corresponding cross-layer features through the pyramid context extraction module, and each is passed through the rectangular self-correction module to obtain intermediate features. The cross-layer features of each level are input into the multi-scale fusion module of the corresponding layer to obtain enhanced features, and the intermediate features of the current level are input into the dynamic interpolation fusion unit together with the enhanced features of the previous level, and the features output by the dynamic interpolation fusion unit are input into the multi-scale fusion module of the next level. The output value of the multi-scale fusion module at the first level from top to bottom is input into the rectangular self-correction module of the second level; the enhanced features of the current level are spliced with the output features of the next level after convolution, and the spliced features are processed by the reparameterized convolution module to obtain the output features of the current level; wherein the enhanced features at the last level from top to bottom are spliced with the enhanced features of the previous level after convolution.
[0041] Specifically, in the step of "inputting the features in the rectangular context fusion pyramid network to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage", the cross-layer features generated by the pyramid context extraction module are used as context descriptors for the constraints in the subsequent intermediate interaction stage and the ultimate optimization stage.
[0042] like Figure 4 As shown, the context fusion pyramid network of this scheme extracts multi-granular context features of the features input into the rectangular context fusion pyramid network through parallel asymmetric adaptive average pooling operations, and uses cascaded rectangular self-correction modules to achieve cross-level multi-granularity context feature channel compression and information aggregation, and then performs separation operations to obtain corresponding cross-layer features, thereby reducing the computational complexity while retaining fine-grained damage features.
[0043] In some specific embodiments, the number of features input into the rectangular context fusion pyramid network RTFPN is 3, the pyramid context extraction module includes 3x3, 5x5, and 7x5 adaptive average pooling in parallel, and each adaptive pooling is cascaded with a rectangular self-correction module, and the calculation formula is as follows: ; Among them, AdaAvgPool is adaptive average pooling, RCA is residual channel attention, γ is a learnable layer scaling parameter, Norm is batch normalization, MLP is a multi-layer perception layer, and Spilt is a separation operation.
[0044] The rectangular self-correction module includes a residual channel attention layer, a batch normalization layer and a multi-layer perception layer connected in sequence. The features input into the rectangular self-correction module are sequentially passed through the residual channel attention layer, the batch normalization layer and the multi-layer perception layer, and then added to the original features input into the rectangular self-correction module to obtain the output features of the rectangular self-correction module.
[0045] In some embodiments, the dynamic interpolation fusion unit includes parallel local perception branches and global association branches, wherein the local perception branches correspond to high-level features, and the global association branches correspond to low-level features. The local perception branches include bilinear interpolation upsampling layers and convolution layers connected in sequence. The high-level features input into the dynamic interpolation fusion unit are sequentially processed by bilinear interpolation upsampling layers and convolution layers to obtain interpolation features. The interpolation features are added to the low-level features to obtain the output features of the current dynamic interpolation fusion unit.
[0046] Specifically, the dynamic interpolation fusion unit of this solution first upsamples the high-level features to the low-level resolution through bilinear interpolation, and applies convolution to the features to extract local details to obtain interpolation features. The interpolation features are added to the low-level features of the global associated branch to obtain the output of the dynamic interpolation fusion unit. The specific formula is as follows: ; Among them, Input1 corresponds to the low-level features of the global association branch input, Input2 corresponds to the high-level features of the local perception branch input, Interpolate() is a bilinear interpolation upsampling used to keep the number of channels aligned, and Conv is a convolutional layer.
[0047] In an embodiment of the present solution, the intermediate features of the current level are used as high-level features, and the enhanced features of the previous level are used as low-level features and are input into the dynamic interpolation fusion unit for processing.
[0048] The multi-scale fusion module of this scheme also includes parallel high-level feature branches and low-level feature branches, wherein the high-level feature branches include convolution layers, activation functions, and bilinear interpolation upsampling layers. The high-level features input into the high-level feature branches are sequentially passed through convolution layers, activation functions, and bilinear interpolation upsampling layers to obtain weighted features. The weighted features are fused with the low-level features according to the weights to obtain the output features of the multi-scale fusion module.
[0049] Specifically, the multi-scale fusion module of this scheme performs convolution operation on high-level features and H-Sigmoid function activation to generate spatial attention weights, and then obtains weight features through bilinear interpolation upsampling. The weight features are fused with low-level features according to the weights to strengthen the features input into the current multi-scale fusion module. The specific formula is as follows: ; Among them, Input1 corresponds to the low-level features of the underlying feature branch input, Input2 corresponds to the high-level features of the high-level feature branch input, Interpolate() is a bilinear interpolation upsampling used to keep the number of channels aligned, Conv is a convolutional layer, and H-Sigmoid is a Hard-sigmoid function.
[0050] In an embodiment of the present solution, the cross-layer features of each level are input into the multi-scale fusion module as high-level features, and the features output by the dynamic interpolation fusion unit of the previous level are input into the multi-scale fusion module as low-level features.
[0051] The dynamic interpolation fusion unit and multi-scale fusion module in the rectangular context fusion pyramid network dynamically adjust the scale transformation process of the feature map through the collaborative optimization of the differentiable bilinear interpolation algorithm and the learning weight coefficient, thereby improving the alignment accuracy of the crack edge.
[0052] In some embodiments, the features output from the rectangular context fusion pyramid network are sequentially input into the decoder after IOU-aware query selection to output the decoding results as output features.
[0053] In step S3, the training set is input into the road damage detection model based on the improved RT-DETR model for training. The loss function and accuracy curve are plotted in real time using TensorBoard. By observing these curves, it is possible to identify whether the model is overfitting or underfitting. According to the changing trend of the curves, the training parameters such as learning rate and batch size are adjusted to ensure that the model can steadily converge to the optimal state and obtain the optimal road damage detection model.
[0054] In some embodiments, after training is completed, the optimal model is tested on the test set, and the model is comprehensively evaluated through six evaluation indicators: recall rate, precision rate, mAP50 value, mAP50-95 value, FPS, and inference time to verify its detection capability in unknown road damage scenarios.
[0055] Embodiment 2 Based on the same concept, this solution provides an application method of a road damage detection model based on an improved RT-DETR model, that is, a road damage detection method based on an improved RT-DETR model, comprising the following steps: The road image is input into the road damage detection model based on the improved RT-DETR model obtained through training in Example 1, and the damage type is output, wherein the damage type is selected from any one of the six damage types of longitudinal cracks, transverse cracks, alligator cracks, oblique cracks, repairs and potholes.
[0056] Embodiment 3 This embodiment also provides an electronic device, referring to Figure 5 , including a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the application method and construction method of the road damage detection model based on the improved RT-DETR model.
[0057] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0058] The memory 404 may include a large-capacity memory 404 for data or instructions. The memory 404 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 402.
[0059] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the application methods and construction methods of the road damage detection model based on the improved RT-DETR model in the above embodiments: Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0060] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0061] The input and output device 408 is used to input or output information. In this embodiment, the input information may be an original road damage image, etc., and the output information may be a road damage detection model based on an improved RT-DETR model, etc.
[0062] Optionally, in this embodiment, the processor 402 may be configured to perform the following steps through a computer program: S1: Obtain original road damage images and annotate the damage type of each original road damage image to form a training set; S2: Construct a road damage detection model based on the improved RT-DETR model: The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network; The cross-stage partial-multi-scale edge feature fusion backbone network includes multiple convolution modules and cross-stage partial-multi-scale edge feature fusion modules in cascade arrangement, and each cross-stage partial-multi-scale edge feature fusion module includes multiple cross-stage feature interaction multi-scale edge feature fusion modules; the rectangular context fusion pyramid network adopts a three-level recursive architecture constructed in sequence by the primary fusion stage, the intermediate interaction stage and the ultimate optimization stage; The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are input into the rectangular context fusion pyramid network together with other hierarchical features after passing through the AIFI module of the RT-DETR basic framework; the features input into the rectangular context fusion pyramid network are used to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage, and the cross-layer features of different levels are strengthened in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are fused using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results; S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train and obtain the road damage detection model based on the improved RT-DETR model.
[0063] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0064] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0065] Embodiments of the present invention can be implemented by computer software, which is executable by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros can be stored in any device readable data storage medium, and they include program instructions for performing specific tasks. Computer program products can include one or more computer executable components configured to perform embodiments when the program is running. One or more computer executable components can be at least one software code or a part thereof. In addition, at this point, it should be noted that any box of the logic flow in the figure can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored in physical media such as memory chips or storage blocks implemented in processors, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and data variants thereof, CDs. Physical media are non-transient media.
[0066] Those skilled in the art should understand that the technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0067] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for constructing a road damage detection model based on an improved RT-DETR model, characterized in that: The following steps are involved: S1: Obtain original road damage images and annotate the damage type of each original road damage image to form a training set; S2: Construct a road damage detection model based on the improved RT-DETR model: The road damage detection model replaces the backbone network and neck network of the RT-DETR basic framework with a cross-stage partial-multi-scale edge feature fusion backbone network and a rectangular context fusion pyramid network; The cross-stage partial-multi-scale edge feature fusion backbone network includes multiple convolution modules and cross-stage partial-multi-scale edge feature fusion modules in cascade arrangement, and each cross-stage partial-multi-scale edge feature fusion module includes multiple cross-stage feature interaction multi-scale edge feature fusion modules; the rectangular context fusion pyramid network adopts a three-level recursive architecture constructed in sequence by the primary fusion stage, the intermediate interaction stage and the ultimate optimization stage; The feature map is input into the cross-stage partial-multi-scale edge feature fusion backbone network for feature extraction and outputs the hierarchical features corresponding to each cross-stage partial-multi-scale edge feature fusion module. The hierarchical features of the last cross-stage partial-multi-scale edge feature fusion module are input into the rectangular context fusion pyramid network together with other hierarchical features after passing through the AIFI module of the RT-DETR basic framework; the features input into the rectangular context fusion pyramid network are used to generate corresponding cross-layer features through the pyramid context extraction module in the primary fusion stage, and the cross-layer features of different levels are strengthened in the intermediate interaction stage using the dynamic interpolation fusion unit and the multi-scale fusion module to generate corresponding enhanced features. The enhanced features are fused using the reparameterized convolution module in the ultimate optimization stage to obtain the output features, and the output features are input into the decoder of the RT-DETR basic framework to output the results; S3: Input the training set into the road damage detection model based on the improved RT-DETR model to train and obtain the road damage detection model based on the improved RT-DETR model.
2. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: The road damage detection model based on the improved RT-DETR model includes a sequentially connected cross-stage partial-multi-scale edge feature fusion backbone network, an AIFI module, a rectangular context fusion pyramid network and a decoder.
3. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: The features input into the cross-stage partial-multi-scale edge feature fusion module are processed by the convolution layer and the separation layer in sequence to obtain multiple separation features. The multiple separation features are input into the multi-level cascade-connected multi-scale edge feature fusion module, and the output features and separation features of each multi-scale edge feature fusion module are added and then convolution is performed to obtain the output features of the current cross-stage partial-multi-scale edge feature fusion module.
4. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: Each multi-scale edge feature fusion module includes multiple layers of edge feature processing layers and a single convolution layer in parallel, wherein each edge feature processing layer includes an adaptive average pooling layer, a double convolution layer, an upsampling layer and an edge enhancer connected in sequence, and the features input to each multi-scale edge feature fusion module are subjected to multi-scale encoding and edge feature processing by the multiple layers of edge feature processing layers to obtain multi-layer edge enhancement features, and the multi-layer edge enhancement features are spliced with the local features extracted by the single convolution layer to obtain spliced features, and the spliced features are convoluted again to obtain the output features of each current multi-scale edge feature fusion module.
5. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 4, characterized in that: The features input to the edge enhancer are average pooled to generate low-frequency features, and then the features input to the edge enhancer are subtracted by element-by-element subtraction based on the low-frequency features to obtain high-frequency features. The high-frequency features are then convolved to obtain edge feature response values. The edge feature response values and the features input to the edge enhancer are added to obtain the edge enhancement features of the edge feature processing layer of the current layer.
6. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 1, characterized in that: The rectangular context fusion pyramid network includes a pyramid context extraction module, a rectangular self-correction module, a dynamic interpolation fusion unit, a multi-scale fusion module and a reparameterized convolution module. The features input into the rectangular context fusion pyramid network generate corresponding cross-layer features through the pyramid context extraction module, and each of them is passed through the rectangular self-correction module to obtain intermediate features. The cross-layer features of each level are input into the multi-scale fusion module of the corresponding layer to obtain enhanced features, and the intermediate features of the current level are input into the dynamic interpolation fusion unit together with the enhanced features of the previous level, and the features output by the dynamic interpolation fusion unit are input into the multi-scale fusion module of the next level, and the output value of the multi-scale fusion module at the first level from top to bottom is input into the rectangular self-correction module of the second level; the enhanced features of the current level are spliced with the output features of the next level after convolution, and the spliced features are processed by the reparameterized convolution module to obtain the output features of the current level; wherein the enhanced features at the last level from top to bottom are spliced with the enhanced features of the previous level after convolution.
7. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 6, characterized in that: The dynamic interpolation fusion unit includes parallel local perception branches and global association branches, wherein the local perception branches correspond to high-level features, and the global association branches correspond to low-level features. The local perception branches include bilinear interpolation upsampling layers and convolution layers connected in sequence. The high-level features input into the dynamic interpolation fusion unit are sequentially processed by bilinear interpolation upsampling layers and convolution layers to obtain interpolation features. The interpolation features are added to the low-level features to obtain the output features of the current dynamic interpolation fusion unit.
8. The method for constructing a road damage detection model based on an improved RT-DETR model according to claim 6, characterized in that: The multi-scale fusion module also includes parallel high-level feature branches and low-level feature branches. The high-level feature branch includes convolution layer, activation function and bilinear interpolation upsampling layer. The high-level features input into the high-level feature branch are successively passed through convolution layer, activation function and bilinear interpolation upsampling layer to obtain weighted features. The weighted features are fused with the low-level features according to the weights to obtain the output features of the multi-scale fusion module.
9. An application method of a road damage detection model based on an improved RT-DETR model, characterized in that: The following steps are involved: The road image is input into the road damage detection model based on the improved RT-DETR model obtained by the construction method of the road damage detection model based on the improved RT-DETR model described in any one of claims 1 to 8, and the damage type is output, wherein the damage type is selected from any one of the six damage types of longitudinal cracks, transverse cracks, crocodile cracks, oblique cracks, repairs and potholes.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a method for constructing a road damage detection model based on an improved RT-DETR model according to any one of claims 1 to 8.
Citation Information
Patent Citations
Road rapid damage detection method and system based on multi-head attention mechanism
CN117132956A
Image tampering positioning method fusing multi-level multi-scale and boundary information
CN117893858A
Road damage detection method based on improved YOLOv8
CN118521869A
Automatic driving target detection method based on improved RT-DETR
CN119314144A
Vehicle detection method and system based on unmanned aerial vehicle and improved YOLO algorithm
CN119516412A
Cited By
Fabricated retaining wall defect identification method and system based on image identification
CN120339285A
Unmanned aerial vehicle target detection method based on feature fusion DINO
CN120544081A
An unmanned aerial vehicle target detection method based on feature fusion DINO
CN120544081B
River bank illegal building identification method based on multi-scale fusion
CN121121499A
Underwater target detection encoder, detection system and method
CN122024028A