A Road Damage Detection Method Based on Diffusion-Based Adaptive Feature Extraction
By adopting a diffusion-based adaptive feature extraction method in road damage detection, using deep residual neural network and denoising diffusion model, the problem of insufficient identification and high false alarm rate in the prior art is solved, and high accuracy and real-time road damage detection is achieved.
Patent Information
- Application Number
- CN202310424939.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-04-20
AI Technical Summary
In the road damage detection, the prior art has problems such as insufficient identification, high false alarm rate, fixed learning query, and difficulty in taking into account real-time and accuracy.
Adaptive feature extraction method based on diffusion is adopted to construct a deep residual neural network under multi-scale convolution with fusion indirect attention, and combined with a denoising diffusion model, adaptive feature extraction and object detection of different input samples are realized.
It improves the accuracy and robustness of road damage detection, enhances the target detection capability in complex road scenarios, and achieves real-time and efficient road damage detection.
Smart Images

Figure CN116630683B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a road damage detection method for adaptive feature extraction. Background Art
[0002] With the rapid development of urban construction and the continuous improvement of the highway network, China's highway transportation industry has begun to transform from a large-scale construction period to a continuous maintenance period, and the problems faced by highway maintenance management have become increasingly important. In order to sense and detect road diseases such as transverse cracks, longitudinal cracks, crocodile cracks, potholes, and spills on the road surface, so as to extend the service life of the road and reduce the overall maintenance cost.
[0003] However, at present, highway maintenance largely relies on manpower rather than machines, the method is relatively rigid, road inspection is somewhat dangerous, with low efficiency and high cost. And some sections have begun to use inspection vehicles equipped with cameras, but the recognition degree of road diseases is low, the classification is not clear enough, and the false alarm rate is relatively high. At the same time, in traditional road damage detection, there is a defect that they rely on a fixed set of learnable queries. Road damage detection needs to balance the issues of real-time and accuracy, and there are problems of multiple types of detection of road diseases and diversification of road scenes.
[0004] Therefore, how to identify road damage in real time with high precision is a technical problem that those skilled in the art need to solve currently. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the present invention provides a road damage detection method based on diffusion for adaptive feature extraction. For the feature extraction of the target object, an improved deep residual neural network is used to adaptively extract the most discriminative features for different input samples, the attention weight matrix is used to ensure the comprehensiveness of feature extraction, and the denoising diffusion model is used to improve the detection accuracy while ensuring the real-time detection, and improve the robustness and generalization ability of target detection under complex roads, which has practical application value.
[0006] The technical solutions adopted by the present invention to solve its technical problems include the following steps:
[0007] Step 1: Construct a deep residual neural network with multi-scale convolution with fused indirect attention, and use this deep residual neural network for feature extraction;
[0008] Step 2: In the training stage, different Gaussian noises are added to the ground truth box to obtain a noise box; the target image is input into the deep residual neural network to extract the deep feature representation and generate the corresponding multi-scale feature map;
[0009] Step 3: Construct an object detection network with sparse candidate boxes. Conditional on the depth features obtained in Step 2, crop the ROI features from the feature maps generated by the deep residual neural network. This object detection network is trained to predict the ground truth boxes without noise.
[0010] Step 4: In the inference stage, generate the object boxes by reversing the learned diffusion process, adjust the noise prior distribution to the learned distribution on the object boxes, and gradually refine the box predictions from the noise boxes to obtain the final result.
[0011] Step 5: After passing the image data to be detected through Steps 1 to 4, obtain the detection result through the optimized model to achieve object detection.
[0012] Preferably, Step 1 is specifically as follows:
[0013] The deep residual neural network with multi-scale convolution with fused indirect attention can automatically perform feature extraction according to the attention weight matrix A to obtain the features of the most discriminative regions. The attention weight matrices A of different layers are shown in Equation (1), and the attention weight vector a between the class label and the image sequence 0 is shown in Equation (2) as follows:
[0014] A = [a 0 ; a1; a2; …; a n (1)
[0015] a 0 = [a 0,0 , a 0,1 , a 0,2 , …, a 0,N (2)
[0016] The single convolutional layer in the residual learning module uses convolutional kernels of three sizes, 1*1, 3*3, and 5*5, at the same level.
[0017] Adopt the method of fusing indirect attention to ensure the comprehensiveness of feature extraction. Indirect attention includes the relationships between various features. The method of fusing indirect attention is shown in Equation (3).
[0018] Z indirect = Z l [Max(a indirect )] (3)
[0019] Find the index of the feature corresponding to the maximum weight in each vector through the indirect attention matrix a indirect . According to these indexes, slice and splice the original feature matrix Z l as the extracted indirect features. The indirect features, as a supplement to the direct features, improve the comprehensiveness of feature extraction.
[0020] Preferably, step 2 is specifically as follows:
[0021] During the training process, first construct a diffusion process from the ground truth boxes to the noise boxes, and then train the model to reverse this diffusion process; first fill some additional boxes into the original ground truth boxes so that all the boxes are totaled to a fixed number N_train, and then add Gaussian noise to the filled ground truth boxes, adding different Gaussian noises to obtain the optimal noise addition strategy.
[0022] Preferably, step 4 is specifically as follows:
[0023] The inference process is a denoising sampling process from the noise to the target boxes, starting from the boxes sampled from the Gaussian distribution and gradually refining the prediction; adopt a box update strategy to restore them by replacing the unexpected boxes with random boxes to obtain the final result.
[0024] The beneficial effects of the present invention are as follows:
[0025] The method of the present invention makes full use of the information of different target objects, can well adapt to different target objects, and can detect the types and damage degrees of the targets with high accuracy. Aiming at the multi-scene problems, detection real-time performance and high accuracy requirements in road damage detection, the object detection task is simulated as a denoising diffusion process from a noise box to a target box, improving the detection accuracy while ensuring the detection real-time performance. The present invention realizes reliable and efficient detection of the state of the highway, and can timely maintain the highway after identifying road diseases, providing a reference for the traffic maintenance department. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the deep residual neural network structure of the present invention.
[0027] Figure 2 It is a schematic diagram of the object detection network structure with sparse candidate boxes of the present invention.
[0028] Figure 3 It is the denoising diffusion model structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described below with reference to the drawings and embodiments.
[0030] In view of the problems in multi-scene road damage detection and the requirements for real-time detection and high accuracy, traditional object detection methods have deficiencies. The present invention proposes a road damage detection method based on a diffusion-based adaptive feature extraction method, which simulates the object detection task as a denoising diffusion process from a noisy box to a target box. In the training stage of the model, the real target box is continuously diffused into a random noise distribution, enabling the model to learn this noise modeling process. In the inference stage, the model gradually refines a set of randomly generated target boxes into the final prediction result in a progressive process, ultimately obtaining a more robust model for accurate object detection.
[0031] As Figure 1 shown, a road damage detection method based on diffusion-based adaptive feature extraction includes the following steps:
[0032] Step 1: Construct a deep residual neural network with multi-scale convolution incorporating fused indirect attention, and use this deep residual neural network for feature extraction;
[0033] Step 2: In the training stage, add different Gaussian noises to the ground truth boxes to obtain noisy boxes. Input the target image into this deep residual neural network. After the target image passes through the residual learning module with multi-scale convolution, extract the deep feature representation and generate the corresponding multi-scale feature maps;
[0034] Step 3: As Figure 2 shown, construct an object detection network with sparse candidate boxes. Using this deep feature as a condition, crop the ROI features from the feature maps generated by the deep residual neural network. This object detection network is trained to predict the ground truth boxes without noise;
[0035] Step 4: In the inference stage, generate the target boxes by reversing the learned diffusion process, which adjusts the noise prior distribution to the learned distribution on the target boxes, gradually refining the box predictions from the noisy boxes to obtain the final result.
[0036] Step 5: After passing the image data to be detected through Steps 1 to 4, give the detection result through the optimized denoising diffusion model to achieve object detection, as Figure 3 shown. Specific embodiments:
[0038] The present invention proposes a real-time and efficient road damage detection system, which simulates the object detection task as a denoising diffusion process from a noisy box to a target box. During the training phase of the model, the real target box is continuously diffused into a random noise distribution, enabling the model to learn this noise modeling process. During the inference phase, the model gradually refines a set of randomly generated target boxes into the final prediction result, ultimately obtaining a more robust model for accurate object detection. The main steps of the present invention are as follows:
[0039] Step 1: Construct a deep residual neural network structure with an adaptive feature extraction module integrating indirect attention, and use this convolutional neural network for feature extraction;
[0040] This step proposes a deep residual neural network structure with an adaptive feature extraction module integrating indirect attention. This network can automatically perform feature extraction according to the attention weight matrix A to obtain the features of the most discriminative regions. The attention weight matrices A of different layers are shown in Equation (1), and the attention weight vector a between the class label and the image sequence 0 is shown in Equation (2).
[0041] A = [a 0 ; a1; a2; …; a n (1)
[0042] a 0 = [a 0,0 , a 0,1 , a 0,2 , …, a 0,N (2)
[0043] For the feature extraction of the target object, the original residual neural network only considered increasing the network depth. In the residual learning module, 3*3 convolutional kernels were used to extract image features, so the extracted features were relatively single, making the expression of image information inaccurate. In order to utilize convolutional kernels of multiple scales, the single convolutional layer in the original residual learning module uses three types of convolutional kernels with sizes of 1*1, 3*3, and 5*5 at the same level. Convolutional kernels of different sizes have different receptive fields. Larger convolutional kernels focus on the extraction of global features, and smaller convolutional kernels can extract more local features. The improved residual neural network is more adaptable to the scale of the object in the image and expands the width of the network, effectively avoiding the occurrence of overfitting caused by the excessive depth of the network.
[0044] Since there are also subtle connections among the features in the image, a method of fusing indirect attention is proposed to ensure the comprehensiveness of feature extraction. Indirect attention contains the relationships among the features, and the features indirectly related to the main object play a huge role in improving the prediction accuracy of the model. The method of fusing indirect attention is shown in Equation (3).
[0045] Z indirect =Z l [Max(a indirect )] (3)
[0046] First, through the indirect attention matrix a indirect find the indices corresponding to the features with the maximum weights in each vector. According to these indices, slice and splice the original feature matrix Z l as the extracted indirect features. The indirect features, as a supplement to the direct features, improve the comprehensiveness of feature extraction.
[0047] Step 2, in the training stage, add different Gaussian noises to the ground truth boxes to obtain noisy boxes. Input the target image into this deep residual neural network. After the target image passes through the residual learning module with multi-scale convolutions, extract the deep feature representation and generate the corresponding multi-scale feature maps;
[0048] During the training process, first construct the diffusion process from the ground truth boxes to the noisy boxes, and then train the model to reverse this process. For modern object detection benchmarks, the number of instances of interest usually varies from image to image. Therefore, first fill some additional boxes into the original ground truth boxes so that all the boxes are totaled to a fixed number N_train. Then add Gaussian noises to the filled ground truth boxes, and different Gaussian noises can be added at this stage to obtain the optimal noise addition strategy.
[0049] Step 3, construct an object detection network with sparse candidate boxes. Conditional on this deep feature, crop the ROI features from the feature maps generated by the deep residual neural network. This detection network is trained to predict the ground truth boxes without noise;
[0050] This step proposes to construct an object detection network with sparse candidate boxes. To better represent some details of the object, introduce learnable proposal features and combine them with the ROI information extracted by the deep residual neural network. The input is an image, a set of proposal boxes, and proposal features. The proposal boxes and proposal features are learnable parameters. After the detection network is trained, it is used to predict the ground truth boxes without noise.
[0051] Step 4, in the inference stage, generate the target boxes by reversing the learned diffusion process. It adjusts the noise prior distribution to the learned distribution on the target boxes, gradually refines the box predictions from the noisy boxes, and obtains the final result.
[0052] The inference process is a denoising sampling process from noise to target bounding boxes. Starting from the bounding boxes sampled from a Gaussian distribution, the model gradually refines its predictions. In the inference stage, an optimal noise addition strategy is used to improve the learning performance and effect of the network. To make the inference better consistent with the training, a bounding box update strategy is adopted. By replacing the unexpected bounding boxes with random bounding boxes to recover them, the final result is obtained.
[0053] Step 5: After the image data to be detected goes through Steps 1 to 4, the optimized denoising diffusion model gives the detection result to achieve object detection.
[0054] The method of the present invention is a road damage detection method with stronger robustness and can adapt to the pavement technical conditions of various roads. It can efficiently and accurately identify road diseases, has strong generalization ability, and at the same time, the method design takes into account the time performance and can process target pictures or videos efficiently and in real time.
Claims
1. A road damage detection method based on diffusion-based adaptive feature extraction, characterized in that It includes the following steps: Step 1: Construct a deep residual neural network with multi-scale convolution under fused indirect attention, and use this deep residual neural network for feature extraction; Step 2: In the training stage, add different Gaussian noises to the ground truth boxes to obtain noisy boxes; input the target image into the deep residual neural network, extract the deep feature representation and generate corresponding multi-scale feature maps; Step 3: Construct an object detection network with sparse candidate boxes, and crop the ROI features from the feature maps generated by the deep residual neural network conditioned on the deep features obtained in Step 2; This object detection network is trained to predict the ground truth boxes without noise; Step 4: In the inference stage, generate the target boxes by reversing the learned diffusion process, adjust the noise prior distribution to the learned distribution on the target boxes, and gradually refine the box predictions from the noisy boxes to obtain the final result; Step 5: After passing the image data to be detected through Steps 1 to 4, obtain the detection result through the optimized model to achieve object detection.
2. The road damage detection method based on diffusion-based adaptive feature extraction according to claim 1, wherein The specific content of Step 1 is as follows: The deep residual neural network under multi-scale convolution with fused indirect attention can automatically perform feature extraction according to the attention weight matrix A to obtain the features of the most discriminative regions; the attention weight matrices A of different layers are shown in Equation (1), and the attention weight vector a between the class label and the image sequence 0 as shown in Equation (2): A = [a 0 ; a1; a2; …; a n (1) a 0 = [a 0,0 , a 0,1 , a 0,2 , …, a 0,N (2) For the single convolutional layer in the residual learning module, convolutional kernels of three sizes, 1*1, 3*3, and 5*5, are used at the same level; Adopt the method of fused indirect attention to ensure the comprehensiveness of feature extraction. The indirect attention contains the relationships between various features. The method of fused indirect attention is shown in Equation (3): Z indirect = Z l [Max(a indirect )] (3) Through the indirect attention matrix a indirect Find the index of the feature corresponding to the maximum weight in each vector. According to these indexes, slice and splice the original feature matrix Z l to obtain the extracted indirect features. The indirect features, as a supplement to the direct features, improve the comprehensiveness of feature extraction.
3. The road damage detection method based on diffusion-based adaptive feature extraction according to claim 1, characterized in that, The specific content of Step 2 is as follows: During the training process, first construct the diffusion process from the ground truth boxes to the noisy boxes, and then train the model to reverse this diffusion process; first fill some additional boxes into the original ground truth boxes so that all the boxes are totaled to a fixed number N_train, and then add Gaussian noise to the filled ground truth boxes, adding different Gaussian noises to obtain the optimal noise addition strategy.
4. The road damage detection method based on diffusion-based adaptive feature extraction according to claim 1, characterized in that The specific content of Step 4 is as follows: The inference process is a denoising sampling process from noise to target boxes, starting from the boxes sampled from the Gaussian distribution and gradually refining the predictions; adopt the box update strategy to restore them by replacing the unexpected boxes with random boxes to obtain the final result.
Citation Information
Patent Citations
Roadside vehicle identification method based on visual sensor
CN111695448A
Aircraft skin surface damage detection method and system based on deep learning
CN114565579A