Road damage detection method, system, device, storage medium and computer equipment
By using feature extraction and attention enhancement processing in the road damage detection model, the problem of low road inspection efficiency in existing technologies has been solved, achieving real-time and accurate road damage detection, and improving detection efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-12-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing road inspection methods are inefficient and fail to detect road damage in a timely manner, affecting road maintenance costs and safety.
A road damage detection model is adopted, which uses a backbone network module for feature extraction, combines a channel spatial attention module for attention enhancement, and uses a post-processing module for road damage prediction to achieve multi-target tracking and recognition.
It improves the accuracy and efficiency of road damage detection, enabling real-time automatic determination of relevant information about road damage objects, reducing manual intervention, and improving detection accuracy and efficiency.
Smart Images

Figure CN115984723B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a road damage detection method, system, device, computer-readable storage medium, and computer equipment. Background Technology
[0002] With the development and improvement of construction machinery, the mileage of roads built in my country is increasing rapidly. However, due to domestic traffic conditions—such as overloaded construction vehicles and high traffic volume—roads are prone to cracks, potholes, and roadbed collapses. The frequent occurrence of these problems primarily results in substantial maintenance costs, and if not addressed promptly, they can even threaten the safety of vehicles and pedestrians. Early maintenance, preventing road cracks from developing into serious problems, can significantly reduce maintenance expenses and extend the lifespan of roads. Therefore, early detection of road cracks is crucial for saving on road maintenance costs and ensuring road safety.
[0003] Existing road inspection methods are limited, with the mainstream approach relying on manual inspections by professionals conducting on-site investigations. Later, advancements in image processing technology revolutionized road inspection with the emergence of semi-automatic inspection systems. These systems primarily involve a collaboration between humans and computers, with the computer handling image acquisition, classification, and storage, while humans make the final decisions.
[0004] Current inspection systems are not very efficient at detecting road damage. Summary of the Invention
[0005] This application provides a method, apparatus, system, computer-readable storage medium, and computer equipment for road damage detection, which can improve the efficiency and accuracy of road damage detection.
[0006] This application provides a method for detecting road damage, including:
[0007] Acquire each frame of road image from real-time captured road video;
[0008] The backbone network module in the road damage detection model is used to extract features from each frame of road image to obtain the feature output results of the backbone network module. The feature output results include multiple different features corresponding to multiple different downsampling factors.
[0009] Multiple different features are input into the channel space attention module of the road damage detection model for attention enhancement processing to obtain multiple different attention enhancement features;
[0010] Multiple different attention-enhanced features are input into the post-processing module of the road damage detection model for road damage prediction, so as to obtain the damage detection boxes of multiple different road damage objects in each frame of road image, the damage probability of the road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0011] This application also provides a method for detecting road damage, including:
[0012] Obtain a training dataset and an initial road damage detection model. The training dataset includes multiple road damage images, training damage detection boxes containing the corresponding training road damage objects in the road damage images, and road damage category labels of the training detection boxes.
[0013] The backbone network module in the initial road damage detection model is used to perform feature extraction processing on each road damage image to obtain the training feature output result of the backbone network module. The training feature output result includes multiple different training features corresponding to multiple different downsampling factors.
[0014] Multiple different training features are input into the channel space attention module of the initial road damage detection model for attention enhancement processing to obtain multiple different training attention enhancement features;
[0015] Multiple different training attention enhancement features are input into the post-processing module of the initial road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0016] Based on the damage detection box, the probability of road damage at the location of the damage detection box, the road damage category at the location of the damage detection box, the training damage detection box, and the road damage category label of the training damage detection box, the network parameters of the initial road damage model are updated to obtain the road damage detection model.
[0017] This application also provides a method for detecting road damage, including:
[0018] Acquire multiple frames of road images from real-time captured road videos;
[0019] A road damage detection model including a channel spatial attention module is used to perform road damage detection processing on the current frame road image in multiple frames of road images to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box containing a road damage object, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box.
[0020] Based on the road damage detection results of the current frame and previous frames, multi-target tracking processing is performed on the road damage object to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object. The damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object are determined.
[0021] Based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each frame of road image, determine whether the tracked target road damage object is valid. If valid, determine the location information of the target road damage object.
[0022] The tracker identifier of the target road damage object, the damage detection frame where the target road damage object is located, the damage probability of the damage detection frame, the road damage category, and the location information are saved.
[0023] This application also provides a road damage detection system, including:
[0024] Unmanned aerial vehicle (UAV) equipment and ground station equipment communicating with the UAV equipment, wherein the UAV equipment is equipped with a camera;
[0025] The flight mission of the UAV is set using the ground station equipment, and the flight mission is sent to the UAV equipment;
[0026] The camera is activated using the drone device, and the drone device is controlled to fly according to the flight mission;
[0027] The drone device acquires multiple frames of road images from real-time road video captured by the camera, and performs road damage detection processing on the multiple frames of road images to obtain road damage identification results in the multiple frames of road images. The road damage identification results include at least one of the following: the tracker identifier corresponding to the target road damage object, the damage detection box where the target road damage object is located, the damage probability of the damage detection box, the road damage category, and the positioning information of the target road damage object.
[0028] The road damage identification results in multiple frames of road images are sent to the ground station equipment so that the road damage identification results can be displayed on the ground station equipment;
[0029] The road damage identification results in the multi-frame road images are obtained according to the method described in any of the above-mentioned methods.
[0030] This application embodiment also provides a road damage detection device, including:
[0031] The first acquisition module is used to acquire each frame of road image from real-time captured road video;
[0032] The backbone processing module is used to perform feature extraction processing on each frame of road image using the backbone network module in the road damage detection model, so as to obtain the feature output result of the backbone network module. The feature output result includes multiple different features corresponding to multiple different downsampling factors.
[0033] The attention processing module is used to input multiple different features into the channel space attention module in the road damage detection model for attention enhancement processing, so as to obtain multiple different attention enhancement features;
[0034] The post-processing module is used to input multiple different attention enhancement features into the post-processing module of the road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0035] This application embodiment also provides a road damage detection device, including:
[0036] The training acquisition module is used to acquire the training dataset and the initial road damage detection model. The training dataset includes multiple road damage images, the training damage detection boxes containing the corresponding training road damage objects in the road damage images, and the road damage category labels of the training detection boxes.
[0037] The training backbone processing module is used to perform feature extraction processing on each road damage image using the backbone network module in the initial road damage detection model, so as to obtain the training feature output result of the backbone network module. The training feature output result includes multiple different training features corresponding to multiple different downsampling factors.
[0038] The training attention processing module is used to input multiple different training features into the channel space attention module in the initial road damage detection model for attention enhancement processing, so as to obtain multiple different training attention enhancement features;
[0039] Multiple different training attention enhancement features are input into the post-processing module of the initial road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0040] The post-training processing module updates the network parameters of the initial road damage model based on the damage detection box, the damage probability of the road damage at the location of the damage detection box, the road damage category at the location of the damage detection box, the training damage detection box, and the road damage category label of the training damage detection box, so as to obtain the road damage detection model.
[0041] The second acquisition module is used to acquire multiple frames of road images from real-time captured road videos;
[0042] The damage detection module is used to perform road damage detection processing on the current frame road image in multiple frames of road images using a road damage detection model including a channel spatial attention module, so as to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box where a road damage object is located, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box.
[0043] The tracking module is used to perform multi-target tracking processing on the road damage object based on the road damage detection results of the road images in the current frame and previous frames, so as to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object, and to determine the damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object;
[0044] The discrimination module is used to determine whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame. If valid, the positioning information of the target road damage object is determined.
[0045] The storage module is used to store the tracker identifier, the damage detection frame where the target road damage object is located, the damage probability of the damage detection frame, and the positioning information.
[0046] This application also provides a computer-readable storage medium storing a computer program adapted for loading by a processor to perform the steps in the road damage detection method as described in any of the above embodiments.
[0047] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the steps in the road damage detection method described in any of the above embodiments by calling the computer program stored in the memory.
[0048] The road damage detection method, apparatus, system, computer-readable storage medium, and computer equipment provided in this application acquire each frame of road images from real-time captured road videos. They then utilize a newly added channel spatial attention module in the road damage detection model to perform attention enhancement processing on multiple different features corresponding to multiple sampling multiples output by the backbone network module of the road damage detection model. This process extracts key features from the multiple different features, enhances the expressive power of key features, weakens the expressive power of unimportant features, and makes the key features among the multiple attention-enhanced features obtained after attention enhancement processing more prominent, thereby improving the accuracy of road damage detection processing. To improve efficiency, the road damage detection results are used to perform multi-target tracking on damaged road objects. This process determines whether the tracked damaged road object and the detected damaged road object in the current frame are the same target. By utilizing the multi-target tracking results, the accuracy of road damage detection across multiple road images is improved. Furthermore, at least two bounding boxes in each road image frame are used to determine the validity of the tracked target damaged road object. If valid, location information is obtained. In this way, relevant information of the target damaged road object is automatically determined in real time, improving the accuracy and efficiency of road damage identification. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic flowchart of the road damage detection method provided in the embodiments of this application.
[0051] Figure 2 A partial network structure diagram of YOLOx-Darknet53 provided in the embodiments of this application.
[0052] Figure 3 This is a schematic diagram of the network structure of the channel spatial attention module provided in an embodiment of this application.
[0053] Figure 4 This is a flowchart illustrating the channel spatial attention module provided in an embodiment of this application.
[0054] Figure 5 This is a schematic diagram of the network structure of the channel attention module provided in an embodiment of this application.
[0055] Figure 6 This is a flowchart illustrating the channel attention module provided in an embodiment of this application.
[0056] Figure 7 This is a schematic diagram of the network structure of the spatial attention module provided in an embodiment of this application.
[0057] Figure 8 This is a flowchart illustrating the spatial attention module provided in an embodiment of this application.
[0058] Figure 9 This is a schematic flowchart of a road damage detection method provided in an embodiment of this application.
[0059] Figure 10 This is another schematic flowchart of the road damage detection method provided in the embodiments of this application.
[0060] Figure 11 This is a schematic diagram of two discrimination boxes provided in an embodiment of this application.
[0061] Figure 12 This is a schematic diagram of the road damage detection system provided in an embodiment of this application.
[0062] Figure 13 This is a schematic diagram of the road damage detection system provided in an embodiment of this application.
[0063] Figure 14 This is another schematic diagram of the road damage detection method provided in the embodiments of this application.
[0064] Figure 15 This is a schematic diagram of the road damage identification results provided in an embodiment of this application.
[0065] Figure 16 This is a schematic diagram of the road damage detection device provided in an embodiment of this application.
[0066] Figure 17 This is a schematic diagram of the road damage detection device provided in an embodiment of this application.
[0067] Figure 18 Another structural schematic diagram of the road damage detection device provided in the embodiments of this application.
[0068] Figure 19 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0069] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0070] This application provides a road damage detection method, apparatus, system, computer-readable storage medium, and computer equipment. Specifically, the road damage detection method of this application can be executed by a computer device, which can be a terminal, server, or other similar device. The terminal can be a smartphone, tablet computer, laptop computer, personal computer (PC), drone, helicopter, or other similar device. The server can be a standalone physical server, a server cluster consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services and cloud databases.
[0071] This application uses a drone as an example of a computer device for illustration, and this will not be repeated below. It is understood that the solution in this application can also be applied to computer devices that are not drones. For example, a drone can be used to acquire real-time road video, which is then sent to a computer device for real-time road detection.
[0072] Before detailing the methods described in the embodiments of this application, let's further analyze current deep learning technologies. Current object detection methods include convolutional neural networks, such as the Fast R-CNN (Fast Region Proposals Convolutional Neural Networks) network model. The Fast R-CNN network model can be used to extract features. When the extracted target features have a high similarity to existing target features, the target is considered to have been detected. Although the Fast R-CNN network model has a low error rate and a low missed detection rate, it is slow as a two-stage detection algorithm and cannot meet the needs of real-time detection scenarios.
[0073] When using drones for road inspection, the real-time performance of the detection algorithm is crucial, as the algorithm's detection efficiency determines the detection distance that the drone can cover in a single flight mission.
[0074] Therefore, the solution in this application embodiment can utilize drones to achieve real-time detection, thereby improving efficiency and accuracy.
[0075] The following will provide a detailed description of a road damage detection method, apparatus, system, computer-readable storage medium, and computer equipment provided in the embodiments of this application. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0076] like Figure 1 The diagram shown is a flowchart of a road damage detection method provided in an embodiment of this application. The method is applied to computer equipment such as drones and includes the following steps.
[0077] 101. Obtain each frame of road image from real-time captured road video.
[0078] A camera is mounted / set up / installed in a computer device such as a drone. Before each flight mission, the camera in the drone is turned on. During flight, the camera is used to capture real-time road video during the flight process. This road video includes multiple frames of road images.
[0079] Because it is filmed in real time, it can capture real-time road video, as well as each frame of the road image within the video.
[0080] 102. The backbone network module in the road damage detection model is used to perform feature extraction processing on each frame of road image to obtain the feature output result of the backbone network module. The feature output result includes multiple different features corresponding to multiple different downsampling factors.
[0081] The road damage model is a one-stage detection model that directly and simultaneously outputs localization and classification results. Therefore, it is more efficient than two-stage detection models and can achieve real-time road detection.
[0082] The road damage detection model includes a backbone network module, a channel spatial attention module, and a post-processing module. The post-processing module can include an enhanced feature extraction module and a prediction module. The backbone network module, typically called the Neck module, is used to extract features. The channel spatial attention module, a new addition in this application, enhances the expressive power of key features while weakening the expressive power of less important features; this will be described in detail later. The enhanced feature extraction module, also known as the Neck module, is generally placed between the backbone module and the prediction module to further improve feature diversity and robustness. The prediction module, also known as the Prediction module, is used to output the detection results.
[0083] The road damage detection module can be a YOLOx target recognition model with an added channel spatial attention module. Here, YOLOx indicates a model in the YOLO series, such as YOLO3, YOLO4, YOLO5, YOLOv3_spp, YOLOx-s, YOLOx-m, YOLOx-l, YOLOx-x, YOLOx-Darknet53, etc. This application uses YOLOx-Darknet53 as an example for illustration in its embodiments.
[0084] like Figure 2 The diagram shown is a schematic of the YOLOx-Darknet53 network model with added channel spatial attention module provided in an embodiment of this application. The backbone network modules, from front to back, include CBL, Res1, Res2, Res8, Res8, Res4, CBL, SPP, and CBL modules.
[0085] Each road image frame (e.g., 640*640*3, each channel being 640*640, for a total of 3 channels) is input into each module of the backbone network of the YOLOx-Darknet53 network model for feature extraction. For example, after processing by the first Res8 module of the backbone network module, the third feature (80*80*256, each channel being 80*80, for a total of 256 channels) is obtained; after processing by the second Res8 module of the backbone network module, the second feature (40*40*512, each channel being 40*40, for a total of 512 channels) is obtained; and after processing by the last CBL module of the backbone network module, the first feature (20*20*1024, each channel being 20*20, for a total of 1024 channels) is obtained.
[0086] The first, second, and third features are multiple different features corresponding to different downsampling factors. Specifically, the third feature corresponds to 8x downsampling, the second feature to 16x downsampling, and the first feature to 32x downsampling. The number of channels for each feature corresponding to a different downsampling factor is different. Therefore, the first, second, and third features not only differ in their downsampling factors but also in the number of channels they correspond to.
[0087] 103. Multiple different features are input into the channel space attention module of the road damage detection model for attention enhancement processing to obtain multiple different attention enhancement features.
[0088] The Convolutional Block Attention Module (CBAM) comprises a Channel Attention Module (CAM) and a Spatial Attention Module (SAM) connected in sequence. The Channel Attention Module is used to extract the important parts of the features, while the Spatial Attention Module is used to extract the more information-rich parts of the features.
[0089] like Figure 3 The diagram shown is a schematic of the channel spatial attention module provided in an embodiment of this application. The channel attention module and the spatial attention module are connected in series via sampling. They are independent of each other but have a clear division of labor, performing attention mechanism operations on the channel and spatial attention mechanism operations respectively.
[0090] Since the backbone network module outputs three different features corresponding to three different downsampling factors—namely, the first feature, the second feature, and the third feature—the channel spatial attention module is copied into three copies: CBAM1, CBAM2, and CBAM3, each corresponding to one of the first, second, and third features, respectively. Figure 2 As shown. In Figure 2 In this context, the CBAM1 module is part of the Neck module. In some embodiments, the Neck module may not include the CBAM1 module.
[0091] In one embodiment, such as Figure 4 As shown, step 103 also includes the following steps, please refer to the details. Figure 3 Come and see Figure 4 , Figure 4 The main function is to use a channel-space attention module to enhance the attention of the input features. Figure 4 This includes steps 201 to 204.
[0092] 201. Each of the multiple different features is input into the channel attention module for channel attention feature extraction processing to obtain the corresponding channel attention feature.
[0093] For example, the first feature among multiple different features is input into the channel attention module in CBAM1 for channel attention feature extraction processing to obtain the channel attention feature of the first feature; the second feature among multiple different features is input into the channel attention module in CBAM2 for channel attention feature extraction processing to obtain the channel attention feature of the second feature; and the third feature among multiple different features is input into the channel attention module in CBAM3 for channel attention feature extraction processing to obtain the channel attention feature of the third feature.
[0094] in, Figure 3 In this context, InputFeature can represent the first, second, or third feature of the input, denoted as F∈R. C*H*W This is represented by the fact that the C*H*W values of the first feature, the second feature, and the third feature are different.
[0095] The structural diagram of the channel attention module is as follows: Figure 5 As shown. Specifically, step 201 may also include Figure 6 Steps 301 to 304 shown are described in conjunction with... Figure 5 See Figure 6 .
[0096] 301. Each of the multiple different features is subjected to two different pooling processes to obtain two different pooled features.
[0097] Each of the multiple different features is subjected to two different pooling processes, for example, global max pooling. Figure 5 MaxPool and global average pooling (in the context of pooling) Figure 5 In AvgPool), two different pooling features are obtained, such as Figure 5 As shown. For example, the input feature F H*W*C After pooling, two different 1×1×C feature maps are obtained, which are two different pooling features.
[0098] 302. Perform multilayer perceptual processing on the two different pooling features to obtain two different perceptual processing features.
[0099] Two different pooled features are input into a small neural network, such as a Multilayer Perceptron (MLP), for multilayer perceptron processing to extract features, resulting in two distinct perceptron features. This MLP has only two layers: the first layer uses ReLU as the activation function and has Cr neurons, while the second layer has C neurons. The two layers share weights. Because the two layers share weights, therefore… Figure 5 The MLP network is also called the SharedMLP network. Since the input to the MLP network is two 1×1×C feature maps, the MLP network will also output two feature maps, that is, two different perceptual processing features.
[0100] The MLP network is used to perform multi-layer perceptual processing on the pooling features corresponding to each channel to extract the features between each channel and the relationship between the features of each channel, highlighting the features of the target between each channel, such as the features of road cracks.
[0101] 303. Perform a first fusion process on two different perceptual processing features to obtain the first fused feature.
[0102] The first fusion process can be an addition process, which involves adding corresponding elements of two different perceptual features one by one (element-wise addition), such as... Figure 5 In the middle operate.
[0103] 304. The first activation function is used to perform the first activation process on the first fused feature to obtain the corresponding channel attention feature.
[0104] The first activation function can be the Sigmoid activation function. The first activation function is used to perform a first activation process on the first fused feature to obtain the corresponding channel attention feature (Channel AttentionM). C That is, the output feature M obtained by using the channel attention module to perform channel attention feature extraction processing. C The first, second, and third features, after being processed by the channel attention module, all yield corresponding output features (channel attention features).
[0105] Among them, the channel attention feature M is obtained by using the channel attention module to perform channel attention feature extraction. C The processing procedure can be shown in formula (1).
[0106]
[0107] In the formula, σ represents the Sigmoid activation function, and W0∈RCr*C W1∈R C*Cr In an MLP network, the weights W0 and W1 are shared weights, meaning that pooling features from the two inputs are shared between each other.
[0108] Input features F are extracted using the channel attention module. H*W*C The system learns the relationships between channels and their weight distributions to highlight the features of targets (such as road cracks) in each channel. The channel attention features obtained by the channel attention module are already output features enhanced by channel attention.
[0109] 202. Based on the channel attention features and the corresponding features of the current input channel attention module, determine the corresponding target channel attention features.
[0110] Based on the channel attention feature corresponding to the first feature and the first feature of the current input channel attention module, the target channel attention feature of the first feature is determined. Similarly, the target channel attention features of the second feature and the target channel attention features of the third feature are also obtained.
[0111] In one embodiment, the channel attention features are multiplied with the corresponding features of the current input channel attention module to obtain the corresponding target channel attention features.
[0112] In one embodiment, the channel attention features are mapped to the same dimension as the corresponding features of the current input channel attention module, for example, by performing a copy mapping. Each element in the mapped channel attention features is then multiplied by the corresponding elements in the corresponding features of the current input channel attention module to obtain the corresponding target channel attention features. For example, the channel attention features corresponding to the first feature are mapped to the same dimension as the first feature, and then the channel attention features corresponding to the first feature are multiplied one-to-one by each element in the first feature, i.e., element-wise multiplication is performed. Figure 3 In operate.
[0113] The process of obtaining the target channel attention features can be represented by the following formula (2).
[0114]
[0115] In the formula, F represents the input feature map, defined as F∈R C*H*W This includes the first feature, the second feature, or the third feature. This involves element-wise multiplication between two feature vectors. F′ represents the feature vector M after channel attention enhancement of F. CPerform with the original input feature map F Feature map after operation.
[0116] The target channel attention features obtained through the channel attention module can highlight the features of the target (road damage object) in each channel and focus on the key / important features in each channel, such as the features of road damage objects like road cracks.
[0117] 203. Input the target channel attention features into the spatial attention module for spatial attention feature extraction processing to obtain the corresponding spatial attention features.
[0118] For example, the target channel attention feature corresponding to the first feature is input into the spatial attention module for spatial attention feature extraction processing to obtain the spatial attention feature corresponding to the first feature. Similarly, the spatial attention features corresponding to the second feature and the spatial attention features corresponding to the third feature can be obtained.
[0119] The network structure diagram of the spatial attention module is shown below. Figure 7 As shown. Step 203 may also include Figure 8 Steps 401 to 404 shown are described in conjunction with... Figure 7 See Figure 8 .
[0120] 401. The target channel attention features are subjected to two different pooling processes to obtain two different pooling features.
[0121] The input to the spatial attention module is the feature M output by the channel attention module. C and the original input features F H*W*C ,conduct The feature map F′ obtained after the element-wise multiplication operation is the target channel attention feature. The target channel attention feature, which includes the key features of each channel extracted by the channel attention module, is then subjected to two different pooling processes: one channel global max pooling and one flat max pooling, to obtain two different pooled features, such as two H×W×1 feature maps. Figure 7 As shown.
[0122] 402. Perform a second fusion process on the two different pooling features to obtain the second fused feature.
[0123] The second fusion process involves concatenation, where two different pooling features are concatenated to obtain a 2H*W*1 feature map, which is the second fused feature. Figure 7 As shown. The second fusion feature includes the more information-rich parts of the key / important features, such as the features of road cracks.
[0124] 403. The second fusion feature is convolved to map it onto a space to obtain spatial convolution features.
[0125] like Figure 7 As shown, the second fused feature is processed by convolutional layer to transform the spatial information of the feature, reduce its dimensionality, and transform it into another space to obtain spatial convolutional feature. The spatial convolutional feature includes the more information-rich parts of important / key features, such as the feature of road cracks.
[0126] 404. By applying the second activation function to the spatial convolution features, the corresponding spatial attention features can be obtained.
[0127] The second activation function can be the Sigmoid activation function. This second activation function is used to perform a second activation process on the second fused feature to obtain the corresponding spatial attention feature (Spatial AttentionM). S In this step, the second activation function is used for activation to obtain the enhanced feature M generated by the spatial attention module. S .
[0128] The corresponding spatial attention features include the spatial attention features corresponding to the first feature, the spatial attention features corresponding to the second feature, and the spatial attention features corresponding to the third feature.
[0129] The process of obtaining the corresponding spatial attention features can be represented by the following formula (3).
[0130]
[0131] In the formula, σ represents the second activation function, such as the Sigmoid function, and f 7*7 This indicates that a 7×7 convolution kernel is used to perform a convolution operation on the concatenated second fused feature. It should be noted that the convolution kernel in formula (3) is just an example and does not constitute a limitation on the convolution kernel. Other convolution kernels can also be used.
[0132] The spatial attention mechanism of the spatial attention module is used to transform the spatial information of the input features into another space, and important features / important information are extracted, such as the information-rich parts of the features of road cracks.
[0133] 204. Based on the spatial attention features and the target channel attention features in the current input spatial attention module, determine the corresponding target spatial attention features, and use the determined target spatial attention features as multiple different attention enhancement features.
[0134] For example, the spatial attention feature corresponding to the first feature and the target channel attention feature corresponding to the first feature in the current input spatial attention module are used to determine the target spatial attention feature corresponding to the first feature. Similarly, the target spatial attention feature corresponding to the second feature and the target spatial attention feature corresponding to the third feature are obtained. The target spatial attention feature corresponding to the first feature, the target spatial attention feature corresponding to the second feature, and the target spatial attention feature corresponding to the third feature are used as multiple different attention enhancement features.
[0135] In one embodiment, the spatial attention features and the target channel attention features in the current input spatial attention module are multiplied together to obtain the corresponding target spatial attention features.
[0136] In one embodiment, the spatial attention features are mapped to the same dimension as the corresponding features currently input to the spatial attention module, for example, by performing a copy mapping. Each element of the mapped spatial attention features is then multiplied by each element of the corresponding features currently input to the spatial attention module to obtain the corresponding target spatial attention features. This involves performing element-wise multiplication, such as... Figure 3 In operate.
[0137] In this way, the spatial attention features are weighted and output, which enhances the feature representation ability of the more information-rich parts of the key features, such as road cracks, and weakens the feature representation ability of unimportant features.
[0138] The process of obtaining the attention features of the target space can be represented by the following formula (4).
[0139]
[0140] F″ is the feature M after spatial attention enhancement of F′. S Perform with feature map F′ The feature map after the operation, F″ feature map, is also the feature map output by the CBAM module, which is the attention-enhanced feature.
[0141] 104. Multiple different attention enhancement features are input into the post-processing module of the road damage detection model for road damage prediction processing to obtain the road damage detection result in each frame of road image. The road damage detection result includes at least one of the following: multiple different road damage objects are located in the damage detection box, the damage probability of the road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0142] The road damage objects are the detected targets, such as road cracks. A damage detection box refers to the annotation of road damage objects as detection boxes in each frame of the road image. Each damage detection box is represented by four parameters: t x t y t w and t h , where t x t y Indicates the offset of the center point of the damage detection frame, t w and t h This indicates the stretching of the damage detection box in width and height. Since there are three branches, each branch can output multiple damage detection boxes, such as three.
[0143] The probability of road damage at the location of the damage detection box refers to the probability that the target in the damage detection box will be damaged. Each damage detection box corresponds to one probability of damage.
[0144] The road damage categories can be set according to the actual situation. For example, to distinguish between damage that is not legitimate, the road damage category can include: road damage present, and damage without justification. Correspondingly, the road damage category should include at least these two situations. For example, the road damage category can also include multiple different categories, such as cracks, potholes, and holes. In YOLOx-Darknet53, each damage detection box can support multiple categories for multi-classification, such as 80 categories.
[0145] Since the post-processing module includes an enhanced feature extraction module and a prediction module, multiple different attention-enhanced features are input into the enhanced feature extraction module for enhanced feature extraction processing to obtain multiple different enhanced features. These multiple different enhanced features are then input into the prediction module for road damage prediction processing to obtain the road damage detection results in each frame of the road image. The road damage detection results include at least one of the following: multiple different road damage objects located in the damage detection boxes, the road damage probability at the location of the damage detection boxes, and the road damage category at the location of the damage detection boxes.
[0146] For details, please refer to the enhanced feature extraction module and prediction module. Figure 2 The network structure shown is Figure 2 Only a portion of the prediction module is shown. Since there are three branches—first feature, second feature, and third feature—the enhancement feature extraction module, after performing enhancement feature extraction, can obtain three different enhancement features, which are then used as the three inputs to the prediction module.
[0147] The prediction module's first input is determined based on the attention enhancement feature corresponding to the first feature. For example, the attention enhancement feature corresponding to the first feature can be directly used as the first input. The second input is determined based on the first input and the attention enhancement feature corresponding to the second feature. For instance, the first input feature undergoes enhancement feature extraction processing, such as convolution and upsampling, and the processed result is concatenated with the attention enhancement feature corresponding to the second feature. This concatenation is then processed through convolution and other steps in the CBL module to obtain the second input. The third input is determined based on the second input and the attention enhancement feature corresponding to the third feature. For instance, the second input undergoes enhancement feature extraction processing, such as convolution and upsampling, and the processed result is concatenated with the attention enhancement feature corresponding to the third feature. This concatenation is then processed through convolution and other steps in the CBL module to obtain the third input.
[0148] Since the three attention-enhancing features all improve the expressive power of key features, the first, second, and third inputs obtained from the three attention-enhancing features further enhance the expressive power of key features such as road damage, weaken the expressive power of unimportant features, further strengthen the key features, improve the prediction accuracy of the prediction module, and improve the accuracy and efficiency of road damage detection.
[0149] In the above embodiments, multiple different features extracted from the road damage detection model at different downsampling multiples are enhanced using the channel attention mechanism and spatial attention mechanism in the channel spatial attention module. This improves the expressive power of important / key features such as road cracks and weakens the expressive power of unimportant features such as background features. This is extremely suitable for road damage detection scenarios, where the target is often small and the background is complex, thus improving the accuracy of detection.
[0150] Figure 9 This is a flowchart illustrating the road damage detection method provided in this application example. The flowchart mainly includes the training process of training the road damage detection model. The training process mainly includes the following steps.
[0151] 501. Obtain the training dataset and the initial road damage detection model. The training dataset includes multiple road damage images, the training damage detection boxes containing the corresponding road damage objects in the road damage images, and the road damage category labels of the training detection boxes.
[0152] The road damage images in the training dataset can be divided into two parts: one part consists of open-source road damage images from the internet, such as images of road cracks, and the other part consists of road damage images captured by drones, such as images of road cracks. After collecting the images for the training dataset, pre-defined software in the field of deep learning object detection, such as LabelImg, is used to annotate the images, obtaining the corresponding XML files, thus completing the construction of the training dataset. To reduce the overfitting phenomenon of the model, data augmentation techniques are used to expand the sample size of the constructed training dataset, such as rotation, adjustment of brightness and contrast, etc. Finally, the augmented training dataset is used to train the initial road damage detection model.
[0153] The initial road damage detection model can be a YOLOx series network model, which is a network model whose network parameters need to be updated.
[0154] In this example, the road damage category labels in the training dataset represent the truly determined road damage category information. The terms used in this embodiment, such as "training road damage object" and "training damage detection box," have the same meaning as those used in the preceding text. The word "training" is added before these terms in this embodiment to distinguish them from their counterparts in the preceding text. The same applies to the terms mentioned below, and will not be repeated hereafter.
[0155] 502. The backbone network module in the initial road damage detection model is used to perform feature extraction processing on each road damage image to obtain the training feature output result of the backbone network module. The training feature output result includes multiple different training features corresponding to multiple different downsampling factors.
[0156] For example, the training feature output includes the third feature corresponding to 8x downsampling, the second feature corresponding to 16x downsampling, and the first feature corresponding to 32x downsampling.
[0157] 503. Multiple different training features are input into the channel space attention module of the initial road damage detection model for attention enhancement processing to obtain multiple different training attention enhancement features.
[0158] Step 503 includes: inputting each of the multiple different training features into the channel attention module for channel attention feature extraction processing to obtain the corresponding training channel attention feature; determining the corresponding training target channel attention feature based on the training channel attention feature and the corresponding training feature currently input into the channel attention module; inputting the training target channel attention feature into the spatial attention module for spatial attention feature extraction processing to obtain the corresponding training spatial attention feature; determining the corresponding training target spatial attention feature based on the training spatial attention feature and the training target channel attention feature currently input into the spatial attention module; and using the determined multiple training target spatial attention features as multiple different training attention enhancement features.
[0159] The step of inputting each of the multiple different training features into the channel attention module for channel attention feature extraction to obtain the corresponding training channel attention feature includes: performing two different pooling processes on each of the multiple different training features to obtain two different training pooling features; performing multilayer perceptron processing on the two different training pooling features to obtain two different training perceptron processing features; performing a first fusion process on the two different training perceptron processing features to obtain a first fusion training feature; and performing a first activation process on the first fusion training feature using a first activation function to obtain the corresponding training channel attention feature.
[0160] The step of inputting the training target channel attention features into the spatial attention module for spatial attention feature extraction to obtain the corresponding training spatial attention features includes: performing two different pooling processes on the training target channel attention features to obtain two different training pooling features; performing a second fusion process on the two different training pooling features to obtain a training second fusion feature; performing convolution processing on the training second fusion feature to map it onto a space to obtain a training spatial convolution feature; and performing a second activation process on the training spatial convolution feature using a second activation function to obtain the corresponding training spatial attention features.
[0161] The step of determining the corresponding training target channel attention feature based on the training channel attention feature and the corresponding training feature of the current input channel attention module includes: mapping the training channel attention feature to the same dimension as the corresponding training feature of the current input channel attention module; and multiplying each element in the mapped training channel attention feature with the corresponding element in the corresponding training feature of the current input channel attention module to obtain the corresponding training target channel attention feature.
[0162] The step of determining the corresponding training target spatial attention feature based on the training spatial attention feature and the training target channel attention feature in the current input spatial attention module includes: mapping the training spatial attention feature to the same dimension as the corresponding training target channel attention feature currently input to the spatial attention module; and multiplying each element in the mapped training spatial attention feature with the corresponding element in the training target channel attention feature in the current input spatial attention module to obtain the corresponding training target spatial attention feature.
[0163] 504. Multiple different training attention enhancement features are input into the post-processing module of the initial road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0164] 505. Based on the damage detection box, the damage probability of road damage at the location of the damage detection box, the road damage category at the location of the damage detection box, the training damage detection box, and the road damage category label of the training damage detection box, update the network parameters of the initial road damage model to obtain the road damage detection model.
[0165] The loss value of the initial road damage model is determined based on the damage detection bounding box, the probability of road damage at the location of the damage detection bounding box, the road damage category at the location of the damage detection bounding box, the training damage detection bounding box, and the road damage category label of the training damage detection bounding box. The network parameters of the initial road damage model are updated based on the loss value until the training stopping condition is met, such as the loss value converges or the number of training epochs reaches a preset limit, at which point training stops and the road damage detection model is obtained.
[0166] This embodiment describes the specific process of training an initial road damage detection model to obtain a road damage detection model.
[0167] Figure 10 This is another flowchart illustrating the road damage detection method provided in this application example. The flowchart mainly involves three parts: a road damage detection model, a multi-object tracking model (multi-object tracking algorithm), and a dual-box algorithm (dual-box model). The flowchart utilizes these three parts to detect road damage and mainly includes the following steps.
[0168] 601, acquire multiple frames of road images from real-time captured road video.
[0169] 602. Using a road damage detection model including a channel spatial attention module, road damage detection processing is performed on the current frame road image in multiple frames of road images to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box where a road damage object is located, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box.
[0170] The channel space attention module is as described above and will not be repeated here.
[0171] Specifically, the channel space attention module in the road damage detection model can be used to perform attention enhancement processing on multiple different features corresponding to multiple downsampling multiples output by the backbone network module. The multiple attention-enhanced features obtained by the attention enhancement processing are then input into the post-processing module in the road damage detection model for road damage prediction processing, so as to obtain the road damage detection results in each frame of road image.
[0172] The road damage detection results of the current frame road image can also be processed according to the above. Figures 1 to 8 The damage detection method described in any of the embodiments is used to determine the damage, and will not be elaborated further.
[0173] 603. Based on the road damage detection results of the road images in the current frame and previous frames, perform multi-target tracking processing on the road damage object to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object, and determine the damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object.
[0174] The multi-target tracking process utilizes a multi-target tracking algorithm. Based on the road damage detection results in the road images of previous frames, the multi-target tracking algorithm predicts the road damage tracking result in the current frame's road image. This road damage tracking result includes at least one damage tracking box containing a road damage object and the tracker identifier corresponding to the road damage object. The road damage detection results and road damage tracking results of the current frame's road image are matched and associated. The road damage object tracked in the current frame and the road damage object detected in the current frame that are successfully matched and associated are identified as the same object, and the same object that is successfully matched and associated is identified as the target road damage object.
[0175] The multi-target tracking algorithm can be implemented using the DeepSORT algorithm, or other multi-target tracking algorithms. This embodiment uses the DeepSORT algorithm as an example for illustration.
[0176] In the DeepSORT algorithm, when tracking targets (road damage objects) in road videos, it uses Kalman filtering to perform optimal estimation of the target in the current frame based on the road damage detection result (target detection result) obtained from the road damage detection model in the k-th frame and the optimal estimate of the road damage detection results in the previous k frames. To verify whether targets are the same, the DeepSORT algorithm uses the correlation between appearance information (also known as appearance description / appearance features) and motion information. The correlation of the target's motion information can be calculated using Mahalanobis distance, while the matching of appearance information can be achieved using minimum cosine distance.
[0177] The multi-target tracking process of the DeepSORT algorithm mainly includes the following:
[0178] First, obtain the damage detection bounding box (also known as the target bounding box) of the road damage object obtained from the road damage detection model.
[0179] Next, a Kalman filter is used to estimate the motion state of the road damage object. For example, based on the acquired target bounding box, the location information of the target (road damage object) in the road image is extracted, and the target is feature-described. For example, an 8-dimensional state space is used to describe the target's motion state. x, y represent the position of the target bounding box in the road image, γ represents the aspect ratio of the target bounding box, and h represents the height of the target bounding box. This indicates the velocity information of the target bounding box.
[0180] Then, target tracking and matching are performed. For example, the road damage detection results of the current frame road image are obtained and the road tracking results estimated / predicted using Kalman filters are matched and associated. If the matching and association are successful, Kalman updates are performed. If the matching and association are unsuccessful, a second matching is performed. The matching and association consists of two parts: the association of motion information and the association of appearance information.
[0181] The motion information association represents the distance between the optimal estimate given by the Kalman filter after Kalman filtering and the target bounding box given by the target detection (road damage detection result). This distance can be calculated using Mahalanobis distance, as shown in Equation (5).
[0182]
[0183] In formula (5), d j Given the bounding box of the j-th target for object detection, y i S is the predicted target bounding box tracked by the tracker for the i-th tracked target. i Let be the covariance matrix.
[0184] Using Mahalanobis distance alone can lead to frequent target replacements. This is because Mahalanobis distance only considers the relationship between the target and the predicted trajectory. To reduce the probability of this phenomenon, it is necessary to use the characteristics of the target to complete the association of the target's appearance information.
[0185] Appearance information association refers to matching and associating objects by calculating the similarity between the features of the target bounding box output by the target detection system and the target bounding box output by the tracker. This similarity is described using the minimum cosine distance. Let r... j To represent the feature description corresponding to the target bounding box of the object detection, using If the feature description corresponding to the target bounding box output by the k-th tracker is used, then the minimum pre-distance can be determined according to formula (6).
[0186]
[0187] By weighted and fused the correlations of motion information and appearance information, the final correlation measurement method can be obtained, leading to the final matching and correlation results. Specifically, as follows... Figure 7 As shown.
[0188] c i,j =λdis (1) (i,j)+(1-λ)dis (2) (i,j) (7)
[0189] Finally, the target tracking status is updated: after successful matching, the tracker is updated using Kalman filtering, which updates all target bounding boxes predicted by the Kalman filter. For trackers and target detection results (road damage detection results) that still fail to match after two matching attempts, the target detection result is initialized as a new target, and a tracker is reassigned to it using the Hungarian algorithm; for the tracker, its lifetime is determined, trackers with a lifetime greater than the threshold are deleted, and those with a lifetime less than the threshold are retained, waiting for the target to reappear.
[0190] By using a multi-target tracking algorithm to track targets, it is possible to distinguish whether the target information (road damage objects) output by the road damage detection model is the same target. At the same time, it will associate the target correlation information between consecutive frames of the road video (including the correlation of motion information and the correlation of appearance information) to ensure that the targets in the video are the same targets, thereby improving the accuracy of road damage detection.
[0191] 604. Based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame, determine whether the tracked target road damage object is valid. If valid, determine the location information of the target road damage object.
[0192] At least two of these bounding boxes involve the binomial bounding box algorithm. This step will describe the binomial bounding box algorithm in detail.
[0193] In this process, when the drone is performing a flight mission, the content of multiple frames of road images in the road video will change along a certain direction as the drone flies. This direction is taken as the first direction, and at least two decision boxes are set along the first direction, and at least two decision boxes do not overlap.
[0194] For example, when a drone is flying, the upper part of each frame of a road image in a road video may be changing (newly added content), while the lower part remains the same as the previous frame. In this case, the content of multiple road images changes along the height direction. Therefore, at least two bounding boxes should be set along the height direction, or at least two bounding boxes should be set horizontally. Figure 11 The diagram shows at least two decision boxes set in the height direction. Correspondingly, the changing content first passes through the first decision box of the at least two decision boxes, and then through the second decision box of the at least two decision boxes, that is, the first decision box is above and the second decision box is below.
[0195] For example, when a drone is in flight, the left side of each frame of a road image in a road video contains changing content (newly added content), while the right side remains the same as the previous frame. In this case, the content of multiple road images changes along the width direction. Therefore, at least two decision boxes should be set along the width direction, or at least two decision boxes should be set vertically. Correspondingly, the changing content first passes through the first decision box of the two decision boxes, and then through the second decision box of the two decision boxes; that is, the first decision box is on the left, and the second decision box is on the right.
[0196] The presence of at least two bounding boxes in each frame of a road image indicates that these two boxes apply to every frame of the road video, and does not imply that at least two bounding boxes are required for each specific road image frame. While a dual-bounding box algorithm can include more boxes for greater accuracy, this increases computational cost. Two boxes, however, can achieve comparable accuracy while minimizing computational complexity. This embodiment uses two bounding boxes as an example.
[0197] Furthermore, it is also necessary to determine the size and position of at least two decision boxes. Specifically, the size of the at least two decision boxes in the first direction is determined based on the image size and scaling factor of the multi-frame road images in the first direction; the size of the at least two decision boxes in the second direction is determined based on the image size of the multi-frame road images in the second direction; and the position of the at least two decision boxes is determined based on the image size of the multi-frame road images in the first direction and the size of the at least two decision boxes in the first direction.
[0198] Wherein, if the first direction is the height direction, the image size of the multi-frame road image in the first direction refers to the image height, and the image size in the second direction refers to the image width. If the first direction is the width direction, the image size of the multi-frame road image in the first direction refers to the image width, and the image size in the second direction refers to the image height. The scaling factor controls the size of at least two decision boxes in the first direction, and the scaling factor is a value between 0 and 1.
[0199] Taking the height direction as an example, the height of at least two decision boxes is determined based on the height of the multi-frame road images in the height direction and the scaling factor. Specifically, the height of the decision boxes can be determined using the first formula in formula (8). The width of the decision boxes in the width direction is determined based on the width of the multi-frame road images. For example, the image width can be determined as the width of the decision boxes. Then, the positions of at least two decision boxes are determined based on the height of the multi-frame road images in the height direction and the height of the at least two decision boxes in the height direction. For example, the positions of at least two decision boxes can be determined using the second formula in formula (8).
[0200]
[0201] In the formula, H is the height of the road image captured by the camera, k is the scaling factor, and R... h P is the height of the decision box (the width is fixed to the width of the image captured by the camera), and P is the position of the decision box in the road image.
[0202] For example, assuming the height H of a multi-frame road image is 1000, and the center is 500, and the scaling factor k is 0.2, then the height R of the decision box... h If the value is 100, the positions P of the two decision boxes are 400 and 600 respectively. After drawing at these two positions, the vertical positions of the two boxes are 300 to 400 and 600 to 700 respectively.
[0203] Assume that at least two discrimination boxes set along the first direction are designated as the first discrimination box and the second discrimination box. Correspondingly, the step of determining whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame includes: determining whether the tracked target road damage object is valid based on whether the detection point corresponding to the damage detection box of the target road damage object appears in the first discrimination box and the second discrimination box, and whether the tracker identifier corresponding to the target road damage object is recorded.
[0204] The detection point corresponding to the damage detection box of the target damaged object is the point within the damage detection box. If the detection point corresponding to the damage detection box of the target road damaged object appears in both the first and second discrimination boxes, and the tracker identifier corresponding to the target road damaged object is recorded in the first discrimination box and also appears in the second discrimination box, then the tracked target road damaged object is determined to be valid; otherwise, further judgment is required.
[0205] Further, the step of determining whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each frame of the determined road image includes: if the detection point corresponding to the damage detection box of the target road damage object is detected to appear in the first discrimination box, check whether the tracker identifier corresponding to the target road damage object is recorded; if not recorded, record the tracker identifier; if the detection point corresponding to the damage detection box of the target road damage object is detected to appear in the second discrimination box, check whether the tracker identifier corresponding to the target road damage object is recorded in the first discrimination box; if so, determine that the tracked target road damage object is valid, and perform statistical counting on the target road damage object; if the tracker identifier corresponding to the target road damage object is recorded in the first discrimination box, but the detection point corresponding to the damage detection box of the target road damage object does not appear in the second discrimination box in multiple consecutive frames, delete the tracker identifier of the target road damage object that did not appear in the second discrimination box from the tracker identifier recorded in the first discrimination box.
[0206] If the target road damage object is determined to be valid, it is statistically counted, and a double-box algorithm is used to avoid double counting. Furthermore, if the target road damage object is determined to be valid, its location information is determined. This can be achieved by directly obtaining the GPS location information of the drone, or by further determining the location information based on the drone's GPS location information, thus facilitating subsequent repair of the target road damage object, such as road cracks.
[0207] 605. Save the tracker identifier of the target road damage object, the damage detection box where the target road damage object is located, the damage probability of the damage detection box, the road damage category and location information.
[0208] All information about valid target road damage objects is stored, such as tracker identifier, damage detection box where the target damage object is located, confidence level, category, location information, etc. The probability of damage to the damage detection box can also be regarded as the confidence level of damage to the damage detection box, and the category is the road damage category, etc.
[0209] In this embodiment, a road damage detection model, a multi-target tracking algorithm, and a dual-box algorithm are used together to identify target road damage objects. The road damage detection model can improve the accuracy and efficiency of detection. Furthermore, multi-target tracking is performed based on the road damage detection model to determine whether the detected road damage objects are the same target object, thereby improving the accuracy of road damage detection. Then, the dual-box algorithm is used to further determine whether the target road damage objects obtained by the multi-target tracking algorithm are valid. If they are valid, statistical counting is performed to further improve the accuracy of road damage detection.
[0210] Figure 12 This is a schematic diagram of a road damage detection system provided in an embodiment of this application. The road damage detection system includes a drone device 11 and a ground station device 12. The drone device 11 communicates with the ground station device 12, such as wirelessly. The ground station device 12 can be a terminal device such as a smartphone, tablet, laptop, personal computer (PC), vehicle terminal, or server device.
[0211] like Figure 13 The diagram shown is a structural schematic of a road damage detection system provided in an embodiment of this application. The UAV device 11 includes a hardware layer and a ROS system layer, which communicate with each other. The hardware layer includes the UAV fuselage and sensors, an onboard deep learning computing platform (the road damage detection method in this embodiment), and an onboard camera. The hardware layer receives information from the ROS system layer and the ground station equipment, and sends its own information to both the ROS system layer and the ground station equipment. The ROS system layer includes a data processing section, a low-level control section, and a communication section. The ROS system layer mainly abstracts the hardware layer, facilitating the control and management of the low-level hardware, as well as the transmission, reception, and processing of data.
[0212] The ground station equipment 12 primarily displays received information and sends control commands to the UAV, such as flight mission commands and flight status acquisition commands. The ground station equipment 12 includes a ground station application (Ground Station APP), through which flight mission settings, real-time video display, and flight status acquisition can be performed.
[0213] Figure 14 This is another schematic flowchart of the road damage detection method provided in this application embodiment. Please refer to the following for details regarding this road damage detection method. Figure 12 and Figure 13 Refer to the road damage detection system shown below, which includes the following steps.
[0214] 701. Use ground station equipment to set the flight mission of the UAV equipment and send the flight mission to the UAV equipment.
[0215] First, turn on the drone and wait for it to perform a self-test and for its system and related functions to start. After startup, connect to the drone wirelessly via the ground station. Then, set the drone's flight mission using the ground station app. This mission includes the start and end points of the flight and / or flight path. Once the flight mission settings are confirmed to be correct, send the mission to the drone via the ground station.
[0216] 702. Use the drone equipment to turn on the camera and control the drone equipment to fly according to the flight mission.
[0217] First, unlock the drone equipment, turn on the camera using the drone equipment, enable the Socket service, communicate data through the Socket service, run the road damage detection method in this embodiment of the application, and start controlling the drone equipment to fly and automatically execute flight tasks.
[0218] 703. Using a drone device to acquire multiple frames of road images from real-time road video captured by a camera, and performing road damage detection processing on the multiple frames of road images to obtain road damage identification results in the multiple frames of road images. The road damage identification results include at least one of the following: the tracker identifier corresponding to the target road damage object, the damage detection box where the target road damage object is located, the damage probability of the damage detection box, the road damage category, and the location information of the target road damage object.
[0219] The road damage detection method in any of the above embodiments can be used to obtain the road damage identification result. Please refer to the content described in the above embodiments for details, which will not be repeated here.
[0220] 704. Send the road damage identification results from multiple road images to the ground station equipment so that the road damage identification results can be displayed on the ground station equipment.
[0221] During a drone's flight mission, upon completing a target detection, the drone sends the detected target information to a ground station app. The app then displays the data, such as marking the location of the target (a valid target road damage object, like a road crack) on the map, displaying the damage detection box containing the target, the number of detected targets, the location information of the targets, and tracker identifiers. This facilitates subsequent repair of the target, such as the road crack. Figure 15 The image shown is a schematic diagram of the road damage identification results provided in an embodiment of this application.
[0222] This embodiment can not only automatically detect targets such as road cracks in real time, improving the efficiency and accuracy of road damage detection, but also allow operators to view the current road damage identification results in real time through the ground station APP, making it more convenient to assist operators in completing the road damage detection task.
[0223] In this embodiment, using drones for road inspection eliminates the time and resource consumption of manual road inspections by driving on the roads. Simply set the drone's flight mission, and after completing the mission, the drone returns to its starting point, allowing the operator to inspect the target road from there. Furthermore, using deep learning for road detection allows machines to perform the detection work, significantly reducing the manual cost of road inspections. Another advantage of using deep learning for road detection is that as more data is collected and the training dataset grows larger, the accuracy of machine detection increases, gradually reducing the need for manual intervention. Moreover, when more types of road damage need to be detected, simply expanding the categories in the training dataset and training the model can handle the detection of new targets.
[0224] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0225] To facilitate better implementation of the road damage detection method of this application, this application also provides a road damage detection device. Please refer to... Figure 16 , Figure 16 This is a schematic diagram of the road damage detection device provided in an embodiment of this application. The road damage detection device 800 may include an acquisition module 801, a backbone processing module 802, an attention processing module 803, and a post-processing module 804.
[0226] The first acquisition module 801 is used to acquire each frame of road image from real-time captured road video.
[0227] The backbone processing module 802 is used to perform feature extraction processing on each frame of road image using the backbone network module in the road damage detection model, so as to obtain the feature output result of the backbone network module. The feature output result includes multiple different features corresponding to multiple different downsampling factors.
[0228] The attention processing module 803 is used to input multiple different features into the channel space attention module in the road damage detection model for attention enhancement processing, so as to obtain multiple different attention enhancement features.
[0229] In one embodiment, the attention processing module 803 is further configured to: input each of a plurality of different features into the channel attention module for channel attention feature extraction processing to obtain a corresponding channel attention feature; determine a corresponding target channel attention feature based on the channel attention feature and the corresponding feature currently input into the channel attention module; input the target channel attention feature into the spatial attention module for spatial attention feature extraction processing to obtain a corresponding spatial attention feature; determine a corresponding target spatial attention feature based on the spatial attention feature and the target channel attention feature currently input into the spatial attention module; and use the determined plurality of target spatial attention features as a plurality of different attention enhancement features.
[0230] The post-processing module 804 is used to input multiple different attention enhancement features into the post-processing module of the road damage model for road damage prediction processing, so as to obtain multiple different road damage objects, the damage probability of the road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0231] Please see Figure 17 , Figure 17 This is a schematic diagram of a road damage detection device provided in an embodiment of this application. The road damage detection device 900 may include a training acquisition module 901, a training backbone processing module 902, a training attention processing module 903, a training post-processing module 804, and an update module 805.
[0232] The training acquisition module 901 is used to acquire a training dataset and an initial road damage detection model. The training dataset includes multiple road damage images, training damage detection boxes containing the corresponding training road damage objects in the road damage images, and road damage category labels of the training detection boxes.
[0233] The training backbone processing module 902 is used to perform feature extraction processing on each road damage image using the backbone network module in the initial road damage detection model, so as to obtain the training feature output result of the backbone network module. The training feature output result includes multiple different training features corresponding to multiple different downsampling factors.
[0234] The training attention processing module 903 is used to input multiple different training features into the channel space attention module in the initial road damage detection model for attention enhancement processing, so as to obtain multiple different training attention enhancement features.
[0235] The post-training processing module 804 inputs multiple different training attention enhancement features into the post-processing module of the initial road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box.
[0236] The update module 805 is used to update the network parameters of the initial road damage model based on the damage detection box, the damage probability of the road damage at the location of the damage detection box, the road damage category at the location of the damage detection box, the training damage detection box, and the road damage category label of the training damage detection box, so as to obtain the road damage detection model.
[0237] Please see Figure 18 , Figure 18 This is another structural schematic diagram of the road damage detection device provided in an embodiment of this application. The road damage detection device 1000 may include a second acquisition module 1001, a damage detection module 1002, a tracking module 1003, a discrimination module 1004, and a storage module 1005.
[0238] The second acquisition module 1001 is used to acquire multiple frames of road images from real-time captured road videos.
[0239] The damage detection module 1002 is used to perform road damage detection processing on the current frame road image in multiple frames of road images using a road damage detection model including a channel spatial attention module, so as to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box where a road damage object is located, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box.
[0240] The tracking module 1003 is used to perform multi-target tracking processing on the road damage object based on the road damage detection results of the road images in the current frame and previous frames, so as to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object, and to determine the damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object.
[0241] The discrimination module 1004 is used to determine whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame. If valid, the positioning information of the target road damage object is determined.
[0242] The storage module 1005 is used to store the tracker identifier, the damage detection frame where the target road damage object is located, the damage probability of the damage detection frame, and the positioning information.
[0243] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0244] Accordingly, embodiments of this application also provide a computer device, which can be a terminal, a server, or, for example, a drone. Figure 19 As shown, Figure 19 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 1100 includes a processor 1101 with one or more processing cores, a memory 1102 with one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 and the memory 1102 are electrically connected.
[0245] The processor 1101 is the control center of the computer device 1100. It connects various parts of the computer device 1100 through various interfaces and lines. By running or loading software programs (computer programs) and / or modules stored in the memory 1102, and calling data stored in the memory 1102, it performs various functions of the computer device 1100 and processes data, thereby monitoring the computer device 1100 as a whole.
[0246] In this embodiment of the application, the processor 1101 in the computer device 1100 will load the instructions corresponding to the processes of one or more applications into the memory 1102 according to the following steps, and the processor 1101 will run the applications stored in the memory 1102 to realize various functions, such as the functions / steps in the road damage detection method in any of the above embodiments.
[0247] The specific implementation of each of the above operations / steps and the beneficial effects they can achieve can be found in the previous embodiments, and will not be repeated here.
[0248] like Figure 19 As shown, the computer device 1100 also includes a camera 1103. Optionally, the computer device 1100 may further include a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the camera 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107. Those skilled in the art will understand that... Figure 19 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0249] The radio frequency circuit 1104 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other computer devices, and to transmit and receive signals with network devices or other computer devices.
[0250] Audio circuitry 1105 can be used to provide an audio interface between a user and a computer device via a speaker and a microphone. Audio circuitry 1105 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 1105, converted back into audio data, and then processed by processor 1101 before being transmitted via radio frequency circuitry 1104 to, for example, another computer device, or output to memory 1102 for further processing. Audio circuitry 1105 may also include an earphone jack to provide communication between peripheral headphones and the computer device.
[0251] The input unit 1106 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0252] Power supply 1107 is used to supply power to various components of computer device 1100. Optionally, power supply 1107 can be logically connected to processor 1101 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 1107 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0253] although Figure 19 As not shown in the diagram, computer device 1100 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.
[0254] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0255] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0256] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of computer programs that can be loaded by a processor to execute the steps in any of the road damage detection methods provided in embodiments of this application.
[0257] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0258] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0259] Since the computer program stored in the storage medium can execute the steps in any of the road damage detection methods provided in the embodiments of this application, the beneficial effects that any of the road damage detection methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0260] The foregoing has provided a detailed description of a road damage detection method, apparatus, system, storage medium, and computer equipment provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for detecting road damage, characterized in that, include: Acquire each frame of road image from real-time captured road video; The backbone network module in the road damage detection model is used to extract features from each frame of road image to obtain the feature output results of the backbone network module. The feature output results include multiple different features corresponding to multiple different downsampling factors. Multiple different features are input into the channel space attention module of the road damage detection model for attention enhancement processing to obtain multiple different attention enhancement features; Multiple different attention enhancement features are input into the post-processing module of the road damage detection model for road damage prediction processing to obtain the road damage detection results in each frame of road image. The road damage detection results include at least one of the following: multiple different road damage objects are located in the damage detection box, the damage probability of the road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box. The channel spatial attention module includes a channel attention module and a spatial attention module connected in sequence; the step of inputting multiple different features into the channel spatial attention module of the road damage detection model for attention enhancement processing to obtain multiple different attention enhancement features includes: Each of the multiple different features is input into the channel attention module for channel attention feature extraction processing to obtain the corresponding channel attention feature; Based on the channel attention features and the corresponding features currently input to the channel attention module, a corresponding target channel attention feature is determined; wherein, the step of determining the corresponding target channel attention feature based on the channel attention features and the corresponding features currently input to the channel attention module includes: mapping the channel attention features to the same dimension as the corresponding features currently input to the channel attention module; multiplying each element in the mapped channel attention features with the corresponding elements in the corresponding features currently input to the channel attention module to obtain the corresponding target channel attention feature; The target channel attention features are input into the spatial attention module for spatial attention feature extraction processing to obtain the corresponding spatial attention features; Based on the spatial attention features and the target channel attention features currently input to the spatial attention module, corresponding target spatial attention features are determined, and the determined target spatial attention features are used as multiple different attention enhancement features. The step of determining the corresponding target spatial attention features based on the spatial attention features and the target channel attention features currently input to the spatial attention module includes: mapping the spatial attention features to the same dimension as the corresponding target channel attention features currently input to the spatial attention module; and multiplying each element in the mapped spatial attention features with the corresponding elements in the target channel attention features currently input to the spatial attention module to obtain the corresponding target spatial attention features.
2. The method according to claim 1, characterized in that, The step of inputting each of the multiple different features into the channel attention module for channel attention feature extraction to obtain the corresponding channel attention feature includes: Each of the multiple different features is subjected to two different pooling processes to obtain two different pooled features; Two different pooling features are processed separately using multilayer perceptron processing to obtain two different perceptron processing features; Two different perceptual processing features are subjected to a first fusion process to obtain the first fused feature; The first fused feature is activated using a first activation function to obtain the corresponding channel attention feature.
3. The method according to claim 1, characterized in that, The step of inputting the target channel attention features into the spatial attention module for spatial attention feature extraction processing to obtain the corresponding spatial attention features includes: The target channel attention features are subjected to two different pooling processes to obtain two different pooling features. Two different pooling features are subjected to a second fusion process to obtain a second fused feature. The second fused feature is convolved to map it onto a space to obtain spatial convolutional features; By applying a second activation function to the spatial convolutional features, the corresponding spatial attention features can be obtained.
4. A method for detecting road damage, characterized in that, include: Obtain a training dataset and an initial road damage detection model. The training dataset includes multiple road damage images, training damage detection boxes containing the corresponding training road damage objects in the road damage images, and road damage category labels of the training detection boxes. The backbone network module in the initial road damage detection model is used to perform feature extraction processing on each road damage image to obtain the training feature output result of the backbone network module. The training feature output result includes multiple different training features corresponding to multiple different downsampling factors. Multiple different training features are input into the channel space attention module of the initial road damage detection model for attention enhancement processing to obtain multiple different training attention enhancement features; Multiple different training attention enhancement features are input into the post-processing module of the initial road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box. Based on the damage detection box, the probability of road damage at the location of the damage detection box, the road damage category at the location of the damage detection box, the training damage detection box, and the road damage category label of the training damage detection box, the network parameters of the initial road damage model are updated to obtain the road damage detection model. The step of inputting multiple different training features into the channel space attention module of the initial road damage detection model for attention enhancement processing to obtain multiple different training attention enhancement features includes: Each training feature from multiple different training features is input into the channel attention module for channel attention feature extraction to obtain the corresponding training channel attention feature; Based on the training channel attention features and the corresponding training features of the currently input channel attention module, a corresponding training target channel attention feature is determined; wherein, the step of determining the corresponding training target channel attention feature based on the training channel attention features and the corresponding training features of the currently input channel attention module includes: mapping the training channel attention features to the same dimension as the corresponding training features of the currently input channel attention module; multiplying each element in the mapped training channel attention features with the corresponding elements in the corresponding training features of the currently input channel attention module to obtain the corresponding training target channel attention feature; The training target channel attention features are input into the spatial attention module for spatial attention feature extraction to obtain the corresponding training spatial attention features; Based on the training space attention features and the training target channel attention features in the current input space attention module, corresponding training target space attention features are determined, and the determined multiple training target space attention features are used as multiple different training attention enhancement features. The step of determining the corresponding training target space attention features based on the training space attention features and the training target channel attention features in the current input space attention module includes: mapping the training space attention features to the same dimension as the corresponding training target channel attention features currently input to the space attention module; and multiplying each element in the mapped training space attention features with the corresponding elements in the training target channel attention features currently input to the space attention module to obtain the corresponding training target space attention features.
5. A method for detecting road damage, characterized in that, include: Acquire multiple frames of road images from real-time captured road videos; A road damage detection model including a channel spatial attention module is used to perform road damage detection processing on the current frame road image in multiple frames of road images to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box containing a road damage object, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box. Based on the road damage detection results of the current frame and previous frames, multi-target tracking processing is performed on the road damage object to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object. The damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object are determined. Based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each frame of road image, determine whether the tracked target road damage object is valid. If valid, determine the location information of the target road damage object. The tracker identifier of the target road damage object, the damage detection frame where the target road damage object is located, the damage probability of the damage detection frame, the road damage category, and the location information are stored. The road damage detection result of the current frame road image is obtained by the method according to any one of claims 1 to 3, or by the road damage detection model trained according to claim 4, which performs road damage detection processing on the current frame road image.
6. The method according to claim 5, characterized in that, The content of multiple road images in the road video changes along a certain direction, which is taken as the first direction. The at least two decision boxes are set along the first direction and do not overlap.
7. The method according to claim 6, characterized in that, The method further includes: The size of the at least two discrimination boxes in the first direction is determined based on the image size and scaling factor of the multi-frame road images in the first direction; The size of the at least two discrimination boxes in the second direction is determined based on the image size of the multi-frame road images in the second direction; The positions of the at least two decision boxes are determined based on the image size of the multi-frame road images in the first direction and the decision box size of the at least two decision boxes in the first direction.
8. The method according to claim 6, characterized in that, At least two decision boxes are set along the first direction, namely a first decision box and a second decision box; The step of determining whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame includes: The validity of the tracked target road damage object is determined by whether the detection point corresponding to the damage detection box of the target road damage object appears in the first discrimination box and the second discrimination box, and whether the tracker identifier corresponding to the target road damage object is recorded.
9. The method according to claim 8, characterized in that, The step of determining whether the tracked target road damage object is valid based on whether the detection point corresponding to the damage detection box of the target road damage object appears in the first discrimination box and the second discrimination box, and whether the tracker identifier corresponding to the target road damage object is recorded, includes: If the detection point corresponding to the damage detection box of the target road damage object is detected to appear in the first discrimination box, check whether the tracker identifier corresponding to the target road damage object is recorded. If not recorded, record the tracker identifier. If the detection point corresponding to the damage detection box of the target road damage object is detected to appear in the second discrimination box, check whether the tracker identifier corresponding to the target road damage object is recorded in the first discrimination box. If so, determine that the tracked target road damage object is valid, and perform statistical counting on the target road damage object. If the tracker identifier corresponding to the target road damage object is recorded in the first discrimination box, but the detection point corresponding to the damage detection box of the target road damage object does not appear in the second discrimination box within multiple consecutive frames, then the tracker identifier of the target road damage object that does not appear in the second discrimination box is deleted from the tracker identifier recorded in the first discrimination box.
10. The method according to claim 5, characterized in that, The step of performing multi-target tracking processing on the road damage object based on the road damage detection results of the current frame and previous frames, to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object, includes: Based on the road damage detection results in the road images of multiple previous frames, a multi-target tracking algorithm is used to predict the road damage tracking result in the current frame road image. The road damage tracking result includes at least one damage tracking box containing a road damage object and the tracker identifier corresponding to the road damage object. Match and associate the road damage detection results and the road damage tracking results of the current frame road image; The road damage object tracked in the current frame and the road damage object detected in the current frame that are successfully matched and associated are identified as the same object, and the same object is identified as the target road damage object.
11. A road damage detection system, characterized in that, include: Unmanned aerial vehicle (UAV) equipment and ground station equipment communicating with the UAV equipment, wherein the UAV equipment is equipped with a camera; The flight mission of the UAV is set using the ground station equipment, and the flight mission is sent to the UAV equipment; The camera is activated using the drone device, and the drone device is controlled to fly according to the flight mission; The drone device acquires multiple frames of road images from real-time road video captured by the camera, and performs road damage detection processing on the multiple frames of road images to obtain road damage identification results in the multiple frames of road images. The road damage identification results include at least one of the following: the tracker identifier corresponding to the target road damage object, the damage detection box where the target road damage object is located, the damage probability of the damage detection box, the road damage category, and the positioning information of the target road damage object. The road damage identification results in multiple frames of road images are sent to the ground station equipment so that the road damage identification results can be displayed on the ground station equipment; The road damage identification results in the multi-frame road images are obtained according to the method described in any one of claims 1 to 10.
12. A road damage detection device, characterized in that, include: The first acquisition module is used to acquire each frame of road image from real-time captured road video; The backbone processing module is used to perform feature extraction processing on each frame of road image using the backbone network module in the road damage detection model, so as to obtain the feature output result of the backbone network module. The feature output result includes multiple different features corresponding to multiple different downsampling factors. The attention processing module is used to input multiple different features into the channel space attention module in the road damage detection model for attention enhancement processing, so as to obtain multiple different attention enhancement features; The post-processing module is used to input multiple different attention enhancement features into the post-processing module of the road damage model for road damage prediction processing, so as to obtain multiple different damage detection boxes where road damage objects are located, the damage probability of road damage at the location of the damage detection box, and the road damage category at the location of the damage detection box. The attention processing module is further used for: Each of the multiple different features is input into the channel attention module for channel attention feature extraction to obtain the corresponding channel attention feature; Based on the channel attention features and the corresponding features currently input to the channel attention module, a corresponding target channel attention feature is determined; wherein, the attention processing module is further configured to: map the channel attention features to the same dimension as the corresponding features currently input to the channel attention module; and multiply each element in the mapped channel attention features with the corresponding elements in the corresponding features currently input to the channel attention module to obtain the corresponding target channel attention feature; The target channel attention features are input into the spatial attention module for spatial attention feature extraction processing to obtain the corresponding spatial attention features; Based on the spatial attention features and the target channel attention features currently input to the spatial attention module, a corresponding target spatial attention feature is determined, and the determined target spatial attention features are used as multiple different attention enhancement features. The attention processing module is further configured to: map the spatial attention features to the same dimension as the corresponding target channel attention feature currently input to the spatial attention module; and multiply each element in the mapped spatial attention feature with the corresponding element in the target channel attention feature currently input to the spatial attention module to obtain the corresponding target spatial attention feature.
13. A road damage detection device, characterized in that, include: The second acquisition module is used to acquire multiple frames of road images from real-time captured road videos; The damage detection module is used to perform road damage detection processing on the current frame road image in multiple frames of road images using a road damage detection model including a channel spatial attention module, so as to obtain the road damage detection result of the current frame road image. The road damage detection result includes at least one damage detection box where a road damage object is located, the damage probability of the road damage at the location of the damage detection box, and the road damage category of the damage detection box. The tracking module is used to perform multi-target tracking processing on the road damage object based on the road damage detection results of the road images in the current frame and previous frames, so as to determine that the road damage object tracked in the current frame and the road damage object detected in the current frame are the same target road damage object, and to determine the damage detection box where the target road damage object is located, the damage probability of the road damage at the location of the damage detection box, the road damage category of the damage detection box, and the tracker identifier corresponding to the target road damage object; The discrimination module is used to determine whether the tracked target road damage object is valid based on the detection point corresponding to the damage detection box of the target road damage object, the tracker identifier, and at least two discrimination boxes in each determined road image frame. If valid, the positioning information of the target road damage object is determined. The storage module is used to store the tracker identifier, the damage detection frame where the target road damage object is located, the damage probability of the damage detection frame, and the positioning information; The road damage detection result of the current frame road image is obtained according to the road damage detection device described in claim 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the steps of the road damage detection method as described in any one of claims 1-10.
15. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the steps of the road damage detection method as described in any one of claims 1-10 by calling the computer program stored in the memory.