Road damage detection method and system
Through the optimized Yolov8 model and drone image acquisition technology, accurate detection and trend prediction of road damage are achieved, and the problems of low detection efficiency and insufficient accuracy in the existing technology are solved, and detection efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202510404228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing road damage detection technology is difficult to take into account both accuracy and calculation volume, and lacks prediction of the development trend of damage, resulting in low detection efficiency and high maintenance costs.
The optimized Yolov8 model is used for road damage detection, and the precise detection and prediction of damage is achieved through feature fusion module, adaptive processing module, attention module and time series analysis, combined with drone image acquisition.
It improves the accuracy and computing efficiency of damage detection, can operate efficiently in resource-constrained scenarios, and can predict damage development trends, optimize maintenance plans, and reduce traffic impacts.
Smart Images

Figure CN120339213A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of road detection and maintenance, and particularly relates to a road damage detection method and system. Background Art
[0002] Nowadays, the number of automobiles has increased rapidly, and the road network has become increasingly complex. As an important national infrastructure, the condition of roads is directly related to economic development and people's daily lives. Good road conditions can not only improve transportation efficiency, enhance connections between regions, but also ensure people's travel safety and maintain the smooth flow of urban traffic; on the contrary, road damage may cause traffic jams, result in economic losses, and even pose serious safety hazards, thus threatening people's lives and property safety. Therefore, it is particularly important to strengthen road damage detection and maintenance.
[0003] Most road damage detection technologies rely on the analysis of static images. Road surface images are obtained through on-vehicle cameras or ground sensors, and then these images are processed and detected. The analysis of static images usually needs to be achieved through image recognition and classification; however, existing image recognition and classification models often have difficulty comprehensively extracting and utilizing feature information in images, resulting in insufficient accuracy of road damage detection, and prone to situations such as positioning errors, classification errors, misidentifications, and missed identifications; while a few high-accuracy recognition and classification models often have a large amount of computation, limited application scenarios, and low efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a road damage detection method and system, which are used to solve the problem that the image recognition and classification models adopted by the road damage detection technology implemented through image recognition and classification in the prior art are difficult to balance the accuracy and computation amount of road damage detection.
[0005] To achieve the above purpose, the present invention provides a road damage detection method, including: inputting a road image to be detected into a damage detection model trained by training samples to obtain the positioning information and category information of the identified damage; The damage detection model before training is obtained by optimizing the Yolov8 model; The optimization process includes: replacing each Bottleneck module in the C2f unit of the neck network in Yolov8 with an optimized feature fusion module; each optimized feature fusion module contains one of the Bottleneck modules and an adaptive processing module for processing its output; in the adaptive processing module, feature dimensionality reduction is performed on the output of the Bottleneck module through the SE attention mechanism, and then the results after the feature dimensionality reduction by the SE attention mechanism are processed respectively through a first branch containing depthwise separable convolution and a second branch containing a lightweight pyramid pooling module and a multi-layer perceptron, and the processing results of the first and second branches are fused to obtain the output of the adaptive processing module.
[0006] Furthermore, the optimization process also includes: adding an attention module between the output of the backbone network and the input of the neck network in Yolov8. In this attention module, feature extraction of different scales is performed on the output of the backbone network in Yolov8 respectively through grouped convolution containing convolutional branches of different scales and then fused. After that, the fused features after fusion are processed through a channel attention mechanism. Then, a spatial attention mechanism is used to process the results after the channel attention mechanism and fuse them with the input of this attention module to obtain the output of the attention module.
[0007] Furthermore, the optimization process also includes: replacing the convolution module for feature extraction in the backbone network of Yolov8 with an optimized feature extraction module; the optimized feature extraction module is obtained by replacing the original Ghost module closest to the output in the part of GhostNetV2 with Stride = 2 with an optimized Ghost module; The optimized Ghost module contains an original Ghost module and a fusion processing module for processing its output; the fusion processing module is used to fuse the output of the original Ghost module in the part of GhostNetV2 with Stride = 2 with the input of the part of GhostNetV2 with Stride = 2 as the output of this optimized Ghost module; Moreover, the convolution in each Bottleneck module in the C2f unit of the backbone network of Yolov8 is replaced with the optimized feature extraction module.
[0008] Further, it further includes: the localization information of the damage recognized from each frame of road image obtained in real time and the feature vectors obtained from all the features extracted in each single-frame road image recognized each time. The damage recognized in each frame of road image is matched through a target tracking algorithm, and according to the matching result and the time sequence corresponding to each frame of road image to be detected, a time series dataset corresponding to each damage is determined; The time series dataset corresponding to each damage stores a time series including the detection time when the damage is first detected, the localization information and category information of the damage recognized each time, the feature vectors obtained from all the features extracted in each single-frame road image recognized each time, and the motion trajectory information of the damage changing over time; From the time series dataset corresponding to a certain damage, the damage features determined by selecting a set number of data groups whose detection time corresponding to the recognized single-frame road image is before the time node to be predicted and is closest to the time node to be predicted are input into the LSTM model trained with time series data samples to predict the damage features corresponding to the time node to be predicted.
[0009] Further, the road images to be detected and the training samples are both obtained by shooting with a drone.
[0010] Further, the convolution branches with different scales in the grouped convolution include: a convolution branch mainly composed of a single 3x3 scale convolution network and a convolution branch mainly composed of two 3x3 scale convolution networks.
[0011] Further, the target tracking algorithm is the Deep SORT algorithm.
[0012] Further, the method for processing the result of the SE attention mechanism feature dimension reduction through the second branch including a lightweight pyramid pooling module and a multi-layer perceptron includes: The result of the SE attention mechanism feature dimension reduction is processed through the lightweight pyramid pooling module, and then input into the multi-layer perceptron to obtain the processing result of the multi-layer perceptron to obtain the output of the second branch.
[0013] The above technical solution of the present invention provides a brand-new road damage detection method, and its beneficial effects include: an optimized Yolov8 model is adopted to implement damage detection. In the optimized Yolov8 model, an optimized feature fusion module adds an adaptive processing module on the basis of the original Bottleneck module; in the adaptive processing module, feature dimensionality reduction is performed through the SE attention mechanism, effectively reducing the computational amount; the expression ability of local features is enhanced through a branch containing depthwise separable convolutions, helping the model to better capture detailed information; different-scale feature pooling and upsampling are performed through a branch containing a lightweight pyramid pooling module, fusing multi-scale context information. At the same time, the multi-layer perceptron in this branch acts on the channel dimension to reassign channel weights, enhancing the feature modeling ability of global information; finally, the two branches are fused to integrate global semantic information and local detailed information, thereby improving the model's representation ability for multi-scale features with different receptive fields. After processing the output of each Bottleneck module and then inputting it into the next module, this not only enables the model to more accurately detect road damage features, but also ensures its efficient operation, thus meeting the requirements of resource-constrained scenarios; that is, it not only enhances the expression ability of the damage detection model network, but also significantly improves the training efficiency and computational efficiency, especially suitable for devices with limited computational resources.
[0014] The present invention also provides a road damage detection system, including a processor, and executable program instructions are stored in the processor, and the executable program instructions are used to be executed to implement the road damage detection method as described above.
[0015] The technical solution of the above road damage detection system of the present invention can achieve the same beneficial effects as the above road damage detection method. Brief Description of the Drawings
[0016] Figure 1 It is a structural principle block diagram of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 2 It is a structural principle block diagram of the adaptive processing module of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 3 It is a structural principle block diagram of the lightweight pyramid pooling module of the adaptive processing module in the embodiment of the road damage detection method of the present invention; Figure 4 It is a structural principle block diagram of the attention module of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 5 It is a structural principle block diagram of the optimized Ghost module of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 6It is the structural principle block diagram of DFC Attention in the optimized Ghost module of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 7 It is the structural principle block diagram of Ghost Module in the optimized Ghost module of the damage detection model in the embodiment of the road damage detection method of the present invention; Figure 8 It is the overall process block diagram of the road damage detection method in the embodiment of the road damage detection method of the present invention. Specific Embodiments
[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0018] Embodiment of Road Damage Detection Method This embodiment provides a technical solution for a road damage detection method. The damage detection model adopted by this road damage detection method enhances global information and local features by fusing spatial attention and channel attention, enabling the model to more accurately detect road damage features; at the same time, lightweight processing is performed to effectively reduce the computational amount and meet the requirements of resource-constrained scenarios.
[0019] The method specifically includes: inputting the road image to be detected into the damage detection model trained by training samples to obtain the localization information and category information of the identified damage; Among them, the damage detection model before training is obtained by optimizing the Yolov8 model; The above optimization process includes: replacing each Bottleneck module in the C2f unit of the neck network in Yolov8 with an optimized feature fusion module; each optimized feature fusion module contains one of the Bottleneck modules and an adaptive processing module for processing its output; in the adaptive processing module, feature dimensionality reduction is performed on the output of the Bottleneck module through the SE attention mechanism, and then the results after the feature dimensionality reduction by the SE attention mechanism are processed respectively through a first branch containing depthwise separable convolution and a second branch containing a lightweight pyramid pooling module and a multi-layer perceptron, and the processing results of the first and second branches are fused to obtain the output of the adaptive processing module.
[0020] This method uses an optimized Yolov8 model to achieve damage detection. In the optimized Yolov8 model, the C2f unit in the original neck network for feature fusion is replaced with an optimized feature fusion module. On the basis of the original Bottleneck module, the optimized feature fusion module adds an adaptive processing module. In the adaptive processing module, feature dimensionality reduction is performed through the SE attention mechanism, effectively reducing the computational amount. The expression ability of local features is enhanced through a branch containing depthwise separable convolutions, helping the model to better capture detailed information. Different-scale feature pooling and upsampling are performed through a branch containing a lightweight pyramid pooling module, integrating multi-scale context information. At the same time, the multi-layer perceptron in this branch acts on the channel dimension to reassign channel weights, enhancing the feature modeling ability of global information. Finally, the two branches are fused to integrate global semantic information and local detailed information, thereby improving the model's representation ability for multi-scale features with different receptive fields. After processing the output of each Bottleneck module and then inputting it into the next module, this not only enables the model to more accurately detect road damage features but also ensures its efficient operation, thus meeting the requirements of resource-constrained scenarios. That is, it not only enhances the expression ability of the damage detection model network but also significantly improves the training efficiency and computational efficiency, especially suitable for devices with limited computational resources.
[0021] Specifically, the method of processing the result of feature dimensionality reduction by the SE attention mechanism through the second branch containing a lightweight pyramid pooling module and a multi-layer perceptron includes: processing the result of feature dimensionality reduction by the SE attention mechanism through the lightweight pyramid pooling module, and then inputting it into the multi-layer perceptron to obtain the processing result of the multi-layer perceptron to obtain the output of the second branch. Refer to Figure 1 , in the optimized Yolov8 model, the C2f_AFIM unit in the neck network is the new unit after replacing each Bottleneck module in the C2f unit of the neck network in Yolov8 with an optimized feature fusion module, replacing the original C2f unit in the neck network; the composition principle of the C2f_AFIM module is as Figure 1 shown in the lower left corner, obtained by connecting an AFIM (Adaptive Feature Integration Module, hereinafter referred to as the adaptive processing module in the embodiment) after each original Bottleneck module in the C2f unit; the composition principle of the AFIM module is referred to Figure 2 , where the part from the input to Re-weight (i.e., the re-weighting operation) is the part of the SE attention mechanism, and then it is divided into two branches; the left side is the first branch, also called the main branch, which contains DSC (i.e., depthwise separable convolution); the right side is the second branch, also called the residual branch, which contains LPPM (i.e., lightweight pyramid pooling module, throughFigure 2 The LPPM part in Figure 2 is represented by the part from 1x1 Conv to 1x1 Conv+Dropout after the LPPM part in Figure 3 . Since its specific structural design is prior art, it will not be elaborated here.
[0022] To further improve the accuracy of road damage detection and ensure efficiency at the same time, in this embodiment, the optimization process further includes: adding an attention module between the output of the backbone network and the input of the neck network in Yolov8. In this attention module, through grouped convolution including convolutional branches of different scales, feature extraction of different scales is respectively performed on the output of the backbone network in Yolov8 and then fused. After that, the fused features after fusion are processed through a channel attention mechanism. Then, after processing the result processed by the channel attention mechanism with a spatial attention mechanism, it is fused with the input of this attention module to obtain the output of the attention module.
[0023] This attention module captures road damage features of different sizes and textures through grouped convolution of different scales. After fusing information of different scales, it uses a channel attention mechanism to highlight important feature channels and suppress secondary feature channels to extract the channel information of the image. The core role of grouped convolution is to achieve efficient multi-branch feature extraction through feature grouping and parameter sharing while keeping the model lightweight. On this basis, the local sensitivity of features is further enhanced through spatial attention. Combining these, the network can simultaneously focus on global semantics and local details, thus obtaining a more sufficient feature representation. Therefore, this attention module has the advantages of spatial attention and channel attention and can effectively improve the performance of road damage detection.
[0024] As Figure 1 shown, in the optimized Yolov8 model, an attention module MESA is added between the output of the backbone network (i.e., after SPPF) and the input of the neck network (i.e., before Upsample). The structure of MESA refers to Figure 4 . Specifically, the convolutional branches of different scales in grouped convolution include: a convolutional branch mainly composed of a single 3x3 scale convolutional network and a convolutional branch mainly composed of two 3x3 scale convolutional networks. The grouped convolution starts from Figure 4 the Group Conv in
[0025] As Figure 4 shown, Add (Element-wise) is element-wise addition, which is a way to fuse the feature extraction results of different convolutional branches; the subsequent channel attention mechanism adopts the ECA attention mechanism, that is, Figure 4 the part from after Add (Element-wise) to Re-weight in Figure 4 ; after the ECA attention mechanism, an operation of fusing the results of average pooling and max pooling and the two poolings is performed to implement the processing of the spatial attention mechanism. For the processing result of the spatial attention mechanism, in this embodiment, through residual connection (that is,
[0026] in Residual Connection), it is fused with the output of the backbone network in Yolov8, that is, it is fused with the input of this attention module, and finally the output of this attention module is obtained. Moreover, in order to improve the effect of feature extraction, in this embodiment, the optimization process further includes: replacing the convolutional module for feature extraction in the backbone network of Yolov8 with an optimized feature extraction module; the optimized feature extraction module is obtained by replacing the original Ghost module closest to the output in the part of GhostNetV2 with Stride = 2 with an optimized Ghost module respectively;
[0027] The optimized Ghost module includes an original Ghost module and a fusion processing module for processing its output; the fusion processing module is used to fuse the output of the original Ghost module in the part of GhostNetV2 with Stride = 2 with the input of the part of GhostNetV2 with Stride = 2 as the output of this optimized Ghost module;
[0028] Figure 1Among them, Ghost ConvV2 is the optimized feature extraction module, and the C2f GhostV2 unit is the new unit obtained by replacing the convolution in each Bottleneck module in the C2f unit of the backbone network in Yolov8 with the optimized feature extraction module. The Ghost Bottleneck V2 in the C2f GhostV2 unit is the new module obtained by replacing the convolution in the Bottleneck module with the optimized feature extraction module. In the optimized feature extraction module, the structural principle of the optimized Ghost module is as shown in Figure 5 . Its structural difference from the existing GhostNetV2 lies in that on the basis of the output of the original Ghost module (i.e., Ghost module) in the part where Stride = 2, an operation of fusing with the input of the part where Stride = 2 in the GhostNetV2 structure is added (i.e., the operation of the Concat+Fusion1x1 Conv part in the part where Stride = 2 in Figure 5 ), and the other parts are the same as the structure of the existing GhostNetV2; Figure 5 For the specific principle example of DFC Attention in Figure 6 , Figure 5 For the specific principle example of the Ghost Module in Figure 7 , since they all belong to the prior art, they will not be elaborated here.
[0029] In addition, considering that road damage is a dynamic and continuously changing process, its occurrence, expansion, and deterioration are often the result of long-term accumulation. Most of the existing detection methods can only reflect the current road conditions and lack the prediction of the damage development trend. This makes it possible that when the road damage has not reached a serious level, the relevant departments may not be able to take maintenance measures in time, resulting in an increase in maintenance costs and an increase in traffic safety hazards. Therefore, the detection method relying solely on static image analysis cannot meet the growing actual needs. Nowadays, in addition to accurately detecting road damage, people also urgently need to accurately predict the future development trend of road damage in order to plan maintenance measures in advance and improve the safety and reliability of road use. Then in this embodiment, referring to Figure 8 , this road damage detection method further includes: according to the location information and category information of the damage identified from each frame of road image obtained in real time, matching each damage identified in each frame of road image through a target tracking algorithm, and determining the time series dataset corresponding to each damage according to the matching result and the time sequence corresponding to each frame of road image to be detected; The time series dataset corresponding to each damage stores a time series including the detection time when the damage was first detected, the location information and category information of each identified damage at that location, the feature vectors obtained respectively from all the extracted features in each single-frame road image identified each time, and the movement trajectory information of the damage at that location over time; From the time series dataset corresponding to a certain damage, select the damage features determined by a set number of data groups whose detection time corresponding to the identified single-frame road image is before the time node to be predicted and is closest to the time node to be predicted, and input them into the LSTM model trained with time series data samples to predict the damage features corresponding to the time node to be predicted.
[0030] In this way, the detection results of road damage can be made dynamic, and through the dynamic detection results, the prediction of the road damage situation at a specific time node can be realized by using the historical data of known road damage, so as to estimate future maintenance requirements in advance, optimize the maintenance plan, reasonably allocate manpower, equipment and materials, minimize the impact on traffic to the greatest extent, and improve the maintenance efficiency and resource utilization rate.
[0031] Among them, the feature vector obtained from all the extracted features in each single-frame road image identified each time is obtained by performing average pooling on all the damage feature vectors in the single frame, which can integrate the information of different damages and avoid distortion caused by extreme values of individual features.
[0032] Specifically, the target tracking algorithm used is the Deep SORT algorithm. Then, according to the location information of the damage identified from each frame of road image obtained in real time and the feature vector obtained from all the extracted features in each single-frame road image identified each time, the specific process of matching each damage identified in each frame of road image by the target tracking algorithm and correspondingly determining the time series dataset corresponding to each damage is as follows: 1) ID marking: In each frame, the Deep SORT algorithm uses the damage bounding box detected by the above-mentioned damage detection model (i.e., the location information of the identified damage) and the appearance features extracted by deep learning (i.e., the location information of the identified damage) to assign a unique ID to each damage. As time goes by, the Deep SORT tracks the damage areas with the same ID through the matching of movement trajectories and appearance features.
[0033] 2) Handling of newly added and disappeared damages: When a new damage area is detected (i.e., damage that fails to match the historical ID in the current frame), the Deep SORT algorithm assigns a new ID to it. This ensures that newly added road damages are recorded independently. If a damage area cannot be matched in consecutive frames, it is considered to have disappeared, and the tracking of it stops.
[0034] 3) Multi-frame tracking and temporal data construction: The sequence of video frames captured regularly provides damage information at different time points. By tracking damage areas with the same ID, the temporal characteristics of the damage areas (such as position changes, area changes, category changes, etc.) can be recorded. Based on these temporal data, the expansion speed, shape changes, and possible causes of the damage can be analyzed.
[0035] 4) Unified data management: The IDs and temporal characteristics of all damage areas are uniformly managed to form a time series dataset. Each record in the dataset includes the initial detection time of a certain damage (i.e., the detection time when the damage was first detected), the location (i.e., the positioning information of the damage each time it is recognized), the category (i.e., the category information of the damage each time it is recognized), the feature vector (i.e., the feature vector obtained based on all the features extracted from the single-frame road image recognized each time), and the trajectory information changing over time (i.e., the time series of the movement trajectory information of the damage changing over time).
[0036] After that, the above time series data of road damages are processed to prevent the existence of missing values, outliers, etc., to ensure the integrity and consistency of the road damage data. Then the time series dataset is standardized to prevent exceeding the standardized range during the above transformation process, which can further accelerate the convergence of the LSTM model used for predicting road damage data and improve the accuracy of model prediction.
[0037] The process example of training an LSTM model for predicting time series road damage data based on the processed dataset is as follows: First, the data needs to be arranged in a format suitable for the input of the LSTM model: the time series is divided into input sequences of a fixed length (such as the damage values at the past 10 time points) and the corresponding prediction targets (such as the damage value at the next time point). Then, the LSTM model is trained, with the input dimension being the number of features of the time series, the hidden layer used to capture time dependencies, and the output layer used to predict future damage values. Finally, the trained model is applied to the test set or the data to be predicted in the new time series dataset, and the future damage changes are predicted through forward propagation.
[0038] In addition, in this embodiment, the road images to be detected and the training samples are both obtained by shooting with a drone. Therefore, compared with the prior art where vehicle-mounted cameras are often used to collect road images, it is more suitable for applications in large-scale and complex scenarios. Moreover, vehicle-mounted cameras capture images at a relatively close ground position with a fixed shooting angle, resulting in limited image information obtained, which affects the training effect of the model. In this embodiment, a drone equipped with a high-resolution camera is used to periodically obtain road images. Since the drone can capture more extensive and multi-angle image information, it can effectively improve the detection coverage rate and accuracy. In addition, in this embodiment, the data corresponding to the road images to be detected or the training samples collected by the drone each time are accompanied by time stamps and geographical location information (such as GPS coordinates) for subsequent time series analysis.
[0039] Specifically, in this embodiment, images are extracted frame by frame from the video captured by the drone, a data set is constructed in chronological order, and the high-resolution images are cropped to adapt to the hardware processing capabilities. The cropped image patches are target-labeled using a labeling tool to mark common road damage categories such as cracks and potholes. Subsequently, the data is preprocessed to ensure that the format is adapted to the detection model. Finally, the data set is divided into a training set (training samples), a validation set, and a test set according to time periods, which are used for training the damage detection model, tuning parameter verification, and evaluation respectively.
[0040] Embodiment of Road Damage Detection System This embodiment provides a technical solution for a road damage detection system, including a processor. The processor stores executable program instructions, and is characterized in that the executable program instructions are used to be executed to implement the road damage detection method in the road damage detection method embodiment as described above.
[0041] The working principle and specific manner of the road damage detection system used in this embodiment have been described in detail in the above road damage detection method embodiment, so they will not be elaborated here.
[0042] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention.
Claims
1. A road damage detection method, characterized in that, Including: Inputting the road image to be detected into the damage detection model trained by training samples to obtain the localization information and category information of the identified damage; The damage detection model before training is obtained by optimizing the Yolov8 model; The optimization process includes: replacing each Bottleneck module in the C2f unit of the neck network in Yolov8 with an optimized feature fusion module; each optimized feature fusion module contains one of the Bottleneck modules and an adaptive processing module for processing its output; in the adaptive processing module, feature dimensionality reduction is performed on the output of the Bottleneck module through the SE attention mechanism, and then the results after the feature dimensionality reduction by the SE attention mechanism are processed respectively through a first branch including depthwise separable convolution and a second branch including a lightweight pyramid pooling module and a multi-layer perceptron, and the processing results of the first and second branches are fused to obtain the output of the adaptive processing module.
2. The road damage detection method according to claim 1, characterized in that The optimization process further includes: adding an attention module between the output of the backbone network and the input of the neck network in Yolov8. In this attention module, feature extraction of different scales is performed on the output of the backbone network in Yolov8 respectively through grouped convolution including convolutional branches of different scales and fused. Then, the fused features after fusion are processed through a channel attention mechanism. After that, the results processed by the channel attention mechanism are processed through a spatial attention mechanism and then fused with the input of this attention module to obtain the output of the attention module.
3. The road damage detection method according to claim 1 or 2, characterized in that, The optimization process further includes: replacing the convolutional module for feature extraction in the backbone network of Yolov8 with an optimized feature extraction module; the optimized feature extraction module is obtained by respectively replacing the original Ghost module closest to the output in the part of GhostNetV2 with Stride = 2 with an optimized Ghost module; The optimized Ghost module contains an original Ghost module and a fusion processing module for processing its output; the fusion processing module is used to fuse the output of the original Ghost module in the part of GhostNetV2 with Stride = 2 and the input of the part of GhostNetV2 with Stride = 2 as the output of this optimized Ghost module; And, replacing the convolution in each Bottleneck module in the C2f unit of the backbone network in Yolov8 with the optimized feature extraction module.
4. The road damage detection method according to claim 1 or 2, characterized in that, Also including: According to the localization information of the damage identified from each frame of road image obtained in real time and the feature vectors obtained from all the features extracted from each single-frame road image identified each time, each damage identified in each frame of road image is matched through a target tracking algorithm. According to the matching results and the time sequence corresponding to each frame of road image to be detected, a time series dataset corresponding to each damage is determined; The time series dataset corresponding to each damage stores a time series including the detection time when the damage was first detected, the location information and category information of each identified damage at that location, the feature vectors obtained respectively from all the extracted features in each single-frame road image identified, and the motion trajectory information of the damage at that location over time; From the time series dataset corresponding to a certain damage, select the damage features determined by a set number of data groups whose detection time corresponding to the identified single-frame road image is before the time node to be predicted and closest to the time node to be predicted, and input them into the LSTM model trained with time series data samples to predict the damage features corresponding to the time node to be predicted.
5. The road damage detection method according to claim 1 or 2, characterized in that, Both the road images to be detected and the training samples are obtained by shooting with drones.
6. The road damage detection method according to claim 2, characterized in that, The convolution branches of different scales in the grouped convolution include: a convolution branch mainly composed of a single 3x3 scale convolution network and a convolution branch mainly composed of two 3x3 scale convolution networks.
7. The road damage detection method according to claim 4, wherein The target tracking algorithm is the Deep SORT algorithm.
8. The road damage detection method according to claim 1 or 2, characterized in that, The method of processing the result after feature reduction by the SE attention mechanism through the second branch including a lightweight pyramid pooling module and a multi-layer perceptron includes: Process the result after feature reduction by the SE attention mechanism through the lightweight pyramid pooling module, and then input it into the multi-layer perceptron to obtain the processing result of the multi-layer perceptron to obtain the output of the second branch.
9. A road damage detection system, including a processor, wherein executable program instructions are stored in the processor, characterized in that The executable program instructions are used to be executed to implement the road damage detection method according to any one of claims 1-8.