Road damage object detection method based on attention mechanism ensemble learning network
By adopting a multi-source image detection method based on attention mechanism and integrated learning in road damage detection, the problem of difficulty in detecting small-scale wear and large-scale damage at the same time in the existing technology is solved, and unified detection and efficient detection of road damage are achieved.
Patent Information
- Application Number
- CN202310371571.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-04-03
AI Technical Summary
The prior art is difficult to detect small-scale wear and large-scale damage of roads simultaneously, and it is difficult to obtain sufficient road damage data in sparsely populated areas or places where natural disasters occur, resulting in unsatisfactory detection results.
The road damage object detection method of the mass source image based on attention mechanism and integrated learning is adopted. The YOLOv5, YOLOv5_SE, and YOLOv5_CA subnets are set in parallel, combined with the attention module and integrated learning, and the small degree of wear and large degree of damage are uniformly marked, and the mass source images such as on-board vehicles and the Internet are used for detection.
It realizes unified inspection of road damage, improves the accuracy and robustness of detection, can effectively detect road damage targets in different regions and conditions, and provides efficient and convenient daily road maintenance and disaster response support.
Smart Images

Figure CN116597270B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and relates to a road damage target detection method, and specifically to a multi-source image road damage target detection method based on an attention mechanism integrated learning network. Background Art
[0002] As basic transportation facilities, roads carry the normal operation of various aspects of social order such as education, medical care, freight, work, and tourism, and promote economic and social development. Convenient, fast, and efficient road damage detection methods are conducive to daily maintenance and disaster repair of roads, and provide strong protection for the health of the road system. According to the different causes, road damage can be divided into two categories: one is the small degree of wear and tear caused by weather, transportation volume, etc., such as cracks and potholes; the other is the sudden large degree of damage due to collapse, landslides, etc., such as collapse and falling rocks.
[0003] At present, in terms of daily road maintenance, due to the time-consuming and labor-intensive manual visual inspection and the expensive 3D scanner equipment, there is a growing number of studies on the use of convenient and efficient image detection methods to detect road diseases. In terms of road disaster response, most studies are based on remote sensing or drone platforms to obtain macroscopic damage information of road networks in overhead images, lacking specific on-site information. Although social media data can provide detailed on-site information for disaster response, its application rarely focuses on road damage.
[0004] With the widespread application and rapid development of deep learning technology in the field of computer vision, deep learning networks can achieve high-precision target detection, semantic segmentation, image classification and other tasks, providing key technical means for road damage detection. At the same time, in order to make the deep learning network focus on important features in the image, plug-and-play attention modules are widely used. The attention module can be inserted into any position in the network to adaptively increase the weight of key features and reduce the interference of redundant features. In addition, single model detection often has limited effects and is prone to misjudgment and omission for complex data. Ensemble learning methods can effectively alleviate this problem, that is, by integrating the detection results of multiple single models, compensating for the defects of single models and obtaining higher accuracy.
[0005] At present, the main problems of road damage detection are as follows:
[0006] (1) Whether it is daily road maintenance or disaster response, the mainstream method is to use images as data sources to train deep learning networks to detect minor wear or major damage, but not both types of damage at the same time. There are similarities between the two, but there is currently no solution to unify the two tasks.
[0007] (2) Although traditional data collection methods such as vehicle-mounted cameras and drones can accurately and completely obtain images of road damage in urban areas, they are difficult to achieve the same results in sparsely populated areas. Natural disasters often occur in remote areas, and road damage data in this area is very scarce.
[0008] (3) The road-related area in the image may only occupy a part, and the other redundant parts interfere with the model detection. Simply applying the existing deep learning model to road damage detection is difficult to achieve ideal results. Moreover, the road conditions in different regions are different, the types and degrees of road damage are different, and the shooting angles, weather, lighting, etc. are not uniform. Deep learning is data-driven, so the generalization ability and robustness of a single deep learning network are difficult to cope with such complex road images. Summary of the invention
[0009] In order to solve the above technical problems, the present invention proposes a multi-source image road damage target detection method based on attention mechanism and ensemble learning. The present invention comprehensively utilizes multi-source images such as vehicle-mounted and Internet images, combines attention mechanism and ensemble learning, and performs road damage target detection.
[0010] The technical solution adopted by the method of the present invention is: a road damage target detection method based on an attention mechanism integrated learning network, comprising the following steps:
[0011] Step 1: Input the image to be detected into the attention mechanism-based integrated learning network for detection;
[0012] The attention mechanism-based integrated learning network is composed of YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks set in parallel;
[0013] The YOLOv5 subnetwork includes two parts: a backbone network and a detection head; the backbone network includes five down-sampling feature extraction blocks and an SPPF module, the first feature extraction block is a convolution layer with a convolution kernel size of 6 and a step size of 2; the second to fifth feature extraction blocks are all a combination of a convolution layer with a convolution kernel size of 3 and a step size of 2 and a C3 module, and the C3 module is composed of a number of serial convolutions and jump connections; a layer of SPPF module is arranged behind the fifth feature extraction block; the detection head includes multi-scale feature aggregation and detection, firstly, the three feature maps extracted by the backbone network are subjected to top-down feature fusion, then the feature enhancement is performed from bottom to top, and finally the feature maps of the three scales are detected;
[0014] The YOLOv5_SE subnetwork includes a backbone network, an SE attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the SE attention module is added before the last layer of the SPPF module in the backbone network to highlight the important features of the channel dimension and suppress the redundant features of the channel dimension; the SE attention module first performs global average pooling on the input feature map along the channel dimension, then obtains the weighted value of the channel through a fully connected layer and an activation function, and finally multiplies it with the input feature map;
[0015] The YOLOv5_CA subnetwork includes a backbone network, a CA attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the CA attention module is added after the last layer of the SPPF module in the backbone network, and is used to highlight the important features at the pixel level and suppress the redundant features at the pixel level; the CA attention module first averagely pools the input feature map along the length and width directions, concatenates them, and then passes through convolution, batch normalization and activation function, and then divides them into feature maps in the length and width directions, respectively passes through convolution and activation function, and finally performs pixel-by-pixel weighted processing on the input feature map;
[0016] Step 2: Perform non-maximum suppression on the output results of the YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks to obtain the final damaged target detection results.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] (1) This paper proposes a method for establishing a training database of road damage images from crowd-source platforms such as the Internet, vehicle-mounted cameras, and drones to overcome the problem of insufficient road damage data. At the same time, by uniformly labeling minor wear and major damage, the two tasks of daily road maintenance and road disaster response are unified.
[0019] (2) The present invention proposes a road damage target detection method that combines attention mechanism and ensemble learning. The attention module is used to allow a single deep learning network to focus on key features, improve the detection ability of a single model, and deal with the problem of redundant areas that are not roads in the image. Then, multiple different improved single models are integrated using ensemble learning to improve generalization ability and robustness, and alleviate complex and diverse image problems.
[0020] (3) After obtaining the integrated learning model combined with the attention mechanism, it is possible to detect road damage targets in multi-source images. This system not only provides an efficient and convenient means of road disease detection for daily road maintenance, but also lays the foundation for quickly obtaining accurate and detailed damage information for road disaster response. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a diagram of an integrated learning network structure based on an attention mechanism according to an embodiment of the present invention;
[0022] Figure 2 This is a flow chart of an integrated learning network training based on an attention mechanism according to an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of object detection and annotation of multi-source images of road damage according to an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments. Obviously, the described embodiments are only examples of a part of the present invention, not all examples. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] Please see Figure 1 The present invention provides a road damage target detection method based on an attention mechanism integrated learning network, comprising the following steps:
[0026] Step 1: Input the image to be detected into the attention mechanism-based integrated learning network for detection;
[0027] The attention mechanism-based integrated learning network of this embodiment is composed of YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks set in parallel;
[0028] The YOLOv5 subnetwork of this embodiment includes two parts: a backbone network and a detection head; the backbone network includes five down-sampling feature extraction blocks and an SPPF module, the first feature extraction block is a convolution layer with a convolution kernel size of 6 and a step size of 2; the second to fifth feature extraction blocks are all a combination of a convolution layer with a convolution kernel size of 3 and a step size of 2 and a C3 module, and the C3 module is composed of several series convolutions and jump connections; a layer of SPPF module is arranged behind the fifth feature extraction block; the detection head includes multi-scale feature aggregation and detection, firstly, the three feature maps extracted by the backbone network are subjected to top-down feature fusion, then the feature enhancement is performed from bottom to top, and finally the feature maps of the three scales are detected;
[0029] The YOLOv5_SE subnetwork of this embodiment is an optimized network based on the YOLOv5 network, including a backbone network, an SE attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the SE attention module is added before the last layer of the SPPF module of the backbone network, and is used to highlight the important features of the channel dimension and suppress the redundant features of the channel dimension; the SE attention module first performs global average pooling on the input feature map along the channel dimension, then obtains the weighted value of the channel through a fully connected layer and an activation function, and finally multiplies it with the input feature map;
[0030] The YOLOv5_CA sub-network of this embodiment is an optimized network based on the YOLOv5 network, including a backbone network, a CA (Coordinate Attention) attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the CA attention module is added after the last layer of the SPPF module of the backbone network, and is used to highlight important features at the pixel level and suppress redundant features at the pixel level; the CA attention module first averagely pools the input feature map along the length and width directions, concatenates them, and then passes through convolution, batch normalization and activation function, and then divides them into feature maps in the length and width directions, respectively passes through convolution and activation function, and finally performs pixel-by-pixel weighted processing on the input feature map;
[0031] Step 2: Perform non-maximum suppression on the output results of the YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks to obtain the final damaged target detection results.
[0032] The non-maximum suppression process implemented in this embodiment includes the following sub-steps:
[0033] Step 2.1: The processing object is the original prediction results of the three sub-networks YOLOv5, YOLOv5_SE, and YOLOv5_CA in the attention mechanism-based integrated learning network, that is, the prediction results of the last three layers of feature maps of each sub-network, a total of nine layers of feature map prediction results; multiply the confidence of each prediction box by the category prediction probability to obtain the probability of each category of the prediction box;
[0034] Step 2.2: Set the threshold Tc. For a single prediction box, when the category probability is greater than Tc, keep the prediction box, otherwise discard the prediction box; when there are multiple category probabilities greater than Tc, take the category corresponding to the maximum value as the category of the prediction box;
[0035] Step 2.3: Set the threshold Ti. For the retained prediction boxes, select the box with the maximum category probability and calculate the IOU values of all other boxes with it. If the value is greater than Ti, it means that the overlap between the two boxes is too high and they are deleted. Otherwise, they are retained.
[0036] Step 2.4: For the remaining prediction boxes, repeat step 2.3 until all prediction boxes are traversed and the IOU values between any two boxes are less than Ti;
[0037] Step 2.5: The remaining prediction box after iteration is the final prediction result.
[0038] In this embodiment, Tc and Ti are set based on empirical values, with Tc being 0.2 and Ti being 0.5.
[0039] Please see Figure 2 The attention mechanism-based integrated learning network of this embodiment is a trained network, and its training process includes the following sub-steps:
[0040] Step 1.1: Obtain a number of road damage crowd-source images and perform pre-processing; crowd-source images come from vehicle-mounted cameras, drones, and the Internet;
[0041] This embodiment obtains multi-source images of road damage from platforms such as vehicle-mounted cameras, drones, and the Internet, removes duplicate images based on the MD5 values of the image files, and then scales all images to a square size of 640 pixels.
[0042] Step 1.2: Manually annotate the road damage multi-source images preprocessed in step 1.1 and establish a sample database;
[0043] like Figure 3 As shown, the annotation format is the target detection box, and the categories include small-scale wear and tear such as cracks and potholes and large-scale damage such as collapse and falling rocks.
[0044] Step 1.3: Input the data in the sample database into the attention mechanism-based ensemble learning network, and train the attention mechanism-based ensemble learning network;
[0045] In this embodiment, the three sub-networks YOLOv5, YOLOv5_SE and YOLOv5_CA are trained separately according to the same hyperparameters; the training hyperparameters include: the image length and width are both 640 pixels, the batch size is 16 images, the entire data set is trained for 100 rounds, the stochastic gradient descent optimizer is used, and the initial learning rate is 0.01.
[0046] The loss function used in the training process of this embodiment consists of three parts: the first is the positioning loss, which calculates the CIOU loss value between the predicted box and the true box; the second is the classification loss, which calculates the binary cross entropy loss function value between the predicted category and the true category; the third is the confidence loss, which calculates the binary cross entropy loss function value of whether the target is contained.
[0047] After obtaining the integrated learning model combined with the attention mechanism, this embodiment can detect road damage targets in multi-source images. Regardless of whether the road image comes from the Internet, vehicle-mounted cameras, drones and other platforms, the model can detect damaged targets in the form of bounding boxes, and has category information of small-degree wear and tear such as cracks and potholes or large-degree damage such as collapse and falling rocks, thereby realizing the unification of the two tasks of daily road maintenance and road disaster response.
[0048] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A road damage target detection method based on an attention mechanism ensemble learning network, characterized in that: The following steps are involved: Step 1: Input the image to be detected into the attention mechanism-based integrated learning network for detection; The attention mechanism-based integrated learning network is composed of YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks set in parallel; The YOLOv5 subnetwork includes a backbone network and a detection head; the backbone network includes five down-sampling feature extraction blocks and an SPPF module, the first feature extraction block is a convolution layer with a convolution kernel size of 6 and a step size of 2; the second to fifth feature extraction blocks are all a combination of a convolution layer with a convolution kernel size of 3 and a step size of 2 and a C3 module, and the C3 module is composed of a number of serial convolutions and jump connections; A layer of SPPF module is arranged behind the fifth feature extraction block; the detection head includes multi-scale feature aggregation and detection, firstly, the three feature maps extracted by the backbone network are subjected to top-down feature fusion, then feature enhancement is performed from bottom to top, and finally the feature maps of the three scales are detected; The YOLOv5_SE subnetwork includes a backbone network, an SE attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the SE attention module is added before the last layer of the SPPF module in the backbone network to highlight the important features of the channel dimension and suppress the redundant features of the channel dimension; the SE attention module first performs global average pooling on the input feature map along the channel dimension, then obtains the weighted value of the channel through a fully connected layer and an activation function, and finally multiplies it with the input feature map; The YOLOv5_CA subnetwork includes a backbone network, a CA attention module and a detection head, wherein the backbone network and the detection head are consistent with the YOLOv5 network; the CA attention module is added after the last layer of the SPPF module in the backbone network, and is used to highlight the important features at the pixel level and suppress the redundant features at the pixel level; the CA attention module first averagely pools the input feature map along the length and width directions, concatenates them, and then passes through convolution, batch normalization and activation function, and then divides them into feature maps in the length and width directions, respectively passes through convolution and activation function, and finally performs pixel-by-pixel weighted processing on the input feature map; Step 2: Perform non-maximum suppression on the output results of the YOLOv5, YOLOv5_SE, and YOLOv5_CA sub-networks to obtain the final damaged target detection results.
2. The method for detecting road damage targets based on an attention mechanism ensemble learning network according to claim 1, characterized in that: In step 1, the attention mechanism-based integrated learning network is a trained network, and its training process includes the following sub-steps: Step 1.1: Obtain a number of road damage crowd-source images and perform pre-processing; the crowd-source images come from vehicle-mounted cameras, drones, and the Internet; Step 1.2: Manually annotate the road damage multi-source images preprocessed in step 1.1 and establish a sample database; Step 1.3: Input the data in the sample database into the attention mechanism-based integrated learning network, and train the attention mechanism-based integrated learning network; The three sub-networks YOLOv5, YOLOv5_SE and YOLOv5_CA are trained separately according to the same hyperparameters; The training hyperparameters include: image length and width are both 640 pixels, batch size is 16 images, the entire dataset is trained for 100 rounds, stochastic gradient descent optimizer is used, and the initial learning rate is 0.
01.
3. The method for detecting road damage targets based on an attention mechanism ensemble learning network according to claim 2, characterized in that: In step 1.1, the preprocessing is to use the MD5 value of the file to deduplicate the image and scale all the images to the same size.
4. The method for detecting road damage targets based on an attention mechanism ensemble learning network according to claim 2, characterized in that: In step 1.2, the manually annotated label format is a target bounding box, that is, the minimum circumscribed rectangular box of the damage instance; the damage categories include small damage such as cracks and potholes and large damage such as collapse and rockfall.
5. The method for detecting road damage targets based on an attention mechanism ensemble learning network according to claim 2, characterized in that: In step 1.3, the loss function used in the training process consists of three parts: the first is the positioning loss, which calculates the CIOU loss value between the predicted box and the true box; the second is the classification loss, which calculates the binary cross entropy loss function value between the predicted category and the true category; the third is the confidence loss, which calculates the binary cross entropy loss function value of whether the target is contained.
6. The method for detecting road damage targets based on an attention mechanism ensemble learning network according to any one of claims 1 to 5, characterized in that: In step 2, the non-maximum suppression process is specifically implemented by including the following sub-steps: Step 2.1: The processing object is the original prediction results of the three sub-networks YOLOv5, YOLOv5_SE, and YOLOv5_CA in the attention mechanism-based integrated learning network, that is, the prediction results of the last three layers of feature maps of each sub-network, a total of nine layers of feature map prediction results; multiply the confidence of each prediction box by the category prediction probability to obtain the probability of each category of the prediction box; Step 2.2: Set the threshold Tc. For a single prediction box, when the category probability is greater than Tc, keep the prediction box, otherwise discard the prediction box; when there are multiple category probabilities greater than Tc, take the category corresponding to the maximum value as the category of the prediction box; Step 2.3: Set the threshold Ti. For the retained prediction boxes, select the box with the maximum category probability and calculate the IOU values of all other boxes with it. If the value is greater than Ti, it means that the overlap between the two boxes is too high and they are deleted. Otherwise, they are retained. Step 2.4: For the remaining prediction boxes, repeat step 2.3 until all prediction boxes are traversed and the IOU values between any two boxes are less than Ti; Step 2.5: The remaining prediction box after iteration is the final prediction result.
Citation Information
Patent Citations
Lightweight road defect detection method based on improved YOLOv5
CN115187583A
Multi-scale road target detection method based on region focusing
CN115690714A