Pavement disease identification method based on YOLO-BSAM
By introducing a double-layer attention module and a spatial attention module in the YOLO model, the problem that the YOLO model is difficult to distinguish disease characteristics and redundant information in complex scenarios is solved, and the accuracy and efficiency of road surface disease detection are achieved.
Patent Information
- Application Number
- CN202411924419.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
When the existing YOLO model deals with road disease detection in complex scenarios, it is difficult to distinguish disease characteristics and redundant information, resulting in a decrease in recognition accuracy and recall in medium and long-range pictures taken by manual inspection.
The road surface disease recognition method based on YOLO-BSAM is adopted, and by introducing the double-layer attention module BAM and the spatial attention module SAM, the attention weight is adjusted, feature extraction ability and object detection performance are enhanced, and adaptive adjustments at different information levels are adaptive.
It improves the detection accuracy of road surface diseases in complex images, enhances the model's detection ability of multiple types of diseases, and improves the detection efficiency and recognition accuracy in actual road detection scenarios.
Smart Images

Figure CN120047788A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of structural health monitoring, and particularly relates to a method for identifying road diseases based on YOLO-BSAM. Background Art
[0002] Road diseases not only affect the appearance of the road surface and driving comfort, but also shorten the overall service life and performance of the road surface. Therefore, early detection of road diseases can reduce the economic cost of road maintenance and ensure the safety of vehicles and drivers on the road. The detection of traditional road diseases mainly relies on manual inspections. However, affected by factors such as the complexity of working conditions and the professionalism of personnel, its detection efficiency is low, the operating cost is high, and the detection window period is long.
[0003] Currently, based on convolutional neural networks, computer vision has been widely used in road disease detection to automatically identify the types and locations of road diseases in images using the network. For example, the common YOLO model relies on convolution to extract disease features and is widely used in machine (inspection vehicle) inspections. It has a good effect when dealing with close-range images with obvious disease features and less background noise. However, due to the receptive field limitation of the convolutional layer in the YOLO model architecture, when dealing with features in different spatial positions and channels, it cannot dynamically adjust the weights, which makes it difficult to distinguish disease features and redundant information in complex scenarios. Therefore, when applied to manual inspections, when faced with manually taken images with a far viewing distance, a complex background, and unclear disease features, the recognition accuracy and recall rate of the YOLO model significantly decrease and it is difficult to meet the detection requirements.
[0004] In view of this, how to make the neural network model meet the requirements of real road images and improve the recognition accuracy of medium and long-distance disease images is an urgent problem to be solved. Summary of the Invention
[0005] In view of the above technical problems in the prior art, the present invention proposes a method for identifying road diseases based on YOLO-BSAM, which can accurately identify road diseases in close-range, medium and long-distance images, and can adapt to road disease images collected by both manual inspections and machine inspections.
[0006] To achieve the above invention object, the present invention adopts the following technical solutions:
[0007] A method for identifying road diseases based on YOLO-BSAM, comprising the following steps:
[0008] S1. Establish a YOLO-BSAM neural network model; the YOLO-BSAM neural network model includes an input, a backbone, a neck, and a head; the BSAM module is inserted at the end of the backbone;
[0009] S2. Use the augmentations.letterbox function to crop, scale, and complete the input image to a unified size to adapt to model calculations. Secondly, use the Mosaic function to enrich data samples through random scaling and arrangement. Finally, use the Focus module to slice the image by channels, splice it in the channel direction, and then perform convolution to shuffle and integrate features to obtain a downsampled feature map.
[0010] S3. The backbone Backbone includes a convolution module, a pooling module, and an attention module, which are responsible for the feature output, feature enhancement, and target feature extraction of the image respectively.
[0011] S4. The neck adopts an improved feature pyramid network PANet structure to enhance the feature extraction ability and target detection performance through a two-way information exchange mechanism.
[0012] S5. The head outputs the detection results.
[0013] Further, as a preferred technical solution of the present invention, in S3, the convolution module uses multiple convolutions to extract the target features in the feature map, namely the edges, textures, and shapes of the target. It loops with a single convolution completed by the cooperation of the Conv module and the C3 module. The single convolution includes: the Conv module uses Conv2d to establish a basic convolution layer to extract the local spatial information of the feature map; the Conv module uses the normalization operation BN to reduce the data gradient and help convergence; the SILU module introduces the Sigmoid non-linear activation function to extract the target feature map; the C3 module divides the target feature map into 2 sub-parts for independent convolution, and then forms the output feature map through cross-channel connection.
[0014] Further, as a preferred technical solution of the present invention, in S3, the pooling module uses fast spatial pyramid pooling SPPF to process different parts of the output feature map through multiple layers of 3×3 convolutions, thereby retaining the multi-scale information of the feature map, further reducing the information loss during the pooling process, and forming the final feature map.
[0015] Further, as a preferred technical solution of the present invention, in S3, the attention module uses the BSAM mechanism to adjust the attention weights, making the network pay more attention to the real disease features in the final feature map, eliminating noise pollution, and thus improving the recognition accuracy of diseases. The BSAM module includes a double-layer attention module BAM and a spatial attention module SAM.
[0016] Further, as a preferred technical solution of the present invention, the BAM module generates attention weights by calculating the relationship between input data, and can significantly improve the calculation efficiency through region screening and refinement processing. Its mechanism is as follows:
[0017]
[0018] Among them, Q is the query matrix used to locate the feature positions; K is the key matrix used to indicate the relationships between positions; V is the value matrix, i.e., the feature values. d k is the normalization constant; the Softmax function calculates the inner product of Q and K for each feature to represent its correlation, and normalizes it to form a probability distribution, thereby expressing the importance of each feature.
[0019] Furthermore, as a preferred technical solution of the present invention, the SAM module emphasizes the features of important regions by generating a spatial attention map; performs average pooling and max pooling on the feature map F' respectively in the channel dimension to obtain two intermediate feature maps and concatenates these two feature maps in the channel dimension, and then convolves them to form a feature map; uses the Sigmoid activation function to generate the spatial attention map M s ; multiplies the spatial attention map M s element-wise with the input feature map F 0 to obtain the final output feature map;
[0020]
[0021] In the formula, f and σ represent the Sigmoid activation function.
[0022] Furthermore, as a preferred technical solution of the present invention, in S4, 1×1 convolutional kernels are used to convolve each level of feature maps to align them in the channel dimension; the top-down path is adopted to transfer the global structure information of the image to the low-level feature maps, and the bottom-up path is adopted to transfer the bottom-level information such as edges and textures to the high-level feature maps, thereby simultaneously improving the model's ability to understand the target background and the ability to detect the target boundaries.
[0023] Furthermore, as a preferred technical solution of the present invention, in S5, non-maximum suppression (NMS) processing is adopted to remove the detected bounding boxes with large overlaps and only retain the bounding box with the highest confidence, thereby ensuring the accuracy and reliability of the detection.
[0024] For the pavement disease recognition method based on YOLO-BSAM described in the present invention, compared with the prior art by adopting the above technical solutions, it has the following technical effects:
[0025] (1) In the present invention, BSAM provides a dual attention module BAM and a spatial attention mechanism SAM. BAM generates adaptive attention weights, extracts only the regions with high relevance to diseases for calculation in regional-level routing, thereby reducing the computational amount, and then uses a dense matrix to enhance disease features in label-level routing, improving the accuracy of disease feature extraction. SAM generates a spatial attention map to screen the disease region features, further enhancing the detection ability of multiple types of diseases. BSAM combines the spatial attention mechanism and the channel attention mechanism, adaptively adjusts the attention weights for different information levels, and realizes simultaneous attention to global structural information and underlying feature information, so that the model can more accurately detect road diseases in complex images.
[0026] (2) Compared with the same type of networks (YOLOv5, YOLOv5+SE, YOLOv5+ECA, YOLOv5+GAM), YOLO-BSAM of the present invention can better adapt to different lighting, perspective, and size scenarios, has high sensitivity to multi-object detection, and strong robustness to environmental noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is the architecture diagram of the YOLO-BSAM model according to the embodiment of the present invention;
[0028] Figure 2 It is the schematic diagram of the BSAM mechanism according to the embodiment of the present invention;
[0029] Figure 3 It is the architecture diagram of the BAM sub-module according to the embodiment of the present invention;
[0030] Figure 4 It is the architecture diagram of the SAM sub-module according to the embodiment of the present invention;
[0031] Figure 5 It is the schematic diagram of the neck PANet structure according to the embodiment of the present invention;
[0032] Figure 6 It is the schematic diagram of the implementation process according to the embodiment of the present invention;
[0033] Figure 7 It is the long-range image set according to the embodiment of the present invention;
[0034] Figure 8 It is the schematic diagram of the YOLO-BSAM disease recognition result according to the embodiment of the present invention;
[0035] Figure 9 It is the P-R curve diagram according to the embodiment of the present invention;
[0036] Figure 10 It is the comparison schematic diagram of YOLO-BSAM and similar models according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The following will further explain and illustrate the present invention in detail with reference to the accompanying drawings, so that those skilled in the art can understand the present invention more deeply and be able to implement it. However, the following is only used to explain the present invention by referring to examples and does not limit the present invention.
[0038] A road surface disease recognition method based on YOLO-BSAM includes the following steps: S1. Establish a YOLO-BSAM neural network model; the YOLO-BSAM neural network model includes an input, a backbone, a neck, and a head; the BSAM module is inserted at the tail of the backbone; as Figure 1 shown; S2. Use the augmentations.letterbox function to crop, scale, and complete the input image to a unified size to adapt to model calculation; secondly, use the Mosaic function to enrich data samples by random scaling and arrangement; finally, use the Focus module to slice the image by channels and splice it in the channel direction, and then perform convolution to shuffle and integrate features to obtain a downsampled feature map; S3. The backbone includes a convolution module, a pooling module, and an attention module; respectively responsible for the feature output, feature enhancement, and target feature extraction of the image; S4. The neck adopts an improved feature pyramid network PANet structure to enhance the feature extraction ability and target detection performance through a two-way information exchange mechanism; S5. The head outputs the detection result.
[0039] The convolution module extracts the target features in the feature map using multiple convolutions, that is, the edges, textures, and shapes of the target; it loops through a single convolution completed by the cooperation of the Conv module and the C3 module; the single convolution includes: the Conv module uses Conv2d to establish a basic convolutional layer to extract the local spatial information of the feature map; the Conv module uses a normalization operation (BatchNormalization, BN) to reduce the data gradient and help convergence; the SILU module introduces a Sigmoid non-linear activation function to extract the target feature map; the C3 module divides the target feature map into 2 sub-independent convolutions, and then forms an output feature map through cross-channel connection to reduce information loss and improve feature expression.
[0040] The pooling module uses fast spatial pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) to process different parts of the output feature map through multiple layers of 3×3 convolutions, thereby retaining the multi-scale information of the feature map, further reducing information loss during the pooling process, and forming the final feature map.
[0041] The attention module uses the BSAM mechanism disclosed in the present invention. By adjusting the attention weights, the network pays more attention to the true disease features in the final feature map, eliminates noise pollution, and thus improves the recognition accuracy of diseases. BSAM includes two sub-modules as Figure 2 shown; the Bilevel Attention Module (BAM) and the Spatial Attention Module (SAM).
[0042] BAM generates attention weights by calculating the relationships between input data, and can significantly improve the calculation efficiency through region screening and refinement processing. Its mechanism is as follows:
[0043]
[0044] Among them, Q is the query matrix for locating feature positions; K is the key matrix for indicating the relationships between positions; V is the value matrix, that is, the feature values. d k is the normalization constant. The Softmax function calculates the inner product of Q and K of each feature to represent its correlation, and normalizes it to form a probability distribution, so as to express the importance of each feature (attention score). BAM adopts a sparse attention strategy, screens important regions (with high attention scores) for calculation, avoids full-range calculation, thereby reducing the calculation amount and improving GPU friendliness.
[0045] The structure of BAM is as Figure 3 . Specifically: DWConv3×3 extracts local spatial features, stabilizes the feature distribution through regularization (LayerNormalization, LN), and ensures the stability of the training process and the rapid convergence of the model. The Bi-level Routing Attention in the bi-level channel attention uses region-level routing to screen important regions, enhances the attention to diseases, and at the same time reduces the amount of computation. Then, label-level attention is used for further refinement processing within the region, and dense matrix operations are used to enhance the target features and improve the adaptability in complex scenarios. The Multi-layer Perceptron further captures the data relationships through non-linear transformation and feature fusion, and enhances the model's expression ability.
[0046] The SAM module further enhances the screening of target features by generating a spatial attention map and emphasizing the features of important regions. Respectively perform average pooling and max pooling on the feature map F' in the channel dimension to obtain two intermediate feature maps and Concatenate these two feature maps in the channel dimension, and then perform convolution to form a feature map. Use the Sigmoid activation function to generate the spatial attention map M s . 4) The spatial attention map Ms Multiply element-wise with the input feature map F 0 to obtain the final output feature map.
[0047]
[0048] where f and σ represent the Sigmoid activation function.
[0049] The structure of SAM is as Figure 4 shown. Weight adjustment (Re-weight) re-adjusts the importance weights of each pixel in the feature map. Channel pooling (Channel Pool) uses global average pooling and global max pooling to obtain intermediate feature maps respectively. 7×7Conv convolves the intermediate feature maps using a 7×7 convolutional kernel. 4) Use the Sigmoid function to adjust the weights (Re-weight) again.
[0050] The neck adopts an improved Feature Pyramid Network (Path Aggregation Network, PANet) structure, enhancing the feature extraction ability and object detection performance through a two-way information exchange mechanism as Figure 5 shown. Use 1×1 convolutional kernels to convolve the feature maps at each level to align them in the channel dimension. Adopt a top-down path to transfer the global structural information of the image to the low-level feature maps, and a bottom-up path to transfer the low-level information such as edges and textures to the high-level feature maps, thereby improving the model's ability to understand the target background and the detection ability of the target boundary simultaneously.
[0051] The head outputs the detection results, which are processed using non-maximum suppression (abbreviated as NMS) to remove the detected bounding boxes with large overlaps and only retain the bounding box with the highest confidence, thereby ensuring the accuracy and reliability of the detection.
[0052] As Figure 6 shown, taking the long-range image set captured by manual inspection as an example, the implementation method of the road surface disease recognition method based on YOLO-BSAM is described in detail;
[0053] Long-range images refer to images in which the disease area accounts for less than 15% of the image and contains a complex background as Figure 7 shown. This working condition contains 6521 original images, and the disease types are potholes, reticular cracks, and strip cracks.
[0054] Use the augmentations.letterbox function to crop, scale, and complete the images to 512×512 pixels; use rotation, symmetry, and scaling to expand the image set to 9876 images.
[0055] Use the Labelme tool to annotate the images, select the disease areas by bounding boxes, and indicate their disease forms, namely potholes, map cracks, and strip cracks. Randomly divide the image set into a training set, a validation set, and a test set at a ratio of 7:2:1, obtaining 6913 images in the training set and 2963 images in the validation set.
[0056] Set the parameters of the YOLO-BSAM model (batch_size = 8, epochs = 200, iou = 0.45, img_size = 512x512, Lr = 0.01, warmup_epochs = 10), and input the training set into the YOLO-BSAM model for training to obtain a road disease recognition model.
[0057] Use the validation set to evaluate the model performance, including average precision, recall rate, accuracy, and computational efficiency.
[0058] Optimize the model performance by adjusting the hyperparameters. The specific hyperparameters include batch_size, epochs, img_size, and warmup_epoches.
[0059] Obtain the optimized model, and use the test set to output the recognition results as Figure 8 The P-R of the detection results is as Figure 9 . The results show that the average precision (AP) of potholes is 82.7%, the AP of map cracks is 62.4%, the AP of strip cracks is 44.1%, and the mean average precision (mAP) of all targets is 63.1%. Compared with similar models as Figure 10 , the precision of YOLO-BSAM is improved by about 10%, the recall rate is improved by 5%, and the average precision is improved by 7%.
[0060] The present invention discloses a road disease recognition method based on YOLO-BSAM, which can accurately recognize road diseases in near-distance, medium-distance, and far-distance images, and can adapt to road disease images collected by both manual inspections and machine inspections at the same time. By using a bilevel spatial attention module (BSAM), the YOLO model can achieve more flexible content perception and computational allocation, filter out redundant information at the wide-area level, and then perform more comprehensive feature filtering in the channel area to ensure that the computational network focuses on areas related to road diseases and improve the disease recognition accuracy. At the same time, BSAM can use sparse matrix operations to reduce redundant calculations, significantly reduce the memory usage, and improve the computational efficiency.
[0061] The results of field tests show that in the inspection work of road diseases, this method significantly reduces the impact of complex environmental noise on the recognition results, and remarkably improves the detection efficiency and recognition accuracy of road disease targets in actual road detection scenarios. It can simultaneously recognize the near-view disease images collected by machines and the medium- and long-view disease images collected manually. The average accuracy for near-distance images reaches 91.8%, and the average accuracy for complex far-distance images reaches 67.6%.
[0062] The specific implementation schemes described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation schemes of the present invention, and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention shall fall within the scope of protection of the present invention.
Claims
1. A road surface disease identification method based on YOLO-BSAM, characterized in that: The following steps are involved: S1. Establish a YOLO-BSAM neural network model. The YOLO-BSAM neural network model includes input, backbone, neck and head. The BSAM module is inserted into the tail of the backbone. S2. Use the augmentations.letterbox function to crop, scale, and complete the input image to a uniform size to adapt to the model calculation; secondly, use the Mosaic function to enrich the data samples by random scaling and arrangement; finally, use the Focus module to slice the image by channel, splice it in the channel direction, and then perform convolution to shuffle and integrate the features to obtain the downsampled feature map; S3, Backbone includes convolution module, pooling module and attention module; they are responsible for image feature output, feature enhancement and target feature extraction respectively; S4 and neck adopt an improved feature pyramid network PANet structure to enhance feature extraction capability and target detection performance through a two-way information exchange mechanism; S5. The head outputs the detection result.
2. The road surface disease identification method based on YOLO-BSAM according to claim 1 is characterized in that: In S3, the convolution module uses multiple convolutions to extract the target features in the feature map, namely the edge, texture and shape of the target; The single convolution performed by the Conv module and the C3 module is circulated; the single convolution includes: the Conv module uses Conv2d to establish a basic convolution layer to extract the local spatial information of the feature map; The Conv module uses the normalization operation BN to reduce data gradients and help convergence; the SILU module introduces the Sigmoid nonlinear activation function to extract the target feature map; the C3 module divides the target feature map into two sub-parts for independent convolution, and then forms an output feature map through cross-channel connections.
3. The road surface disease identification method based on YOLO-BSAM according to claim 2 is characterized in that: In S3, the pooling module uses fast spatial pyramid pooling SPPF to process different parts of the output feature map through multiple layers of 3×3 convolutions, thereby retaining the multi-scale information of the feature map, thereby reducing the information loss in the pooling process and forming the final feature map.
4. The road surface disease identification method based on YOLO-BSAM according to claim 3 is characterized in that: In S3, the attention module uses the BSAM mechanism to adjust the attention weight so that the network pays more attention to the real disease features in the final feature map and removes noise pollution, thereby improving the recognition accuracy of the disease; The BSAM module consists of a double-layer attention module BAM and a spatial attention module SAM.
5. The road surface disease identification method based on YOLO-BSAM according to claim 4 is characterized in that: The BAM module generates attention weights by calculating the relationship between input data, which can significantly improve computational efficiency through region screening and refinement. Its mechanism is as follows: Where Q is the query matrix, used to locate the feature position; K is the key matrix, used to indicate the relationship between positions; V is the value matrix, i.e., the eigenvalue. k is a normalization constant; the Softmax function calculates the inner product of Q and K of each feature to express its correlation, and normalizes it to form a probability distribution, thereby expressing the importance of each feature.
6. The road surface disease identification method based on YOLO-BSAM according to claim 4 is characterized in that: The SAM module generates a spatial attention map to emphasize important regional features; it performs average pooling and maximum pooling on the feature map F' in the channel dimension to obtain two intermediate feature maps and The two feature maps are concatenated in the channel dimension and then convolved to form a feature map; the Sigmoid activation function is used to generate the spatial attention map M s ; The spatial attention map M s Multiply element-by-element with the input feature map F0 to obtain the final output feature map; Where f, σ represents the Sigmoid activation function.
7. The road surface disease identification method based on YOLO-BSAM according to claim 1, characterized in that: In S4, a 1×1 convolution kernel is used to convolve the feature maps at each level to align them in the channel dimension; a top-down path is used to transfer the global structural information of the image to the low-level feature map, and a bottom-up path is used to transfer the underlying information such as edges and textures to the high-level feature map, thereby improving the model's ability to understand the target background and detect the target boundary at the same time.
8. The road surface disease identification method based on YOLO-BSAM according to claim 1, characterized in that: In S5, non-maximum suppression (NMS) processing is used to remove detected bounding boxes with large overlaps and only retain the bounding boxes with the highest confidence, thereby ensuring the accuracy and reliability of detection.