Road disease identification and grading early warning method based on DynAM-YOLO11
By introducing the DynAM-YOLO11 network model and combining it with multi-scale local enhancement and dynamic parameter adaptation modules, the problems of low road damage identification accuracy and untimely warning in existing technologies are solved. Efficient and accurate road damage identification and graded warning are achieved, improving the efficiency and safety of urban road management.
Patent Information
- Application Number
- CN202510622238.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies for road hazard identification and early warning have problems such as low detection accuracy, single function, high cost, poor environmental adaptability, and untimely early warning, making it difficult to meet the high-frequency and high-precision detection needs of modern urban road networks.
The DynAM-YOLO11 network model is introduced. By introducing the DynAM module into the YOLO11 network, the fine-grained feature extraction capability is enhanced. Combined with the multi-scale local enhancement module, the dynamic parameter adaptation module, and the cross-channel attention fusion module, the recognition accuracy of complex damage is improved and timely warnings are provided.
It improves the detection accuracy and environmental adaptability of road hazard identification, realizes low-cost and high-efficiency road hazard graded warning, optimizes maintenance efficiency, and improves road safety and urban traffic management capabilities.
Smart Images

Figure CN120823474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic management and smart city technology, and specifically to a road hazard identification and graded early warning method based on DynAM-YOLO11. Background Art
[0002] In the process of new urbanization and the digital transformation of transportation infrastructure, real-time perception and intelligent assessment of road health have become strategic requirements for the modernization of urban governance systems. Traditional manual inspections for road damage identification suffer from low efficiency and subjectivity. Early warning systems rely primarily on manual inspections, regular road surface assessments, and monitoring with specialized equipment, but these still suffer from high costs and poor timeliness. Traditional identification and early warning methods struggle to meet the high-frequency, high-precision inspection demands of modern urban road networks.
[0003] With the continuous development of deep learning, road damage identification methods based on deep learning have developed rapidly: for example, deep learning algorithms such as convolutional neural networks (CNN) are used to automatically analyze road surface images, and network structures such as YOLO (You Only Look Once) achieve efficient and real-time detection. However, they are insufficient in processing fine-grained features, enhancing spatial information, and detecting multi-scale targets. For example, see application number CN202411427707.0, which discloses a road damage detection method based on YOLO-V5. By adding a detection layer, the model's ability to recognize small targets and densely damaged areas is improved. This method realizes the automated detection of road damage and can output damage type and coordinate information, significantly improving detection efficiency. Although the system uses deep learning technology and improves traditional methods, the invention is not technologically advanced enough, the detection precision and accuracy are not high, and the function is single, only realizing disease detection, lacking graded warning and PCI quantitative evaluation, and cannot achieve comprehensive road health status monitoring.
[0004] Therefore, there is an urgent need to provide a road disease identification and graded warning system that can improve detection accuracy while combining with the warning mechanism to provide timely graded warning information on road diseases, provide support for urban traffic management and road maintenance, and solve the problems of high equipment cost, poor environmental adaptability, and untimely warning in existing technologies. Summary of the Invention
[0005] In order to overcome the shortcomings of the existing technical problems, the purpose of the present invention is to provide a road damage identification and graded warning method based on DynAM-YOLO11. The DynAM module is introduced into the YOLO11 network to ensure the speed of inference and warning while enhancing the fine-grained feature extraction capability and improving the recognition accuracy of complex damage; and timely warning of identified road damage is provided.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A road hazard identification and graded warning method based on DynAM-YOLO11 includes the following steps:
[0008] (1) Collect road condition images of different types of roads in various environments, pre-process the images, and annotate the road damage types and coordinates to construct a data set;
[0009] (2) Constructing a DynAM-YOLO11 network model for identifying road damage, wherein the DynAM-YOLO11 network model uses the YOLO11 model as a basic network and introduces a DynAM module into the neck network to enhance the local perception capability of minor damage; and training the DynAM-YOLO11 network model using the data set from step (1);
[0010] (3) Labelme is used to mark the road surface area that needs to be inspected, and the trained DynAM-YOLO11 network model is used to identify the road surface area that needs to be inspected, segment the road damage area, and obtain an image with road damage conditions; the road surface area that needs to be inspected and the road damage area are covered with masks respectively, and the areas of the road surface area that needs to be inspected and the road damage area at the pixel level are calculated, thereby calculating the road damage rate DR and the road damage index PCI;
[0011] (4) Based on the Highway Technical Condition Assessment Standard, road damage is graded using PCI values, and the level of damage in the road surface area that needs to be inspected is determined, and corresponding warnings are issued.
[0012] In the present invention, the road disease types include cracks, block cracks, longitudinal cracks, transverse cracks, subsidence, bulges, potholes, looseness, oil overflow, and repairs.
[0013] In the present invention, the DynAM module introduces a multi-scale local enhancement module, a dynamic parameter adaptation module, and a cross-channel attention fusion module based on SimAM. The DynAM module works as follows:
[0014] First, the original features generate basic attention based on the attention calculation logic of the mean and variance of SimAM, and the dynamic parameter adaptive module adjusts the smoothing strength of the attention calculation to generate spatial attention;
[0015] Secondly, the original features generate local attention through a multi-scale local enhancement module;
[0016] Then, the original features are processed through global average pooling and multi-layer perceptron in the cross-channel attention fusion module to generate channel attention, and the local attention, spatial attention and channel attention are fused to obtain fused attention.
[0017] Finally, the original features are enhanced by fusing attention to enhance important areas and suppress unimportant areas to obtain enhanced features; the enhanced features are finally combined with the original features through residual connections to output the final results.
[0018] In the present invention, the dynamic parameter adaptation module generates a dynamic smoothing coefficient λ through global average pooling and a fully connected layer, which is used to adjust the smoothing strength of the attention calculation. The original features generate basic attention based on SimAM and spatial attention with the dynamic smoothing coefficient λ.
[0019] In the present invention, the global average pooling compresses the feature map into 1x1, and then generates the λ value through two convolutional layers, captures the channel importance through global average pooling, and dynamically adjusts the denominator smoothing strength;
[0020] The mathematical expression for dynamic lambda generation is:
[0021] λ=λ min +(λ max -λ min )·σ(W2δ(W1(AvgPool(x)))
[0022] Among them, λ min ,λ max are the minimum and maximum values of the smoothing coefficient respectively;
[0023] σ is the Hardsigmoid function, with an output range of [0,1];
[0024] δ is the ReLU activation function;
[0025] W1 and W2 are learnable parameter matrices respectively;
[0026] AvgPool() is the global average pooling operation;
[0027] x is the original input feature map.
[0028] In the present invention, the multi-scale local enhancement module realizes multi-scale perception of local context through the design of depth-wise separable convolution plus channel remapping. The multi-scale local enhancement module generates local attention. The specific mathematical expression is as follows:
[0029] LocalAtt=σ(Conv1x1(ReLU(BN(DWConv3x3(x)))))
[0030] Where: x is the original feature map of the input;
[0031] DWConv is the depthwise convolution part of the depthwise separable convolution;
[0032] BN: Batch normalization stabilizes training;
[0033] ReLU: introduces nonlinearity;
[0034] Conv1x1: channel remapping;
[0035] σ is the Hardsigmoid function, with an output range of [0,1];
[0036] LocalAtt is local attention.
[0037] In the present invention, the working process of the cross-channel attention fusion module is as follows:
[0038] First, the original features are processed through global average pooling and multi-layer perceptron to generate channel attention. The mathematical expression for channel attention calculation is as follows:
[0039] ChannelAtt=σ(MLP((AvgPool(x)))
[0040] Among them, ChannelAtt is channel attention;
[0041] σ is the Hardsigmoid function;
[0042] MLP is a multi-layer perceptron;
[0043] AvgPool is average pooling;
[0044] x is the original input feature;
[0045] Then, the channel attention is expanded to the same shape as the enhanced spatial attention. The mathematical expression of channel attention expansion is as follows:
[0046] ChannelAtt expanded =ChannelAtt·1 H×W
[0047] Among them, ChannelAtt expanded is the expanded channel attention; H is the height; W is the width;
[0048] Finally, the spatial attention, local attention and the multiplied channel attention are fused through multiplication interaction to enhance the ability to perceive details and express features. Finally, the fused attention is generated through 1x1 convolution and Sigmoid function. The formula is as follows:
[0049] FusedAtt=σ(W f ·(SpatialAtt☉LocalAtt☉ChannelAtt expanded )+b f )
[0050] Among them, FusedAtt is fused attention; ⊙ represents element-by-element multiplication.
[0051] In the present invention, the calculation formula of the road surface damage rate DR is:
[0052]
[0053] Among them, A i is the area of the i-th type of road damage area, A is the road surface area that needs to be detected, i represents the type of road damage, ω i is the weight or conversion factor of the i-th type of road damage, i0 is the total number of damage types;
[0054] The calculation formula of the pavement damage index PCI is:
[0055]
[0056] Among them, a0 and a1 are model parameters, which are determined according to the road surface type.
[0057] In the present invention, the PCI value is used to classify road damage as follows:
[0058] When PCI ≥ 85, the road surface is in excellent condition, with little damage and no need for repair;
[0059] When 70≤PCI<85, the road surface is in good condition with minor damage and preventive maintenance is required;
[0060] When 55≤PCI<70, there is some damage and appropriate repair is required;
[0061] When 40≤PCI<55, the damage is serious and should be repaired in time;
[0062] When PCI is less than 40, the damage is extremely serious and large-scale repair or reconstruction work must be carried out immediately.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1. The present invention introduces the spatially weighted slicing DynAM module. Based on SimAM, this module incorporates a multi-scale local enhancement module, a dynamic parameter adaptation module, and a cross-channel attention fusion module. It designs a multi-dimensional feature fusion mechanism to enhance fine-grained feature extraction capabilities, improve the recognition accuracy of complex damage, and construct a lightweight inference engine to achieve a breakthrough improvement in detection speed. While ensuring the speed of inference and warning, it also improves detection accuracy and environmental adaptability.
[0065] 2. This invention combines road hazard identification with early warning, quantitatively evaluating road hazard conditions and categorizing them into early warning levels based on the Highway Technical Condition Assessment Standard (JTG5210-2018). Accurately monitoring road hazard conditions and issuing timely early warnings in complex urban environments, this low-cost, high-efficiency technology improves road safety, optimizes maintenance efficiency, and enhances the ability of municipal infrastructure systems to address road hazard issues, providing a strong guarantee for the safe operation of urban roads. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 Flow chart of the method of the present invention.
[0067] Figure 2 Schematic diagram of the DynAM-YOLO11 network structure of the present invention.
[0068] Figure 3 Schematic diagram of the multi-scale local context module structure of the present invention.
[0069] Figure 4 Schematic diagram of the dynamic parameter adaptive module of the present invention.
[0070] Figure 5 Schematic diagram of the cross-channel attention fusion module introduced in this invention.
[0071] Figure 6 This is a road surface disease identification effect diagram detected by the present invention.
[0072] Figure 7 This is a rendering of the pavement cracks detected by the present invention.
[0073] Figure 8 The user interface design diagram of the road disease identification and PCI index calculation and graded early warning of the present invention.
[0074] Figure 9 The present invention is a complex scene to be detected on a rainy day and affected by reflection.
[0075] Figure 10 The DynAM-YOLO11 detection effect diagram of the present invention in a complex scene affected by reflections on a rainy day.
[0076] Figure 11The YOLO11-seg detection effect diagram of the present invention in a complex scene affected by reflections on a rainy day. DETAILED DESCRIPTION
[0077] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0078] The present invention discloses a road hazard identification and graded early warning method based on DynAM-YOLO11, which is as follows:
[0079] (1) Data collection and preprocessing
[0080] High-definition cameras are strategically placed along different road types, such as highways, urban arterials, and rural roads, based on road length and traffic flow. Road surface images are regularly collected at different times of day, including during different light intensities during the day, at night, and in unusual weather conditions (such as after rain or snow), to capture road conditions in a variety of environments. Images are annotated using the LabelMe annotation tool (a tool for image annotation, primarily used to create labeled data for computer vision tasks). This data includes information such as past damage types and damage coordinates, providing a richer sample for model training and helping the model learn the development patterns and characteristic changes of road damage. The collected raw images are then preliminarily screened to remove blurry, overexposed, or heavily occluded images. After screening, LabelMe is used to annotate road damage areas, clearly identifying damage types (such as cracks and potholes) and recording the location and extent of the damage, thus forming a high-quality annotated dataset.
[0081] In terms of data augmentation, random cropping technology is used to set the cropping size within a reasonable range to ensure that the cropped area can contain sufficient crack or pothole feature information. This allows the model to grasp the manifestation of damaged areas at different positions and sizes. At the same time, the range of rotation angles is reasonably set. Considering that the direction of cracks or potholes may vary depending on the shooting angle, moderate rotation operations help the model learn the characteristics of cracks or potholes in different directions, thereby improving the model's rotation invariance. In addition, by setting the flip angle, the model can learn the characteristics of cracks or potholes in various directions, further improving the model's adaptability to image changes. Finally, random simulation technology is used to adjust brightness and contrast (changing in the HSV color space) to simulate cracks or potholes under different lighting conditions, thereby increasing the diversity of the dataset.
[0082] (2) Building the DynAM-YOLO11 network model and training
[0083] DynAM-YOLO11 uses the YOLO11 model as its basic network and is based on the SimAM attention mechanism. However, SimAM also has some shortcomings, which mainly include: limited local feature capture capability, and the global receptive field design is insensitive to small targets; the fixed smoothing coefficient (λ) cannot adapt to complex scenarios; insufficient interaction between channels, ignoring channel semantic associations; high implicit computing costs, which affect the inference speed; limited feature fusion capabilities, and lack of effective combination of multi-dimensional features.
[0084] To address the above shortcomings, the present invention uses the YOLO11 model as the basic network and introduces the DynAM module into the neck network of the YOLO11 model. The DynAM module is based on SimAM and introduces a dynamic parameter adaptation module, a multi-scale local enhancement module, and a cross-channel attention fusion module into SimAM; residual connections are used to retain the original features, and dynamic adjustments and feature fusion are performed in many aspects. The dynamic parameter adaptation module dynamically generates hyperparameters in the attention calculation, enabling the module to adaptively adjust according to the input features. Multi-scale local enhancement enhances local contextual information through multi-scale depthwise separable convolution, improving the ability to capture details. Cross-channel attention combines spatial and channel attention through the cross-channel attention fusion mechanism to enhance the expressiveness of features. By integrating these components, the DynAM module significantly improves the performance of visual tasks while maintaining its lightweight.
[0085] First, the original features generate basic attention based on the attention calculation logic of the mean and variance of SimAM, and then the smoothing strength of the attention calculation is adjusted through the dynamic parameter adaptive module to generate spatial attention; the original features generate local attention through the multi-scale local enhancement module; the original features generate channel attention through the global average pooling and multi-layer perceptron (MLP) of the cross-channel attention fusion module, and the local attention, spatial attention and channel attention are fused to obtain fused attention; the original features are element-wise multiplied by the fused attention to enhance important areas and suppress unimportant areas to obtain enhanced features; finally, the enhanced features are combined with the original features through the residual connection to output the final result. The mathematical expression of the residual connection is as follows:
[0086] Output=identity*FusedAtt+identity
[0087] Among them, Output means output;
[0088] identity represents the original characteristics;
[0089] FusedAtt represents the fused attention weight.
[0090] Through the above steps, the DynAM-YOLO11 module achieves local feature enhancement, dynamic smoothing coefficient generation, channel attention fusion, and feature preservation of residual connections. These improvements significantly enhance the model's ability to capture detailed features while maintaining computational efficiency.
[0091] The specific steps to build the DynAM-YOLO11 network model are as follows:
[0092] (1) Dynamic parameter adaptation module: Generates basic attention based on the attention calculation logic of SimAM's mean and variance. The original SimAM uses a fixed smoothing coefficient λ (default 1e-4) and cannot adapt to different scenarios. The energy function of the original SimAM is:
[0093]
[0094] Where: z i is the value of the i-th neuron in the feature map.
[0095] μ is the local mean (context) of the feature map.
[0096] S 2 is the local variance of the feature map.
[0097] λ is a smoothing factor (to prevent the denominator from being zero).
[0098] The specific process is as follows Figure 4 As shown in Figure 1, the dynamic parameter adaptation module generates a dynamic smoothing coefficient λ through global average pooling and fully connected layers. This is used to adjust the smoothing strength of the attention calculation, enabling the module to adapt to the input features. Global average pooling compresses the feature map into 1x1, and then generates the λ value through two convolutional layers. Global average pooling captures channel importance and dynamically adjusts the denominator smoothing strength. The mathematical expression for dynamic λ generation is:
[0099] λ=λ min +(λ max- λ min )·σ(W2δ(W1(AvgPool(x)))
[0100] Among them, λ min ,λ max are the minimum and maximum values of the smoothing coefficient, respectively.
[0101] σ is the Hardsigmoid function (output range [0,1]).
[0102] δ is the ReLU activation function.
[0103] W1 and W2 are both learnable parameter matrices.
[0104] AvgPool() is a global average pooling operation.
[0105] x is the original input feature map.
[0106] DynAM dynamically generates λ through a lightweight dynamic parameter adaptive module, automatically increasing λ for simple backgrounds (to suppress noise) and reducing λ for complex scenes (to enhance detail response), thereby enhancing local features and improving the ability to capture details. It then generates spatial attention SpatialAtt based on the basic attention and dynamic parameter λ generated by the attention calculation logic of SimAM's mean and variance.
[0107] (2) Multi-scale local enhancement module
[0108] The multi-scale local enhancement module achieves multi-scale perception of local context at a low computational cost through the lightweight design of "depthwise separable convolution + channel remapping", achieving a balance between accuracy and efficiency in road damage detection tasks.
[0109] The specific process is as follows Figure 3 As shown in the figure, although the module is named "multi-scale", it actually indirectly achieves multi-scale perception by combining different operations: depthwise separable convolution is implemented through grouped convolution (groups = channels), achieving channel-independent convolution, reducing the number of parameters and keeping the number of channels unchanged. Each group of channels performs a separate 3x3 convolution. The local module extracts high-frequency details (crack edges) through 3x3 depthwise convolution, and then obtains channel-level global information through 1x1 convolution, establishing cross-channel associations. The implicit global context improves the ability to capture high-frequency information such as road disease edges and textures at small scales (i.e., local details). Local attention is generated through the multi-scale local enhancement module. The specific mathematical expression is as follows:
[0110] LocalAtt=σ(Conv1x1(ReLU(BN(DWConv3x3(x)))))
[0111] Where: x is the original feature map of the input.
[0112] DWConv is the depthwise convolution part of the depthwise separable convolution.
[0113] BN: Batch Normalization stabilizes training.
[0114] ReLU: Introduces nonlinearity.
[0115] Conv1x1: channel remapping.
[0116] σ is the Hardsigmoid function (output range [0,1]).
[0117] LocalAtt is local attention.
[0118] The original features are extracted from local spatial features through depthwise separable convolution, and BN and ReLU are used for normalization, accelerated convergence and introduction of nonlinear activation. Finally, they are output through channel remapping and Sigmoid activation function, and the feature response strength is adjusted according to the importance of local context as local attention.
[0119] (3) Cross-channel attention fusion module. First, the original features are processed through global average pooling and multi-layer perceptron (MLP) to generate channel attention weights. The mathematical expression for channel attention weight calculation is as follows:
[0120] ChannelAtt=σ(MLP(AvgPool(x)))
[0121] Among them, ChannelAtt is channel attention; σ is the Hardsigmoid function (output range [0,1]); MLP is the multi-layer perceptron; AvgPool is the average pooling; x: the original feature map of the input.
[0122] Then, the channel attention is expanded to the same shape as the enhanced spatial attention. The mathematical expression of channel attention expansion is as follows:
[0123] ChannelAtt expanded =ChannelAtt·1 H×W
[0124] Among them, ChannelAtt expanded is the expanded channel attention; H is the height; W is the width.
[0125] Finally, the spatial attention, local attention and the multiplied channel attention are fused through multiplication interaction to enhance the ability to perceive details and express features; finally, the fused attention is generated through 1x1 convolution and Sigmoid function. The formula is as follows:
[0126] FusedAtt=σ(W f ·(SpatialAtt☉LocalAtt☉ChannelAtt expanded )+b f )
[0127] Among them, σ is the Hardsigmoid function, W f , b f are the weight and bias of 1x1 convolution respectively, ⊙ represents element-by-element multiplication, FusedAtt is fused attention, SpatialAtt is spatial attention, LocalAtt is local attention, ChannelAtt expanded is the expanded channel attention
[0128] (3) Road condition detection and calculation of road technical condition index
[0129] Based on the warning standards in the Highway Technical Condition Assessment Standard (JTG5210-2018), and taking into account factors such as road age and traffic volume, warning areas are defined on the map. The labelme tool is used to mark warning boxes in key areas to clearly define the road surface to be inspected. The system then overlays the road damage areas identified by DynAM-YOLO11 with the road surface areas to be inspected.
[0130] Based on the Highway Technical Condition Assessment Standard (JTG5210-2018), this paper introduces the pavement damage rate (DR) and the pavement damage index (PCI) as evaluation indicators to comprehensively and accurately assess road damage conditions. The calculation formula for the pavement damage rate (DR) is:
[0131]
[0132] Among them, A i is the area of the i-th type of road damage area, A is the road surface area that needs to be detected, i represents the type of road damage, ω i is the weight or conversion coefficient of the i-th type of road disease (wherein, the conversion coefficient ω for the automated detection of road diseases such as cracks, block cracks, subsidence, bulges, potholes, and looseness is i =1.0, the conversion factor ω for automated detection of longitudinal and transverse cracks i 2.0; Conversion coefficient ω for automated detection of road damage such as oil spills and repairs i is 0.2); i0 is the total number of road damage types;
[0133] The calculation formula of the pavement damage index PCI is:
[0134]
[0135] Among them, a0 and a1 are model parameters, which are determined according to the pavement type. For example, a0 is 15.00 and a1 is 0.412 for asphalt pavement; a0 is 10.66 and a1 is 0.461 for cement concrete pavement.
[0136] Calculate the ratio of road damage area to road surface area that needs to be inspected and determine the warning level. Read road damage from the specified file, set the road surface area that needs to be inspected and create a mask. Then create a binary mask (an image used to mark a specific area in an image) of cracks or potholes and warning areas with the same size as the image. After obtaining the mask of the crack and pothole areas in the warning area, use the calculate_area function to calculate the number of pixels with a median value of 1 in the mask, and obtain the area of the crack and pothole areas and the road surface area that needs to be inspected respectively. Then use the calculate_area function to calculate the area of the intersection area. Finally, use a specific image analysis algorithm to accurately calculate the relationship between the road damage area and the road surface area that needs to be inspected at the pixel level to calculate the road damage rate DR and the road damage condition index PCI.
[0137] Road damage conditions are assessed by calculating DR and PCI values. A higher PCI value indicates better road damage; conversely, a lower PCI value indicates worse road damage. The DR value directly reflects the proportion of road damage area within the inspected or surveyed area. A higher DR value indicates more severe road damage.
[0138] (IV) Road hazard classification and early warning
[0139] Different warning thresholds are set based on the PCI level. When PCI ≥ 85, the road surface is in excellent condition, with little damage and no need for repair. When PCI ≤ 70 < 85, the road surface is in good condition with minor damage, and preventive maintenance is appropriate. When PCI ≤ 55 < 70, some damage has occurred and appropriate repair is required. When PCI ≤ 55 < 55, the damage is severe and prompt repair is required. When PCI < 40, the damage is extremely severe and large-scale repair or reconstruction must be carried out immediately. If the PCI or DR value reaches or exceeds a certain warning threshold, the warning level is adjusted accordingly, providing a more reliable basis for road maintenance decisions.
[0140] Once an early warning is triggered, the road maintenance management platform pushes the early warning information to the relevant road maintenance departments and staff, and sends a text message notification at the same time. The early warning information includes the specific location of the road disease (determined by the camera position and image coordinates), the type of damage, the PCI assessment level, etc. The road maintenance department initiates the corresponding handling plan according to the early warning level. In the case of a mild early warning, maintenance personnel are arranged to conduct on-site inspections and record in the near future; in the case of a moderate early warning, a small repair team is organized, the corresponding repair materials and equipment are prepared, and repairs are carried out as soon as possible; in the case of a severe early warning, traffic control measures are immediately implemented, large-scale construction equipment and professional personnel are deployed, and emergency repair work is carried out to ensure road traffic safety and normal traffic.
[0141] Implementation Case 1
[0142] This example uses a Windows 11 operating system, an NVIDIA RTX 4090 GPU, and Python, integrated with the PyTorch deep learning framework. The constructed DynAM-YOLO11 network model is loaded into the configured training environment, and the network model parameter file is initialized. The experimental training parameters are as follows: 200 training iterations, 8 batches, 1 object category, an initial learning rate of 0.01, the Auto optimizer, and a uniform image size of 640*640.
[0143] To ensure effective dataset utilization and objective evaluation of algorithm performance, this case used a publicly available online dataset for annotation and data augmentation. After data preprocessing, the dataset was divided into a training set and a validation set in an 8:2 ratio. This dataset included 700 images in the pothole damage category, including 560 training images and 140 validation images.
[0144] In order to intuitively verify the effectiveness of the improved model in this paper, Figure 6 and Figure 7 The renderings of potholes and cracks identified by the model of the present invention are respectively displayed, which can more accurately detect road defects and accurately determine their locations and categories, thereby more effectively responding to challenges in actual road applications.
[0145] To comprehensively evaluate the improved model's performance in road defect detection, we selected several mainstream YOLO models for comparative analysis. This comprehensive comparison was conducted while ensuring consistent experimental conditions and datasets. The experimental results are shown in Table 1.
[0146] Table 1 Comparison of training indicators of different models
[0147]
[0148]
[0149] In terms of detection accuracy, the DynAM-YOLO11 model achieved an outstanding mAP50 value of 78.1%, surpassing both YOLO11-seg (75.0%) and YOLOv8-seg (76.1%). Comparative experimental data shows that DynAM-YOLO11 demonstrates superior overall performance in object detection tasks, outperforming both YOLO11-seg and YOLOv8-seg models in both precision (77.6%) and recall (73.8%). In particular, DynAM-YOLO11 achieved a 3.6% improvement in recall over YOLOv8-seg, demonstrating a lower missed detection rate. While YOLO11-seg outperformed DynAM-YOLO11 in accuracy (85.0%), its recall (63.7%) and mAP50 (75.0%) were 10.1% and 3.1% lower, respectively. It is worth noting that while YOLOv8-seg's recall (70.2%) and mAP50 (76.1%) are slightly higher than YOLO11-seg, they are still lower than DynAM-YOLO11. Overall, DynAM-YOLO11 demonstrates advantages in detection accuracy and object coverage through improvements to its model structure and training strategy.
[0150] In the actual application of road damage detection, the diversity and complexity of damage types place higher demands on the generalization ability of the detection model. The present invention enhances environmental adaptability by fusing DynAM with the YOLO11 neck network, while more accurately identifying these subtle differences, thereby improving detection accuracy. In addition, the present invention introduces the PCI indicator and early warning mechanism to quantitatively evaluate road damage. The application diagram of the present invention's road damage identification and PCI indicator calculation and graded early warning is shown in the figure below. Figure 8 As shown in the figure, after selecting the trained model, it provides fast and accurate road disease identification and graded warning functions in multiple scenarios such as images, videos, and cameras, which to a certain extent solves the problems of poor real-time performance and delayed manual detection.
[0151] Field tests have verified that this solution significantly improves road damage identification under typical urban road conditions in terms of accuracy, recall, map50, map50-95, and other indicators compared to previous models. By accurately calculating PCI and DR values, reasonably setting warning thresholds, and conducting detailed assessments of different pavement materials and damage types, the present invention achieves comprehensive and accurate assessment and early warning of road damage.
[0152] In general, the present invention improves the detection accuracy of road damage, assesses the road damage status more comprehensively and accurately, and can still issue early warnings in time on rainy days. It has both timeliness and environmental adaptability, providing strong support for road maintenance decisions.
[0153] Implementation Case 2
[0154] Detection object: potholes on the road.
[0155] Environmental parameters
[0156] Light conditions: cloudy.
[0157] Road interference: reflections from accumulated water, dynamic vehicles, and debris.
[0158] Target features: small potholes, irregular edges, and water artifacts.
[0159] Case Description
[0160] In complex urban road scenes, the detection differences between DynAM-YOLO and YOLO11-seg are significant:
[0161] Anti-interference ability
[0162] DynAM-YOLO in complex environments with water reflection and gravel interference ( Figure 10 ) still maintains high confidence detection, and its multi-scale feature fusion mechanism effectively distinguishes real potholes from reflective artifacts; while YOLO11-seg( Figure 11 ) shows obvious misjudgment of similar scenes, indicating that traditional segmentation networks are sensitive to optical noise.
[0163] In summary, DynAM-YOLO11 significantly improves detection robustness in complex weather conditions, such as rainy and cloudy conditions, through dynamic parameter adjustment, multi-scale local perception, and cross-dimensional feature fusion. Core improvements create a synergistic effect: Dynamic Lambda addresses illumination variations, local enhancement extracts occluded objects, channel fusion suppresses reflections, and residual connections protect weak signals. Actual tests demonstrate that this solution maintains high accuracy even in harsh low-contrast, high-noise environments, providing a reliable solution for all-weather road detection in intelligent transportation systems.
[0164] The above is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A road hazard identification and graded warning method based on DynAM-YOLO11, characterized in that: The following steps are involved: (1) Collect road condition images of different types of roads in various environments, pre-process the images, and annotate the road damage types and coordinates to construct a data set; (2) Constructing a DynAM-YOLO11 network model for identifying road damage, wherein the DynAM-YOLO11 network model uses the YOLO11 model as a basic network and introduces a DynAM module into the neck network to enhance the local perception capability of minor damage; and training the DynAM-YOLO11 network model using the data set from step (1); (3) Use labelme to mark the road surface area that needs to be inspected, use the trained DynAM-YOLO11 network model to identify the road surface area that needs to be inspected, segment the road damage area, and obtain an image with road damage; Perform masking on the road surface area to be inspected and the road damage area respectively, calculate the area of the road surface area to be inspected and the road damage area at the pixel level, and thus calculate the road damage rate DR and road damage index PCI; (4) Based on the Highway Technical Condition Assessment Standard, road damage is graded using PCI values, and the level of damage in the road surface area that needs to be inspected is determined, and corresponding warnings are issued.
2. The road hazard identification and graded warning method based on DynAM-YOLO11 according to claim 1 is characterized in that: The types of road damage include cracks, block cracks, longitudinal cracks, transverse cracks, subsidence, bulges, potholes, looseness, oil spills, and repairs.
3. The road hazard identification and graded early warning method based on DynAM-YOLO11 according to claim 1 is characterized in that: The DynAM module is based on SimAM and introduces a multi-scale local enhancement module, a dynamic parameter adaptation module, and a cross-channel attention fusion module. The DynAM module works as follows: First, the original features generate basic attention based on the attention calculation logic of the mean and variance of SimAM, and the dynamic parameter adaptive module adjusts the smoothing strength of the attention calculation to generate spatial attention; Secondly, the original features generate local attention through a multi-scale local enhancement module; Then, the original features are processed through global average pooling and multi-layer perceptron in the cross-channel attention fusion module to generate channel attention, and the local attention, spatial attention and channel attention are fused to obtain fused attention. Finally, the original features are enhanced by fusing attention to enhance important areas and suppress unimportant areas to obtain enhanced features; the enhanced features are finally combined with the original features through residual connections to output the final results.
4. The road hazard identification and graded early warning method based on DynAM-YOLO11 according to claim 3 is characterized in that: The dynamic parameter adaptation module generates a dynamic smoothing coefficient λ through global average pooling and a fully connected layer to adjust the smoothing strength of the attention calculation. The original features generate basic attention based on SimAM and spatial attention based on the dynamic smoothing coefficient λ.
5. The road hazard identification and graded early warning method based on DynAM-YOLO11 according to claim 4 is characterized in that: The global average pooling compresses the feature map into 1x1, and then generates the λ value through two convolutional layers. The global average pooling captures the channel importance and dynamically adjusts the denominator smoothing strength. The mathematical expression for dynamic lambda generation is: λ=λ min +(λ max -l min )·σ(W2δ(W1(AvgPool(x))) Among them, λ min ,λ max are the minimum and maximum values of the smoothing coefficient respectively; σ is the Hardsigmoid function, with an output range of [0,1]; δ is the ReLU activation function; W1 and W2 are learnable parameter matrices respectively; AvgPool is the global average pooling operation; x is the original input feature map.
6. The road hazard identification and graded warning method based on DynAM-YOLO11 according to claim 3 is characterized in that: The multi-scale local enhancement module realizes multi-scale perception of local context through the design of depth-wise separable convolution plus channel remapping. The multi-scale local enhancement module generates local attention. The specific mathematical expression is as follows: LocalAtt=σ(Conv1x1(ReLU(BN(DWConv3x3(x))))) Where: x is the original feature map of the input; DWConv is the depthwise convolution part of the depthwise separable convolution; BN: Batch normalization stabilizes training; ReLU: introduces nonlinearity; Conv1x1: channel remapping; σ is the Hardsigmoid function, with an output range of [0,1]; LocalAtt is local attention.
7. The road hazard identification and graded warning method based on DynAM-YOLO11 according to claim 3 is characterized in that: The working process of the cross-channel attention fusion module: First, the original features are processed through global average pooling and multi-layer perceptron to generate channel attention. The mathematical expression for channel attention calculation is as follows: ChannelAtt=σ(MLP(AvgPool(x))) Among them, ChannelAtt is channel attention; σ is the Hardsigmoid function; MLP is a multi-layer perceptron; AvgPool is average pooling; x is the original input feature; Then, the channel attention is expanded to the same shape as the enhanced spatial attention. The mathematical expression of channel attention expansion is as follows: ChannelAtt expanded =ChannelAtt·1 H×W Among them, ChannelAtt expanded is the expanded channel attention; H is the height; W is the width; Finally, the spatial attention, local attention and the multiplied channel attention are fused through multiplication interaction to enhance the ability to perceive details and express features. Finally, the fused attention is generated through 1x1 convolution and Sigmoid function. The formula is as follows: FusedAtt=σ(W f ·(SpatialAtt⊙LocalAtt⊙ChannelAtt expanded )+b f ) Among them, SpatialAtt is spatial attention; FusedAtt is fused attention; ⊙ represents element-by-element multiplication.
8. The road hazard identification and graded early warning method based on DynAM-YOLO11 according to claim 1 is characterized in that: The calculation formula of road surface damage rate DR is: Among them, A i is the area of the i-th type of road damage area, A is the road surface area that needs to be detected, i represents the type of road damage, ω i is the weight or conversion factor of the i-th type of road damage, i0 is the total number of damage types; The calculation formula of the pavement damage index PCI is: Among them, a0 and a1 are model parameters, which are determined according to the road surface type.
9. The road hazard identification and graded warning method based on DynAM-YOLO11 according to claim 1 is characterized in that: The PCI value is used to classify road damage as follows: When PCI ≥ 85, the road surface is in excellent condition, with little damage and no need for repair; When 70≤PCI<85, the road surface is in good condition with minor damage and preventive maintenance is required; When 55≤PCI<70, there is some damage and appropriate repair is required; When 40≤PCI<55, the damage is serious and should be repaired in time; When PCI is less than 40, the damage is extremely serious and large-scale repair or reconstruction work must be carried out immediately.
Citation Information
Patent Citations
Pavement damage condition detection method and device based on YOLO-V5
CN119399128A