Dynamic multi-scale fusion detection method and device for road damage under complex streetscape and storage medium

By constructing a multi-scale hybrid extended convolutional group fusion GFB module and a multi-branch feature fusion MBFF module, combined with dynamic serpentine convolution, the multi-scale and nonlinear problems of road damage detection under complex street scenes are solved, and high-precision road damage detection is achieved.

CN120953188APending Publication Date: 2025-11-14HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511026597.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously and effectively capture millimeter-level microcracks and meter-level structural damage in complex streetscapes, and the variable crack orientations result in insufficient detection accuracy.

Method used

A multi-scale hybrid extended convolutional group fusion GFB module and a multi-branch feature fusion MBFF module are constructed. The DMSFNet network is constructed by using a dynamic serpentine convolution adaptive capture curve mode and combining cross-stage feature fusion.

Benefits of technology

It improves the detection accuracy and stability of road damage under complex street scenes, effectively captures multi-scale damage, especially slender cracks, and reduces background interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953188A_ABST
    Figure CN120953188A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic multi-scale fusion detection method and device for road damage under a complex street scene and a storage medium, and the method comprises the steps: constructing a multi-scale hybrid extension convolution group fusion module GFB based on a YOLOv11 network architecture; constructing a multi-branch feature fusion module MBFF; and constructing a novel backbone network structure to obtain a DMSFNet model. The detection device comprises a memory, a processor and a computer program which is stored on the memory and can run on the processor, and the steps of the fusion detection method can be realized when the computer program is loaded to the processor. A program instruction capable of being read and operated is stored in the storage medium, and when the program instruction is read and operated, the steps in the fusion detection method can be executed. According to the method, the detection precision of large-scale changing road damage and long and thin cracks in the complex streetscape is effectively improved through adaptive adjustment of the receptive field and multi-scale feature fusion, and the method can be applied to improving the road maintenance efficiency of urban road management departments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target detection technology in computer vision, specifically to a dynamic multi-scale fusion detection method, device, and storage medium for road damage in complex street scenes. Background Technology

[0002] Roads are vital infrastructure connecting regions and facilitating communication, playing a crucial role in driving economic development. However, with the surge in traffic volume and changing climate conditions, road damage has also emerged. This damage affects driving comfort and safety and can even lead to traffic accidents. Therefore, effectively detecting road images is essential for timely road repair.

[0003] Early road damage detection typically relied on manual identification and image annotation, which significantly increased labor and material costs. Subsequently, several automated detection methods emerged, employing laser scanning, infrared thermal imaging, and ground-penetrating radar for road damage detection. However, multi-sensor road detection is prohibitively expensive. Due to the powerful feature extraction capabilities of convolutional neural networks (CNNs) on images, they have been widely applied to road damage detection. Generally, deep learning-based road damage detection methods can be categorized into three types: image classification-based methods, semantic segmentation-based methods, and object detection-based methods.

[0004] Image classification-based road damage detection methods primarily use a simple CNN architecture with customized connection layers to classify road damage in specific scenarios. However, this method struggles to accurately determine the location, extent, and shape of damage in real-world scenes. Image segmentation-based road damage detection methods aim to divide the input image into multiple semantically meaningful, non-overlapping regions to extract pixel-level information about defects, including their location, shape, and area. While semantic segmentation can accurately depict defect regions, pixel-level segmentation and quantization are time-consuming. Compared to the two methods mentioned above, object detection-based methods can not only determine the presence of defects in an image but also precisely pinpoint their location.

[0005] While the methods described above have achieved good performance in detecting road damage, the diversity and complexity of background interference in complex street scenes still present some significant challenges:

[0006] 1) Road damage varies greatly in scale, making it difficult for existing models to effectively capture both millimeter-level microcracks and meter-level structural damage simultaneously, leading to missed detections of small targets or blurred boundaries of large targets. For example, the Faster R-CNN feature extraction network lacks sufficient resolution to capture the details of small targets. RetinaNet addresses the class imbalance problem by introducing Focal Loss, improving the detection rate of small targets. However, its feature pyramid network may lose boundary information when processing large targets due to the feature fusion method, thus affecting the accuracy of boundary measurements.

[0007] 2) Cracks typically exhibit nonlinearity and variable orientation, making it difficult for conventional convolution kernels to adapt to their diverse paths, leading to boundary positioning errors or breakage. Therefore, we propose a dynamic multi-scale fusion detection method for road damage in complex street scenes. Summary of the Invention

[0008] To address the technical challenges of detecting road damage in complex streetscape environments, such as large scale differences and complex, variable nonlinear crack directions, this invention provides a dynamic multi-scale fusion detection method, device, and storage medium for road damage in complex streetscapes. It constructs a multi-scale hybrid extended convolutional group fusion GFB module to improve the accuracy of cross-scale damage detection; and a multi-branch feature fusion MBFF module that adaptively captures curve patterns through dynamic serpentine convolution and obtains more detailed features through cross-stage feature fusion. This invention provides a dynamic multi-scale fusion detection method for road damage in complex streetscapes that reduces background interference, improves the ability to adapt to multi-scale road damage, and enhances the detection accuracy of slender cracks, effectively solving the aforementioned problems.

[0009] This invention is achieved through the following technical solution:

[0010] A dynamic multi-scale fusion detection method for road damage in complex street scenes includes the following steps:

[0011] Step 1: Obtain the keyframe illumination correction dataset D4 from the road damage video stream;

[0012] Step 2: Based on the YOLOv11 network architecture, construct a multi-scale hybrid extended convolutional group fusion GFB module, and perform multi-scale feature fusion in parallel through multiple convolutional layers with different kernel sizes;

[0013] Step 3: Construct a multi-branch feature fusion (MBFF) module. Through a three-branch cross-stage feature extraction structure, obtain more detailed curve-type crack information features, construct a novel backbone network structure, and obtain a dynamic multi-scale fusion network DMSFNet for complex street scene environments.

[0014] Step 4: Perform complex street scene road damage detection and bounding box regression on the D4 keyframe illumination correction dataset of the road damage video stream on the DMSFNet detection model, and output the predicted data set R of the road damage detection results.

[0015] Furthermore, the operation method of step 1 is as follows: define a pre-acquired real complex street scene road damage video stream dataset D1, perform keyframe processing on the road damage video stream dataset D1 to obtain a road damage video stream keyframe dataset D2; perform data cleaning, data standardization, and data augmentation on the road damage video stream keyframe dataset D2 to obtain a road damage video stream keyframe augmentation dataset D3 containing illumination noise; perform illumination correction on the images in the road damage video stream keyframe augmentation dataset D3 containing illumination noise to obtain a road damage video stream keyframe illumination correction dataset D4.

[0016] Furthermore, the operation of performing illumination correction on the images in the road damage video stream keyframe enhancement dataset D3 containing illumination noise in step 1 to obtain the road damage video stream keyframe illumination correction dataset D4 includes the following steps:

[0017] Step 1.1: Use the LabelImg software to annotate the images in the D3 keyframe augmentation dataset of road damage video stream containing illumination noise, marking the location and type of road damage;

[0018] Step 1.2: Input the road damage image due to uneven illumination;

[0019] Step 1.3: Use OpenCV to read the road damage image and convert it to floating-point format;

[0020] Step 1.4: Estimate the illumination components and apply large kernel Gaussian blur to the road damage image to simulate uneven illumination.

[0021] Step 1.5: Calculate the reflection component, i.e. the damage and texture of the road damage, and separate the reflection component using logarithmic domain subtraction;

[0022] Step 1.6: Linearly stretch the reflection component to the range [0, 255] and convert it to an 8-bit image;

[0023] Step 1.7: Use cv2.convertScaleAbs() to adjust the contrast alpha and brightness beta;

[0024] Step 1.8: Improve the contrast of the damaged area to obtain the road damage video stream keyframe illumination correction dataset D4.

[0025] Furthermore, the specific operational steps of step 2 include:

[0026] Step 2.1: Obtain the shallow road damage feature X1 of Stage 1 through 1×1 convolution. Input the shallow road damage feature X1 of Stage 1 into two parallel depthwise separable convolutions to obtain the multi-scale road damage feature X of Stage 2. 21 X 22 The two parallel depthwise separable convolution kernels are 3×3 and 5×5 respectively, and are used to convert the Stage 2 road damage features X. 21 X 22 Along the channels, road damage features are spliced ​​together to obtain Stage 3 cross-scale feature fusion road damage feature X3. The channel number of Stage 3 cross-scale feature fusion road damage feature X3 is split into four groups to obtain Stage 4 decoupled road damage feature X. 41 X 42 The two sets of features from Stage 4 are subjected to 1×1 depthwise separable convolutions to obtain the refined road damage features X from Stage 5. 51 X 52 The features refined in Stage 5 are recombined by 1×1 convolution to obtain the deep road damage feature X6 in Stage 6. Finally, the deep road damage feature X6 and the shallow road damage feature X1 are added bitwise.

[0027] Step 2.2: The GFB module performs multi-scale feature fusion in parallel using multiple convolutional layers with different kernel sizes to extract deep and rich multi-scale road damage information. The specific process is shown in the following formula:

[0028] F shallow =f 1×1 (X)

[0029]

[0030] Where: X represents the initial features of the input; f 1×1 (·) represents a 1×1 convolution; and These represent 1×1 depthwise separable convolution, 3×3 depthwise separable convolution, and 5×5 depthwise separable convolution, respectively; Concat(·) indicates that the acquired features are... and Connecting the data along the channel dimension allows for the aggregation of road damage information at different scales; Split(·) represents the channel grouping operation, performing channel-level feature transformation on the extracted road damage information; F s1 F s2 This represents the two groups of road damage feature maps after regrouping; F shallow F represents the shallow road damage characteristics obtained. gfbThis represents the deep road damage features acquired by the GFB module.

[0031] Furthermore, the specific operational steps of step 3 include:

[0032] Step 3.1: The Multi-Branch Feature Fusion (MBFF) module contains three parallel processing branches and a feature concatenation and fusion layer. The first branch of the MBFF module uses an improved C2f module. The bottleneck of the C2f module is replaced with an inverse residual convolution module to extract feature information at different road damage scales and obtain richer road damage feature information. The specific calculation process of the first branch is as follows:

[0033]

[0034] Wherein: F mid F1, F2, ..., F represents the shallow features extracted by a 1×1 convolution; n+2 Indicates the segmented features; IRB i (·) represents the inverted residual module; i represents the i-th inverted residual module; n represents the number of inverted residual modules;

[0035] Step 3.2: The second branch of the MBFF module uses concatenated depthwise separable convolutions and dynamic serpentine convolutions; firstly, depthwise separable convolutions are used to extract initial damage features, and then dynamic serpentine convolutions are used to further obtain damage features of slender bending cracks; the specific calculation process of the second branch is as follows:

[0036]

[0037] in: Represents a 3×3 dynamic serpentine convolution; F n+3 This indicates the output feature of the second branch;

[0038] Step 3.3: The third branch of the MBFF module uses a 1×1 bypass convolution to improve the module's ability to identify road damage in complex street scenes; the specific calculation process of the third branch is as follows:

[0039]

[0040] in: This represents a 1×1 depthwise separable convolution; F n+4 Indicates the output characteristics of the third branch;

[0041] Step 3.4: The semantic information of the three branch features is fused and dimensionality reduced through cascading operations and 1×1 convolutions. The specific calculation process is as follows:

[0042] F mbff =f 1×1 (Concat(F1,F2,…,Fn+2 ,F n+3 ,F n+4 ));

[0043] Wherein: F mbff This indicates the output characteristics of the MBFF module.

[0044] Furthermore, the specific operational steps of step 4 include:

[0045] Step 4.1: Apply the DMSFNet model to the D4 keyframe illumination correction dataset of the road damage video stream for target detection;

[0046] Step 4.2: The road damage images are used to extract three-level feature maps through the backbone network architecture of DMSFNet, and after bidirectional fusion by the path aggregation network, an enhanced multi-scale feature set F is output;

[0047] Step 4.3: At each location in the feature set F, the bounding box parameter set B0, the confidence score set S0, and the class probability distribution set C0 are generated simultaneously through the decoupling head;

[0048] Step 4.4: Perform conditional checks on all predicted bounding boxes. If S0 > θ s θ s If the confidence threshold is reached, the predicted bounding box is retained; otherwise, it is discarded.

[0049] Step 4.5: Multi-class DIoU-NMS processing, grouping the retained prediction boxes by category: calculate the DIoU between each pair of boxes in the same category, considering the IoU variant of the distance between the center points;

[0050] Step 4.6: If DIoU > θ0, where θ0 is the overlap threshold, retain the box with the highest S0 score and suppress the remaining overlapping boxes; if DIoU ≤ θ0, retain all boxes.

[0051] Step 4.7: Output the final prediction dataset R of the detection results.

[0052] A dynamic multi-scale fusion detection device for road damage under complex street scenes includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded into the processor, it can implement the steps of the above method.

[0053] A dynamic multi-scale fusion detection storage medium for road damage under complex street scenes, wherein the storage medium stores program instructions that can be read and run, and the program instructions, when read and run, can effectively execute the steps in the above method.

[0054] Beneficial effects

[0055] The present invention proposes a dynamic multi-scale fusion detection method, device, and storage medium for road damage under complex street scenes, which, compared with the prior art, has the following advantages:

[0056] This invention constructs a multi-scale hybrid extended convolutional group fusion (MBFF) module to improve cross-scale detection accuracy. By combining convolutional kernels with different dilation rates, the same convolutional layer can perceive damage at different scales, capturing multi-scale information from local details to global structure. Features are segmented through grouping operations, and depthwise separable convolutions are used to reorganize the channels of features, fusing cross-channel information for multi-scale features. This reduces computational cost while enhancing the expressive power of features at each scale. Using the constructed MBFF module, dynamic serpentine convolution, through its unique serpentine kernel design, can more flexibly capture curves in images. By learning offsets to adjust the shape and position of the convolutional kernel, it can dynamically change according to the shape and boundary information of the target, thus better fitting the curve structure and improving the feature extraction capability for fine cracks. Richer road damage features are obtained through residual connections and cross-stage feature fusion. The DMSFNet detection model of this invention combines multi-scale hybrid extended convolution and channel grouping fusion mechanism with dynamic serpentine convolution adaptability, which improves the ability to extract road damage and irregular crack features with large scale changes. Moreover, the DMSFNet detection model can still maintain high detection accuracy and stability when affected by the surrounding street scene environment. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of the overall process of the method in this invention.

[0058] Figure 2 This is a diagram of the DMSFNet network structure in this invention.

[0059] Figure 3 This is a structural diagram of GFB in this invention.

[0060] Figure 4 This is a structural diagram of the MBFF in this invention.

[0061] Figure 5 This is a comparison chart of the detection results in an embodiment of the present invention.

[0062] Figure 6 This is a comparison chart of performance parameters in this invention. Detailed Implementation

[0063] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. The described embodiments are merely some embodiments of the present invention, and not all embodiments. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the design concept of the present invention should fall within the protection scope of the present invention.

[0064] Example 1:

[0065] A dynamic multi-scale fusion detection method for road damage in complex street scenes includes the following steps:

[0066] Step 1: Define a pre-acquired road damage video stream dataset D1 under real complex street scenes. Perform keyframe processing on the road damage video stream dataset D1 to obtain a road damage video stream keyframe dataset D2. Perform data cleaning, data standardization, and data augmentation on the road damage video stream keyframe dataset D2 to obtain a road damage video stream keyframe augmentation dataset D3 containing illumination noise. Perform illumination correction on the images in the road damage video stream keyframe augmentation dataset D3 containing illumination noise to obtain a road damage video stream keyframe illumination correction dataset D4.

[0067] The steps involved in performing illumination correction on images in the keyframe enhancement dataset D3 of the road damage video stream containing illumination noise to obtain the road damage video stream keyframe illumination correction dataset D4 are as follows:

[0068] Step 1.1: Use the LabelImg software to annotate the images in the D3 keyframe augmentation dataset of road damage video stream containing illumination noise, marking the location and type of road damage;

[0069] Step 1.2: Input the road damage image due to uneven illumination;

[0070] Step 1.3: Use OpenCV to read the road damage image and convert it to floating-point format;

[0071] Step 1.4: Estimate the illumination components and apply large kernel Gaussian blur to the road damage image to simulate uneven illumination.

[0072] Step 1.5: Calculate the reflection component, i.e. the damage and texture of the road damage, and separate the reflection component using logarithmic domain subtraction;

[0073] Step 1.6: Linearly stretch the reflection component to the range [0, 255] and convert it to an 8-bit image;

[0074] Step 1.7: Use cv2.convertScaleAbs() to adjust the contrast alpha and brightness beta;

[0075] Step 1.8: Improve the contrast of the damaged area to obtain the road damage video stream keyframe illumination correction dataset D4.

[0076] Step 2: Based on the YOLOv11 network architecture, construct a multi-scale hybrid extended convolutional group fusion (GFB) module, which performs multi-scale feature fusion in parallel through multiple convolutional layers with different kernel sizes; specific operation steps include:

[0077] Step 2.1: Obtain the shallow road damage feature X1 of Stage 1 through 1×1 convolution. Input the shallow road damage feature X1 of Stage 1 into two parallel depthwise separable convolutions to obtain the multi-scale road damage feature X of Stage 2. 21 X 22 The two parallel depthwise separable convolution kernels are 3×3 and 5×5 respectively, and are used to convert the Stage 2 road damage features X. 21 X 22 Along the channels, road damage features are spliced ​​together to obtain Stage 3 cross-scale feature fusion road damage feature X3. The channel number of Stage 3 cross-scale feature fusion road damage feature X3 is split into four groups to obtain Stage 4 decoupled road damage feature X. 41 X 42 The two sets of features from Stage 4 are subjected to 1×1 depthwise separable convolutions to obtain the refined road damage features X from Stage 5. 51 X 52 The features refined in Stage 5 are recombined by 1×1 convolution to obtain the deep road damage feature X6 in Stage 6. Finally, the deep road damage feature X6 and the shallow road damage feature X1 are added bitwise.

[0078] Step 2.2: The GFB module performs multi-scale feature fusion in parallel using multiple convolutional layers with different kernel sizes to extract deep and rich multi-scale road damage information. The specific process is shown in the following formula:

[0079] F shallow =f 1×1 (X)

[0080]

[0081] Where: X represents the initial features of the input; f 1×1 (·) represents a 1×1 convolution; and These represent 1×1 depthwise separable convolution, 3×3 depthwise separable convolution, and 5×5 depthwise separable convolution, respectively; Concat(·) indicates that the acquired features are... and Connecting the data along the channel dimension allows for the aggregation of road damage information at different scales; Split(·) represents the channel grouping operation, performing channel-level feature transformation on the extracted road damage information; F s1 F s2 This represents the two groups of road damage feature maps after regrouping; F shallow F represents the shallow road damage characteristics obtained. gfb This represents the deep road damage features acquired by the GFB module.

[0082] Step 3: Construct a multi-branch feature fusion (MBFF) module. Through a three-branch cross-stage feature extraction structure, obtain more detailed curve-type crack information features, construct a novel backbone network structure, and obtain the dynamic multi-scale fusion network DMSFNet for complex street scene environments. Specific operation steps include:

[0083] Step 3.1: The Multi-Branch Feature Fusion (MBFF) module contains three parallel processing branches and a feature concatenation and fusion layer. The first branch of the MBFF module uses an improved C2f module. The bottleneck of the C2f module is replaced with an inverse residual convolution module to extract feature information at different road damage scales and obtain richer road damage feature information. The specific calculation process of the first branch is as follows:

[0084]

[0085] Wherein: F mid F1, F2, ..., F represents the shallow features extracted by a 1×1 convolution; n+2 Indicates the segmented features; IRB i (·) represents the inverted residual module; i represents the i-th inverted residual module; n represents the number of inverted residual modules.

[0086] Step 3.2: The second branch of the MBFF module uses concatenated depthwise separable convolutions and dynamic serpentine convolutions; firstly, depthwise separable convolutions are used to extract initial damage features, and then dynamic serpentine convolutions are used to further obtain damage features of slender bending cracks; the specific calculation process of the second branch is as follows:

[0087]

[0088] in: Represents a 3×3 dynamic serpentine convolution; F n+3 This indicates the output feature of the second branch.

[0089] Step 3.3: The third branch of the MBFF module uses a 1×1 bypass convolution to improve the module's ability to identify road damage in complex street scenes; the specific calculation process of the third branch is as follows:

[0090]

[0091] in: This represents a 1×1 depthwise separable convolution; F n+4 Indicates the output characteristics of the third branch;

[0092] Step 3.4: The semantic information of the three branch features is fused and dimensionality reduced by cascading operations and 1×1 convolution. The specific calculation process is shown as follows.

[0093] F mbff =f 1×1 (Concat(F1,F2,…,F n+2 ,F n+3 ,F n+4 ));

[0094] Wherein: F mbff This indicates the output characteristics of the MBFF module.

[0095] Step 4: Perform complex street scene road damage detection and bounding box regression on the D4 keyframe illumination correction dataset of the road damage video stream using the DMSFNet detection model, and output the predicted data set R of the road damage detection results; the specific operation steps include:

[0096] Step 4.1: Apply the DMSFNet model to the D4 keyframe illumination correction dataset of the road damage video stream for target detection;

[0097] Step 4.2: The road damage images are used to extract three-level feature maps through the backbone network architecture of DMSFNet, and after bidirectional fusion by the path aggregation network, an enhanced multi-scale feature set F is output;

[0098] Step 4.3: At each location in the feature set F, the bounding box parameter set B0, the confidence score set S0, and the class probability distribution set C0 are generated simultaneously through the decoupling head;

[0099] Step 4.4: Perform conditional checks on all predicted bounding boxes. If S0 > θ s θ s If the confidence threshold is reached, the predicted bounding box is retained; otherwise, it is discarded.

[0100] Step 4.5: Multi-class DIoU-NMS processing, grouping the retained prediction boxes by category: calculate the DIoU between each pair of boxes in the same category, considering the IoU variant of the distance between the center points;

[0101] Step 4.6: If DIoU > θ0, where θ0 is the overlap threshold, retain the box with the highest S0 score and suppress the remaining overlapping boxes; if DIoU ≤ θ0, retain all boxes.

[0102] Step 4.7: Output the final prediction dataset R of the detection results.

[0103] Please see the appendix Figure 2 This is a diagram of the DMSFNet network structure in this embodiment. The overall architecture of the DMSFNet model in this embodiment is divided into four modules, including the Input layer, the Backbone network, the Neck structure, and the Head output layer.

[0104] Please see the appendix Figure 3 To accurately capture multi-scale features and large aspect ratio impairments, a multi-scale hybrid extended convolutional group fusion GFB module is constructed. Specifically, the GFB employs multiple convolutional layers with different kernel sizes for parallel processing and adaptively adjusts the receptive field according to the target's scale. Figure 3 As shown, firstly, parallel 3×3 and 5×5 depthwise separable convolutions are used to extract spatial features of multi-scale damage. Secondly, the feature maps extracted by different depthwise separable convolutions are concatenated through channels to enrich semantic information and enhance feature representation capabilities. Then, the features are segmented through grouping operations, and clustered using 1×1 depthwise separable convolutions. Finally, a residual structure is used to fuse the rich road damage features through 1×1 convolutions to generate a unified feature map.

[0105] Please see the appendix Figure 4 To accurately capture the characteristics of slender cracks under complex street scene conditions, a multi-branch feature fusion (MBFF) module was constructed, such as... Figure 4 As shown, this module consists of three branches. The first branch of MBFF extracts more road damage information at different scales by replacing the regular convolutions in C2f with an inverse residual module. The second branch of MBFF uses depthwise separable convolutional groups and dynamic serpentine convolutions for feature extraction. First, the input depthwise separable convolutional group, composed of three different convolutional kernels, is used to enhance the ability to extract detailed features and capture long-range dependencies between the target and the background. Then, by adaptively adjusting the shape and position of the convolutions, the obtained features are passed to the dynamic serpentine convolutions to flexibly capture the features of thin and curved cracks. The third branch of MBFF uses 1×1 bypass convolutions to improve the recognition of complex patterns and the efficiency of information flow.

[0106] To address the challenges of detecting large-scale variations in road damage and the difficulty in detecting fine cracks in complex streetscape environments, a DMSFNet network model was constructed. The network structure parameters are shown in Table 1.

[0107] Table 1 Network Structure Parameters

[0108]

[0109]

[0110] To determine the optimal model parameters for the DMSFNet model in complex street scene video stream data with road damage, a case study is used to identify the optimal parameters. The example provides the training and validation sets for training the model, as well as the training method.

[0111] Case 1:

[0112] In this case study, a pre-acquired real-world complex street scene road damage video stream dataset D1 was used. Keyframe processing was performed on the road damage video stream dataset D1 to obtain a road damage video stream keyframe dataset D2. Data cleaning, standardization, and augmentation were then performed on the road damage video stream keyframe dataset D2 to obtain a road damage video stream keyframe augmentation dataset D3. Illumination correction was then performed on the road damage video stream keyframe images containing illumination noise to obtain a road damage video stream keyframe illumination correction dataset D4 consisting of 13,973 images.

[0113] The illumination correction dataset D4 was labeled using the LabelImg tool. The class Class was defined as follows: Class = {0: D00, 1: D10, 2: D20, 3: D40}, where 0-3 are numbers, and D00, D10, D20 and D40 represent longitudinal cracks, transverse cracks, alligator cracks and pits, respectively. Each target object in each image is labeled with a bounding box and assigned a specific classification label. Each target object in each image will have a category label and a set of corresponding position coordinates. The label information of the target objects in each image is saved as a txt file with the same name as the image to obtain the label dataset L1. Let L1 = {id, Ox, Oy, width, high}, where id, Ox, Oy, width, and high are the category number, center point x-axis coordinate, center point y-axis coordinate, width, and height, respectively. The road damage video stream keyframe illumination correction dataset D4 is divided into training set V1 and validation set V2 in an 8:2 ratio. The training set V1 has 11,178 data points, and the validation set V2 has 2,795 data points.

[0114] The training set V1 and validation set V2 are input into DMSFNet and YOLOv11 for training and testing. The optimal weight files for the model parameters of DMSFNet and YOLOv11 during the training process are obtained. The optimal weight file for the DMSFNet model parameters is P1, and the optimal weight file for the YOLOv11 model parameters is P2. These files are saved separately.

[0115] By comparing the training and validation results, the optimal model parameters were obtained by fine-tuning the model parameters to be 300 training epochs, batch size of 8, initial learning rate of 0.01, minimum learning rate of 0.01, and weight decay coefficient of 0.0005.

[0116] This case study uses a comprehensive evaluation metric to assess performance, including precision (P), recall (R), mAP50, and mAP50-95. These metrics have different focuses: precision (P) measures the number of positive samples that meet the criteria for a positive sample, while recall (R) measures the number of accurately predicted positive samples. Mean accuracy (AP) is the average precision across different confidence thresholds, calculated using the Intersection over Union (IoU) to measure the overlap between detected and ground truth targets. AP is presented using a precision-recall curve. Mean accuracy (mAP) is the average of AP across all classes and is a crucial metric for evaluating the performance of an object detection system. mAP is calculated at a confidence threshold; for example, mAP50 represents the mAP value at a 50% confidence threshold. Calculating mAP values ​​within the confidence threshold range of 50% to 95% evaluates the model's robustness and accuracy. Specifically, this includes:

[0117] Precision rate P:

[0118]

[0119] Recall rate R:

[0120]

[0121] Average accuracy mAP50:

[0122]

[0123]

[0124] In the above formula, TP refers to the number of samples that are positive and are predicted as positive, FP refers to the number of samples that are negative and are predicted as positive, FN refers to the number of samples that are positive and are predicted as negative, M is the number of target classes detected, and AP(i) is the AP of the i-th target class.

[0125] Comparative experiments were conducted to verify the optimization effects of the constructed GFB and MBFF modules, and to further evaluate the performance of the DMSFNet model. Compared with YOLOv11, the method proposed in this invention can reduce background interference in complex street scene environments, improve the ability to adapt to multi-scale road damage, and enhance the detection accuracy of slender cracks.

[0126] This case study was tested on three publicly available road damage datasets: SVRDD, RDD2020, and USRDD.

[0127] SVRDD is the first public dataset for road damage detection based on street view images. It contains 8,000 street view images, covering a variety of urban road types and conditions. This invention selects 7,000 images for training and 1,000 images for testing.

[0128] The RDD2020 dataset contains 26,620 images of road damage from Japan, India, and the Czech Republic. This invention uses 18,500 images from the original training set for training and 2,541 images for testing.

[0129] The USRDD dataset is the US road damage dataset from the 2022 Crowd-Based Road Damage Detection Challenge. This invention uses 4000 images from the original training set for training and 805 images for testing.

[0130] Please refer to Tables 2, 3, and 4, which compare the method of this invention with several mainstream object detection algorithms, including Faster-R-CNN, RetinaNet, FCOS, RT-DETR, YOLOv5, YOLOX, YOLOv8, and YOLOv11, on three publicly available datasets. Experiments show that the method of this invention significantly outperforms other detection methods when handling slender cracks and road damage with large-scale variations.

[0131] Table 2

[0132]

[0133]

[0134] Table 3

[0135]

[0136] Table 4

[0137]

[0138] Table 2 shows the performance of different damage categories in the SVRDD dataset using different methods, with mAP@0.5. It can be observed that: 1) In this embodiment, the mAP@0.5 for all damage categories is superior to other comparative models, exceeding the optimal comparative model YOLOv11 by 10.6%; 2) The mAP@0.5 for transverse cracks, longitudinal patches, and transverse patches is 16.8%, 24.4%, and 7.4% higher than YOLOv11, respectively, indicating that the DMSFNet detection model in this embodiment has achieved significant improvement in handling slender cracks. Potholes in the SVRDD dataset are small and may be confused by the surrounding environment, but the DMSFNet detection model still achieves 75.4% and 74.8% accuracy in the pothole and manhole cover categories, respectively, demonstrating that the method of this invention can more accurately detect road damage, especially small cracks and potholes, in complex street scene environments.

[0139] Table 3 shows the performance of different damage categories on different methods using the RDD2020 dataset. It can be observed that: 1) The method of this invention outperforms other comparative models in terms of mAP@0.5 for all damage categories, while RetinaNet and FCOS have the worst detection performance with mAP@0.5 below 50%; 2) The detection results of D00, D10, and D20 are 3.4%, 6.6%, and 2.7% higher than YOLOX, respectively, demonstrating that the method of this invention has significant advantages in detecting road damage and cracks with large-scale variations.

[0140] Table 4 shows the performance of different methods on different damage categories in the USRDD dataset. It can be observed that: 1) The method of this invention achieves an mAP@0.5 of 69.2%, significantly outperforming all other models. YOLOv11's overall mAP@0.5 is 54.3%, still significantly lower than DMSFNet; 2) In the D00, D10, D20, and D40 categories, DMSFNet's mAP@0.5 is 7.6%, 15.8%, 7.8%, and 26.1% higher than YOLOv11, respectively, demonstrating that the method of this invention can more effectively and accurately capture and identify various damage types.

[0141] Please see the appendix Figure 5 The figures below show a comparison of the application effects in this case study, where: (a) is the original image, (b), (c), (d), and (e) are the results of other mainstream detection models, and (f) is the result of the DMSFNet detection model of this invention. Figure 5The comparison shows that the DMSFNet detection model can accurately detect both fine and large-scale cracks even under the interference of tree projections, while other detection models tend to miss or falsely detect them. Therefore, the method of this invention can effectively reduce background interference, adapt to multi-scale damage, and improve the accuracy of detecting fine cracks in complex street scene environments.

[0142] Please see the appendix Figure 6 The figure shows a comparison of the performance parameters of this invention. To verify the effectiveness of GFB and MBFF, this invention removed the GFB and MBFF modules from the DMSFNet detection model, and conducted three experiments on the SVRDD dataset. Figure 6 The performance curve of the baseline model consistently falls below that of the DMSFNet detection model proposed in this invention. Experimental results demonstrate that the proposed GFB and MBFF modules effectively improve feature representation and detection accuracy in DMSFNet.

[0143] Example 2:

[0144] A dynamic multi-scale fusion detection device for road damage under complex street scenes includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded into the processor, it can implement the steps of the method described in Embodiment 1.

[0145] Example 3:

[0146] A dynamic multi-scale fusion detection storage medium for road damage under complex street scenes, wherein the storage medium stores program instructions that can be read and run, and the program instructions, when read and run, can effectively execute the steps in the method described in Embodiment 1.

[0147] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered by the present invention.

Claims

1. A dynamic multi-scale fusion detection method for road damage under complex street scenes, characterized in that: Including the following steps: Step 1: Obtain the keyframe illumination correction dataset D4 from the road damage video stream; Step 2: Based on the YOLOv11 network architecture, construct a multi-scale hybrid extended convolutional group fusion GFB module, and perform multi-scale feature fusion in parallel through multiple convolutional layers with different kernel sizes; Step 3: Construct a multi-branch feature fusion (MBFF) module. Through a three-branch cross-stage feature extraction structure, obtain more detailed curve-type crack information features, construct a novel backbone network structure, and obtain a dynamic multi-scale fusion network DMSFNet for complex street scene environments. Step 4: Perform complex street scene road damage detection and bounding box regression on the D4 keyframe illumination correction dataset of the road damage video stream on the DMSFNet detection model, and output the predicted data set R of the road damage detection results.

2. The dynamic multi-scale fusion detection method for road damage under complex street scenes according to claim 1, characterized in that: The operation method of step 1 is as follows: Define a pre-acquired real complex street scene road damage video stream dataset D1; perform keyframe processing on the road damage video stream dataset D1 to obtain a road damage video stream keyframe dataset D2; perform data cleaning, data standardization, and data augmentation on the road damage video stream keyframe dataset D2 to obtain a road damage video stream keyframe augmentation dataset D3 containing illumination noise; perform illumination correction on the images in the road damage video stream keyframe augmentation dataset D3 containing illumination noise to obtain a road damage video stream keyframe illumination correction dataset D4.

3. The dynamic multi-scale fusion detection method for road damage under complex street scenes according to claim 2, characterized in that: The operation of performing illumination correction on the images in the road damage video stream keyframe enhancement dataset D3 containing illumination noise in step 1 to obtain the road damage video stream keyframe illumination correction dataset D4 includes the following steps: Step 1.1: Use the LabelImg software to annotate the images in the D3 keyframe augmentation dataset of road damage video stream containing illumination noise, marking the location and type of road damage; Step 1.2: Input the road damage image due to uneven illumination; Step 1.3: Use OpenCV to read the road damage image and convert it to floating-point format; Step 1.4: Estimate the illumination components and apply large kernel Gaussian blur to the road damage image to simulate uneven illumination. Step 1.5: Calculate the reflection component, i.e. the damage and texture of the road damage, and separate the reflection component using logarithmic domain subtraction; Step 1.6: Linearly stretch the reflection component to the range [0, 255] and convert it to an 8-bit image; Step 1.7: Use cv2.convertScaleAbs() to adjust the contrast alpha and brightness beta; Step 1.8: Improve the contrast of the damaged area to obtain the road damage video stream keyframe illumination correction dataset D4.

4. The dynamic multi-scale fusion detection method for road damage under complex street scenes according to claim 1, characterized in that: The specific steps of step 2 include: Step 2.1: Obtain the shallow road damage feature X1 of Stage 1 through 1×1 convolution. Input the shallow road damage feature X1 of Stage 1 into two parallel depthwise separable convolutions to obtain the multi-scale road damage feature X of Stage 2. 21 X 22 The two parallel depthwise separable convolution kernels are 3×3 and 5×5 respectively, and are used to convert the Stage 2 road damage features X. 21 X 22 Along the channels, road damage features are spliced ​​together to obtain Stage 3 cross-scale feature fusion road damage feature X3. The channel number of Stage 3 cross-scale feature fusion road damage feature X3 is split into four groups to obtain Stage 4 decoupled road damage feature X. 41 X 42 The two sets of features from Stage 4 are subjected to 1×1 depthwise separable convolutions to obtain the refined road damage features X from Stage 5. 51 X 52 The features refined in Stage 5 are recombined by 1×1 convolution to obtain the deep road damage feature X6 in Stage 6. Finally, the deep road damage feature X6 and the shallow road damage feature X1 are added bitwise. Step 2.2: The GFB module performs multi-scale feature fusion in parallel using multiple convolutional layers with different kernel sizes to extract deep and rich multi-scale road damage information. The specific process is shown in the following formula: F shallow =f 1×1 (X) Where: X represents the initial features of the input; f 1×1 (·) represents a 1×1 convolution; and These represent 1×1 depthwise separable convolution, 3×3 depthwise separable convolution, and 5×5 depthwise separable convolution, respectively; Concat(·) indicates that the acquired features are... and Connecting the data along the channel dimension allows for the aggregation of road damage information at different scales; Split(·) represents the channel grouping operation, performing channel-level feature transformation on the extracted road damage information; F s1 F s2 This represents the two groups of road damage feature maps after regrouping; F shallow F represents the shallow road damage characteristics obtained. gfb This represents the deep road damage features acquired by the GFB module.

5. The dynamic multi-scale fusion detection method for road damage under complex street scenes according to claim 1, characterized in that: The specific steps of step 3 include: Step 3.1: The Multi-Branch Feature Fusion (MBFF) module contains three parallel processing branches and a feature concatenation and fusion layer. The first branch of the MBFF module uses an improved C2f module. The bottleneck of the C2f module is replaced with an inverse residual convolution module to extract feature information at different road damage scales and obtain richer road damage feature information. The specific calculation process of the first branch is as follows: Wherein: F mid F1, F2, ..., F represents the shallow features extracted by a 1×1 convolution; n+2 Indicates the segmented features; IRB i (·) represents the inverted residual module; i represents the i-th inverted residual module; n represents the number of inverted residual modules; Step 3.2: The second branch of the MBFF module uses concatenated depthwise separable convolutions and dynamic serpentine convolutions; firstly, depthwise separable convolutions are used to extract initial damage features, and then dynamic serpentine convolutions are used to further obtain damage features of slender bending cracks; the specific calculation process of the second branch is as follows: in: Represents a 3×3 dynamic serpentine convolution; F n+3 This indicates the output feature of the second branch; Step 3.3: The third branch of the MBFF module uses a 1×1 bypass convolution to improve the module's ability to identify road damage in complex street scenes; the specific calculation process of the third branch is as follows: in: This represents a 1×1 depthwise separable convolution; F n+4 Indicates the output characteristics of the third branch; Step 3.4: The semantic information of the three branch features is fused and dimensionality reduced through cascading operations and 1×1 convolutions. The specific calculation process is as follows: F mbff =f 1×1 (Concat(F1,F2,…,F n+2 ,F n+3 ,F n+4 )); Wherein: F mbff This indicates the output characteristics of the MBFF module.

6. The dynamic multi-scale fusion detection method for road damage under complex street scenes according to claim 1, characterized in that: The specific steps in step 4 include: Step 4.1: Apply the DMSFNet model to the D4 keyframe illumination correction dataset of the road damage video stream for target detection; Step 4.2: The road damage images are used to extract three-level feature maps through the backbone network architecture of DMSFNet, and after bidirectional fusion by the path aggregation network, an enhanced multi-scale feature set F is output; Step 4.3: At each location in the feature set F, the bounding box parameter set B0, the confidence score set S0, and the class probability distribution set C0 are generated simultaneously through the decoupling head; Step 4.4: Perform conditional checks on all predicted bounding boxes. If S0 > θ s θ s If the confidence threshold is reached, the predicted bounding box is retained; otherwise, it is discarded. Step 4.5: Multi-class DIoU-NMS processing, grouping the retained prediction boxes by category: calculate the DIoU between each pair of boxes in the same category, considering the IoU variant of the distance between the center points; Step 4.6: If DIoU > θ0, where θ0 is the overlap threshold, retain the box with the highest S0 score and suppress the remaining overlapping boxes; if DIoU ≤ θ0, retain all boxes. Step 4.7: Output the final prediction dataset R of the detection results.

7. A dynamic multi-scale fusion detection device for road damage under complex street scenes, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into the processor, it can implement the steps of the method described in any one of claims 1-6.

8. A dynamic multi-scale fusion detection storage medium for road damage under complex street scenes, wherein the storage medium stores program instructions that can be read and executed, characterized in that: When the program instructions are read and run, they can effectively execute the steps of the method described in any one of claims 1-6.