Road defect detection optimization method based on YOLOv11s

By optimizing the multi-scale feature fusion and loss function of the YOLOv11s model, the accuracy and efficiency of the detection of road cracks by aerial photography by drone is improved, the problem of insufficient detection accuracy in traditional methods is solved, and efficient road disease detection is achieved.

CN120339235APending Publication Date: 2025-07-18NANJING TECH UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510438863.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology has insufficient pavement crack detection accuracy in drone aerial photography scenarios, traditional methods are low in efficiency and high leakage detection rate, making it difficult to apply on a large scale.

Method used

Based on YOLOv11s, the road defect detection method is optimized by designing the multi-scale feature fusion module BiFPN_Concat, replacing the C3k2 module with C3k2_DFF, and using the Inner_CIoU loss function, the YOLOv11s model is optimized and the detection performance is improved.

Benefits of technology

The detection accuracy has been improved, and the average accuracy of the improved model in the actual environment is increased by 1.9%, the accuracy rate is increased by 3.3%, and the recall rate is increased by 2%, meeting the requirements of road defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339235A_ABST
    Figure CN120339235A_ABST
Patent Text Reader

Abstract

The invention discloses a road defect detection optimization method based on YOLOv11s, which is applied to the technical field of unmanned aerial vehicle road inspection, and comprises the steps of constructing a road defect target detection model, designing a multi-scale feature fusion module BiFPNConcat in a YOLOv11s neck network, fusing a BiFPN network, and improving object detection and segmentation performance. In the neck network, a C3k2DFF module is adopted to replace a C3k2 module, and the problem that local features of different scales may lose information in the fusion process is solved; according to the method, a CIoU loss function is replaced with an InerCIoU loss function, the defects of a bounding box regression method are overcome, the detection capability is further improved, the road defect detection performance of the model is improved by optimizing a YOLOv11s model framework, and the detection precision of an algorithm model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and specifically relates to an optimized method for road defect detection based on YOLOv11s Background Art

[0002] As a common but seriously underestimated disease, pavement cracks pose a risk of expansion and more serious pavement damage, threatening driving safety. Targeted improvement of object detection technology and its application on an unmanned aerial vehicle (UAV) platform to detect pavement cracks can provide important technical support for the high-efficiency and intelligentization of road inspection

[0003] Currently, there is little application research on pavement crack detection in the scenario of UAV aerial photography, and the detection accuracy of small cracks is insufficient. Traditional road defect detection mainly relies on manual detection and core drilling physical detection methods, which generally have problems such as low efficiency, high missed detection rate, strong subjectivity, and difficulty in large-scale application. Therefore, it is of great significance to apply deep learning to the inspection of pavement cracks to provide strong support for timely discovery of pavement diseases and efficient inspection, and it has strong adaptability and versatility Summary of the Invention

[0004] 1. Technical problems to be solved

[0005] In view of the above technical problems, the present invention provides an optimized method for road defect detection based on YOLOv11s, starting from the algorithm network structure, aiming to improve the detection performance

[0006] 2. Technical solutions

[0007] An optimized method for road defect detection based on YOLOv11s, characterized by comprising the following steps

[0008] Step 1: Make a road defect detection dataset. The open-source dataset UAV-PDD2023 is used in the experiment, and the training set: validation set: test set is divided in a ratio of 8:1:1. The test set does not participate in the training process

[0009] Step 2: Build a road defect object detection model based on the YOLOv11s algorithm. The target defect detection model includes a backbone feature extraction network, a neck fusion network, and a head detection network

[0010] Step 3: Optimize the YOLOv11s neck fusion network. Design a multi-scale feature fusion module BiFPN_Concat, integrate it into the BiFPN network to improve object detection and segmentation performance, and use the C3k2_DFF module to replace the C3k2 module to solve the problem that local features of different scales may lose information during the fusion process

[0011] Step 4: Use the Inner_CIoU loss function to replace the CIoU loss function to make up for the deficiencies of the bounding box regression method and further improve the detection ability;

[0012] Step 5: Use the completed road defect dataset to train the improved YOLOv11s detection model to obtain a trained road defect target detection model;

[0013] Step 6: Input the road defect picture to be detected into the trained target detection model to obtain the road defect detection result;

[0014] Step 7: Evaluate the model, and evaluate the model using the mean average precision MAP, precision Precision, and recall Recall.

[0015] 3. Beneficial effects:

[0016] (1) In this method, YOLOv11s is optimized for road defect detection. By designing a multi-scale feature fusion module BiFPN_Concat in the neck network of YOLOv11s and integrating it into the BiFPN network to generate feature pyramids with different resolutions, it helps to improve object detection and segmentation performance.

[0017] (2) In this method, YOLOv11s is optimized for road defect detection. The C3k2_DFF module is used to replace the C3k2 module in the neck network to solve the problem of information that may be lost during the fusion of local features at different scales.

[0018] (3) In this method, YOLOv11s is optimized for road defect detection. The Inner_CIoU loss function is used to replace the CIoU loss function to make up for the deficiencies of the bounding box regression method and further improve the detection ability.

[0019] (4) The average precision (mAP@0.5) of the improved algorithm model is 89.8%, which is 1.9% higher than the average precision (mAP@0.5) of the original YOLOv11s. The precision Precision reaches 90.5%, which is 3.3% higher than the original version, and the recall Recall reaches 93%, which is 2% higher than the original version, indicating that the algorithm can meet the requirements of road defect detection in the actual production environment. Description of the drawings

[0020] Figure 1 It is the network structure diagram after the optimization of the YOLOv11s detection algorithm in the present invention;

[0021] Figure 2 It is the explanatory diagram of the DFF module structure in the present invention;

[0022] Figure 3 This is the detailed diagram of the BiFPN bidirectional feature pyramid architecture in the present invention;

[0023] Figure 4 This is the detailed diagram of the C3k_DFF module in the present invention;

[0024] Figure 5 This is the detailed diagram of the Bottleneck_DFF module in the present invention;

[0025] Figure 6 This is the MAP curve diagram before the optimization algorithm in the specific embodiment of the present invention;

[0026] Figure 7 This is the MAP curve diagram after the optimization algorithm in the specific embodiment of the present invention;

[0027] Figure 8 This is the road defect detection effect diagram of YOLOv11s in the specific embodiment; Detailed implementation manners

[0028] The present invention will be specifically described below with reference to the accompanying drawings.

[0029] A road defect detection optimization method based on YOLOv11s, characterized by comprising the following steps:

[0030] Step 1: Make a road defect detection data set. The open-source data set UAV-PDD2023 is used in the experiment, and the training set: validation set: test set is divided in a ratio of 8:1:1. The test set does not participate in the training process;

[0031] Step 2: Build a road defect target detection model based on the YOLOv11s algorithm. The target defect detection model includes a backbone feature extraction network, a neck fusion network, and a head detection network;

[0032] Step 3: Optimize the YOLOv11s neck fusion network, design a multi-scale feature fusion module BiFPN_Concat, integrate it into the BiFPN network to improve object detection and segmentation performance, and use the C3k2_DFF module to replace the C3k2 module to solve the problem of possible information loss of local features at different scales during the fusion process;

[0033] Step 4: Use the Inner_CIoU loss function to replace the CIoU loss function to make up for the deficiency of the bounding box regression method and further improve the detection ability;

[0034] Step 5: Use the divided road defect data set to train the improved YOLOv11s detection model to obtain a trained road defect target detection model;

[0035] Step 6: Input the road defect image to be detected into the trained object detection model to obtain the road defect detection result;

[0036] Step 7: Evaluate the model; evaluate the model using the mean average precision MAP, precision, and recall.

[0037] Furthermore, Step 1 specifically includes: The defect types in the open-source UAV-PDD2023 dataset include 6 typical pavement cracks, namely longitudinal crack, transverse crack, alligator crack, oblique crack, repair, and potholes, under different weather conditions and construction qualities. The training set: validation set: test set is divided in a ratio of 8:1:1. The test set does not participate in the training process. Among them, the training set contains 1952 images, and the validation set and test set contain 244 images each. The annotation information of the dataset is converted from XML format to TXT format for running in the YOLOv11s model based on the PyTorch framework. The input image specification of YOLOv11s is a resolution of 640*640 for each image.

[0038] Furthermore, Step 3 specifically includes: Incorporate the BiFPN network into the Concat module of the original YOLOv11s neck network, design the multi-scale feature fusion module BiFPN_Concat to improve object detection and segmentation performance, and replace the C3k2 module with the C3k2_DFF module to solve the problem of possible loss of information of local features at different scales during the fusion process;

[0039] As shown in the Figure 3 attachment, in this method, different-resolution feature pyramids are generated by combining BiFPN. By combining the top-down and bottom-up paths, features of different resolutions are effectively fused. At the same time, learnable weights are introduced to weight the fused features, significantly improving the effect and efficiency of feature extraction. BiFPN introduces a weighted fusion mechanism, adding an additional weight to each input feature to let the network learn the importance of each feature;

[0040] The core of BiFPN lies in weighted feature fusion, and the formula for this weighted fusion is as follows:

[0041]

[0042] In the above formula, O is the output value, representing the result after weighted fusion; w iDenotes the learnable weight associated with the i-th input feature, where i represents the index number of the i-th feature map of the input; ∈ is a small constant, usually set to 0.0001, used to avoid the case of a zero denominator, thus ensuring numerical stability; I i Is the input value, representing the i-th input feature map; the core idea of weighted feature fusion is to use the learned weight w i To weight different input features I i And ensure that the sum of these weights is 1 through normalization, thereby obtaining an output feature O that fuses all the input feature information, enabling BiFPN to effectively fuse features from different levels and improve the performance of tasks such as object detection.

[0043] Furthermore, as shown in the appendix Figure 2 The embedded DFF module mainly contains two input features And After being processed by 1×1×1 convolution respectively and then added together, and generating the spatial weight W through the Sigmoid activation function sp The input features are first concatenated (Concat), then passed through global average pooling (AVGPool), convolution (Conv) and Sigmoid operations to generate the channel weight W ch The fused features Adjust the channel information through 1×1×1 convolution and perform element-wise multiplication with the spatially attention-weighted features to obtain the final fused features This mechanism can enhance the ability to retain details of local features, and at the same time combine global information to improve the segmentation accuracy and robustness of the model.

[0044] Furthermore, as shown in Figure 1 In order to better fuse multi-scale features, make full use of global information, and improve the performance of image segmentation, the C3k2 module used in the YOLOv11s model is fused with the dynamic feature fusion (DFF) module. DFF aims to adaptively fuse multi-scale local feature maps based on global information, and select important features during the fusion process through a dynamic mechanism to address the deficiencies in feature fusion of the above-mentioned existing technologies. Among them, the C3k module and the Bottleneck module in the original C3k2 module are fused with the DFF module into the C3k2_DFF module, and C3k_DFF and Bottleneck_DFF are called in the cases corresponding to Ture and False;

[0045] The C3k module extracts features through a configurable convolution kernel size (default is 3x3), expands the receptive field, and captures complex spatial features. By combining the C3k module in the C3k2 module of the original YOLOv11s neck network with the DFF module into the C3k_DFF module, the feature fusion ability is further enhanced. The specific steps are as follows: In the C3k_DFF module, by integrating the DFF module, first, the input features and the features of the skip connection are concatenated together, and the number of channels is doubled. Global average pooling and convolution operations are used to generate attention weights to weight the concatenated features. Then, the weighted features are compressed to the target number of channels through convolution operations. Finally, two 1x1 convolutions are used to generate additional attention maps, which are multiplied by the compressed features to further enhance feature fusion. Through the above method, the C3k_DFF module can effectively extract and fuse multi-scale features, improving the feature expression ability and detection performance of the model;

[0046] As Figure 4 shown, CBS refers to a combined module that includes convolution, batch normalization, and activation function. The input features of the C3k_DFF module first enter the first CBS layer, and the output features are used for feature extraction and fusion through a series of Bottleneck_DFF modules. The output of the Bottleneck_DFF module is concatenated with the output of another parallel CBS layer. The concatenated features are integrated through the third CBS layer, and the output after concatenation is used as the final output features. Bottleneck_DFF*n indicates that there are n Bottleneck_DFF modules, and these modules are connected in series, with the n value set to 3;

[0047] The Bottleneck module first compresses the number of channels of the input feature map through a dimensionality reduction operation using a 1x1 convolution, then uses a 3x3 convolution kernel for feature extraction while keeping the spatial dimension of the feature map unchanged, and finally expands the number of channels back to the original size through a 1x1 convolution. The Bottleneck module is combined with the residual connection to form a skip connection, which directly adds the input features to the output features to alleviate the problem of gradient disappearance. The Bottleneck module in the C3k2 module of the original YOLOv11s neck network is combined into the Bottleneck_DFF module. As Figure 5As shown, CBS refers to a combined module that includes Convolution, Batch Normalization, and Activation. The specific steps of the Bottleneck_DFF module are as follows: First, the input feature x is passed through the first convolutional operation to obtain the hidden layer feature y1. At the same time, the hidden layer feature y1 is passed through the second convolutional operation to obtain the output feature y2. Then, the sizes of the input channel number c1 and the output channel number c2 are compared. If they are equal, a residual connection is performed to add the input feature x and the output feature y2. If they are not equal, y2 is directly used as the output. Finally, the input feature x and the output feature y2 are passed to the DFF module for further feature fusion. According to the above steps, the Bottleneck_DFF module can not only extract features through 2 layers of convolution but also further enhance feature fusion through the DFF module, improving the feature expression ability and the performance of the model.

[0048] Furthermore, in step four, the Inner_CIoU loss function is used instead of the CIoU loss function to make up for the deficiencies of the bounding box regression method. The Inner-CIoU loss function is a variant of the CIoU loss function. The Inner_CIoU loss function calculates the IoU loss through an auxiliary bounding box and introduces a scale factor ratio to control the scale size of the auxiliary bounding box for calculating the loss, ensuring that the model can learn the position information of the target during training. The introduced scale factor ratio ratio is set to 0.7, ensuring that the width and height of the bounding box are 70% of the width and height of the original bounding box respectively. The formulas for each part of the Inner_CIoU loss function are as follows:

[0049]

[0050]

[0051] α = (1 - IoU) + v;

[0052]

[0053] L Inner-CIoU = L CIoU + IoU - IOU inner ;

[0054] Where: L represents the loss function, B is the predicted bounding box, B gt is the true bounding box, |B ∩ B gt | is the area of the intersection of the two bounding boxes, |B ∪ B gt | is the area of the union of the two bounding boxes, ρ 2 (b, b gt ) is the predicted bounding box b and the true bounding box bgt The Euclidean distance between them, c is the diagonal length of the smallest closed enclosure containing the two bounding boxes. α is a weight factor used to balance the influence of different terms. v is a term related to the aspect ratio, used to penalize the inconsistent aspect ratios between the predicted bounding box and the ground truth bounding box. inter calculates the area of the intersection region of the two bounding boxes, and union calculates the area of the union region of the two bounding boxes. Specific embodiments:

[0056] This experiment training was carried out in the Ubuntu 16.04.7 LTS and CUDA 11.7 environment, using the Pytorch deep learning framework. The Pytorch version is torch-2.6.0+cu124; the Ultralytics version is 8.3.99, and the Python version is 3.10.16; the GPU configuration is: two NVIDIA GeForce RTX 2080Ti, with a total of 22G video memory; the CPU model: Intel(R) Xeon(R) CPU E5-2678 v3 @ 2.50GHz; the input image specification of YOLOv11s is that the resolution of each image is 640*640, the number of iterations epoch is set to 400 times, the number of threads workers is set to 8, and the training batch batch is set to 16.

[0057] This embodiment uses the open-source UAV-PDD2023 dataset. The dataset image specification is 2592*1944, which is obtained by splitting 4K camera images taken from a height of 30 meters. The defect types include 6 typical pavement cracks: Longitudinal crack, Transverse crack, Alligator crack, Oblique crack, Repair, and Potholes under different weather conditions and construction qualities. The training set: validation set: test set is divided in the ratio of 8:1:1, and the test set does not participate in the training process.

[0058] Among them, the training set contains 1952 pictures, the validation set and the test set contain 244 pictures. The annotation information of the dataset is converted from XML format to TXT format for running in the YOLOv11s model based on the PyTorch framework. The input image specification of YOLOv11s is that the resolution of each image is 640*640.

[0059] Through the training set and the validation set, the optimized YOLOv11s road defect detection model is trained and validated, and the following metrics are used to evaluate the model: Mean Average Precision (MAP), Precision, Recall;

[0060] Mean Average Precision (MAP) is an important metric for evaluating the overall performance of object detection models. Precision represents the proportion of true target samples among all the boxes predicted as positive samples (targets) by the model, and Recall represents the proportion of samples correctly identified by the model among all real targets. In multi-class object detection tasks, by calculating the Average Precision (AP) for each class and taking the average, a comprehensive performance evaluation metric MAP is obtained. The specific formula is as follows:

[0061]

[0062]

[0063]

[0064]

[0065] Among them, TP (True Positives) represents the number of correctly identified targets, FP (False Positives) represents the number of samples that misclassify the background as a target, FN represents False Negative, that is, the number of samples that the model fails to correctly predict as positive, and FP represents False Positive, that is, the number of samples that the model incorrectly predicts as positive. MAP is the mean of the average accuracies of all classes, and C is the number of classes. Through these metrics, the effectiveness and efficiency of the improved YOLOv11s road defect detection model in practical applications can be comprehensively evaluated.

[0066] The above detailed steps are mainly the implementation methods of the road defect detection method of the present invention and the process of the network model processing pictures. During the training process, the model processes the input picture information in the aforementioned manner. After training, the trained weight and parameter files will be obtained for model deployment and testing.

[0067] To verify the effectiveness of the improved algorithm for the road defect dataset in this paper, a series of experiments were conducted, such as Figure 6 、 7As shown, the MAP curve graphs before and after the improvement are plotted to more intuitively compare the precision performance of each algorithm for different categories. The average precision (MAP) of the improved algorithm model is 89.8%, which is a 1.9% increase compared to 87.9% of the original YOLOv11s. Moreover, except for the fluctuations of 0.005% and 0.002% in alligator cracks and oblique cracks respectively among the six categories, the detection accuracy of other categories has been significantly improved. Among them, the precision of longitudinal cracks reaches 92%, the precision of transverse cracks reaches 94.6%, the precision of alligator cracks reaches 96.8%, the precision of oblique cracks reaches 92.2%, the precision of repair reaches 98.3%, and the precision of potholes reaches 64.5%. The overall detection precision shows an upward trend, indicating that the algorithm can meet the requirements of road defect detection in the actual production environment.

[0068] As attached Figure 8 shown, the improved model is used to detect road images containing 6 different defects.

[0069] Although the present invention has been disclosed above in preferred embodiments, they are not used to limit the present invention. Anyone skilled in this art can make various changes or modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be defined by the protection scope of the claims of this application.

Claims

1. An optimized method for road defect detection based on YOLOv11s, characterized in that: It includes the following steps: Step 1: Make a road defect detection dataset. The open-source dataset UAV-PDD2023 is used in the experiment. The training set, validation set, and test set are divided at a ratio of 8:1:

1. The test set does not participate in the training process; Step 2: Build a road defect object detection model based on the YOLOv11s algorithm. The object defect detection model includes a backbone feature extraction network, a neck fusion network, and a head detection network; Step 3: Optimize the YOLOv11s neck fusion network. Design a multi-scale feature fusion module BiFPN_Concat and integrate it into the BiFPN network to improve object detection and segmentation performance. Replace the C3k2 module with the C3k2_DFF module to solve the problem that local features of different scales may lose information during the fusion process; Step 4: Use the Inner_CIoU loss function instead of the CIoU loss function to make up for the deficiencies of the bounding box regression method and further improve the detection ability; Step 5: Use the completed road defect dataset to train the improved YOLOv11s detection model to obtain a trained road defect object detection model; Step 6: Input the road defect picture to be detected into the trained object detection model to obtain the road defect detection result; Step 7: Evaluate the model. The model is evaluated using the mean average precision MAP, precision, and recall.

Citation Information

Cited By

  • Single-plant-scale tree positioning and identifying method

    CN120747757A

  • Unmanned aerial vehicle aerial photography road defect detection method and system based on improved YOLOv11

    CN121353964A

  • A Method and System for Road Defect Detection Based on Improved YOLOv11 Aerial Photography

    CN121353964B