Vehicle-mounted lightweight road crack detection method

By optimizing the YOLO11 architecture, introducing dual convolution and feature refinement modules, combined with SEFN module, the accuracy and real-time problems of existing road crack detection methods in complex scenarios are solved, and efficient and accurate road crack detection is achieved.

CN120279002APending Publication Date: 2025-07-08CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510531648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing road crack detection methods are insufficient in complex scenarios, excessive computing resource consumption, and poor real-time performance. Traditional methods are prone to noise interference and missed detection. When deployed on edge devices, deep learning methods are large in computing and slow inference speed, making it difficult to take into account both high efficiency and high precision.

Method used

By introducing the DFS-YOLO11 model, the convolutional neural network is optimized, the computational complexity is reduced and the feature expression ability is enhanced. Combined with the spatial attention mechanism and the efficient layer aggregation network, the crack detection accuracy and real-timeness of the model in complex backgrounds are improved.

Benefits of technology

It realizes efficient and real-time road crack detection on edge devices, improves detection accuracy and efficiency in complex backgrounds, adapts to environments with limited computing resources, reduces processing delays, and meets real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279002A_ABST
    Figure CN120279002A_ABST
Patent Text Reader

Abstract

The invention relates to the field of road crack detection, in particular to a vehicle-mounted lightweight road crack detection method. According to the method, the dual convolution structure, the FRM module and the SEFN module are introduced, so that the calculation complexity and the memory requirement are effectively reduced, and the reasoning speed is increased. Double convolution optimizes performance and calculation efficiency through a parallel convolution kernel, FRM reinforces crack detection precision by fusing local and global information, and SEFN realizes low-cost multi-scale feature fusion by using a multi-branch structure. The innovations not only improve the real-time detection capability of the model on edge equipment, but also ensure efficient operation in an environment with limited computing resources. The lightweight design of the YOLO11 model not only improves the reasoning speed, but also can be quickly deployed on the edge equipment, reduces the processing delay and adapts to the calculation limitation of the edge equipment, so that the accuracy and efficiency of road crack detection are improved, and efficient and accurate road surface condition evaluation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of road crack detection, and particularly to a vehicle-mounted lightweight road crack detection method. Background Art

[0002] Existing road crack detection methods mainly rely on traditional digital image processing techniques and deep learning-based methods.

[0003] Traditional digital image processing techniques usually rely on geometric features, texture information, edge features, etc. of images, and use algorithms such as edge detection, morphological processing, and threshold segmentation to detect cracks. Through these methods, obvious crack edges or crack regions can be found in the image, but they are easily affected by factors such as noise and illumination changes, and the detection accuracy for complex backgrounds or fine cracks is relatively low. Traditional methods have a small amount of calculation and are suitable for real-time processing, but they often cannot handle complex scenarios or cracks with more details.

[0004] Deep learning-based road crack detection methods have developed rapidly in recent years and have become the key technology to solve the deficiencies of traditional image processing methods. Although significant progress has been made in accuracy, there are still some key problems. First, deep learning models usually require a large amount of computing resources and memory, which makes their deployment on edge devices very difficult. Due to the computing power, storage space, and power consumption limitations of edge devices, the scale and complexity of deep learning models often cannot meet the requirements of these devices, resulting in serious constraints on real-time performance and deployment efficiency. When performing road crack detection, traditional deep learning models require high-performance hardware support, such as powerful GPUs and large-capacity memory, which are not available in many edge devices. To solve this problem, the computing efficiency is often improved by lightweight model structures, but the lightweight models may not be able to fully extract high-level features, affecting the overall performance. Second, most of the existing lightweight networks rely on traditional convolutional neural network architectures and cannot provide sufficient robustness in complex environments while improving computing efficiency.

[0005] With the increase in the service life of roads, affected by various factors such as bad weather, natural disasters, and traffic pressure, diseases such as cracks, potholes, and bumps often appear on the road surface, seriously affecting the safety and service life of the road. These diseases not only increase the probability of traffic accidents but also may bring huge economic losses and threats to driving safety. Therefore, it is crucial to regularly inspect and maintain roads. However, with the continuous growth of traffic flow and the normalization of overloaded road operation, traditional road maintenance work is facing increasing pressure. There are still certain problems in the existing maintenance mechanism and management system, and the improvement of maintenance technology level lags behind relatively, resulting in low road maintenance efficiency and difficulty in meeting the growing demand.

[0006] The rapid development of urban transportation has made it difficult for traditional manual inspection methods to keep up with maintenance requirements. Traditional manual inspection methods are not only inefficient and costly, but also pose significant safety hazards. Due to relying on manual experience, inspection results are often highly subjective and vulnerable to human factors, thus affecting the accuracy of decision-making.

[0007] Existing road crack detection technologies mainly face two core problems: Traditional image processing methods rely on algorithms such as edge detection and threshold segmentation. Although the computational cost is small, they are vulnerable to complex backgrounds, uneven illumination, and noise interference, making it difficult to capture fine or occluded cracks, resulting in a high miss rate; Deep learning methods, although improving accuracy, rely on a large amount of labeled data and high-performance GPUs. The model has a large number of parameters and slow inference speed, facing bottlenecks such as poor real-time performance, high memory occupancy, and excessive power consumption when deployed on edge devices. Existing lightweight models, although reducing the computational cost by compressing the network, significantly weaken the feature extraction ability. In scenarios with complex road surface textures and variable crack morphologies, the detection accuracy drops sharply, making it difficult to balance high efficiency and high precision. This imbalance contradiction of "accuracy - efficiency - resources" severely restricts the automation process of road maintenance, resulting in high manual inspection costs and being unable to meet the real-time monitoring requirements. Summary of the Invention

[0008] To solve the problems of insufficient accuracy, excessive consumption of computing resources, and poor real-time performance of existing road crack detection methods in complex scenarios, the present invention provides a vehicle-mounted lightweight road crack detection method. The disclosed DFS-YOLO11 model has been improved in the following three aspects: First, by improving the traditional convolutional module into a DualConv module, the present invention significantly reduces the computational complexity and memory consumption of the model, thereby improving the inference speed. Second, to make up for the possible loss of feature information caused by lightweight design, a Feature Refinement Module (FRM) is integrated after each feature extraction module, which enhances the feature expression ability of the model in complex road environments and improves the detection accuracy of different types of cracks. Finally, the SEFN (SPFF-ELAN Fusion Network) module is introduced to replace the SPPF module in the original network, further enhancing the model's ability to identify cracks of different widths and depths. The method mainly includes:

[0009] S1: Introduce a dual convolutional structure to improve the traditional convolutional structure in the convolutional neural network model, and use the improved convolutional neural network model to convert the obtained original image data into multi-scale feature maps;

[0010] S2: After the C3K2 module of the YOLO11 backbone network, add the FRM module. Use the FRM module to process the multi-scale feature maps, fuse the local features and global context information of the original image, and obtain the key features of the original image. The key features include the width, depth, and texture changes of road cracks;

[0011] S3: Incorporate channel and spatial attention mechanisms in the YOLO11 network to optimize the feature extraction process. By screening and emphasizing the key features in the image, obtain the most critical features;

[0012] S4: Introduce the SEFN module in the YOLO11 network to splice and fuse the most critical features output by all branches of the YOLO11 network to obtain the final feature map;

[0013] S5: Through the improvements in steps S1 - S4, obtain the DFS-YOLO11 network. Input the acquired road images into the DFS-YOLO11 network to obtain the localization and classification of cracks in the target road.

[0014] A computer device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0015] A computer-readable storage medium stores a computer program. When the program is executed by a processor, it implements the steps of the above method.

[0016] A computer program product includes a computer program or instruction. When the program or instruction is executed by a processor, it implements the steps of the above method.

[0017] A lightweight intelligent unmanned vehicle system based on the DFS-YOLO11 network deploys the DFS-YOLO11 network on the edge device of the unmanned vehicle. This network is used for road crack detection to capture road surface information in real-time; the system has two controllers. The main controller is responsible for deep learning tasks and executes the motion tasks of the unmanned vehicle through the sub-controller; in the data acquisition stage, send Bluetooth instructions through the mobile phone APP to control the unmanned vehicle to start. The camera and lidar perform data acquisition and feedback to the main and sub-controllers to form a closed-loop control and achieve the function of autonomous obstacle avoidance.

[0018] The beneficial effects brought by the technical solution provided by the present invention are as follows: By optimizing the YOLO11 architecture and adopting a lightweight design strategy, the present invention reduces the computational amount and the number of parameters of the network, enabling the model to operate efficiently on edge devices with limited computing resources and meeting the real-time detection requirements. At the same time, by introducing a feature refinement module and fusing local features with global context information, the present invention remedies the problem of low detection accuracy after lightweight processing of the model, enhances the model's ability to extract crack features in complex backgrounds, especially in the performance of fine cracks and diverse scenarios. Finally, the present invention combines the advantages of spatial pyramid pooling and an efficient layer aggregation network to propose a new network structure, forming a more comprehensive feature representation and improving the model's detection effect on cracks of different sizes. In addition, the lightweight design of the YOLO11 model not only improves the inference speed but also enables rapid deployment on edge devices, reduces processing latency, adapts to the computing limitations of edge devices, and thus improves the accuracy and efficiency of road crack detection. To achieve the detection task, the intelligent unmanned vehicle system designed by the present invention provides a carrier for the crack detection model. The system can use the on-board sensors to autonomously cruise and avoid obstacles, and obtain road surface images in real time, which are processed by the lightweight crack detection model integrated into the system, thereby realizing efficient and accurate road condition assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0020] Figure 1 is a flowchart of a vehicle-mounted lightweight road crack detection method in an embodiment of the present invention;

[0021] Figure 2 is a structural diagram of the improved DFS-YOLO11 network in an embodiment of the present invention;

[0022] Figure 3 is a structural diagram of the FRM module in an embodiment of the present invention;

[0023] Figure 4 is a structural diagram of the SEFN module in an embodiment of the present invention;

[0024] Figure 5 is a composition diagram of the intelligent unmanned vehicle system in an embodiment of the present invention;

[0025] Figure 6 is a hardware structural diagram of the intelligent unmanned vehicle system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the drawings.

[0027] Embodiment 1

[0028] Please refer to Figure 1 , Figure 1 which is a flowchart of a vehicle-mounted lightweight road crack detection method in an embodiment of the present invention, specifically including:

[0029] S1: Introduce a dual convolution structure (DualConv) to improve the traditional convolution structure in the convolutional neural network model, and use the improved convolutional neural network model to convert the acquired original image data into a multi-scale feature map;

[0030] S2: Add an FRM module after the C3K2 module of the YOLO11 backbone network, use the FRM module to process the multi-scale feature map, fuse the local features and global context information of the original image, and obtain the key features of the original image, where the key features include the width, depth, and texture changes of road cracks;

[0031] S3: Incorporate channel and spatial attention mechanisms in the YOLO11 network to optimize the feature extraction process, and obtain the most critical features by screening and emphasizing the key features in the image;

[0032] S4: Introduce an SEFN (SPFF-ELAN Fusion Network) module in the YOLO11 network to splice and fuse the most critical features output by all branches of the YOLO11 network to obtain the final feature map;

[0033] S5: Through the improvements in steps S1 - S4, obtain the DFS-YOLO11 network, input the acquired road image into the DFS-YOLO11 network, and obtain the localization and classification of cracks in the target road.

[0034] YOLO11 is the latest product in the series of real-time object detectors released by the Ultralytics team. Based on the YOLO series, YOLO11 further optimizes the accuracy, speed, and efficiency of object detection through major improvements in architecture and training methods. Compared with YOLO, YOLO11 further improves network performance and flexibility, making YOLO11 an ideal choice for various computer vision tasks such as object detection and tracking, instance segmentation, image classification, and pose estimation. The core of the YOLO architecture consists of three basic components. First, the backbone network, as the main feature extractor, uses a convolutional neural network model to convert the original image data into a multi-scale feature map. Second, the neck component, as an intermediate processing stage, uses specialized layers to aggregate and enhance feature representations at different scales. Finally, the head component, as a prediction mechanism, generates the final output based on the optimized feature map for object localization and classification. The improved DFS-YOLO11 network structure is as Figure 2 shown, and the specific improvements are as follows:

[0035] a) Dual convolutional structure

[0036] In the optimization of deep neural network models, lightweight design is one of the key objectives, especially applicable to models deployed on resource-constrained devices. A lightweight network refers to a network that, in its design, adopts a series of techniques and methods to reduce the computational amount, the number of parameters, and the storage requirements of the network, so that the network can operate efficiently on devices with limited resources while maintaining as good performance as possible. The present invention proposes a lightweight improvement method for a feature extraction module, which improves the traditional convolutional structure by introducing a dual convolutional structure (DualConv). DualConv applies two parallel convolutional kernels at the same position. One convolutional kernel is responsible for capturing shallow features, and the other delves into extracting complex features, effectively balancing performance and computational efficiency.

[0037] This structure reduces the computational cost and memory usage of the deep neural network model, while ensuring the model performance and providing richer feature representations. The core advantage of the DualConv structure lies in generating richer feature maps using fewer computational resources. These feature maps are concatenated after expansion to form the final output feature map. Replacing the traditional convolutional structure with the DualConv structure significantly reduces the computational cost in the model feature extraction stage and achieves the design goal of a lightweight model.

[0038] b) Feature Refinement Module (FRM)

[0039] The present invention proposes a Feature Refinement Module (FRM), as Figure 3 shown. This module is integrated after the C3K2 module of the YOLO11 backbone network, aiming to optimize the road crack detection performance. FRM enhances the model's sensitivity to road cracks, especially key features such as the width, depth, and texture changes of cracks, by fusing local features and global context information, which is crucial for accurately identifying and classifying different types of road cracks. FRM uses efficient depthwise separable convolution technology to extract local features, which reduces the computational burden while maintaining sensitivity to key spatial information, helping to accurately capture crack details. In addition, FRM uses average pooling and max pooling operations in the horizontal and vertical directions to effectively capture feature information at different scales.

[0040] The present invention also integrates channel and spatial attention mechanisms, optimizing the feature extraction process. By screening and emphasizing key features in the image, it focuses the model's computational resources on the most critical features, thereby improving the accuracy and efficiency of crack detection. FRM adopts residual connection technology, directly connecting the input to the output, effectively alleviating the problem of gradient disappearance, promoting the deep training of the network, and retaining the original feature information. This design strategy not only improves the learning ability of the model but also significantly enhances the overall performance of road crack detection by integrating it into the YOLO11 network.

[0041] c) SEFN

[0042] In the field of object detection, for the effective identification of cracks, especially the accurate extraction and fusion of multi-scale features, are crucial for ensuring the detection of various crack sizes and shapes. Fast Spatial Pyramid Pooling (SPPF) uses fixed-size pooling kernels to examine the input feature map and extract multi-level features, thus significantly enhancing the model's ability to recognize cracks of various sizes. SPPF usually relies on basic concatenation operations to integrate features from different pooling layers, which limits the potential of leveraging complementary information between multi-scale features and may thus affect the overall performance of the model. Efficient Layer Aggregation Networks (ELANs) adopt a multi-branch parallel structure, achieving the effective extraction and fusion of features at different levels under the condition of low computational cost, showing significant advantages in feature fusion.

[0043] To integrate the advantages of SPPF and ELANs and further optimize the model's detection ability for cracks of various sizes, the present invention proposes a new model, SEFN (SPFF-ELAN Fusion Network), as Figure 4 shown. This module draws on the idea of multiple branches in parallel in ELAN to extract multi-scale features and cleverly uses the SPPF module as the specific implementation of the branches. Each SPPF module has pooling kernels of different sizes, which can extract feature information at different levels. Larger pooling kernels capture the overall contour and distribution features of the crack, while smaller pooling kernels focus on the texture and local details of the crack. Finally, the SEFN module splices and fuses the outputs of all branches, integrating feature information at different levels in this simple and efficient way to form a more comprehensive feature representation and improving the detection effect of the model on cracks of different sizes.

[0044] By introducing a dual convolutional structure, an FRM module, and an SEFN module, the present invention effectively reduces the computational complexity and memory requirements and improves the inference speed. The dual convolution optimizes the performance and computational efficiency through parallel convolutional kernels. The FRM enhances the crack detection accuracy by fusing local and global information. The SEFN uses a multi-branch structure to achieve low-cost multi-scale feature fusion. These innovations not only improve the real-time detection ability of the model on edge devices but also ensure efficient operation in an environment with limited computing resources.

[0045] Embodiment 2

[0046] A computer device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.

[0047] Embodiment 3

[0048] A computer-readable storage medium stores a computer program. When the program is executed by a processor, the steps of the above method are implemented.

[0049] Embodiment 4

[0050] A computer program product includes a computer program or instruction. When the program or instruction is executed by a processor, the steps of the above method are implemented.

[0051] Embodiment 5

[0052] Based on the DFS-YOLO11 network, the present invention designs and builds an intelligent unmanned vehicle system based on Jetson Nano B01. The composition of the system is as Figure 5 shown. The intelligent unmanned vehicle system refers to a system that senses the surrounding environment, processes data in real time, and makes driving decisions without human intervention to complete specified tasks by carrying sensors and an efficient computing platform. The lightweight DFS-YOLO11 network after YOLO11 is deployed on the edge device Jetson Nano B01, and this network is used for road crack detection and can capture road surface information in real time. Edge deployment means deploying the computing and data processing capabilities on "edge" devices close to the data source or users, rather than sending all data to a remote server or cloud for processing. The system has two controllers. The main controller Jetson Nano B01 is responsible for deep learning tasks and executes the motion tasks of the unmanned vehicle through the secondary controller STM32F103C8T6. In the data acquisition stage, Bluetooth instructions are sent through a mobile phone APP to control the unmanned vehicle to start, and the camera and lidar collect data, and a closed-loop control is formed to achieve the function of autonomous obstacle avoidance. The trained network model and code are embedded in Jetson Nano B01, and its operating environment is configured and a self-start command is added.

[0053] This system is a cascade control system, and the hardware is built as Figure 6 shown. The main loop uses a Jetson nano B01 controller. Since this controller has a high cost and is equipped with a small GPU, it is mainly used to execute deep learning tasks and indirectly control the secondary controller stm32f103c8t6. The secondary controller executes the motion tasks of the unmanned vehicle. Sensors such as lidar and encoders output feedback signals to the main and secondary controllers to form a closed-loop control and achieve the function of autonomous obstacle avoidance.

[0054] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A vehicle-mounted lightweight road crack detection method, characterized in that Including: S1: Introduce a dual convolutional structure to improve the traditional convolutional structure in the convolutional neural network model, and use the improved convolutional neural network model to convert the obtained original image data into a multi-scale feature map; S2: Add an FRM module after the C3K2 module of the YOLO11 backbone network, use the FRM module to process the multi-scale feature map, fuse the local features and global context information of the original image, and obtain the key features of the original image. The key features include the width, depth, and texture changes of road cracks; S3: Integrate channel and spatial attention mechanisms in the YOLO11 network to optimize the feature extraction process, and obtain the most critical features by screening and emphasizing the key features in the image; S4: Introduce an SEFN module in the YOLO11 network to splice and fuse the most critical features output by all branches of the YOLO11 network to obtain the final feature map; S5: Through the improvements in steps S1 - S4, obtain the DFS-YOLO11 network, input the obtained road image into the DFS-YOLO11 network, and obtain the localization and classification of cracks in the target road.

2. The on-vehicle lightweight road crack detection method according to claim 1, wherein, In S1, the dual convolutional structure applies two parallel convolutional kernels to the same position in the convolutional neural network model. One convolutional kernel is responsible for capturing shallow features, and the other convolutional kernel is used to deeply extract complex features.

3. The on-vehicle lightweight road crack detection method according to claim 1, characterized in that In S2, the FRM module adopts a residual connection technique to directly connect the input to the output, effectively alleviating the problem of gradient disappearance, promoting the deep training of the YOLO11 network, and at the same time retaining the original feature information of the image.

4. The vehicle-mounted lightweight road crack detection method according to claim 1, wherein In S4, the SEFN module has pooling kernels of different sizes for extracting feature information at different levels. The larger pooling kernel is used to further capture the overall contour and distribution features of the cracks, while the smaller pooling kernel is used to further extract the texture and local details of the cracks.

5. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes a computer program to implement the steps of the vehicle-mounted lightweight road crack detection method described in any one of claims 1 - 4.

6. A computer-readable storage medium, characterized in that, There is a stored computer program, which, when executed by the processor, implements the steps of the vehicle-mounted lightweight road crack detection method described in any one of claims 1 - 4.

7. A computer program product, characterized in that, Including a computer program or instruction, which, when executed by the processor, implements the steps of the vehicle-mounted lightweight road crack detection method described in any one of claims 1 - 4.

8. A lightweight intelligent unmanned vehicle system based on the DFS-YOLO11 network, characterized in that, Deploy the DFS-YOLO11 network on the edge device of the unmanned vehicle. This network is used for road crack detection and captures road surface information in real time. The system has two controllers. The main controller is responsible for deep learning tasks and executes the motion tasks of the unmanned vehicle through the sub-controller. In the data acquisition stage, send Bluetooth instructions through the mobile phone APP to control the unmanned vehicle to start, and the camera and lidar perform data acquisition and feedback to the main and sub-controllers to form a closed-loop control to achieve the function of autonomous obstacle avoidance.

Citation Information

Cited By

  • Safe wearing vision automatic detection method, device and system based on edge end equipment

    CN120747873A