Road crack analysis method based on improved YOLO algorithm

By improving the YOLO algorithm, constructing data sets and training the target detection network, the problems of insufficient detection accuracy and difficulty in applying on mobile devices in the prior art are solved, and efficient, accurate and intelligent detection of road cracks is achieved.

CN120014500APending Publication Date: 2025-05-16CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Application Number
CN202510488637.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing road crack detection methods have problems such as insufficient detection accuracy, high cost, large parameters, insufficient multi-scale feature processing capabilities, and difficulty in applying on mobile terminal devices.

Method used

The improved YOLO algorithm is used to construct the data set and train the object detection network, use StarNet as the backbone network, the C2f-Star module is used for feature extraction, the GN-YoloHead module is used for detection head design, and Focaler-MPD-WIoUv3 is used in the loss function.

Benefits of technology

It realizes efficient, accurate and intelligent detection of road cracks, reduces computing resource consumption, is suitable for running on mobile devices and edge computing devices, and improves the ability to identify multi-scale cracks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014500A_ABST
    Figure CN120014500A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a road crack analysis method based on an improved YOLO algorithm, and the method comprises the steps: constructing a data set; constructing a target detection network; training the target detection network by adopting a data set to obtain a target recognition model; and acquiring a to-be-detected crack image from the road surface by using the unmanned aerial vehicle, and inputting the to-be-detected crack image into the target recognition model for target recognition. According to the method, a backbone network of YOLOv8 is replaced by StarNet, so that the calculation burden is remarkably reduced, and the detection speed is improved; a C2f-Star module is introduced to enhance multi-scale feature extraction and suppress interference of irrelevant features; a GN-YoloHead detection head is used for optimizing computing resource management; a traditional loss function is replaced by a Focaler-MPD-WIoUv3 loss function, so that the detection performance is improved, and the method is specially used for identifying the multi-scale crack of the road.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a road crack analysis method based on an improved YOLO algorithm. Background Art

[0002] Cracks are considered to be the most common form of road damage and are one of the key points of road technical condition detection. They not only directly reflect the decline in road service performance, but may also cause a series of serious problems such as roadbed settlement and water damage. Therefore, it is particularly important to quickly identify and accurately detect road cracks, which not only helps to extend the service life of roads, but also has important practical significance in ensuring traffic safety. However, traditional road crack detection mostly relies on manual operation, which has shortcomings such as dangerous working environment, low detection efficiency and high dependence on subjective experience. When necessary, traffic routes need to be closed, which brings potential safety problems to detection personnel. Therefore, traditional manual detection methods can no longer meet the needs of highway maintenance in my country. At the same time, multi-scale road cracks also have a certain impact on detection accuracy. At present, automatic classification and recognition of cracks in collected road images has become the mainstream detection method in this field.

[0003] At present, the advanced technical means used for road crack detection mainly include: A Chinese invention with application number 202411226865.X, entitled "A method, device, system, and storage medium for identifying pavement cracks", belongs to the field of image data processing technology, and includes obtaining historical pavement images, constructing a data set, establishing a convolutional neural network model for pavement crack identification based on YOLOv8, using the convolutional neural network model to identify pavement disease images; and using non-maximum suppression to filter the crack boundary box output by the convolutional neural network model. This invention effectively solves the problems of low recognition efficiency and low detection target accuracy of existing methods.

[0004] A Chinese invention with application number 202410878110.1, entitled A method for identifying small targets in pavement cracks based on a lightweight algorithm, belongs to the field of image data processing technology, including constructing a data set containing pavement crack images of different resolutions; constructing a lightweight optimization model based on the YOLOv5s algorithm, using MobileNetV2 as the backbone network to reduce computational complexity, and embedding a fusion domain attention module to enhance image features and adaptively assign feature weights. After the model parameters are set, input images for training to obtain training weights; and identify small targets in pavement cracks based on the evaluated algorithm model. This invention improves the recognition accuracy and computational speed of the improved network model in small target detection tasks.

[0005] These technologies have all made innovations in the field of road crack recognition, but they still have some limitations. Although the YOLOv8-based method improves recognition efficiency and accuracy, it may consume a lot of computing resources when processing large-scale data, and may not be effective in identifying multi-scale cracks. Secondly, although the lightweight algorithm-based method performs well in small target recognition, it may sacrifice accuracy when facing larger or more complex cracks. These methods may not perform well in real-time processing and response to multiple crack types, especially in applications that need to adapt to a wide range of environmental conditions. Summary of the invention

[0006] In view of the problems of existing road disease detection methods, such as insufficient detection accuracy, high cost, large parameters, insufficient multi-scale feature processing capabilities, and difficulty in application on mobile terminal devices, the present invention aims to provide a road crack analysis method based on an improved YOLO algorithm to achieve efficient, accurate and intelligent detection of cracks; The technical solution adopted by the present invention is a road crack analysis method based on an improved YOLO algorithm, the method comprising: Constructing a data set, using a drone to obtain images of road surfaces containing multi-scale cracks, preprocessing the images and marking the cracks; all images with marked cracks constitute a data set; Construct a target detection network, select YOLOv8 including backbone network, neck network, head network and loss calculation module as the basic recognition network, use StarNet as the backbone network, the neck network includes C2f-Star module, and use GN-YoloHead module based on group normalization and shared convolution in the detection stage. In the target position loss function, use Focaler-MPD-WIoUv3 loss function to replace the original loss function; get the target detection network; The target detection network is trained using the data set consisting of the images of the annotated cracks to obtain a target recognition model; The unmanned aerial vehicle is used to obtain the crack images to be detected from the road surface, and the crack images are input into the target recognition model for target recognition.

[0007] Furthermore, the backbone network of the target detection network is a five-layer structure, including a first layer, a second layer, a third layer, a fourth layer and a fifth layer; The first layer includes the Stem layer of the custom convolutional batch normalization module; The second layer includes a downsampling layer and a C2f-Star module; The third layer includes a downsampling layer and a C2f-Star module; The fourth layer includes a downsampling layer and three C2f-Star modules; The fifth layer includes a downsampling layer and a C2f-Star module.

[0008] Furthermore, the StarNet includes a convolutional layer, a C2f-Star layer, a batch normalization layer, an activation function, a global average pooling layer, and a fully connected layer; The C2f-Star layer includes a depth-separable convolutional layer, a fully connected layer, a batch normalization layer, and a star operation; wherein the star operation is specifically: (1) in, is the index channel, Indicates the number of input channels, are the coefficients of each channel, is the weight of a single output channel, To represent a single input feature element, is transposed; The process of star calculation produces Different items: (2) No. Layer star operation output The recursion is expressed as: (3) in, is the network width, It is the comprehensive parameter matrix of each layer in StarNet.

[0009] Furthermore, the GN-YoloHead module includes a group normalization convolution layer, a prediction box convolution layer, a classification convolution layer and a scaling layer; the specific working mode of the GN-YoloHead module is as follows: the crack feature map of each scale passes through a 1×1 group normalization convolution layer, and then the feature map is further processed by two 3×3 group normalization convolution layers, and the weights are shared between the two layers. Then, the feature map of each scale is processed by the prediction box convolution layer and the classification convolution layer with shared weights; wherein, the prediction box convolution layer is used to predict the bounding box of the target, and the classification convolution layer is used to classify the object in each prediction box; finally, the output of each scale adjusts the size of the prediction box through a scaling layer to adapt to targets of different scales.

[0010] Furthermore, the Focaler-MPD-WIoUv3 loss function includes Wise-IoU, MPD-IoU and Focaler-IoU loss functions: (5).

[0011] Furthermore, the bounding box regression loss of the Wise-IoU loss function Specifically: (6) (7) (8) in, Used to measure the abnormality of the prediction box. is the gradient gain, and To adjust the hyperparameters of the loss function; Indicates the horizontal and vertical coordinate values ​​of the prediction box, are the horizontal and vertical coordinate values ​​of the target frame, The values ​​correspond to the height and width of the two boxes respectively. is the height and width of the overlapping area of ​​the two frames, is the area of ​​the non-overlapping part of the two boxes, is the height and width of the prediction box, The height and width of the target box.

[0012] Furthermore, the bounding box loss function MPD-IoU It is used to optimize the matching degree between the predicted frame and the real frame, and improve the positioning accuracy of targets with significant size differences. MPD-IoU The loss function is specifically: (9) MDP-IoU Bounding Box Regression Loss Function Defined as: (10) in, is the distance from the lower left corner of the real box to the lower left corner of the predicted box, d 2 is the distance from the upper right corner of the real box to the upper right corner of the predicted box, The values ​​correspond to the height and width of the two boxes respectively. is the area of ​​the non-overlapping part of the two boxes, The height and width of the overlapping area of ​​the two boxes.

[0013] Furthermore, the Focaler-IoU is used to increase the loss weight of difficult-to-classify samples, so that the model pays more attention to the difficult-to-identify areas; specifically: (11) in, After reconstruction focaler-IoU Loss function, IoU For the original IoU value, , For the lowest IoU Threshold, indicating the invalid area; For the highest IoU Threshold, indicating high-quality prediction area; by adjusting and The value of Focusing on different regression samples, the loss is defined as follows: (12).

[0014] The application of the technical solution of the present invention has the following beneficial effects: 1) The StarNet network is used as the backbone network of the YOLOv8 algorithm, downsampling is achieved through the convolution layer, and the C2f-Star module is applied for feature extraction. In order to improve computational efficiency, batch normalization is used to replace layer normalization and is placed after the deep convolution layer to facilitate fusion in the inference stage. The network structure iteratively uses multiple star blocks for feature extraction, avoiding complex structures or hyperparameter settings, optimizing the balance between the model's precision and recall, and effectively reducing the amount of computation, thereby achieving real-time detection of cracks; 2) Use the C2f-Star module to replace the original C2f module. The C2f-Star module can achieve the same or better recognition effect on the basis of reducing the computational complexity. It improves the performance of the model in feature layer processing and enhances the effect of extracting key features from multi-scale complex data; 3) The introduction of the GN-YoloHead detection head fully utilizes the advantages of group normalization (GroupNorm) and shared convolution to effectively fuse feature information at different levels, improving the model's ability to detect targets of different sizes. In addition, due to the efficient management of computing resources, the detection head is particularly suitable for use in application scenarios with limited computing power but high real-time requirements, such as mobile devices and edge computing devices; 4) Using Focaler-MPD-WIoUv3 in the loss function of the YOLOv8 algorithm can enhance the model performance in different aspects, especially more accurately detecting small cracks or irregular cracks. Especially on unbalanced data sets such as crack detection, it can effectively improve the detection rate of cracks.

[0015] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 is a flowchart of the steps of the road crack identification method in this embodiment; Figure 2 This is a schematic diagram of the StarNet structure in this embodiment; Figure 3 Schematic diagram of the structure of C2f module and C2f-Star module in this embodiment, (a) is C2f module, (b) is C2f-Star module; Figure 4 Schematic diagram of the depthwise separable convolution structure in this implementation; Figure 5 GN-YoloHead network structure diagram in this implementation mode; Figure 6 Schematic diagram of the group normalization structure in this implementation mode; Figure 7 2 is a schematic diagram of the frame regression model structure in an embodiment of the present invention; Figure 8 It is a flow chart of the target detection network in the implementation mode of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] like Figure 1-Figure 8 As shown, this embodiment discloses a road crack analysis method based on an improved YOLO algorithm, comprising the following steps: S1: Construct a dataset. Use drones to obtain images of road surfaces containing multi-scale cracks. Based on the initial crack images, introduce a variety of image enhancement techniques to process them and annotate the cracks. All images with annotated cracks constitute the dataset.

[0020] In this embodiment, for the multi-scale cracks, a single rectangular frame is uniformly used for marking.

[0021] S2: Construct target detection network (such as Figure 8 As shown in Figure 1, YOLOv8 is selected as the basic recognition network. The traditional YOLOv8 consists of a backbone network, a neck network, a head network and a loss calculation module. The backbone network, the neck network, the head network and the loss calculation module are improved respectively. Specifically, the lightweight network StarNet is used as the backbone network, and the C2f module (as shown in Figure 1) in the YOLOv8 network structure is replaced by the C2f module (as shown in Figure 1). Figure 3 (a)) is replaced by the C2f-Star module (such as Figure 3 (b)), and in the detection stage, the GN-YoloHead module based on group normalization and shared convolution is adopted. In the target position loss function, the Focaler-MPD-WIoUv3 loss function is used to replace the original loss function to obtain the target detection network.

[0022] In this embodiment, the backbone network is mainly responsible for extracting basic and deep feature representations from the input image. This process involves multi-level convolution operations to capture the spatial hierarchy of the image, from low-level edges and textures to high-level semantic information. In order to ensure the effectiveness and consistency of feature extraction, the C2f-Star module is adopted, which is optimized and can perform feature learning stably and efficiently. The neck network, as a bridge between the backbone network and the detection or segmentation head, has the core task of fusing and enhancing the multi-scale feature maps from the backbone network. By integrating features at different levels, the neck network can enhance the robustness to changes in object size and improve the model's detection ability for targets of different scales. Similarly, the neck network also uses the C2f-Star module to maintain feature processing logic consistent with the backbone network, thereby ensuring the consistency and coherence of feature expression in the entire network architecture.

[0023] Using the same C2f-Star module in the backbone and neck networks not only helps to simplify the network design, but also promotes seamless feature transfer, allowing the two parts to work together more harmoniously, ultimately improving the overall performance of visual tasks.

[0024] like Figure 2 As shown, the lightweight network StarNet includes a convolutional layer, a C2f-Star layer, a batch normalization layer, an activation function, a global average pooling layer, and a fully connected layer; In this embodiment, the lightweight network StarNet effectively captures high-dimensional features and nonlinear features in low-dimensional space. In a single-layer neural network, StarNet combines the weight matrix and bias into one entity, represented as ,Right now is the comprehensive parameter matrix of each layer in StarNet, which combines the weight matrix and bias term into a unified entity, where represents the weight term, represents the bias term; accordingly, the input vector The new input matrix is ​​expanded to contain a constant term (usually 1) ,Through this arrangement, StarNet implements the star operation, which is expressed as , Specifically, define is the weight of a single output channel and can be easily extended to accommodate multiple output channels , and handle multiple feature elements To simplify the analysis, this implementation first focuses on the scenario with a single input and a single output, as shown in formula (1): (1) Where: is the index channel, Indicates the number of input channels, are the coefficients of each channel, is a single input feature element, which is the input matrix A component of: (2) The star operation process produces different terms, as shown in Equation (2). In addition, by stacking multiple layers, the hidden dimension can be recursively increased to approach infinity. Assuming the network width , the output of the previous star operation is expressed as (3), then Layer star operation output It can be expressed recursively as: (3) like Figure 3 As shown in (b), the C2f-Star layer includes a depthwise separable convolutional layer (DW-Conv), a fully connected layer (FC), a batch normalization layer (BN), and a star operation; In this embodiment, the crack image is firstly processed by a depth-separable convolutional layer (e.g. Figure 4 The image is processed by a 10-th order convolutional layer (as shown in the figure), followed by batch normalization to normalize the output, thereby improving the training efficiency and stability of the model. The image is then passed to a fully connected layer with dimension expansion, and nonlinear features are introduced using the ReLU6 activation function. As the process continues, the image is split in the network, with one part continuing to be processed by another combination of a fully connected layer and a depthwise separable convolutional layer, while the other part is integrated with the output of the first fully connected layer through a star operation to finally generate the output of the model.

[0025] In a specific embodiment, the depthwise separable convolution decomposes a convolution operation in a conventional convolution structure into two parts: a 3×3 depthwise convolution and a 1×1 pointwise convolution. The depthwise convolution performs convolution operations independently for each input channel, while the pointwise convolution performs weighted combination of the generated feature layers in the depth direction to finally form a new feature layer. q dpt for: (4) Where: are the number of input feature channels and the number of output feature channels, respectively.

[0026] like Figure 5 As shown, the lightweight GN-YoloHead module includes a group normalization (GroupNorm) convolution layer, a prediction box convolution layer, a classification convolution layer and a scaling layer; In this implementation, the crack feature map of each scale passes through a 1×1 group normalized convolution layer (GN_Conv 1×1), and then the feature map is further processed by two 3×3 group normalized convolution layers (GN_Conv 3×3), which share weights. Then, the feature map of each scale passes through a prediction box convolution layer (Conv_Box) and a classification convolution layer (Conv_Cls) with shared weights. The Conv_Box layer is responsible for predicting the bounding box (position and size) of the target, while the Conv_Cls layer is used to classify the object in each prediction box. Finally, the output of each scale passes through a Scale module to adjust the size of the prediction box to accommodate targets of different scales.

[0027] In some specific implementations, the feature map is obtained from the original annotated crack map through continuous processing of the downsampling layers and the C2f-Star module in the backbone network and the neck network.

[0028] In a specific embodiment, the group normalization (e.g. Figure 6 The channels are divided into several groups, the mean and variance are calculated in each group, and then all the data in the group are normalized according to the calculated results.

[0029] The shared convolution reduces training parameters through a weight sharing mechanism, thereby improving computational efficiency and reducing the risk of overfitting. This mechanism can effectively extract local features and maintain spatial structure. Its invariance enhances the model's adaptability and generalization performance to translation transformations, and also supports the processing of multi-channel inputs, thereby further improving the model's expressiveness.

[0030] The Focaler-MPD-WIoUv3 loss function includes Wise-IoU, MPD-IoU and Focaler-IoU loss functions; In this embodiment, the Focaler-MPD-WIoUv3 loss function combines three strategies: Wise-IoU adjusts the loss weight to cope with the category imbalance of the data set, which is particularly suitable for the imbalance between crack and non-crack areas; MPD-IoU optimizes the matching degree between the predicted box and the real box, and improves the positioning accuracy of targets with significant size differences; Focaler-IoU increases the loss weight of difficult-to-classify samples, so that the model pays more attention to difficult-to-identify areas.

[0031] In a specific embodiment, the Wise-IoU (WIoU) loss function is a bounding box regression loss with a dynamic non-monotonic focusing mechanism. It evaluates the quality of the anchor box by "outliers" to avoid over-penalizing geometric factors in low-quality labeled data. The present invention uses WIoUv3 with a two-layer attention mechanism and a dynamic non-monotonic mechanism. Its WIoUv3 bounding box regression loss function The expression is as follows: (6) (7) (8) In the formula, It is used to measure the abnormality of the prediction box. The lower the value, the higher the quality of the anchor box. The non-monotonic focal loss function can convert small gradient gains Assigned to prediction boxes with high outliers, thereby effectively suppressing the adverse gradient effects brought by low-quality training samples; and is the hyperparameter for adjusting the loss function. The meanings of other parameters are as follows Figure 7 As shown, and Indicates the horizontal and vertical coordinate values ​​of the prediction box, and are the horizontal and vertical coordinate values ​​of the target frame, and The values ​​correspond to the height and width of the two boxes respectively. and is the height and width of the overlapping area of ​​the two boxes, is the area of ​​the non-overlapping part of the two boxes, and is the height and width of the prediction box, and The height and width of the target box.

[0032] The MPD-IoU can be calculated based on the minimum point distance between points, taking into account the overlap area, center point distance, and width and height deviations. The meanings of the parameters are as follows: Figure 7 As shown, The distance from the lower left corner of the real box to the lower left corner of the predicted box 、 is the distance from the upper right corner of the real box to the upper right corner of the predicted box: (9) MDPIoU bounding box regression loss function Defined as: (10) The formula of Focaler-IoU is as follows: (11) In the formula, For the reconstructed Focaler-IoU, IoU is the original IoU value, , by adjusting and The value of Focusing on different regression samples, the loss is defined as follows: (12) In some specific implementations, the three loss functions of Focaler-MPD-WIoUv3 are not fused by a simple weighted summation method, but a stacked fusion method is adopted. First, MPD-IoU is used as the basic IoU calculation method. This loss function takes into account the distance measurement between the predicted box and the true box, and optimizes the positional relationship between the predicted box and the true box; then WIoU is applied as a modulation mechanism on top of MPD-IoU to adjust the loss weight to deal with the category imbalance problem in the data set; finally, Focaler-IoU is another modulation mechanism, which adjusts the sample weight according to the difficulty and increases the weight of difficult-to-classify samples; the fusion process of Focaler-MPD-WIoUv3 is specifically as follows: (5) S3: Using the data set to train the target detection network to obtain a target recognition model; S4: Use the UAV to obtain the crack image to be detected from the road surface and input it into the target recognition model for target recognition.

[0033] In some possible implementations, the backbone network has a five-layer structure, including a first layer, a second layer, a third layer, a fourth layer, and a fifth layer; The first layer includes the Stem layer of the custom convolutional batch normalization module; The second layer includes a downsampling layer and a C2f-Star module; The third layer includes a downsampling layer and a C2f-Star module; The fourth layer includes a downsampling layer and three C2f-Star modules; The fifth layer includes a downsampling layer and a C2f-Star module; In some specific implementations, the backbone network and the neck network both use C2f-Star modules with the same structure, but their respective functional positioning is different. In the backbone network, the C2f-Star module works in conjunction with the downsampling layer, mainly undertaking the task of preliminary extraction of image features. It generates feature maps of different resolutions through layer-by-layer deepening convolution operations, capturing visual features from low-level to high-level, thereby providing rich and diverse feature representations for subsequent processing. In contrast, in the neck network, the C2f-Star module focuses on deep fusion and enhancement of different scale features from the backbone network. This process aims to integrate multi-level feature information, strengthen cross-scale semantic associations, and ultimately generate high-quality multi-scale feature maps for the detection head to perform accurate target detection or classification. Although the C2f-Star modules in the backbone network and the neck network have the same structural design, this consistency ensures the coherence and stability of the feature processing logic; however, their positions in the network determine their unique functional roles, namely feature extraction and feature fusion enhancement. In this way, the present implementation can effectively maintain the consistency of feature expression, while optimizing the task performance at different stages, and ultimately improve the overall performance of the model.

[0034] Example 1 This embodiment discloses a road crack analysis method based on an improved YOLO algorithm. The method of this embodiment adopts the StarNet network as the backbone network of the YOLOv8 algorithm, optimizes the balance between the model's precision and recall, effectively reduces the amount of calculation, and thus realizes real-time detection of cracks; uses the C2f-Star module to replace the original C2f module, improves the performance of the model in feature layer processing, and enhances the effect of extracting key features from multi-scale complex data; introduces the GN-YoloHead detection head to improve the efficiency of the model in computing resource management; adopts Focaler-MPD-WIoUv3 to be applied to the loss function of the YOLOv8 algorithm, improves the adaptability of the model to unbalanced crack data sets, and avoids additional computational burden. The method of this embodiment solves the problem that the existing algorithm may not perform well in real-time processing and coping with multiple crack types, especially in applications that need to be widely adapted to different environmental conditions, which makes it difficult to further improve the recognition accuracy of disease images.

[0035] In order to verify the effectiveness of the method of the present invention, based on the above-mentioned embodiment, a comparative test is conducted between the method of the present invention and several commonly used target detection algorithms: The drone aerial images acquired during disease detection on multiple roads were systematically summarized to form an initial dataset of pavement cracks containing 1,150 multi-scale features. The diversity of actual road projects is reflected in the complex background and different lighting conditions (including bright and dark) when the drones are shooting. In order to further enrich the dataset and expand the image background information, a variety of image enhancement techniques were introduced on the basis of the initial dataset, including random rotation, cropping, horizontal flipping, vertical flipping, color temperature adjustment, noise addition, and blur processing, expanding the dataset to 2,000 images.

[0036] The dataset is divided into training set, validation set and test set in a ratio of 8:1:1, and compared with 7 commonly used target detection algorithms. The experimental results are shown in Table 1.

[0037] Table 1 Comparison of different target detection algorithms , mAP: The full name is Mean Average Precision, which refers to the average accuracy of all categories. It is a commonly used evaluation indicator in the field of target detection and is used to measure the performance of target detection algorithms.

[0038] GFLOPs: The full name is floating point operations. GFLOPs is the number of floating point operations that can be performed per second. It is an important indicator for measuring the size of the target detection model.

[0039] Note: The YOLOv5 and YOLOv7 used in Table 1 are different versions of the same algorithm.

[0040] It can be seen from Table 1 that the method of the present invention can achieve higher recognition accuracy, and the mAP can reach 62.4%, which is 2.4%, 4.0%, 2.1%, 6.7%, 4.9%, 19.9%, and 12.9% higher than the YOLOv11, YOLOv10, YOLOv8, YOLOv7, YOLOv5, Faster RCNN, and SSD comparison models, respectively. It can be seen that compared with other algorithms, the method of the present invention has the best road multi-scale crack image detection accuracy. At the same time, the GFLOPs of the present invention is 4.5G, which is 1.8G, 3.9G, 3.6G, 8.5G, 11.4G, 110.9G, and 248.1G less than the YOLOv11, YOLOv10, YOLOv8, YOLOv7, YOLOv5, Faster RCNN, and SSD comparison models, respectively. The present invention is the most lightweight road multi-scale crack image detection model, which meets the requirements of lightweight disease detection and is more suitable for road multi-scale crack detection tasks.

[0041] In addition, this embodiment also discloses a computer device, including a memory and a processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the above-mentioned road multi-scale crack disease identification method when executing the computer program.

[0042] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the computer device.

[0043] The computer device may be a computing device such as a mobile phone, a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may also include an input / output device, a network access device, a bus, etc.

[0044] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0045] This embodiment discloses a road crack analysis method based on an improved YOLO algorithm. The StarNet network is used as the backbone network of the YOLOv8 algorithm to optimize the balance between the model's precision and recall, effectively reducing the amount of calculation, thereby realizing real-time detection of cracks; the C2f-Star module is used to replace the original C2f module, thereby improving the performance of the model in feature layer processing and enhancing the effect of extracting key features from multi-scale complex data; the GN-YoloHead detection head is introduced to improve the efficiency of the model in computing resource management; Focaler-MPD-WIoUv3 is used in the loss function of the YOLOv8 algorithm to improve the adaptability of the model to unbalanced crack data sets and avoid additional computational burdens. This solves the problem that the existing algorithms may not perform well in real-time processing and coping with multiple crack types, especially in applications that need to be widely adapted to different environmental conditions, which makes it difficult to further improve the recognition accuracy of disease images.

[0046] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0047] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A road crack analysis method based on an improved YOLO algorithm, characterized in that: The method includes: Constructing a data set, using a drone to obtain images of road surfaces containing multi-scale cracks, preprocessing the images and marking the cracks; all images with marked cracks constitute a data set; Construct a target detection network, select YOLOv8 including backbone network, neck network, head network and loss calculation module as the basic recognition network, use StarNet as the backbone network, the neck network includes C2f-Star module, and use GN-YoloHead module based on group normalization and shared convolution in the detection stage. In the target position loss function, use Focaler-MPD-WIoUv3 loss function to replace the original loss function; get the target detection network; The target detection network is trained using the data set consisting of the images of the annotated cracks to obtain a target recognition model; The unmanned aerial vehicle is used to obtain the crack images to be detected from the road surface, and the crack images are input into the target recognition model for target recognition.

2. A road crack analysis method based on an improved YOLO algorithm according to claim 1, characterized in that: The backbone network of the target detection network is a five-layer structure, including the first layer, the second layer, the third layer, the fourth layer and the fifth layer; The first layer includes the Stem layer of the custom convolutional batch normalization module; The second layer includes a downsampling layer and a C2f-Star module; The third layer includes a downsampling layer and a C2f-Star module; The fourth layer includes a downsampling layer and three C2f-Star modules; The fifth layer includes a downsampling layer and a C2f-Star module.

3. The road crack analysis method based on the improved YOLO algorithm according to claim 1, characterized in that: The StarNet includes a convolutional layer, a C2f-Star layer, a batch normalization layer, an activation function, a global average pooling layer, and a fully connected layer; The C2f-Star layer includes a depth-separable convolutional layer, a fully connected layer, a batch normalization layer, and a star operation; wherein the star operation is specifically: (1) in, is the index channel, Indicates the number of input channels, are the coefficients of each channel, is the weight of a single output channel, To represent a single input feature element, is transposed; The process of star calculation produces Different items: (2) No. Layer star operation output The recursion is expressed as: (3) in, is the network width, It is the comprehensive parameter matrix of each layer in StarNet.

4. The road crack analysis method based on the improved YOLO algorithm according to claim 1, characterized in that: The GN-YoloHead module includes a group normalization convolution layer, a prediction box convolution layer, a classification convolution layer and a scaling layer; the specific working mode of the GN-YoloHead module is as follows: the crack feature map of each scale passes through a 1×1 group normalization convolution layer, and then the feature map is further processed by two 3×3 group normalization convolution layers, and the weights are shared between the two layers. Then, the feature map of each scale is processed by the prediction box convolution layer and the classification convolution layer with shared weights; wherein the prediction box convolution layer is used to predict the bounding box of the target, and the classification convolution layer is used to classify the object in each prediction box; finally, the output of each scale adjusts the size of the prediction box through a scaling layer to adapt to targets of different scales.

5. The road crack analysis method based on the improved YOLO algorithm according to claim 1, characterized in that: The Focaler-MPD-WIoUv3 loss function includes Wise-IoU, MPD-IoU and Focaler-IoU loss functions: (5)。 6. A road crack analysis method based on improved YOLO algorithm according to claim 5, characterized in that: The bounding box regression loss of the Wise-IoU loss function Specifically: (6) (7) (8) in, Used to measure the abnormality of the prediction box. is the gradient gain, and To adjust the hyperparameters of the loss function; Indicates the horizontal and vertical coordinate values ​​of the prediction box, are the horizontal and vertical coordinate values ​​of the target frame, The values ​​correspond to the height and width of the two boxes respectively. is the height and width of the overlapping area of ​​the two frames, is the area of ​​the non-overlapping part of the two boxes, is the height and width of the prediction box, The height and width of the target box.

7. The road crack analysis method based on the improved YOLO algorithm according to claim 5, characterized in that: The bounding box loss function MPD-IoU It is used to optimize the matching degree between the predicted frame and the real frame, and improve the positioning accuracy of targets with significant size differences. MPD-IoU The loss function is specifically: (9) MDP-IoU Bounding Box Regression Loss Function Defined as: (10) in, The distance from the lower left corner of the real box to the lower left corner of the predicted box , is the distance from the upper right corner of the real box to the upper right corner of the predicted box, The values ​​correspond to the height and width of the two boxes respectively. is the area of ​​the non-overlapping part of the two boxes, The height and width of the overlapping area of ​​the two boxes.

8. The road crack analysis method based on the improved YOLO algorithm according to claim 5, characterized in that: The Focaler-IoU is used to increase the loss weight of difficult-to-classify samples, so that the model pays more attention to the difficult-to-identify areas; specifically: (11) in, After reconstruction focaler-IoU Loss function, IoU is the original IoU value, , For the lowest IoU Threshold, indicating the invalid area; For the highest IoU Threshold, indicating high-quality prediction area; by adjusting and The value of Focusing on different regression samples, the loss is defined as follows: (12)。

Citation Information

Patent Citations

  • Road surface crack small target identification method based on lightweight algorithm

    CN118781325A

  • Pavement crack identification method, device and system, and storage medium

    CN119131553A

Cited By

  • OCT image splicing method

    CN120355571A

  • OCT image stitching method

    CN120355571B

  • Method and device for collecting road image by unmanned aerial vehicle, electronic equipment and medium

    CN120416663A

  • Pavement crack detection method based on Star-YOLO11

    CN120635703A

  • Multi-scale road crack detection method and device in extreme weather and medium

    CN120747078A