Pavement crack detection method based on Star-YOLO11

By introducing StarBlock, CPCA attention mechanism and SMAFPN structure into pavement crack detection through the Star-YOLO11 method, the problems of low automation and large model parameters in the existing technology are solved, and efficient and accurate pavement crack detection is achieved, which is suitable for lightweight mobile devices.

CN120635703AActive Publication Date: 2025-09-12Jiangxi Jiaotong Maintenance Technology Group Co., Ltd.

Patent Information

Application Number
CN202510742678.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing pavement crack detection technology has a low degree of automation and its detection results rely on human subjective factors, leading to misjudgments and deviations. It is unable to carry out timely maintenance of highway diseases. In addition, the model parameters are large, making it difficult to detect multiple crack diseases in real time on lightweight mobile devices.

Method used

A pavement crack detection method based on Star-YOLO11 is adopted. By introducing StarBlock in the backbone network, the CPCA attention mechanism in the feature refinement layer, and the SMAFPN structure in the neck network, traditional and complex data augmentation strategies are combined to improve the model's feature extraction and generalization capabilities.

Benefits of technology

The efficiency and accuracy of pavement crack detection have been improved, the model has become more lightweight, and it can detect various crack diseases in real time on mobile devices, enhancing the detection effect of small target cracks and the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635703A_ABST
    Figure CN120635703A_ABST
Patent Text Reader

Abstract

The invention relates to a pavement crack detection method based on Star-YOLO11. The method comprises the following steps: acquiring a pavement image to be detected; a to-be-detected pavement image is input into the Star-YOLO11 model, a pavement crack detection result is output, the Star-YOLO11 model is obtained by improving the YOLO11 model and training the YOLO11 model through a training set, and the improved YOLO11 model comprises the steps of introducing StarBlock into a backbone network, introducing an attention mechanism into a feature refining layer, connecting feature information based on a pyramid structure in a neck network, and obtaining a pavement crack detection result. The training set comprises pavement crack images containing crack labels. According to the method, the Star-YOLO11 model based on YOLO11n is designed in combination with a pavement crack detection scene, algorithm support can be provided for lightweight mobile equipment, the pavement crack detection efficiency and precision are effectively improved, and effective support is provided for pavement detection equipment deployment and practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pavement crack detection, and in particular to a pavement crack detection method based on Star-YOLO11. Background Art

[0002] At present, highway disease survey still relies on manual methods, and the degree of automation is not high. The detection results will cause many misjudgments and deviations due to human subjective factors. The reliability and detection efficiency are low, and highway diseases cannot be maintained in a timely manner.

[0003] The YOLO (You Only Look Once) algorithm is a single-stage object detection network framework. Compared to traditional multi-stage detection algorithms, including Faster R-CNN and Mask R-CNN, all computations are performed in a single forward pass. This makes YOLO highly advantageous in real-time detection, maintaining high speed while ensuring comparable accuracy to multi-stage detection algorithms. Since its introduction, YOLO has been enthusiastically adopted by researchers. After numerous iterations, YOLO has continuously improved its detection accuracy and speed, making significant contributions to the field of object detection. Su Weiguo et al. applied the YOLOv3 deep learning algorithm to a road crack detection model and validated its effectiveness. Qiu et al. compared the performance of YOLOv2, YOLOv3, and YOLOv4-tiny in pavement crack detection. Xing et al. proposed an improved YOLOv5 algorithm, called EMG-YOLO, for road crack detection on edge computing devices. By optimizing the model structure and loss function, they improved the accuracy and efficiency of crack detection. Zhang et al. proposed an improved YOLOv5 algorithm named USSC-YOLO. By introducing the ShuffleNetV2 block, the coordinate attention (CA) mechanism, and the SwinTransformer block, it improved the accuracy and robustness of multi-scale crack detection in complex backgrounds. Chen Xiuxian et al. introduced the attention mechanism module CBAM, the decoupled decoupling head, and the improved αDIoU loss function to the Yolov5s algorithm, effectively improving the accuracy and efficiency of pavement crack detection. Bai Feng et al. improved AC-YOLO based on YOLOv8s for pavement crack detection in drone aerial photography. By introducing a dynamic large convolution kernel attention mechanism, a multi-scale feature fusion strategy, and an improved loss function, they improved the accuracy and efficiency of pavement crack detection. Su et al. made significant improvements to the YOLO framework and proposed a new algorithm named MOD-YOLO. By retaining the original feature information, expanding the receptive field to the global level, and incorporating the attention mechanism, this algorithm improves the accuracy and real-time performance of crack detection in civil infrastructure. Youwai et al. proposed YOLO9tr, which incorporates a partial attention mechanism. While using a smaller number of model parameters, it maintains high accuracy and fast inference in road damage detection. Meng et al. improved the lightweight nature of the YOLOv8 model in pavement crack detection scenarios through knowledge distillation and stochastic learning strategies. Mulyanto et al. proposed a lightweight, fast ground crack detection method based on YOLOv8, which reduced model parameters while improving accuracy, F1 score, and FPS.Pei et al. introduced a new feature extraction and fusion method, a dynamic snake convolution (DSC-C2f) module, and a coordinate attention mechanism to propose a YOLO-RDD road defect detection model based on YOLOv8. This model improves the accuracy and real-time performance of detecting road cracks and other defects. Although Transformer-related models hold great promise in the detection field, their large number of parameters remains a significant limitation.

[0004] In summary, the application of deep learning technology to pavement crack detection is highly significant. Current research primarily focuses on improving the accuracy of detecting a single pavement defect, while ignoring the inaccuracy of crack detection when multiple cracks coexist. Furthermore, the resulting models have increasingly large parameters, making them difficult to deploy on lightweight mobile devices for real-time detection. To enhance the real-time and accurate detection of multiple cracks and other defects, this paper provides a pavement crack detection method based on Star-YOLO11. Summary of the Invention

[0005] The purpose of this invention is to provide a pavement crack detection method based on Star-YOLO11, which effectively improves the efficiency and accuracy of pavement crack detection and provides effective support for the equipment deployment and practical application of pavement detection.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] The pavement crack detection method based on Star-YOLO11 includes:

[0008] Acquire a road surface image to be detected;

[0009] The road surface image to be detected is input into the Star-YOLO11 model, and a road surface crack detection result is output, wherein the Star-YOLO11 model is obtained by improving the YOLO11 model and training with a training set, wherein the improved YOLO11 model includes introducing StarBlock in the backbone network, introducing an attention mechanism in the feature refinement layer, and connecting feature information based on a pyramid structure in the neck network, and the training set includes road surface crack images containing crack labels.

[0010] Optionally, obtaining the training set includes: obtaining an original pavement crack image dataset, performing traditional data enhancement and complex data enhancement on the original pavement crack images, and obtaining the enhanced images as the training set, wherein the traditional data enhancement includes center cropping, vertical flipping, and horizontal flipping operations, and the complex data enhancement includes adding noise, brightness and saturation, grayscale, and environmental factors.

[0011] Optionally, introducing StarBlock into the backbone network includes: replacing the C3k2 module with StarBlock in the backbone network, wherein the third layer extraction block for extracting pavement crack features stacks three layers of StarBlock, and using one layer of StarBlock in each of the remaining layers of extraction blocks.

[0012] Optionally, introducing the attention mechanism in the feature refinement layer includes: adding a CPCA attention mechanism to the end of the original C2PSA structure in the feature refinement layer, wherein the CPCA attention mechanism adopts a multi-scale depth-separable convolution module to form spatial attention.

[0013] Optionally, the CPCA attention mechanism uses a multi-scale depth-separable convolution module to form spatial attention, including:

[0014] Aggregate spatial information from feature maps using average pooling and max pooling operations through channel attention, feed the spatial information into a shared MLP, and add them together to generate a channel attention map;

[0015] Obtain channel priors by element-wise multiplication of input features and channel attention maps;

[0016] The channel prior is input into the deep convolution module, and the multi-scale features are added after three branches to generate a spatial attention map.

[0017] Optionally, the pyramid structure is obtained by adding two connection blocks and multi-branch connection information between the bottom layer and the upper layer on the basis of PAFPN, including several surface auxiliary fusion modules and several high-level auxiliary fusion modules;

[0018] The surface auxiliary fusion module is used to connect the information transmitted from the backbone network to the neck network;

[0019] The advanced auxiliary fusion module is used to connect the fusion of different scale features in the neck network.

[0020] Optionally, the information transmitted from the surface auxiliary fusion module to the neck network through the backbone network includes:

[0021] Align and connect the output feature map of the upper layer of the backbone network and the output feature map of the auxiliary fusion module of the next layer with the input feature map of the current layer of the backbone network through downsampling and upsampling respectively;

[0022] The input feature map of the current layer of the backbone network, the output feature map of the previous layer in the aligned backbone network, and the feature map of the next layer of the surface auxiliary fusion module are concatenated and fused through the Concat module to output the output feature map of the current layer of the surface auxiliary fusion module.

[0023] Optionally, the advanced auxiliary fusion module connects the fusion of different scale features in the neck network, including:

[0024] Down-sample the output feature map of the previous layer's surface auxiliary fusion module and the output feature map of the previous layer's advanced auxiliary fusion module, up-sample the output feature map of the next layer's surface auxiliary fusion module, and align and connect them with the output feature map of the current layer's surface auxiliary fusion module;

[0025] The aligned connected output feature map of the previous layer of surface auxiliary fusion module, the output feature map of the previous layer of advanced auxiliary fusion module, the output feature map of the next layer of surface auxiliary fusion module and the output feature map of the current layer of surface auxiliary fusion module are spliced ​​and fused through the Concat module.

[0026] The beneficial effects of the present invention are as follows: To achieve the task of pavement crack detection, a new pavement crack detection method based on Star-YOLO11 is proposed to address the problems of multiple types of defects, difficulty in extracting features for different types of defects, and slow defect recognition speed in highway pavement defect detection. By replacing the C3k2 module with a multi-layer StarBlock in the feature extraction backbone network, the implicit dimension space of the feature is expanded, which can capture more effective information. Then, a channel prior convolutional attention mechanism is introduced in the feature refinement layer, combined with the PSA self-attention mechanism to form a CPC-C2PSA structure. At the end of the feature refinement layer, attention is strengthened on crack defects with a small pixel proportion in the image. Finally, based on the multi-auxiliary branch feature pyramid structure idea, a neck deep and shallow layer connection is constructed to ensure that feature fusion can simultaneously obtain deeper and shallower layer information, thereby improving the model's detection effect on small and medium-sized target cracks. To ensure good robustness of the experimental dataset and improve the generalization ability of the model, this paper designs a dataset enhancement method that combines traditional data augmentation with complex data augmentation. Through multiple comparative experiments and ablation experiments, the effectiveness of the dataset enhancement method and the Star-YOLO11 algorithm is verified. This invention can meet the actual pavement crack detection task and provide strong support for road inspection. Future work can focus on deploying the model to mobile devices such as mobile phones or drones for pavement crack detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 This is a structural diagram of Star-YOLO11 according to an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of a data set enhancement strategy according to an embodiment of the present invention;

[0030] Figure 3 A comparison diagram of different backbone network architectures according to an embodiment of the present invention;

[0031] Figure 4 This is a structural diagram of the CPC-Attention working principle of an embodiment of the present invention;

[0032] Figure 5 This is a structural diagram of the CPC-C2PSA fusion attention module according to an embodiment of the present invention;

[0033] Figure 6 Figures 1 and 2 are network structure diagrams of PAFPN, BiFPN, and SMAFPN according to embodiments of the present invention, wherein (a) is a PAFPN structure diagram, (b) is a biFPN structure diagram, and (c) is a SMAFPN overall structure diagram.

[0034] Figure 7 1 is a structural diagram of SAF and AAF according to an embodiment of the present invention;

[0035] Figure 8 These are five example images of categories in the dataset of an embodiment of the present invention, where (a) is a longitudinal crack, (b) is a transverse crack, (c) is a network of cracks, (d) is a pothole, and (e) is a repair.

[0036] Figure 9 Schematic diagram of some enhancement effects of an embodiment of the present invention, where (a) is the original image, (b) is vertical flip + shadow, (c) is vertical flip + snow, (d) is horizontal flip + brightness, and (e) is center cropping + Gaussian blur;

[0037] Figure 10 This is a curve diagram of the dynamic change of the training loss value according to an embodiment of the present invention;

[0038] Figure 11 This is a curve diagram of the dynamic change of the validation set loss value of an embodiment of the present invention;

[0039] Figure 12 This is a dynamic change curve diagram of average detection accuracy according to an embodiment of the present invention;

[0040] Figure 13 The Star-YOLO11 detection result indicators for each category in the embodiment of the present invention;

[0041] Figure 14 This is a graph showing comparative experimental results of different algorithms according to an embodiment of the present invention;

[0042] Figure 15A comparison of heatmaps of different algorithms according to an embodiment of the present invention, including (a) the original image, (b) the YOLO11 heatmap, and (c) the Star-YOLO11 heatmap.

[0043] Figure 16 : This figure compares the detection images of different algorithms in an embodiment of the present invention, where (a) is the original image, (b) is the YOLO11 detection image, and (c) is the Star-YOLO11 detection image. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] This embodiment provides a pavement crack detection method based on Star-YOLO11, including:

[0047] Acquire a road surface image to be detected;

[0048] The road image to be detected is input into the Star-YOLO11 model, and the road crack detection result is output. The Star-YOLO11 model is obtained by improving the YOLO11 model and training it with a training set. The improved YOLO11 model includes introducing StarBlock in the backbone network, introducing the attention mechanism in the feature refinement layer, and connecting feature information based on a pyramid structure in the neck network. The training set includes road crack images with crack labels. The Star-YOLO11 model structure is as follows: Figure 1 shown.

[0049] Specifically, to ensure the dataset effectively simulates real-world environments, a dataset augmentation strategy was designed to ensure sufficient training volume for the model, improving generalization capabilities. To improve the detection performance of road crack detection equipment, a Star-YOLO11 model was constructed. The large number of C3k2 modules used in the backbone network resulted in mediocre detection performance and a large number of parameters. This approach proposes replacing C3k2 modules with multiple StarBlocks, combined with downsampling convolutions to form a StarStage block. This element-by-element multiplication improves image feature extraction while still reducing parameter count. Given the significant pixel ratio between road cracks and the background, a third CPCA attention mechanism was added to the existing two-layer PSA attention in the feature refinement layer to enhance focus on small and less noticeable cracks. To address the problem of the neck failing to adaptively integrate high-level semantic information with low-level spatial information, a SMAFPN was proposed, which incorporates two layers of C3k2 modules and uses multiple branches to strengthen the connection between the neck and the StarBlock backbone.

[0050] Furthermore, obtaining a training set includes: obtaining an original pavement crack image dataset, performing traditional data enhancement and complex data enhancement on the original pavement crack images, and obtaining the enhanced images as a training set, wherein traditional data enhancement includes center cropping, vertical flipping, and horizontal flipping operations, and complex data enhancement includes adding noise, brightness and saturation, grayscale, and environmental factors.

[0051] Specifically, in order to enable the model to achieve a sufficient amount of training, the data set needs to have enough high-quality samples. Data enhancement technology reduces the time for manual annotation of images. By transforming existing data, such as rotating, scaling, and cropping, the number of training data sets can be expanded. Data enhancement methods are divided into traditional data enhancements that operate on image size and position, and complex data enhancements that simulate environmental changes on images. However, the data enhancement methods used in existing studies often only use traditional data enhancement or complex data enhancement, and cannot combine the advantages of the two enhancement methods well. Therefore, this embodiment proposes a new data set enhancement strategy, which adds Gaussian noise, ISO noise, brightness and saturation, grayscale, etc. that are more commonly used in complex data enhancement on the basis of center cropping, vertical flipping, and horizontal flipping operations commonly used in traditional data enhancement. In addition, considering that the model may be greatly affected by weather in actual applications, five types of environmental factors, namely fog, rain, snow, sunlight, and shadow, are added, totaling 12 enhancement methods. After permutation and combination, up to 27 enhancement strategies can be formed, such as Figure 2As shown, each image must be able to undergo one traditional data augmentation and one complex data augmentation, so that each image in the dataset has more complexities, which can further enrich the dataset and enhance the generalization ability of the model.

[0052] Furthermore, StarBlock is introduced into the backbone network, including replacing the C3k2 module with StarBlock in the backbone network, wherein the third layer extraction block for extracting pavement crack features stacks three layers of StarBlock, and one layer of StarBlock is used in each of the remaining layers of extraction blocks.

[0053] Specifically, the feature extraction backbone network is an important stage in YOLO11 that continuously transitions from low-level information to high-level information. The more information it can extract, the better the subsequent detection effect will be. However, due to the limited model size, the feature extraction block C3k2 used in the original YOLO11 cannot extract more information within the limited model size. Therefore, based on the idea of ​​element-by-element multiplication in StarBlock, the StarBlock stacked structure is used to replace the original C3k2 module to increase the accuracy of model detection. Depthwise separable convolution is used before and after multiplication instead of ordinary 1*1 convolution to reduce the number of computational parameters of the backbone network.

[0054] The structure of StarBlock is very simple and efficient. Its core idea is element-by-element multiplication, which is usually written as Where W represents the weight matrix and B represents the bias. For convenience, the weight matrix and bias are combined into one entity W = [WB] T , and X = [X 1] T Thus we get By fusing the features of two linear transformations by element-by-element multiplication, in the single output channel transformation and single element input scenario, Where d is the number of input channels, which is extended to scenarios with multiple output channels and multi-element inputs. This can be rewritten by the following steps:

[0055]

[0056] Where i, j are channel subscripts, and α is the coefficient of each sub-item:

[0057]

[0058] After rewriting, we can get the combination of (d+2)(d+1) / 2 different sub-terms as shown in Equation 4, except for α (d+1,:) x d+1Each subterm of x is nonlinearly related to x, indicating that they are separate implicit dimensions, which can be obtained by element-wise multiplication in d-dimensional space. Therefore, StarBlock can significantly amplify the feature dimension without incurring any additional computational overhead within a single layer. Compared with C3k2, it can obtain more dimensional information in the pavement crack detection task.

[0059] In this embodiment, the number of stacked layers of StarBlock is designed in the backbone network, and the improved backbone network is named Star-s50. Compared with the original backbone network of YOLO11 and the Starnet-s1 and s2 backbone networks provided in the StarBlock related literature, Star-s50 has a lighter architecture and higher feature map extraction capability. Only three layers of StarBlock are stacked in the third layer of extraction blocks that focus on extracting road crack features, and one layer of StarBlock is used in the extraction of relative image input and feature map output to retain a larger receptive field. In subsequent experiments, it can significantly improve the information observation capability of the backbone network and keep the model lightweight. With the same number of parameters, it can have better detection capabilities for complex scenes than the C3k2 module, such as Figure 3 shown.

[0060] Furthermore, introducing the attention mechanism in the feature refinement layer includes: adding a CPCA attention mechanism at the end of the original C2PSA structure in the feature refinement layer, wherein the CPCA attention mechanism adopts a multi-scale depth-separable convolution module to form spatial attention.

[0061] Furthermore, the CPCA attention mechanism uses multi-scale depth-separable convolutional modules to form spatial attention, including:

[0062] Through channel attention, average pooling and maximum pooling operations are used to aggregate spatial information from the feature map, input the spatial information into the shared MLP, and add them together to generate a channel attention map;

[0063] Obtain channel priors by element-wise multiplication of input features and channel attention maps;

[0064] The channel prior is input into the deep convolution module, and the multi-scale features are added after three branches to generate a spatial attention map.

[0065] Specifically, the CPCA (Channel Prior Convolutional Attention) attention mechanism is a new type of attention that combines channel attention and spatial attention. Compared with SE attention that only focuses on the channel dimension and completely ignores the spatial dimension information, and CBAM attention that only generates single spatial information in the spatial dimension, CPCA uses multi-scale depth-separable convolution modules to form spatial attention. It can dynamically allocate attention weights in the channel and spatial dimensions, effectively extract spatial relationships while retaining channel priors, and has a good ability to focus on information channels and important areas. Its structure is as follows Figure 4 shown.

[0066] By adopting the strategy of channel attention first and then spatial attention in series, as shown in formulas (6)-(7), the channel attention part adopts average pooling and maximum pooling operations to gather spatial information from the feature map, and then inputs it into the shared MLP (multi-layer perceptron), which is added to generate a channel attention map, see formula (8), and the channel prior is obtained by element-by-element multiplication of the input feature and the channel attention map. Subsequently, the channel prior is input into the deep convolution module, and the multi-scale features are added after three branches to generate a spatial attention map, see formula (9), where F represents the feature map, F c represents the channel prior refined feature map, Represents the final feature map, CA represents channel attention, SA represents spatial attention, σ represents Sigmoid activation function, Branch represents the three branches in spatial attention, and finally the channel mixing result and the channel prior are element-wise multiplied to obtain the final refined feature.

[0067]

[0068] The original YOLO11 designed a C2PSA structure, which nested two layers of the PSA partial self-attention mechanism proposed by YOLOv10. It has significant global modeling capabilities. However, since self-attention cannot consider the local structure of the data, it cannot effectively focus on target information with many local features such as road cracks. Therefore, based on the end of the original C2PSA structure, the CPCA attention mechanism is added to form the CPC-C2PSA model structure to strengthen the end with the weakest information, such as Figure 5 shown.

[0069] Furthermore, the pyramid structure is obtained by adding two connection blocks and multi-branch connection between the bottom layer and the upper layer on the basis of PAFPN, including several surface auxiliary fusion modules and several high-level auxiliary fusion modules;

[0070] The surface auxiliary fusion module is used to connect the information from the backbone network to the neck network;

[0071] An advanced auxiliary fusion module for fusing features of different scales within the concatenated neck network.

[0072] Furthermore, the surface auxiliary fusion module connects to the backbone network and transmits the following information to the neck network:

[0073] Align and connect the output feature map of the upper layer of the backbone network and the output feature map of the auxiliary fusion module of the next layer with the input feature map of the current layer of the backbone network through downsampling and upsampling respectively;

[0074] The input feature map of the current layer of the backbone network, the output feature map of the previous layer in the aligned backbone network, and the feature map of the next layer of the surface auxiliary fusion module are concatenated and fused through the Concat module to output the output feature map of the current layer of the surface auxiliary fusion module.

[0075] Furthermore, the advanced auxiliary fusion module connects the fusion of different scale features in the neck network including:

[0076] Down-sample the output feature map of the previous layer's surface auxiliary fusion module and the output feature map of the previous layer's advanced auxiliary fusion module, up-sample the output feature map of the next layer's surface auxiliary fusion module, and align and connect them with the output feature map of the current layer's surface auxiliary fusion module;

[0077] The aligned connected output feature map of the previous layer of surface auxiliary fusion module, the output feature map of the previous layer of advanced auxiliary fusion module, the output feature map of the next layer of surface auxiliary fusion module and the output feature map of the current layer of surface auxiliary fusion module are spliced ​​and fused through the Concat module.

[0078] Specifically, since the PAFPN used in the neck of the original YOLO11 tends to merge feature maps of homogeneous scales, it lacks comprehensive processing and fusion of multi-scale information of different resolution layers, and cannot effectively and adaptively integrate high-level semantic information and underlying spatial information at the same time. Therefore, a SMAFPN (Star Multi-Branch Auxiliary Feature Pyramid Network) based on MAFPN connected to the StarBlock in the backbone network is proposed. SMAFPN does not use the RepHELAN module designed by MAFPN but uses the C3k2 module in the original YOLO11 to suit the network structure of YOLO11. The connection form is divided into two parts, namely the surface assisted fusion (SAF, Superficial Assisted Fusion) in the first half, which is mainly responsible for connecting the information passed to the neck by the backbone network, and the advanced auxiliary fusion (AAF, Advanced Assisted Fusion) in the second half, which mainly connects the fusion of features of different scales in the neck. Figure 6 (a)-(c) show the structural comparison between the improved feature fusion structure and the unimproved PAFPN and the currently popular feature fusion module biFPN in the YOLO architecture. Based on PAFPN, biFPN adds an extraction block to the P3 layer, introducing the same-layer information in the backbone network into the extraction block of the previous stage of the detection head. SMAFPN adds two connection blocks and multi-branch connections between the bottom and high layers on the basis of PAFPN, thereby retaining the shallow information of StarBlock and fusing it with the deep information of the neck, thereby improving the detection effect of small target road cracks.

[0079] like Figure 7 and Figure 6 As shown in (c), the connection of SAF mainly consists of four parts, including the input feature map P of the current layer of the backbone network n Output feature map P′ of the auxiliary fusion module of the current surface layer n , the feature map P output by the previous layer of StarStage (each layer of the improved backbone network is called a layer of StarStage) n-1 Through downsampling and the current layer input feature map P of the backbone network n Perform alignment connection and output feature map P′ of the next layer surface auxiliary fusion module n+1 After upsampling and the input feature map P of the current layer of the backbone network n Align the connection, and then output the feature map P of the current layer of the backbone network n , the StarStage output feature map P of the previous layer in the backbone network n-1 , the next layer of surface auxiliary fusion module output feature map P′n+1 The three inputs are concatenated and fused through the Concat module, integrating deep information with features from the same level and high-resolution shallow layers within the backbone. This preserves rich localization details to enhance the network's spatial representation and improve the ability to detect smaller objects. The output is shown in Equation (10), where U(·) represents the upsampling operation, Down represents a 3×3 downsampling convolution with a batch normalization layer, δ represents the siLU activation function, and C represents a 1×1 convolution that controls the number of channels.

[0080] P′ n =Concat(δ(C(Down(P n-1 )),P n ,U(P′ n+1 )#(10)

[0081] AAF further enhances the interactive utilization of feature layer information by integrating multi-scale information at the neck. It mainly includes five parts, in addition to the output feature map P′ of the current surface auxiliary fusion module. n , the current layer advanced auxiliary fusion module output feature map P″ n And the output feature map P″ of the previous advanced auxiliary fusion module n-1 In addition, the next layer of surface auxiliary fusion module output feature map P′ is added n+1 And the output feature map P′ of the upper surface auxiliary fusion module n-1 . n+1 Output feature map P′ through upsampling and current layer surface auxiliary fusion module n Perform alignment connection, P′ n-1 and P″ n-1 By downsampling and feature map P′ n Align the connection and then connect it with the feature map P′ n The concatenation and fusion are performed together through the Concat module, and the output result is shown in formula (11). At this time, the Medium output can merge the information from four different layers, which significantly enhances the performance of medium-sized targets.

[0082]

[0083] The specific experimental process of the method of this embodiment is provided below:

[0084] 1. Dataset preparation:

[0085] The experiment uses the public dataset RDD2022 China dataset, which contains five types of road cracks captured by drones and non-motor vehicles equipped with smart cameras. These include longitudinal cracks (Longitudinal Crack), transverse cracks (Transverse Crack), alligator cracks (Alligator Crack), potholes (Pothole), and repairs (Repair). Figure 8 As shown in (a)-(e).

[0086] The data was divided into training set, validation set and test set in the ratio of 8:1:1, which contains 3502 training sets, 437 validation sets and 439 test sets. The training set was then enhanced using the established data set enhancement strategy. Each image was enhanced by center cropping, vertical flipping, horizontal flipping, and Gaussian noise, ISO noise, brightness saturation, rain, fog, snow, shadow, sunlight, and grayscale. The following diagram shows the effect of some enhancements. Figure 9 As shown in (a)-(e).

[0087] The expanded training set contains 17,494 images, which can improve the quantity and quality of labels for different categories and improve the generalization ability of the model. The types and quantities of labels in the enhanced dataset are shown in Table 1.

[0088] Table 1

[0089]

[0090] 2. Training environment:

[0091] To verify the effectiveness of the method proposed in this example, Pycharm was selected as the programming environment, using the Windows 10 operating system, Python 3.10 as the programming language, Pytorch 2.3.0 as the deep learning framework, and Cuda version 11.8. The hardware test environment for training the model was an 11th Gen Intel (R) Core (TM) i7-11700 @ 2.50 GHz processor with 16 GB of memory and an NVIDIA RTX 4060 Ti GPU (16 GB of video memory). The specific training hyperparameter settings are shown in Table 2.

[0092] Table 2

[0093]

[0094] 3. Evaluation indicators:

[0095] In order to comprehensively evaluate the model, the evaluation indicators used include precision (P), which reflects the model's ability to distinguish negative samples; recall (R), which reflects the model's ability to identify positive samples; mean average precision (mAP), which measures the accuracy of the model. mAP50 means that when the intersection of the true box and the predicted box is greater than 50, the model prediction is considered correct; F1 score (F1-Score), which measures the overall performance and stability of the model; parameters, which are used to measure the size and complexity of the model; and GFLOPs, which represents billions of floating-point operations per second and is used to measure the performance of the computing device when performing floating-point operations. The specific calculation formula is as follows (12):

[0096]

[0097] 4. Training process analysis:

[0098] The training process of deep learning models can be dynamically monitored through two core indicators: loss function and average detection precision (mAP). The lower the loss value and the higher the average detection precision, the better the model performance. Figure 10 and Figure 11 As shown in the training and validation loss curves, the 200 iterations are roughly divided into three phases: rapid decline, gradual convergence, and steady-state convergence. Initially, the loss is high because the model struggles to accurately fit the probability distribution of the predicted box. As training progresses, the probability distribution of the predicted box is gradually optimized. In the initial rounds, the loss shows an exponentially rapid decline, followed by a gradual convergence phase where the decline slows, ultimately reaching a flat steady-state convergence, indicating that the model has reached a good training state.

[0099] Figure 12 The average detection accuracy curve verified during the model training process is shown. There is not much fluctuation overall, and the shape is a smooth curve, which illustrates the effectiveness of the algorithm architecture and the reliability of the dataset. The average detection accuracy curve finally converges, proving that the model training has been completed.

[0100] 5. Test result analysis:

[0101] To verify the effectiveness of the dataset enhancement strategy designed in this embodiment, the unenhanced training set and the enhanced training set were used to train the original YOLO11n model and then tested and compared on the same test set. The comparison of the test results before and after the training set enhancement is shown in Table 3. It can be seen that all indicators are improved. The use of the dataset enhancement strategy proposed in this embodiment can significantly improve the generalization ability of the model and has high robustness.

[0102] Table 3

[0103]

[0104] To verify the effectiveness of the Star-s50 backbone network, an experimental comparison of backbone network replacement was conducted based on YOLO11n. Both were trained using the enhanced training set and tested using the same test set. The comparison of test results for different backbone network replacements is shown in Table 4.

[0105] Table 4

[0106]

[0107] It can be seen that the Star-s50 backbone network proposed in this embodiment achieves 88.3% on mAP50, which is better than the other backbone networks. Subsequently, based on YOLO11n using Star-s50, an experimental comparison of different attention mechanism fusions is conducted to demonstrate the superiority of the CPC-C2PSA attention structure fused in this embodiment. The comparison of different attention fusion test results is shown in Table 5.

[0108] Table 5

[0109]

[0110] Table 5 shows that the model incorporating the CPCA attention mechanism achieves a mAP50 of 90%, far exceeding other attention mechanisms. To verify the effectiveness of the neck structure improvement achieved by feature fusion, we conducted controlled variable experiments comparing SMAFPN with PAFPN and biFPN based on YOLO11n using star-s50 and CPC-C2PSA. Table 6 shows the comparison of the test results for different feature fusion neck structures.

[0111] Table 6

[0112]

[0113] As can be seen from Table 6, the feature fusion structure of SMAFPN can effectively improve the accuracy of the model to 90.3%. When using biFPN, the accuracy is only slightly better than PAFPN, and far behind the SMAFPN structure used in this embodiment. Finally, to verify the effectiveness of Star-YOLO11, an ablation experiment was conducted based on YOLO11n. "√" represents the use of this module. The ablation experiment test is shown in Table 7. The accuracy, recall rate and average detection accuracy of the method proposed in this embodiment in the detection results of each category are shown in Table 7. Figure 13 shown.

[0114] Table 7

[0115]

[0116] As can be seen from Table 7, all three modules can effectively improve the detection accuracy of the model. The Star-s50 backbone network significantly reduces the number of model parameters of YOLO11, and the mAP50 is greatly improved, which is significantly better than the original YOLO11 backbone network. The introduction of the CPCA attention mechanism can effectively improve the accuracy of the model, deepen the focus on small target cracks and inconspicuous cracks, and combine it with Star-s50 to make the mAP50 as high as 90%, and it can still improve by nearly one point on mAP50-95. The introduction of SMAFPN allows the model to pay attention to the interaction between shallower information and deeper information. Although it does not have much advantage over the other two models, the sum of the three can make P as high as 89.9%, mAP as high as 90.3%, and mAP50-95 can also break through to 62.7%. At the same time, in order to verify the superiority of the improved model compared with other mainstream algorithms, Star-YOLO11 is compared with other classic algorithms on the same dataset. The comparison test results of different algorithms are shown in the figure below. Figure 14 As shown in the figure, Star-YOLO11 performs well in all key metrics compared to similarly scaled network models. With the exception of recall, all other metrics are at the top of the list. Furthermore, in addition to improved detection performance, Star-YOLO11 also offers significant advantages over the original YOLO11 in terms of algorithm model size and real-time computing speed. A comparison of these model data is shown in Table 8.

[0117] Table 8

[0118]

[0119] After comparing the data during model testing, we found that the improved Star-YOLO11 model has reduced its model size and the overall number of parameters has been reduced by 18.8%. In theory, it can be well applied to the algorithm deployment of lightweight mobile devices. The model F1 score has also been slightly improved. Although its FPS is slightly lower than that of the original model, it still remains above 200. The detection speed is far ahead of the current models and can well meet actual detection needs.

[0120] 6. Visual Analysis:

[0121] After analyzing the experimental data, in order to intuitively reflect the improvement effect of the method proposed in this paper and observe the degree of optimization of the model by each part of the improvement, a visual analysis of the key links of the experiment was made. In order to verify the information capture and feature extraction capabilities of the model, as well as the optimization of the receptive field by the channel prior convolution attention, a heat map comparison of the model before and after improvement was performed, as shown in the figure below. Figure 15 (a)-(c). Figure 15The difference in color intensity in different areas in the image directly reflects the significant difference in the attention paid to road cracks before and after the model improvement. The Star-YOLO11 model shows a higher global coverage in the crack perception range and a significantly enhanced focus on crack targets. To verify the detection effect of the model in actual engineering applications, the detection performance of the model before and after improvement on the test set is compared. The detection results are as follows: Figure 16 (a)-(c), the improved algorithm significantly improves the problems of missed detection and detection accuracy of the original YOLO11 model.

[0122] By designing the Star-s50 feature extraction backbone formed by stacking layers of StarBlocks, it replaces the C3k2 extraction block used in the original YOLO11 backbone, improving the accuracy of pavement crack detection while increasing the lightweightness of the model. To address the problem that YOLO11's original PSA attention cannot efficiently focus on target information with many local features such as pavement cracks, a channel prior convolutional attention mechanism (CPC-Attention) is introduced, which multiplies the input features by the channel attention map element-by-element. Finally, to address the problem that the original YOLO11's neck cannot adaptively integrate high-level semantic information and low-level spatial information, a SMAFPN feature pyramid structure connected to Star-50 and using C3k2 as the extraction block is proposed based on MAFPN. A dataset augmentation strategy combining traditional and complex data augmentation is designed, and training and testing are performed on a dataset including five types of pavement cracks. Test results show that the Star-YOLO11 model achieves 89.9% accuracy, a 3.5% improvement over the original model; mAP reaches 90.3%, a 2.6% increase; and the F1 score reaches 85.8%, a 0.5% increase. The model size is reduced by 18.8% and the FPS reaches 225.73. The model is very fast and can be effectively used for pavement crack detection.

[0123] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. The pavement crack detection method based on Star-YOLO11 is characterized by: include: Acquire a road surface image to be detected; The road surface image to be detected is input into the Star-YOLO11 model, and a road surface crack detection result is output, wherein the Star-YOLO11 model is obtained by improving the YOLO11 model and training with a training set, wherein the improved YOLO11 model includes introducing StarBlock in the backbone network, introducing an attention mechanism in the feature refinement layer, and connecting feature information based on a pyramid structure in the neck network, and the training set includes road surface crack images containing crack labels.

2. The pavement crack detection method based on Star-YOLO11 according to claim 1 is characterized in that: Obtaining the training set includes: obtaining an original pavement crack image dataset, performing traditional data enhancement and complex data enhancement on the original pavement crack images, and obtaining the enhanced images as the training set, wherein the traditional data enhancement includes center cropping, vertical flipping, and horizontal flipping operations, and the complex data enhancement includes adding noise, brightness and saturation, grayscale, and environmental factors.

3. The pavement crack detection method based on Star-YOLO11 according to claim 1, characterized in that: Introducing StarBlock into the backbone network includes: replacing the C3k2 module with StarBlock in the backbone network, wherein the third layer extraction block for extracting pavement crack features stacks three layers of StarBlock, and each of the remaining layers of extraction blocks uses one layer of StarBlock.

4. The pavement crack detection method based on Star-YOLO11 according to claim 1, characterized in that: Introducing the attention mechanism in the feature refinement layer includes: adding a CPCA attention mechanism to the end of the original C2PSA structure in the feature refinement layer, wherein the CPCA attention mechanism adopts a multi-scale depth-separable convolution module to form spatial attention.

5. The pavement crack detection method based on Star-YOLO11 according to claim 4 is characterized in that: The CPCA attention mechanism uses a multi-scale depth-separable convolution module to form spatial attention, including: Aggregate spatial information from feature maps using average pooling and max pooling operations through channel attention, feed the spatial information into a shared MLP, and add them together to generate a channel attention map; Obtain channel priors by element-wise multiplication of input features and channel attention maps; The channel prior is input into the deep convolution module, and the multi-scale features are added after three branches to generate a spatial attention map.

6. The pavement crack detection method based on Star-YOLO11 according to claim 1, characterized in that: The pyramid structure is obtained by adding two connection blocks and multi-branch connection between the bottom layer and the upper layer on the basis of PAFPN, including several surface auxiliary fusion modules and several high-level auxiliary fusion modules; The surface auxiliary fusion module is used to connect the information transmitted from the backbone network to the neck network; The advanced auxiliary fusion module is used to connect the fusion of different scale features in the neck network.

7. The pavement crack detection method based on Star-YOLO11 according to claim 6, characterized in that: The information transmitted from the surface auxiliary fusion module to the backbone network and into the neck network includes: Align and connect the output feature map of the upper layer of the backbone network and the output feature map of the auxiliary fusion module of the next layer with the input feature map of the current layer of the backbone network through downsampling and upsampling respectively; The input feature map of the current layer of the backbone network, the output feature map of the previous layer in the aligned backbone network, and the feature map of the next layer of the surface auxiliary fusion module are concatenated and fused through the Concat module to output the output feature map of the current layer of the surface auxiliary fusion module.

8. The pavement crack detection method based on Star-YOLO11 according to claim 6, characterized in that: The advanced auxiliary fusion module connects the fusion of different scale features in the neck network and includes: Down-sample the output feature map of the previous layer's surface auxiliary fusion module and the output feature map of the previous layer's advanced auxiliary fusion module, up-sample the output feature map of the next layer's surface auxiliary fusion module, and align and connect them with the output feature map of the current layer's surface auxiliary fusion module; The aligned connected output feature map of the previous layer of surface auxiliary fusion module, the output feature map of the previous layer of advanced auxiliary fusion module, the output feature map of the next layer of surface auxiliary fusion module and the output feature map of the current layer of surface auxiliary fusion module are spliced ​​and fused through the Concat module.

Citation Information

Patent Citations

  • Pavement crack detection method based on deep learning

    CN116503336A

  • Road crack detection method, medium and product

    CN118941526A

  • Segmentation method for rail defect detection based on improved YOLOv10 and SETR

    CN119671959A

  • Light unmanned aerial vehicle aerial image small target detection method and system based on TFA-YOLO11

    CN119810409A

  • Lightweight mobile phone screen defect detection method based on improved YOLOv10

    CN119904449A

Cited By

  • Improved YOLO11-based method for detecting breakage of hair braid of oil pumping unit in low-light scene

    CN121330456A

  • Efficient concrete crack detection method, device and system

    CN121921306A