Pavement crack detection method based on Yolov8 model
Through the pavement crack detection method based on the Yolov8 model, using technologies such as PKIblock, GELAN module and SimSPPF, the problems of low efficiency and low accuracy of existing pavement crack detection are solved, and efficient and accurate crack automatic detection and real-time maintenance support are achieved.
Patent Information
- Application Number
- CN202510966048.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-21
AI Technical Summary
Existing pavement crack detection methods are inefficient and low-precision. Manual detection poses safety risks. The equipment costs of detection vehicles are high, and cross-lane detection is difficult. As a result, road maintenance data is not updated in a timely manner, making it difficult to meet the needs of modern road maintenance for efficient and accurate detection.
A pavement crack detection method based on the Yolov8 model is adopted. By building the YOLOv8-RCI model, the PKIblock multi-scale convolution kernel, GELAN module, EMA attention mechanism and SimSPPF are introduced to perform data enhancement and crack feature extraction, realizing automatic crack detection and instance segmentation.
It achieves efficient and accurate detection of pavement cracks, provides crack location and mask information, supports real-time maintenance decisions, reduces detection costs and manpower input, and reduces traffic disruptions.
Smart Images

Figure CN120823178A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of crack detection, and in particular relates to a pavement crack detection method based on the Yolov8 model. Background Art
[0002] In recent years, with the continuous urbanization in my country, the length of roads has increased significantly. Roads are an integral part of the transportation system, and a country's highway infrastructure reflects its economic and modernization status. However, with the continuous expansion of the road network, the task of pavement maintenance has become increasingly demanding. Road cracks are the most common pavement defect, and the causes of cracks are complex and diverse. These cracks not only threaten driving safety but can also cause property damage and even casualties. Therefore, research on pavement crack detection is particularly important. According to the road network plan released by the Ministry of Transport, my country's highways have entered the maintenance phase.
[0003] Currently, the primary method for detecting pavement cracks relies on a combination of manual inspection and multi-functional road inspection vehicles. However, traditional manual inspection methods are inefficient and lack high accuracy. Furthermore, manual inspections may require traffic closures, which not only disrupts normal traffic but also increases safety risks for operators, making them ineffective for large-scale, long-term road inspections. While road inspection vehicles can automatically collect road surface data at a certain speed, improving inspection efficiency, their limited inspection range, high equipment costs, and difficulty performing cross-lane inspections have limited their widespread use. Furthermore, insufficient inspection frequency results in delayed road maintenance data updates, making it difficult to accurately monitor the development of pavement defects. Typically, only limited sampling inspections can be conducted annually. However, sampling inspections cannot meet the efficient and accurate inspection requirements of modern road maintenance. Therefore, a flexible, low-cost, and efficient data collection solution is needed to improve the efficiency of road inspection and maintenance.
[0004] With the continuous development of drone technology, the field of road crack detection is increasingly using drones as a data collection method. Currently, methods for road crack detection have been proposed, including ultrasonic detection, ground-penetrating radar scanning, vibration signal analysis, and digital image processing. These methods involve mounting equipment on vehicles to capture road surface image data. These methods then combine image recognition algorithms to analyze crack characteristics, enabling automated crack detection. Furthermore, digital image processing technology eliminates the need for road closures, improving detection speed and accuracy while reducing manpower and traffic disruption. Combining this with intelligent algorithms such as deep learning can provide a more informed basis for decision-making in road maintenance. Summary of the Invention
[0005] In view of this, the object of the present invention is to provide a pavement crack detection method based on the Yolov8 model.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: The present invention provides a pavement crack detection method based on the Yolov8 model, comprising the following steps: Constructing a pavement crack segmentation dataset: We collected road crack images using drones and, combined with public datasets, performed data enhancement processing on the images using geometric and color transformations. We also used Gaussian filtering to denoise noisy images and used Labelme to annotate the cracks. Based on the YOLOv8-Seg model, we improved it by introducing the PKIblock multi-scale convolution kernel, the generalized efficient layer aggregation network GELAN module, the EMA attention mechanism, and replacing the spatial pyramid pooling layer SPPF with SimSPPF to construct the road crack identification model YOLOv8-RCI. The trained YOLOv8-RCI model is used to perform crack detection and instance segmentation on UAV images, and output crack location and mask information.
[0007] Preferably, the process of constructing the road crack identification model YOLOv8-RCI includes: The PKIblock multi-scale convolution kernel is embedded after the SPPF pyramid pooling layer to enhance the multi-scale crack feature extraction capability; The C2f module in the original YOLOv8-Seg model is replaced by the generalized efficient layer aggregation network GELAN module to reduce the number of model parameters and optimize the feature transfer path to reduce computational costs; Add an EMA attention mechanism module to the output of the four C2f modules in the Neck part to focus on the key crack areas; The spatial pyramid pooling layer SPPF is replaced with SimSPPF, which consists of SimConv and a 5×5 maximum pooling layer to improve the model detection speed.
[0008] Preferably, the PKIblock multi-scale convolution kernel enhances the features of the central area by combining global average pooling and 1×1 strip convolution, thereby effectively capturing long-range contextual information, where: The PKIblock multi-scale convolution kernel consists of a PKI module and a CAA module. The PKI module extracts local information through a small convolution kernel. After the extraction is completed, a set of parallel deep convolution layers are used to capture contextual information across multiple scales. After the PKI module extracts the features, it uses 1×1 convolution to fuse the local features and context information to represent the relationship between different channels; CAA uses average pooling operations to extract global context information and capture the dependencies between distant pixels. It then fuses the global information with local features through 1×1 convolution to establish relationships between modeling channels.
[0009] Preferably, the GELAN module achieves model compression and performance improvement by integrating the feature map segmentation and reorganization mechanism of the CSPNet architecture with the gradient path optimization strategy of the ELAN module; The feature map segmentation and reorganization mechanism of the CSPNet architecture includes: dividing the input feature map into two branches. The first branch is directly transferred across stages, and the second branch extracts features through local dense blocks. Finally, the features of the two branches are recombined to eliminate gradient redundancy and reduce memory bandwidth usage; The gradient path optimization strategy of the ELAN module includes: building a dual-branch structure in the network layer. The first branch adjusts the channel dimension through a 1×1 convolution kernel, and the second branch optimizes the gradient propagation path through a hierarchical convolution module consisting of a 1×1 convolution and four stacked 3×3 convolutions to shorten the longest gradient path and suppress the gradient vanishing problem of deep networks.
[0010] Preferably, the EMA attention mechanism is embedded in the four C2f level outputs in the Neck module, and feature enhancement is achieved through the collaboration of multi-channel grouping and two-way semantic extraction, specifically including: The channel dimension partitioning module divides the input feature map channels into multiple subgroups and performs batch dimension reshaping to establish cross-group associations; The spatial semantic extraction branch uses global average pooling to compress spatial information and generate spatial weights; The channel semantic extraction branch performs dimensionality reduction through a 1×1 convolution kernel; The dual-branch features are fused through a joint activation mechanism and then element-wise added to the original input to generate enhanced features. The incentive mechanism dynamically allocates channel importance weights through the normalization layer and the Softmax function, and the adjustment mechanism implements feature calibration based on 2D global average pooling.
[0011] Preferably, SimSPPF replaces the two traditional convolutional layers in the original SPPF module with SimConv modules; The SimConv module consists of a convolutional layer Conv, a batch normalization layer BN, and a SiLU activation function; The three-level cascaded maximum pooling operation is retained to ensure the ability to capture multi-scale features.
[0012] The beneficial effects of the present invention are: 1. This paper introduces the PKIblock multi-scale convolution kernel, enhancing the model's ability to extract multi-scale features. The GELAN module improves model accuracy and addresses the issue of increased parameter count. The EMA attention mechanism further enhances accuracy and better focuses on crack features without increasing model parameters. Finally, the SimSPPF algorithm improves model detection speed and performance while maintaining the same number of parameters.
[0013] 2. This paper uses the YOLOv8-RCI model to realize the full process automation from crack detection to quantitative analysis. During the detection process, crack information will be displayed in real time, including crack width and crack area information, to provide support for subsequent pavement crack maintenance.
[0014] 3. In the present invention, the use of geometric changes will not change the pixel values of the image, and color transformation can generate new samples by adjusting the pixel values of the image.
[0015] Other advantages, objectives and features of the present invention will be described in the following description and will be apparent to those skilled in the art to some extent, or those skilled in the art can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to make the purpose, technical solutions and beneficial effects of the invention clearer, the present invention is described with the following drawings: Figure 1 This is a flow chart of the steps of a pavement crack detection method based on the Yolov8 model of the present invention; Figure 2 Schematic diagram of the traditional YOLOv8-Seg network structure of the present invention; Figure 3 This is a schematic diagram of the PKIBlock module structure of the present invention; Figure 4 This is a schematic diagram of the Backbone structure after the KIBlock module is integrated into the present invention; Figure 5 This is a schematic structural diagram of the GELAN of the present invention; Figure 6 This is a schematic diagram of the Backbone structure after the GELAN module is integrated into the present invention; Figure 7 Schematic diagram of the EMA attention mechanism structure of the present invention; Figure 8 Schematic diagram of three fusion methods of the EMA attention module of the present invention; Figure 9 This is a schematic diagram of the improved YOLOv8-RCI structure of the present invention. DETAILED DESCRIPTION
[0017] like Figures 1 to 9 As shown, the present invention provides a pavement crack detection method based on the Yolov8 model, comprising the following steps: Constructing a pavement crack segmentation dataset: We collected road crack images using drones and, combined with public datasets, performed data enhancement processing on the images using geometric and color transformations. We also used Gaussian filtering to denoise noisy images and used Labelme to annotate the cracks. Based on the YOLOv8-Seg model, we improved it by introducing the PKIblock multi-scale convolution kernel, the generalized efficient layer aggregation network GELAN module, the EMA attention mechanism, and replacing the spatial pyramid pooling layer SPPF with SimSPPF to construct the road crack identification model YOLOv8-RCI. The trained YOLOv8-RCI model is used to perform crack detection and instance segmentation on UAV images, and output crack location and mask information.
[0018] The working principle and beneficial effects of the above technical solution are as follows: To address the shortcomings of road crack segmentation datasets, this invention combines public datasets and utilizes drone image acquisition to construct a road crack segmentation dataset. In recent years, image segmentation technology has been increasingly used in crack detection. This method can extract crack morphology with high accuracy. To address the low efficiency and accuracy of crack segmentation algorithms, this invention builds the road crack identification model YOLOv8-RCI based on the instance segmentation algorithm YOLOv8-Seg through four improvements. The introduction of the PKIblock multi-scale convolutional kernel enhances the model's ability to extract multi-scale features. To further improve model accuracy and address the issue of increased parameter count, the Generalized Efficient Layer Aggregation Network (GELAN) module is introduced to replace the C2f module in the original model. The EMA attention mechanism is also incorporated, further improving accuracy and better focusing on crack features without increasing model parameters. Finally, the SPPF is replaced with SimSPPF, which increases model detection speed and slightly improves model performance while maintaining the same number of parameters. These four improvements enhance the model's detection and segmentation performance, and the model meets the requirements of real-time detection.
[0019] In a specific embodiment, the present invention selects to improve the road crack identification model YOLOv8-RCI of YOLOv8-Seg. Compared with other YOLO algorithms, YOLOv8 (YouOnlyLookOnceVersion8) is a more mature and stable version, achieving a better balance in terms of parameter quantity and detection accuracy. YOLOv8's segmentation model provides five models of different sizes to choose from, namely YOLOv8n-seg, YOLOv8s-seg, YOLOv8m-seg, YOLOv8l-seg, and YOLO8x-seg. In order to meet the needs of real-time detection and ensure that the model is sufficiently lightweight, the present invention selects the YOLOv8n-seg model for improvement; YOLOv8 is a real-time object detection algorithm that improves upon YOLOv5. Its most significant innovation is the introduction of the AnchorFree mechanism, which replaces traditional anchor-based methods. Traditional anchor-based methods rely on preset anchor boxes of varying sizes and shapes to assist with object regression and classification. This approach is inefficient due to the numerous hyperparameters required. YOLOv8, by introducing the AnchorFree mechanism, eliminates the need for anchor box configuration, reducing time and computing power requirements while also avoiding missed or duplicate detections caused by improper anchor box configuration.
[0020] YOLOv8-Seg is an advanced single-stage instance segmentation model that extends and optimizes the YOLOv8 object detection framework. YOLOv8-Seg utilizes the Dynamic Focus (DF) loss function, an improvement that effectively addresses class imbalance and enables the model to perform better on multi-category instance segmentation tasks. The model also incorporates a module that enhances feature extraction capabilities and improves the model's accuracy in identifying object boundaries through cross-stage feature fusion.
[0021] In instance segmentation, YOLOv8-Seg adopts a novel mask generation strategy, which draws on and improves on the design principles of YOLOACT. The model generates accurate instance segmentation results by linearly combining the learned prototype mask with the predicted mask coefficients. This approach not only improves segmentation accuracy but also maintains high computational efficiency, enabling YOLOv8-Seg to excel in real-time detection.
[0022] The network structure of YOLOv8-Seg is: input layer, backbone network (Backbone), feature fusion network (Neck), head network (Head) and loss function layer. The input layer is responsible for receiving and processing the original image data; the backbone network uses convolutional neural networks to extract as much feature information as possible from the data; the feature fusion network uses upsampling and downsampling to fuse different features; the head network is responsible for generating the final detection box and segmentation mask; the loss function layer will analyze all losses in training, such as classification loss, positioning loss and segmentation loss, so as to better guide the subsequent training process of the model. The network structure of YOLOv8-Seg is as follows: Figure 2 shown.
[0023] In a specific embodiment, the process of building the road crack identification model YOLOv8-RCI includes: The PKIblock multi-scale convolution kernel is embedded after the SPPF pyramid pooling layer to enhance the multi-scale crack feature extraction capability; The C2f module in the original YOLOv8-Seg model is replaced by the generalized efficient layer aggregation network GELAN module to reduce the number of model parameters and optimize the feature transfer path to reduce computational costs; Add an EMA attention mechanism module to the output of the four C2f modules in the Neck part to focus on the key crack areas; The spatial pyramid pooling layer SPPF is replaced with SimSPPF, which consists of SimConv and a 5×5 maximum pooling layer to improve the model detection speed.
[0024] In a specific embodiment, the PKIblock multi-scale convolution kernel enhances the features of the central region by combining global average pooling and 1×1 strip convolution, thereby effectively capturing long-range contextual information, where: The PKIblock multi-scale convolution kernel consists of a PKI module and a CAA module. The PKI module extracts local information through a small convolution kernel. After the extraction is completed, it uses a set of parallel deep convolution layers to capture contextual information across multiple scales. After the PKI module extracts the features, it uses 1×1 convolution to fuse the local features and context information to represent the relationship between different channels; CAA uses average pooling operations to extract global context information and capture the dependencies between distant pixels. It then fuses the global information with local features through 1×1 convolution to establish relationships between modeling channels.
[0025] The working principle and beneficial effects of the above technical solution are as follows: when performing multi-scale target detection tasks under the interference of complex backgrounds, During the feature extraction process, the module is difficult to effectively retain key feature information, and some detailed features may be lost, which will directly affect the accuracy of the detection results. It also has a major disadvantage for the detection task of small-scale targets. The feature information of small targets will gradually decay in continuous convolution operations as the network depth increases, which will cause the feature identification to become blurred, and ultimately cause a decrease in detection accuracy. This situation is mainly due to the loss of detailed information in the deep network, and the continuous reduction in the resolution of the feature map will also make it difficult to effectively capture and retain the discriminant features of small targets. This is The module has limitations in handling target detection tasks in complex scenarios. This paper optimizes the model's ability to extract multi-scale features by adding the PKIBlock module. PKINet (PolyKernelInceptionNetwork) combines global average pooling and Strip convolution is used to enhance the features of the central area, which can effectively capture long-distance contextual information. This method is different from the traditional method that relies on large convolution kernels or dilated convolution. PKINet uses multiple deep convolution kernels of different sizes to extract multi-scale texture features from different receptive fields without dilation. This enables PKINet to flexibly obtain multi-scale information without significantly increasing the computational complexity, and improve the model's ability to extract diverse features. Each PKINet consists of a PKI module and a CAA module. The structures of PKINet, PKIBlock and CAA are as follows: Figure 3 shown.
[0026] The PKI module extracts local information using a small convolution kernel. After extraction, it uses a set of parallel deep convolutional layers to capture contextual information across multiple scales. This design enables the PKI module to effectively capture rich contextual features from different scales while maintaining focus on local details. The mathematical expression of the PKI module is shown as follows: Where: Local features extracted by convolution; ——No. Second-rate Contextual features extracted by deep convolution; ——Stage input.
[0027] Because the PKI module does not use dilated convolution, it will not generate overly sparse feature representations. The convolution of fused local features with contextual information, characterizing the relationship between different channels and improving the richness and accuracy of feature expression. The mathematical expression is shown in Eq.
[0028] Where: ——Output features.
[0029] CAA (Context Anchor Attention) attention mechanism is a context modeling method that can capture the contextual dependencies between distant pixels in an image while enhancing the feature representation capability of the central area. CAA uses average pooling to extract global context information and capture the dependencies between distant pixels; then Convolution fuses global information with local features, establishing effective relationships between modeling channels. The CAA attention mechanism does not use dilated convolutions, nor does it generate sparse features. This reduces computational complexity while improving the ability to fuse multi-scale features. Through this design, CAA can enhance the model's focus on the target area, reduce background interference, and significantly improve the model's performance in tasks such as object detection and semantic segmentation in complex scenarios. Its mathematical expression is shown in the formula: Where: for ,have Then use two depth-wise separable convolutions as the formula for the standard large kernel depth-wise convolution as shown in the formula: Where: ——Width is Depthwise separable convolution; ——Width is Depthwise separable convolution.
[0030] The feature maps generated by these convolution operations contain rich contextual information, which is then fused with the original features to capture multi-scale contextual relationships. Through this mechanism, the CAA module significantly improves the model's detection and segmentation capabilities for small objects.
[0031] The PKIblock module is integrated into the backbone part of the model, specifically after the SPPF pyramid pooling layer. The backbone structure is as follows Figure 4 shown.
[0032] In a specific embodiment, the GELAN module achieves model compression and performance improvement by integrating the feature map segmentation and reorganization mechanism of the CSPNet architecture with the gradient path optimization strategy of the ELAN module; The feature map segmentation and reorganization mechanism of the CSPNet architecture includes: dividing the input feature map into two branches. The first branch is directly transferred across stages, and the second branch extracts features through local dense blocks. Finally, the features of the two branches are recombined to eliminate gradient redundancy and reduce memory bandwidth usage; The gradient path optimization strategy of the ELAN module includes: building a dual-branch structure in the network layer. The first branch adjusts the channel dimension through a 1×1 convolution kernel, and the second branch optimizes the gradient propagation path through a hierarchical convolution module consisting of a 1×1 convolution and four stacked 3×3 convolutions to shorten the longest gradient path and suppress the gradient vanishing problem of deep networks.
[0033] The working principle and beneficial effects of the above technical solution are as follows: after integrating the PKIBlock module, the accuracy of the model is greatly improved, but the number of parameters and computational complexity are also increased. Therefore, the present invention introduces the Generalized Efficient Layer Aggregation Network (GELAN) module to replace the original model. module, thereby significantly reducing the number of model parameters and lowering the computational cost.
[0034] GELAN (Generalized Efficient Layer Aggregation Network) is an innovative network architecture that combines the splitting and reassembling concepts of CSPNet (Cross Stage Partial Network) and introduces the layered convolution processing method of ELAN (Efficient Layer Aggregation Network) in each part. This design can comprehensively improve the overall performance of the model because GELAN significantly improves the inference speed and accuracy while maintaining lightweight. The network structure diagram of GELAN is shown below. Figure 5 shown.
[0035] CSPNet is an innovative convolutional neural network backbone architecture that not only reduces computational resources and memory usage but also improves the network's learning capabilities. CSPNet primarily addresses the problem of duplicate gradient information in traditional network optimization by integrating feature maps at different network stages, thereby enhancing gradient diversity. CSPNet's architecture includes local dense blocks and local transition layers. These designs not only extend the gradient path but also balance the computational burden across layers, effectively reducing memory bandwidth requirements and enabling the model to operate efficiently even in resource-constrained environments.
[0036] CSPNet also replaces traditional dense connections with spaced connections. This adjustment improves parameter utilization efficiency and the network's learning ability. This enables CSPNet to achieve higher model accuracy while maintaining low computational complexity. It is particularly suitable for tasks that require efficient feature extraction, such as object detection and image classification.
[0037] like Figure 5 As shown in Figure 3, the cross-stage hierarchical structure of CSPNet enhances network performance while maintaining high efficiency and lightweight through innovative design.
[0038] ELAN is a highly efficient neural network architecture designed to improve the efficiency of gradient propagation during the training of deep models. ELAN accurately analyzes the shortest and longest gradient paths within the network and employs an optimized layer-by-layer aggregation strategy. This significantly shortens gradient propagation paths, avoiding the extended gradient paths associated with excessive transition layers in traditional networks. This strategy not only mitigates the potential risk of vanishing or exploding gradients but also enhances network stability as network depth increases.
[0039] ELAN's flexible design achieves a good balance between accuracy and computational cost. By adopting a stacking strategy within the computational block, the stability of the gradient path is ensured, so that even if the network depth increases, the model can still efficiently perform gradient propagation and parameter optimization. In addition, ELAN's design also gives the network greater adaptability, enabling rapid convergence in various tasks while maintaining efficient computational performance. Its network structure is as follows: Figure 5 As shown in the figure, the network structure of the ELAN module consists of two branches: the first branch is connected through a single The convolution kernel adjusts the number of channels; the second branch first passes through a The convolution module changes the number of channels and then passes through four The convolution module performs feature extraction.
[0040] The present invention integrates the GELAN module into the Backbone part of the model, replacing Module. The specific structure is as follows Figure 6 shown.
[0041] In a specific embodiment, the EMA attention mechanism is embedded in the four C2f level outputs of the Neck module, and feature enhancement is achieved through multi-channel grouping and two-way semantic extraction. Specifically, it includes: The channel dimension partitioning module divides the input feature map channels into multiple subgroups and performs batch dimension reshaping to establish cross-group associations; The spatial semantic extraction branch uses global average pooling to compress spatial information and generate spatial weights; The channel semantic extraction branch performs dimensionality reduction through a 1×1 convolution kernel; The dual-branch features are fused through a joint activation mechanism and then element-wise added to the original input to generate enhanced features. The incentive mechanism dynamically allocates channel importance weights through the normalization layer and the Softmax function, and the adjustment mechanism implements feature calibration based on 2D global average pooling.
[0042] The working principle and beneficial effects of the above technical solution are as follows: In the field of computer vision, especially in object detection tasks, the attention mechanism is considered a key technology for improving model performance. The attention mechanism enhances the model's recognition ability in specific tasks by focusing on key areas in the image. The attention mechanism is derived from human visual attention and draws on human attention logic. Its core logic is to assign different importance weights to different parts of the input data, enabling the model to more effectively focus on information relevant to the current task. For example, in person recognition, the attention mechanism allows the model to assign greater weight to facial features, which are key to person recognition. By dynamically adjusting weights, the attention mechanism enhances the model's ability to capture long-range dependencies and key information.
[0043] Inspired by the human brain's attention mechanism, the attention mechanism in neural networks enables models to automatically and selectively focus on specific parts of input data, helping them complete various complex tasks more efficiently. This mechanism learns a set of weight parameters to measure the relative importance of each input and then applies weighted processing to different parts of the input data based on these weights, achieving selective attention to key information. The introduction of the attention mechanism enhances the ability of neural networks to process complex data, improving model performance and robustness.
[0044] In the field of deep learning, attention mechanisms are mainly divided into two categories: channel attention and spatial attention. (Squeeze-and-Excitation) network, which is the channel attention method, The network extracts key channel information by explicitly modeling the interaction between channels, enhancing the model's attention to important features. (ConvolutionalBlockAttentionModule) and SpatialGroup-wiseEnhance focuses more on leveraging the semantic dependencies between the spatial and channel dimensions in feature maps. Spatial attention methods divide the channel dimension into multiple sub-feature groups, optimizing the spatial distribution of these sub-features and improving the model's expressive power.
[0045] While dimensionality reduction and grouping strategies can improve model performance, they increase computational overhead, potentially leading to higher latency and resource consumption. To address this, this paper incorporates an efficient multi-scale attention mechanism, EMA (Efficient Multi-Scale Attention), into the model. The EMA attention mechanism enhances the model's ability to integrate multi-scale features while maintaining efficient computation.
[0046] The two core mechanisms of the EMA attention mechanism are the incentive mechanism and the adjustment mechanism. The incentive mechanism prioritizes input features, differentially marking data based on its importance to the task at hand, enabling efficient screening of key information. This mechanism helps the model more accurately capture task-relevant features and reduces interference from redundant information. The adjustment mechanism further optimizes model performance by adjusting the weights of each component and normalizing them using the Softmax function to ensure that the sum of the weights for each row of data is 1.
[0047] Compared to traditional attention mechanisms, the EMA attention mechanism not only avoids the information loss caused by dimensionality reduction but also achieves efficient fusion of multi-scale features through the synergy of incentive and adjustment mechanisms. This design reduces computational complexity and improves the model's performance in complex tasks. It is particularly effective in scenarios requiring multi-scale feature extraction, such as object detection and image segmentation. Figure 7 The main structure of the EMA attention mechanism is shown.
[0048] When processing the input feature map, the EMA attention mechanism divides the channels into multiple groups, each of which contains a certain number of channels. The grouped channels are rearranged and converted into batch dimensions. The convolution kernel performs dimensionality reduction on the channel. Branch, EMA adopts Global average pooling is used to encode global spatial information; If a branch is added, its dimensions are directly adjusted to match those of the other branches. Finally, these processed features are fused in the joint activation mechanism of the channel features. The reduced feature map is added to the original feature map to generate a new feature map. The 2D global average pooling operation formula in this process is shown as follows: Where: ——the number of input channels; 、 ——The height and width of the input feature; ——Global average pooling; ——Input features at the c-th channel.
[0049] The EMA attention mechanism groups channels, enabling connections between channels within the same group. It also reshapes each channel into a batch dimension, enabling connections between channels in different groups. This cross-channel connection helps the model more effectively build deep visual representations.
[0050] To explore the effectiveness of the EMA attention mechanism in the YOLOv8-Seg model, this paper incorporates the EMA attention mechanism into different locations within the YOLOv8-Seg model. Experiments were conducted under the same conditions to verify the effectiveness of different approaches in improving model performance. The model resulting from integrating the EMA attention mechanism into the SimSPPF of the YOLOv8-Seg Backbone component was named YOLOv8-EMA1. When the EMA attention mechanism was incorporated into the Neck component of the YOLOv8-Seg model, the resulting models were named YOLOv8-EMA2 and YOLOv8-EMA3, respectively, depending on the location of the integration.
[0051] Figure 8 The three network structures incorporating the EMA attention mechanism are shown, clearly illustrating how the EMA module is embedded at different locations and how it integrates with the original network structure. This method allows us to compare the impact of the EMA attention mechanism at different locations on model performance.
[0052] The table below shows the experimental results of models embedding the attention mechanism at three different locations. The experimental results show that the introduction of the attention mechanism optimizes model performance to varying degrees and improves the overall accuracy of the model.
[0053] Comparison of EMA attention mechanism results at different positions The model that integrates the EMA module in Backbone has improved mAP@0.5(B), mAP@0.5(M) and Mask(R) indicators compared to the original network model. Compared with all other models, Mask(R) is the highest. Compared with the original model, the model From 0.802 to Increased from 0.657 to 0.669; an increase of 1.8% and 1.2% respectively.
[0054] The overall improvement of adding the EMA attention mechanism to the Neck part is better than adding the module to the Backbone part. For the two models that add the EMA attention mechanism to the Neck part, the better effect is the method of adding the EMA module to the four C2f modules, that is, YOLOv8-EMA2; compared with the original algorithm, AP@0.5 (B) increased from 0.802 to 0.823, a 2.1% improvement in accuracy; mAP@0.5 (M) increased from 0.657 to 0.670, a 1.3% improvement in accuracy. The YOLOv8-EMA2 model is superior to or equal to the YOLOv8-EMA3 model in terms of accuracy, recall, and average precision.
[0055] After comparing the results, YOLOv8-EMA2 has the best effect, and the improvement strategy of YOLOv8-EMA2 will be adopted in subsequent model improvements.
[0056] In a specific embodiment, SimSPPF replaces two traditional convolutional layers in the original SPPF module with SimConv modules; The SimConv module consists of a convolutional layer Conv, a batch normalization layer BN, and a SiLU activation function; The three-level cascaded maximum pooling operation is retained to ensure the ability to capture multi-scale features.
[0057] The working principle and beneficial effects of the above technical solution are as follows: Spatial Pyramid Pooling (SPP) is a technology widely used in deep learning models. Its core goal is to solve the problem of non-fixed input image size. SPP performs multi-scale pooling operations on the output feature maps of the convolutional layer, converting feature maps of different sizes into fixed dimensions, allowing the model to process input images of any size. Traditional convolutional neural networks (CNNs) require input images to have a fixed size so that they can be processed in the fully connected layer. However, in actual applications, the sizes of images often vary. In order to meet the input requirements of the model, it is usually necessary to force the image to be scaled, but this method may cause geometric deformation of the image and loss of feature information, affecting the performance of the model.
[0058] SPP can significantly improve model performance in tasks such as image classification and object detection because it preserves multi-scale spatial information. The SPP layer performs pooling operations on feature maps at different scales, generating multiple fixed-size feature vectors that are then concatenated into a unified feature representation. This approach not only avoids geometric distortion caused by image scaling but also enhances the model's ability to capture features at different scales within the image. For example, features of small objects can be extracted using smaller pooling regions, while features of larger objects can be extracted using larger pooling regions, thus achieving balanced processing of objects at multiple scales.
[0059] The Spatial Pyramid Pooling Network (SPPNet), proposed in existing technology, is a successful application of this technology. SPPNet introduces an SPP layer before the fully connected layer, effectively addressing the issue of inconsistent input image sizes. This approach enables the network to generate fixed-length feature representations without changing the original scale of the input image, improving the model's robustness and accuracy. The introduction of SPP not only simplifies the image preprocessing process but also provides a more flexible and efficient solution for processing multi-scale images.
[0060] In SPPNet, the input image first passes through a convolutional layer to extract features, which are then processed with normalization and activation functions to produce a set of feature maps. The SPP layer then performs max pooling on these feature maps at multiple scales to extract spatial information at different scales. These pooled results are flattened and concatenated to form a fixed-length feature vector. This approach allows SPPNet to preserve the multi-scale information of the input image while ensuring that the length of the output features is independent of the input image size. This design makes SPPNet a flexible and efficient architecture, capable of accommodating input images of any size and generating stable feature representations without losing critical information. This feature has led to its widespread application in computer vision tasks, including image classification and object detection, providing an effective solution for deep learning models to handle diverse inputs. The emergence of SPPNet has significantly improved model performance and provided new insights into the design of subsequent network structures. For example, in modern deep learning architectures, SPP layers are often integrated into more complex networks to further enhance the model's ability to handle multi-scale features.
[0061] In YOLOv8, the pyramid layer structure draws on the design ideas of Feature Pyramid Network (FPN) and Path Aggregation Network (PANet). This structure is The structure can realize the fusion of multi-scale features and improve the performance of target detection. The SpatialPyramidPooling-Fast architecture integrates contextual information, combining detailed information from high-resolution feature maps with semantic information from low-resolution feature maps. This allows the model to adapt to the detection needs of both small and large objects. High-level features are upsampled and fused with low-level features, enhancing the ability to capture detailed information. Low-level features, on the other hand, are downsampled to convey semantic information, forming a bidirectional information flow, both top-down and bottom-up.
[0062] In order to further improve the inference efficiency of the model, this paper adopts SimSPPF (Simplified SpatialPyramidPooling-Fast) to replace the original structure. SimSPPF is an improved structure that has been optimized in the network structure. This improvement can significantly speed up the calculation speed while maintaining or even improving the performance of the model. The core improvement of SimSPPF is to replace the two traditional convolutional layers in the SPPF module with Module. The SimConv module consists of a convolutional layer (Conv), a batch normalization layer (BN), and This combination simplifies the calculation process and enhances the efficiency of feature extraction. imSPPF retains the SPPF The two-dimensional maximum pooling operation is used to ensure that the ability to capture multi-scale features is not affected.
[0063] By introducing the SimConv module, SimSPPF effectively accelerates model convergence while reducing computational complexity. This improvement is particularly suitable for scenarios requiring efficient inference, such as real-time object detection or video processing. SimSPPF's design not only optimizes computing resource utilization but also further enhances model stability and generalization capabilities through the combination of batch normalization and the ReLU activation function. Experiments demonstrate that SimSPPF reduces inference time while maintaining high accuracy, providing a more efficient solution for deploying deep learning models in real-world applications.
[0064] To demonstrate SimSPPF's superior computational speed compared to SPPF, a validation dataset was created for experimental purposes. Code was used to display the FPS during model recognition. All other conditions remained the same, and tests were conducted using both SPPF and SimSPPF, with the average value taken after five experiments. The table below shows the operating speeds and parameter counts of the two architectures. It can be seen that the SimSPPF module takes less computation time than the SPPF module, while maintaining the same parameter count. Furthermore, test results show that the SimSPPF module achieves a 13.7% improvement in FPS and a 9.5% improvement in detection speed compared to the SPPF module. This demonstrates that this module is more suitable for use in lightweight models, and therefore, this invention incorporates the SimSPPF module as a strategy for lightweight model improvement.
[0065] Module calculation time and parameter quantity In a specific embodiment, the road crack identification model YOLOv8-RCI is constructed in the following way: Compared to common detection targets, road cracks occupy a much lower percentage of pixels in an image, making crack feature extraction significantly more difficult. Furthermore, because crack feature information is easily lost in deep networks, the model is prone to false or missed detections. This not only affects detection accuracy but also places higher demands on the precision of the recognition algorithm. Therefore, this paper introduces the PKIBlock multi-scale convolution kernel to enhance the model's ability to extract multi-scale features, thereby better capturing the subtle characteristics of cracks.
[0066] To further improve model accuracy and address the issues associated with increased parameter count, this paper introduces a Generalized Efficient Layer Aggregation Network (GELAN) module to replace the C2f module in the original model. The GELAN module optimizes the model's network structure, significantly improving accuracy while reducing the number of model parameters and computational complexity, thereby increasing model efficiency.
[0067] This invention also introduces an EMA attention mechanism. By focusing on key areas of crack characteristics, the EMA mechanism further improves the model's detection accuracy without adding additional parameters. To further optimize model performance, the present invention replaces the spatial pyramid pooling module (SPPF) with the SimSPPF module. The SimSPPF module preserves multi-scale spatial information while increasing the model's speed, making it more efficient when processing large-scale image data.
[0068] By leveraging these four improvements, we successfully developed a road crack identification model, YOLOv8-RCI, that balances detection performance and speed. This model improves detection accuracy and reduces computational resource consumption compared to the original model, significantly enhancing both segmentation performance and runtime speed. Figure 9 The complete structure of the YOLOv8-RCI model is shown, clearly showing the synergy between the modules and the optimized overall network architecture. Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. A pavement crack detection method based on the Yolov8 model, characterized in that: The following steps are involved: Constructing a pavement crack segmentation dataset: We collected road crack images using drones and, combined with public datasets, performed data enhancement processing on the images using geometric and color transformations. We also used Gaussian filtering to denoise noisy images and used Labelme to annotate the cracks. Based on the YOLOv8-Seg model, we improved it by introducing the PKIblock multi-scale convolution kernel, the generalized efficient layer aggregation network GELAN module, the EMA attention mechanism, and replacing the spatial pyramid pooling layer SPPF with SimSPPF to construct the road crack identification model YOLOv8-RCI. The trained YOLOv8-RCI model is used to perform crack detection and instance segmentation on UAV images, and output crack location and mask information.
2. A pavement crack detection method based on Yolov8 model according to claim 1, characterized in that: The process of building the road crack identification model YOLOv8-RCI includes: The PKIblock multi-scale convolution kernel is embedded after the SPPF pyramid pooling layer to enhance the multi-scale crack feature extraction capability; The C2f module in the original YOLOv8-Seg model is replaced by the generalized efficient layer aggregation network GELAN module to reduce the number of model parameters and optimize the feature transfer path to reduce computational costs; Add an EMA attention mechanism module to the output of the four C2f modules in the Neck part to focus on the key crack areas; The spatial pyramid pooling layer SPPF is replaced with SimSPPF, which consists of SimConv and a 5×5 maximum pooling layer to improve the model detection speed.
3. A pavement crack detection method based on Yolov8 model according to claim 2, characterized in that: The PKIblock multi-scale convolution kernel enhances the features of the central area by combining global average pooling and 1×1 strip convolution, thereby effectively capturing long-range contextual information, where: The PKIblock multi-scale convolution kernel consists of a PKI module and a CAA module. The PKI module extracts local information through a small convolution kernel. After the extraction is completed, a set of parallel deep convolution layers are used to capture contextual information across multiple scales. After the PKI module extracts the features, it uses 1×1 convolution to fuse the local features and context information to represent the relationship between different channels; CAA uses average pooling operations to extract global context information and capture the dependencies between distant pixels. It then fuses the global information with local features through 1×1 convolution to establish relationships between modeling channels.
4. A pavement crack detection method based on Yolov8 model according to claim 1, characterized in that: The GELAN module achieves model compression and performance improvement by integrating the feature map segmentation and reorganization mechanism of the CSPNet architecture with the gradient path optimization strategy of the ELAN module; The feature map segmentation and reorganization mechanism of the CSPNet architecture includes: dividing the input feature map into two branches. The first branch is directly transferred across stages, and the second branch extracts features through local dense blocks. Finally, the features of the two branches are recombined to eliminate gradient redundancy and reduce memory bandwidth usage; The gradient path optimization strategy of the ELAN module includes: building a dual-branch structure in the network layer. The first branch adjusts the channel dimension through a 1×1 convolution kernel, and the second branch optimizes the gradient propagation path through a hierarchical convolution module consisting of a 1×1 convolution and four stacked 3×3 convolutions to shorten the longest gradient path and suppress the gradient vanishing problem of deep networks.
5. The pavement crack detection method based on the Yolov8 model according to claim 1, characterized in that: The EMA attention mechanism is embedded in the four C2f level outputs of the Neck module. It achieves feature enhancement through multi-channel grouping and dual-path semantic extraction. Specifically, it includes: The channel dimension partitioning module divides the input feature map channels into multiple subgroups and performs batch dimension reshaping to establish cross-group associations; The spatial semantic extraction branch uses global average pooling to compress spatial information and generate spatial weights; The channel semantic extraction branch performs dimensionality reduction through a 1×1 convolution kernel; The dual-branch features are fused through a joint activation mechanism and then element-wise added to the original input to generate enhanced features. The incentive mechanism dynamically allocates channel importance weights through the normalization layer and the Softmax function, and the adjustment mechanism implements feature calibration based on 2D global average pooling.
6. A pavement crack detection method based on Yolov8 model according to claim 1, characterized in that: SimSPPF replaces the two traditional convolutional layers in the original SPPF module with the SimConv module; The SimConv module consists of a convolutional layer Conv, a batch normalization layer BN, and a SiLU activation function; The three-level cascaded maximum pooling operation is retained to ensure the ability to capture multi-scale features.
Citation Information
Cited By
Multi-level concrete crack detection method based on unmanned aerial vehicle
CN121147227A
A multi-level concrete crack detection method based on a drone
CN121147227B
Long-distance mouth breathing state detection method and device for preschool children
CN121459411A
Bridge crack identification method based on improved YOLOv8-seg model
CN121640215A
Ground penetrating radar image crack identification method and system based on frequency domain enhanced YOLOv11
CN121837191A