Road crack target detection method based on improved YOLOv8
By introducing an efficient multi-scale attention mechanism and reconstructing multi-scale convolution in YOLOv8, and replacing the loss function with WIoU, the problem of insufficient accuracy in multi-scale crack detection is solved, the detection accuracy and recall rate are improved, and it is suitable for road crack detection.
Patent Information
- Application Number
- CN202510629048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
AI Technical Summary
When detecting multi-scale crack targets, the YOLOv8 algorithm has insufficient detection accuracy and there are false detection and missed detection.
The efficient multi-scale attention mechanism EMA is introduced in the backbone part of YOLOv8, and the multi-scale convolution MSConv is reconstructed, and the loss function is replaced by WIoU, enhancing the model's capture ability of multi-scale features and bounding box regression accuracy.
The accuracy and recall of multi-scale crack detection of the model is improved, the computational complexity is reduced, the model's generalization ability in complex backgrounds is enhanced, and the accuracy and efficiency balance is achieved.
Smart Images

Figure CN120525840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection, and in particular to a road crack target detection method based on improved YOLOv8. Background Art
[0002] Road crack detection is a key technology widely used in industry. Its primary goal is to accurately locate and identify various cracks in road surfaces, such as longitudinal cracks, transverse cracks, and lattice cracks, through automated methods. These cracks are common road surface issues, impacting road safety and driver safety.
[0003] With the rapid development of computer vision and deep learning, steel surface defect detection is increasingly adopting automated methods based on image processing and machine learning. There are two main types of deep learning-based methods: one is a two-stage algorithm, represented by RCNN and Faster-RCNN, which first generates region proposals and then classifies samples. The other is a one-stage algorithm, represented by the YOLO series and SSD. These algorithms do not generate region proposals but directly generate detection categories and locations, achieving final detection results in a single stage.
[0004] The YOLO series of algorithms has been widely used in industry due to its excellent detection performance and speed. YOLOv8 boasts several new features compared to previous versions, including a C2f structure with richer gradient flows, a mainstream decoupling head, and a transition from anchor-based to anchor-free. However, the YOLOv8 algorithm still lacks accuracy when detecting multi-scale cracks, with false and missed detections occurring frequently. Summary of the Invention
[0005] (1) Technical problems solved
[0006] In response to the shortcomings of the existing technology, the present invention provides a road crack target detection method based on an improved YOLOv8, which solves the problems raised in the above background technology while appropriately increasing the number of parameters and reducing the detection speed.
[0007] (2) Technical solution
[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0009] A road crack target detection method based on improved YOLOv8 is characterized by comprising the following steps:
[0010] S1, obtain the public road damage dataset RDD2022, which divides road damage into cracks and potholes. Cracks are further divided into three categories: longitudinal cracks, transverse cracks, and network cracks;
[0011] S2, divide the RDD2022 dataset into training set, test set and validation set according to the ratio of 8:1:1;
[0012] S3, build the road crack target detection network YOLOv8-MEW based on the improved YOLOv8;
[0013] The improvement process of YOLOv8-MEW in S3 includes:
[0014] First, an efficient multi-scale attention mechanism EMA is introduced into the Bottleneck of the last C2f in the backbone part of the target detection network to form C2f_EMA;
[0015] Then reconstruct the multi-scale convolution MSConv and replace all traditional Convs of the original YOLOv8 network with MSConv;
[0016] Finally, WIoU is used to replace CIoU as the loss function of the network;
[0017] S4, import the training set and validation set into YOLOv8-MEW for training to obtain a road crack target detection model;
[0018] S5, use the trained road crack target detection model to detect road cracks on the test set.
[0019] Furthermore, the C2f_EMA includes: the EMA in C2f_EMA adopts two one-dimensional global average pooling operations in the 1×1 branch, encodes the channels along the length and width spatial directions respectively, and models cross-channel interaction information to connect the encoded features, share the 1×1 convolution kernel so that it does not reduce the dimension on this branch, and decompose the output into two vectors in the length and width directions. The two-dimensional binomial distribution on the linear convolution is then fitted by the Sigmoid nonlinear activation function, and finally cross-channel interaction is achieved by multiplication aggregation channel attention, and a 3×3 convolution is used in the 3×3 branch to capture multi-scale feature representation.
[0020] Furthermore, the multi-scale convolution MSConv includes: reconstructing a new multi-scale convolution MSConv using standard convolution Conv and deep convolution DWConv; the specific steps are: the input features are first subjected to a 3x3 convolution kernel Conv for preliminary feature extraction, and then further feature extraction is performed through parallel 3x3 convolution kernel DWConv, 3x3 convolution kernel Conv and 5x5 convolution kernel DWConv, while setting a residual structure; the features extracted in the first and second steps are fused together and then passed through a BN normalization layer and SiLU activation function, and finally passed through a 1x1 convolution kernel Conv to adjust the number of output channels. The DWC onv comes from depthwise separable convolution, a convolution operation commonly used in neural networks, mainly used to capture the spatial features of input information. The difference between DWConv and Conv is that it does not require each input channel to be convolved with all filters. Instead, each input channel is convolved with a separate filter to generate a corresponding output feature map, which greatly reduces the number of parameters. The reconstructed MSConv combines the advantages of DWConv and Conv, while reducing redundant calculations, it improves the model's multi-scale feature extraction capability and detection accuracy. In addition, the residual structure can construct a richer gradient flow, which is conducive to solving the gradient disappearance problem in deep networks.
[0021] Furthermore, the WIoU includes: a dynamic non-monotonic focusing mechanism that can adaptively adjust the loss weight according to the overlap quality of the predicted box and the true box, ensuring that the gradient of high-quality samples dominates the optimization direction, using "outlier degree" instead of intersection-over-union to evaluate the quality of the predicted box, quantifying the degree of deviation of the predicted box, and suppressing the negative impact of low-quality samples; WIoU is defined as follows:
[0022] S u =wh+w gt h gt -W i H i
[0023]
[0024] L WIoU =rR WIoU L IoU
[0025]
[0026] Where IoU is the intersection over union ratio between the predicted box and the real box, L IoU is the intersection-over-union loss function; w, h are the width and height of the prediction box, (x, y) are the center coordinates of the prediction box; w gt ,h gtis the width and height of the real frame, (x gt ,y gt ) is the center coordinate of the real frame; R WIoU is the penalty term, R WIoU is the WIoU loss function; W i and H i is the width and height of the overlapping part of the predicted box and the real box; W g 、H g is the width and height of the minimum enclosing box, in order to prevent R WIoU Producing gradients that hinder convergence, W g 、H g Separated from the computational graph, “*” indicates the operation; β is the outlier degree, is the dynamic average IoU value; r is the gradient gain, α and δ are learning parameters; S u It is the connected union of the smallest enclosing box and the center point.
[0027] (3) Beneficial effects
[0028] Compared with the existing technology, the present invention provides a road crack target detection method based on improved YOLOv8, which has the following beneficial effects:
[0029] This paper optimizes YOLOv8n for multi-scale target crack detection by reconstructing the multi-scale convolution MSConv: using standard convolution Conv with different convolution kernels and deep convolution DWConv combined with residual structure to reconstruct the multi-scale convolution MSConv, and replacing the traditional convolution of the original YOLOv8 network with the reconstructed MSConv. While reducing redundant calculations, the multi-scale feature extraction capability of the model is enhanced.
[0030] This paper optimizes YOLOv8n for multi-scale target crack detection by integrating EMA in the last C2f of the backbone part of the target detection network: EMA is an efficient multi-scale attention mechanism. Through the collaborative design of multi-scale feature extraction and dynamic attention weight allocation, it enhances the model's capture of local details and global information while maintaining a low computational cost, thereby improving the model's detection accuracy.
[0031] This paper optimizes YOLOv8n for multi-scale crack detection by replacing the original CIoU loss function with WIoU. WIoU strengthens the loss weight of small objects through a dynamic focusing mechanism and improves the bounding box regression accuracy of irregular cracks through geometric decoupling. Its gradient smoothing strategy addresses the fluctuation problem of CIoU in low-overlap scenarios, while suppressing noise interference through a quality-aware mechanism, enhancing the model's generalization in complex backgrounds. Compared to the limitations of CIoU's fixed aspect ratio penalty, WIoU significantly improves the detection recall, convergence speed, and interference resistance of small cracks while maintaining the same computational efficiency, making it more suitable for the multi-scale feature balance required by lightweight models. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a flow chart of a road crack target detection method based on improved YOLOv8 in the present invention;
[0033] Figure 2 This is a network structure diagram of a road crack target detection method based on improved YOLOv8 in the present invention;
[0034] Figure 3 This is a structural diagram of C2f_EMA of the present invention;
[0035] Figure 4 It is the structural diagram of Bottleneck_EMA of the present invention;
[0036] Figure 5 This is the structural diagram of the efficient multi-scale attention EMA of the present invention;
[0037] Figure 6 This is the structural diagram of the multi-scale convolution MSConv of the present invention;
[0038] Figure 7 This is the convergence performance diagram of YOLOv8 of the present invention;
[0039] Figure 8 This is the convergence performance diagram of the improved YOLOv8-MEW of the present invention;
[0040] Figure 9 This is the detection effect diagram of YOLOv8-MEW of the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] Example
[0043] like Figure 1 As shown, an embodiment 1 of the present invention provides a flowchart of a road crack target detection method based on an improved YOLOv8, which specifically includes the following steps:
[0044] S1. Obtain the public road damage dataset RDD2022. This dataset categorizes road damage into cracks and potholes. Cracks are further divided into three types: longitudinal cracks, transverse cracks, and reticular cracks. The RDD2022 dataset defines the crack type labels as D00 for longitudinal cracks, D10 for transverse cracks, and D20 for reticular cracks.
[0045] S2: Select 10,000 images from the RDD2022 dataset and divide them into training set, test set, and validation set in a ratio of 8:1:1.
[0046] The training set, test set and validation set in S2 are specifically as follows: In the target detection task based on YOLOv8, the training set, validation set and test set are three key data sets for model development and evaluation. The training set is used to train the model so that the model learns to extract features from the input image and identify the target, and adjust the weights. The validation set is used to regularly evaluate the model performance and adjust the hyperparameters during training to avoid overfitting. The test set is used to evaluate the generalization ability of the model after the final model training is completed. The specific operation of dividing the data set in S2 is: write a python script file to create summary files train.txt and val.txt that store the absolute path of the picture and the label position and category by line, and finally put the label tags and jpg pictures in the training set and test set in the same directory.
[0047] S3: Build the road crack target detection model YOLOv8-MEW based on the improved YOLOv8. The network structure is as follows Figure 2 As shown. The operation in S3 is: first, add an efficient multi-scale attention mechanism EMA to the end of the Bottleneck inside the last C2f in the backbone part of the target detection network to form C2f_EMA. The structure of C2f_EMA is as follows Figure 3 As shown in the figure, the specific operation is: add EMA at the end of the Bottleneck inside C2f to form Bottleneck_EMA. The structure of Bottleneck_EMA is as follows: Figure 4 As shown, the structure of the efficient multi-scale attention EMA in Bottleneck_EMA is as follows Figure 5As shown in the figure, the EMA specifically involves two one-dimensional global average pooling operations in the 1×1 branch, encoding channels along the length and width directions, respectively. Cross-channel interaction information is modeled to connect the encoded features, and a shared 1×1 convolution kernel is used to avoid dimensionality reduction in this branch. The output is decomposed into two vectors in the length and width directions. A 2D binomial distribution is then fitted to the linear convolution using a Sigmoid nonlinear activation function. Finally, cross-channel interaction is achieved through multiplicative aggregation of channel attention. A single 3×3 convolution is used in the 3×3 branch to capture multi-scale feature representations.
[0048] Then, we reconstruct the multi-scale convolution MSConv and replace all the traditional Convs in the original YOLOv8 network with MSConv. The traditional convolution has a single receptive field, poor multi-scale feature extraction effect, and redundant parameters. The reconstructed MSConv uses standard convolution Conv and deep convolution DWConv combined with a multi-scale parallel structure to optimize computational efficiency. The combination of residual structure and feature fusion improves the model's feature expression ability and generalization performance, making it suitable for road crack target detection tasks that require multi-scale perception. The structure of the MSconv is as follows: Figure 6 As shown in the figure, the input features are first processed through a 3x3 convolution kernel (Conv) for preliminary feature extraction. They are then further extracted through parallel 3x3 convolution kernels (DWConv), 3x3 convolution kernels (Conv), and 5x5 convolution kernels (DWConv), while a residual structure is also set. The features extracted in the first and second steps are fused together, then passed through a batch normalization layer and a SiLU activation function, and finally passed through a 1x1 convolution kernel (Conv) to adjust the number of output channels.
[0049] Finally, we propose to replace the original network loss function CIoU with WIoU. WIoU specifically includes: a dynamic non-monotonic focusing mechanism that can adaptively adjust the loss weight based on the overlap quality between the predicted box and the true box, ensuring that the gradient of high-quality samples dominates the optimization direction; and using "outlier degree" instead of intersection-over-union to evaluate the quality of the predicted box, quantifying the degree of deviation of the predicted box and suppressing the negative impact of low-quality samples. WIoU is defined as follows:
[0050] S u =wh+w gt h gt -W i H i
[0051]
[0052] L WIoU =R WIoU L IoU
[0053]
[0054] Where IoU is the intersection over union ratio between the predicted box and the real box, L IoU is the intersection-over-union loss function; w, h are the width and height of the prediction box, (x, y) are the center coordinates of the prediction box; w gt ,h gt is the width and height of the real frame, (x gt ,y gt ) is the center coordinate of the real frame; R WIoU is the penalty term, R WIoU is the WIoU loss function; W i and H i is the width and height of the overlapping part of the predicted box and the real box; W g 、H g is the width and height of the minimum enclosing box, in order to prevent R WIoU Producing gradients that hinder convergence, W g 、H g Separated from the computational graph, “*” indicates the operation; β is the outlier degree, is the dynamic average IoU value; r is the gradient gain, α and δ are learning parameters; S u It is the connected union of the smallest enclosing box and the center point.
[0055] S4: Evaluate the training of the improved YOLOv8-MEW road crack detection model. The operation in S4 is to visualize the training process and the final training results through Tensorboard, such as Figure 7 and Figure 8The figures show the convergence performance of YOLOv8n and YOLOv8-MEW on the RDD2002 dataset. The figures include curves for the bounding box regression loss, confidence loss, and classification loss on the training and validation sets, as well as convergence curves for precision, recall, mAP50, and mAP50-95. The convergence performance graphs for the two models show that YOLOv8-MEW achieves lower bounding box regression and classification loss values than the original YOLOv8n model, and converges more smoothly. YOLOv8 achieves an mAP50 of 63%, while the improved YOLOv8-MEW achieves 66.5%. The YOLOv8-MEW network parameters, computational complexity (GFLOPs), and inference speed (FPS) are compared with the original YOLOv8 network. Compared to the original YOLOv8 network, this model improves mAP50 by 5.5% without significantly increasing the number of parameters or model size. This significantly improves model detection performance while achieving a good balance between accuracy and efficiency. The improved model achieves an FPS of 168, meeting the requirements for subsequent mobile deployment. Figure 9 This is the detection effect diagram of YOLOv8-MEW. In the figure, D00 represents that the detected crack is a longitudinal crack, D10 represents that the detected crack is a transverse crack, and D20 represents that the detected crack is a mesh crack.
[0056] The model training environment of the present invention is: the CPU uses Intel (R) Xeon (R) Platinum 8362, the GPU uses RTX3090, the running memory uses 24GB, the operating system is Ubuntu 20.04, and the PyTorch 2.0.0 deep learning framework and the Python 3.8 programming language are used.
[0057] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A road crack target detection method based on improved YOLOv8, characterized by: The steps include: S1, obtain the public road damage dataset RDD2022, which divides road damage into cracks and potholes. Cracks are further divided into three categories: longitudinal cracks, transverse cracks, and network cracks; S2, divide the RDD2022 dataset into training set, test set and validation set according to the ratio of 8:1:1; S3, build the road crack target detection network YOLOv8-MEW based on the improved YOLOv8; The improvement process of YOLOv8-MEW in S3 includes: First, an efficient multi-scale attention mechanism EMA is introduced into the Bottleneck of the last C2f in the backbone part of the target detection network to form C2f_EMA; Then reconstruct the multi-scale convolution MSConv and replace all traditional Convs of the original YOLOv8 network with MSConv; Finally, WIoU is used to replace CIoU as the loss function of the network; S4, import the training set and validation set into YOLOv8-MEW for training to obtain a road crack target detection model; S5, use the trained road crack target detection model to detect road cracks on the test set.
2. The road crack target detection method based on improved YOLOv8 according to claim 1, characterized in that: The C2f_EMA includes: the EMA in C2f_EMA adopts two one-dimensional global average pooling operations in the 1×1 branch, encodes the channels along the length and width spatial directions respectively, and models cross-channel interaction information to connect the encoded features, sharing the 1×1 convolution kernel so that it does not reduce the dimension on this branch, and the output is decomposed into two vectors in the length and width directions. The two-dimensional binomial distribution on the linear convolution is then fitted by the Sigmoid nonlinear activation function, and finally cross-channel interaction is achieved by multiplication aggregation channel attention, and a 3×3 convolution is used in the 3×3 branch to capture multi-scale feature representation.
3. The road crack target detection method based on improved YOLOv8 according to claim 1, characterized in that: The multi-scale convolution MSConv includes: reconstructing a new multi-scale convolution MSConv using standard convolution Conv and deep convolution DWConv; the specific steps are: the input features are first subjected to a 3x3 convolution kernel Conv for preliminary feature extraction, and then further feature extraction is performed through parallel 3x3 convolution kernel DWConv, 3x3 convolution kernel Conv and 5x5 convolution kernel DWConv, while setting a residual structure; the features extracted in the first and second steps are fused together and then passed through a BN normalization layer and SiLU activation function, and finally passed through a 1x1 convolution kernel Conv to adjust the number of output channels. The DWCon v comes from depthwise separable convolution, a convolution operation commonly used in neural networks, mainly used to capture the spatial features of input information. The difference between DWConv and Conv is that it does not require each input channel to be convolved with all filters. Instead, each input channel is convolved with a separate filter to generate a corresponding output feature map, which greatly reduces the number of parameters. The reconstructed MSConv combines the advantages of DWConv and Conv, while reducing redundant calculations, it improves the model's multi-scale feature extraction capability and detection accuracy. In addition, the residual structure can construct a richer gradient flow, which is conducive to solving the gradient disappearance problem in deep networks.
4. The road crack target detection method based on improved YOLOv8 according to claim 1, characterized in that: The WIoU includes: a dynamic non-monotonic focusing mechanism that can adaptively adjust the loss weight according to the overlap quality of the predicted box and the true box, ensuring that the gradient of high-quality samples dominates the optimization direction, using "outlier degree" instead of intersection-over-union to evaluate the quality of the predicted box, quantifying the degree of deviation of the predicted box, and suppressing the negative impact of low-quality samples; WIoU is defined as follows: S u =wh+w gt h gt -W i H i L WIoU =rR WIoU L IoU Where IoU is the intersection over union ratio between the predicted box and the real box, L IoU is the intersection-over-union loss function; w, h are the width and height of the prediction box, (x, y) are the center coordinates of the prediction box; w gt ,h gt is the width and height of the real frame, (x gt ,y gt ) is the center coordinate of the real frame; R WIoU is the penalty term, R WIoU is the WIoU loss function; W i and H i is the width and height of the overlapping part of the predicted box and the real box; W g 、H g is the width and height of the minimum enclosing box, in order to prevent R WIoU Producing gradients that hinder convergence, W g 、H g Separated from the computational graph, "*" represents the operation; β is the outlier degree, is the dynamic average IoU value; r is the gradient gain, α and δ are learning parameters; S u It is the connected union of the smallest enclosing box and the center point.