Pavement crack detection method based on improved YOLOV8

By improving the YOLOV8 algorithm, the feature extraction is enhanced by using the C2f_DCNV2 module and the CoordAtt attention mechanism, combined with the EIOU loss function, the accuracy and robustness of road crack detection are solved, and efficient crack detection and positioning are achieved.

CN120375064AInactive Publication Date: 2025-07-25YANCHENG INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510459904.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing pavement crack detection methods have poor results in small-sized crack detection and are not robust enough in complex backgrounds, making it difficult to meet the needs of road maintenance and management.

Method used

The improved YOLOV8 algorithm is adopted to enhance the feature extraction capability of the backbone network through the C2f_DCNV2 module, and a CoordAtt attention mechanism and EIOU regression loss function are introduced to improve the model's ability to capture cross-channel and spatial features and output target detection results.

Benefits of technology

Accurate classification and positioning of road cracks is achieved, improving the work efficiency of maintenance personnel and reducing retrieval costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375064A_ABST
    Figure CN120375064A_ABST
Patent Text Reader

Abstract

The invention provides a pavement crack detection method based on improved YOLOV8, and the method comprises the following steps: S10, obtaining to-be-processed pavement crack data, carrying out the preprocessing of the data, and carrying out the enhancement of the preprocessed data in a training stage; s20, adopting a C2fDCNV2 module to improve a backbone network of the model, carrying out feature extraction on data and fusing the network, enhancing the ability of a convolutional neural network to capture cross-channel and spatial important features by a CoordAtt attention mechanism, and keeping relatively low calculation and parameter cost at the same time; s30, the improved YOLOv8 algorithm is trained; s40, carrying out image preprocessing on the detection image and outputting a target detection result; according to the method, the problem of a traditional method for detecting the pavement cracks through manual detection is solved, the category of each pavement crack in the image can be accurately detected, the specific position of the pavement crack can be positioned, similar cracks can be retrieved for maintenance personnel, the maintenance personnel can be helped to maintain the pavement cracks more conveniently, and the maintenance efficiency is improved. Therefore, the working efficiency of maintenance personnel is improved, and the retrieval cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and more specifically, to a road crack detection method based on improved YOLOV8. Background Art

[0002] Road transportation, as an important infrastructure in modern society, provides strong support for people's travel, material circulation, and national economic development. Through continuous technological innovation in highway construction, from early dirt roads and gravel roads to asphalt roads, cement roads, and then to highways, the quality, safety, and transportation efficiency of roads have been gradually improved. However, with the extension of the service life of roads and the increase in traffic flow, various damages inevitably occur on roads, among which cracks are the most common. Cracks not only affect the flatness, stability, and durability of roads, but also may bring poor driving experiences and safety hazards to drivers, and increase maintenance costs and resource consumption. Therefore, timely detection and repair of cracks are crucial for road maintenance and management.

[0003] Due to the complexity and diversity of road surface cracks, it is difficult to fully cover all situations. At the same time, existing methods also face some challenges, such as poor detection effects on small-sized cracks and insufficient robustness in complex backgrounds. In view of the above problems, an improved YOLOV8 road crack detection method is designed to improve the efficiency and accuracy of road crack detection algorithms. Summary of the Invention

[0004] To solve the above technical problems, the present invention accurately determines the category of each crack detection in the image through target detection technology, retrieves similar cracks for maintenance personnel, and helps maintenance personnel maintain road cracks more conveniently, thereby improving the work efficiency of maintenance personnel and reducing retrieval costs. The specific technical solutions are as follows:

[0005] A road crack detection method based on improved YOLOV8, comprising the following steps:

[0006] Obtain the data to be processed for road cracks, preprocess the data, and enhance the preprocessed data during the training phase;

[0007] Use the C2f_DCNV2 module to improve the backbone network of the model, extract and fuse features from the data, and the CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially while maintaining low computational and parameter costs;

[0008] Train the improved YOLOv8 algorithm;

[0009] Preprocess the detection image and output the target detection result.

[0010] As a preferred technical solution of the present application, the steps of obtaining the data of the road surface cracks to be processed, preprocessing the data, and enhancing the preprocessed data in the training stage include:

[0011] Sort out the obtained road surface crack detection data;

[0012] Randomly read four pictures from the data in the training stage, and perform flipping, scaling, color gamut transformation, and random cropping operations on the obtained four pictures;

[0013] Arrange the processed photos with the first picture placed in the upper left, the second picture placed in the lower left, the third picture placed in the lower right, and the fourth picture placed in the upper right. Use a matrix method to intercept the fixed areas of the four pictures and splice them into a new picture containing a series of bounding boxes and label information;

[0014] Divide the preprocessed data set into a training set, a validation set, and a test set according to a given ratio, and the ratio of the training set, the validation set, and the test set is 7:2:1.

[0015] As a preferred technical solution of the present application, in the step of improving the backbone network of the model using the C2f_DCNV2 module, extracting features from the data and fusing the network, while the CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially and maintains a low computational and parameter cost, add the C2f_DCNV2 module. The C2f_DCNV2 module enhances the receptive field and feature extraction ability by integrating the deformable convolution DCNv2 in the C2f module of the backbone network to improve the performance of the model in object detection and image segmentation tasks.

[0016] As a preferred technical solution of the present application, the steps of improving the backbone network of the model using the C2f_DCNV2 module, extracting features from the data and fusing the network, while the CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially and maintains a low computational and parameter cost include:

[0017] Input the data into the improved backbone network, and perform feature extraction through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module in sequence;

[0018] The CoordAtt attention mechanism embeds spatial coordinate information (in the height and width directions) into the channel attention mechanism, enabling the model to focus on important features at different positions. Among them, CoordAtt is used to make up for the defect caused by the traditional attention mechanism ignoring the spatial structure of the feature map;

[0019] The EIOU regression loss function is used to solve the problems existing in CIOU. Based on the penalty term of CIOU, the influence factors of the aspect ratio of the candidate box and the actual target area are split, and the length and width values of the candidate box and the actual target area are calculated separately. The expression formula of the EIOU regression loss function is as follows:

[0020]

[0021] As a preferred technical solution of this application, the c ω and c h are the length and width of the minimum bounding rectangle of the closed candidate box and the actual area. The formula contains three parts: the overlap loss L IOU between the candidate box and the actual area, the center distance loss L dis between the candidate box and the actual area, and the aspect ratio loss L asp between the candidate box and the actual area. loU is the anchor box loss function, and its result is obtained by dividing the overlapping part of the two regions of the predicted box and the true box by the union part of the two regions; b, b gt are the center points of the predicted box and the true box respectively; ρ 2 (b, b gt ) is the Euclidean distance between the center points of the predicted box and the true box; c is the diagonal distance of the smallest closure region that can contain both the predicted box and the true box; α is a balancing parameter that does not participate in the gradient calculation; υ is a parameter used to measure the aspect ratio consistency; w, h are the width and height of the predicted box respectively; w gt , h gt are the width and height of the true box respectively. The EIOU loss function calculates the length and width of the target box separately, solves the problem of large errors in the horizontal and vertical directions of the GIOU loss function, and can improve the convergence speed and regression accuracy.

[0022] As a preferred technical solution of this application, the way of feature extraction in the step of inputting data into the improved backbone network and passing through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module in sequence is as follows:

[0023] Given an intermediate feature map F ∈ R C×H×WAs the input, first perform global max pooling and average pooling on the input by channel. Feed the two resulting one-dimensional vectors after pooling into a fully connected layer, add the results after the operation, and generate a one-dimensional channel attention M C ∈R C ×1×1 Then, multiply the channel attention by the input elements to obtain the feature map F' adjusted by the channel attention. Next, perform global max pooling and average pooling on F' spatially. Concatenate the two resulting two-dimensional vectors after pooling and perform a convolution operation to finally generate a two-dimensional spatial attention M S ∈R 1×H×W Then, multiply the spatial attention by F' element-wise.

[0024] As a preferred technical solution of the present application, in the step of inputting data into the improved backbone network and sequentially performing feature extraction through a Conv module, a Conv module, a C2f_DCNV2 module, a Conv module, a C2f_DCNV2 module, a Conv module, a C2f_DCNV2 module, a Conv module, a C2f_DCNV2 module, and an SPPF module, the implementation process of the C2f_DCNV2 module includes the following steps:

[0025] Feature transformation, perform feature transformation on the input data through two convolutional layers (cv1 and cv2) and extract features at different levels and degrees of abstraction;

[0026] Branch processing, divide the input data into two branches for processing. One branch is directly passed to the output, and the other branch is processed through multiple Bottleneck modules to increase the non-linearity and representational ability of the network;

[0027] Feature fusion, concatenate the features of different branches in the channel dimension to enrich the feature expression ability.

[0028] As a preferred technical solution of the present application, the CoordAtt attention mechanism embeds spatial coordinate information (in the height and width directions) into the channel attention mechanism, enabling the model to focus on important features at different positions. Among them, the steps for CoordAtt to make up for the defects caused by the traditional attention mechanism ignoring the spatial structure of the feature map include:

[0029] Input feature decomposition and global pooling, used to process the input feature map and perform global pooling in the height and width directions respectively;

[0030] Information aggregation and encoding, used to perform feature encoding on the feature vectors in the height and width directions through a shared convolutional layer;

[0031] Generate directional attention weights, which are used to perform convolutions on the aggregated feature vectors in the height and width directions respectively to generate two different directional attention weights;

[0032] Weight application and feature recalibration, which are used to apply the generated height and width directional attention weights back to the original feature map respectively and recalibrate the features in each direction.

[0033] As a preferred technical solution of the present application, the steps of training the improved YOLOv8 algorithm include:

[0034] Input the paths of the training set and the validation set and the path of the initial training weight file into the parameters required for network model training to perform training, and obtain the improved yolov8 model after training is completed.

[0035] As a preferred technical solution of the present application, the steps of performing image preprocessing on the detection image and outputting the object detection result include:

[0036] Output the object recognition results, including four results of longitudinal cracks, transverse cracks, reticular cracks, and potholes. It is written in Python language, and the correctness of the code is verified on a PC using PyCharm, and a camera is connected for image analysis to perform detection experiments to verify the effectiveness of the method.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] By adding the C2f_DCNV2 module in front of the SPPF layer in the backbone network of the yolov8 network, it brings a stable performance improvement. Then, the CoordAtt attention mechanism is introduced to endow the model with attention to important features at different positions. By using the EIOU loss function to calculate the length and width of the target box separately, it can solve the problem of large errors of the GIOU loss function in the horizontal and vertical directions, and can improve the convergence speed and regression accuracy. The present invention can accurately judge the category of each pavement crack in the image, locate the specific position of the pavement crack, retrieve similar cracks for maintenance personnel, and help maintenance personnel maintain pavement cracks more conveniently, thereby improving the work efficiency of maintenance personnel and reducing the retrieval cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is the structural flow chart of the present invention;

[0040] Figure 2 It is the schematic diagram of the C2f_DCNV2 module of the present invention;

[0041] Figure 3 It is the schematic diagram of the CoordAtt attention mechanism of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0043] Embodiment:

[0044] As Figures 1 to 3 shown, this embodiment provides a road surface crack detection method based on improved YOLOV8, including the following steps: S10, S20, S30, and S40.

[0045] Specifically, S10: Obtain the data to be processed for road surface cracks, preprocess the data, and enhance the preprocessed data during the training phase.

[0046] In this process, first collect a large number of road surface crack pictures as the data to be processed, then preprocess the collected data to be processed, and then enhance the preprocessed data during the training phase.

[0047] Furthermore, S10 further includes the following steps:

[0048] S11: Organize the obtained road surface crack detection data;

[0049] After the road surface crack detection data is collected and obtained, the obtained road surface crack data can be organized.

[0050] S12: Randomly read four pictures from the data during the training phase, and perform flipping, scaling, color gamut transformation, and random cropping operations on the obtained four pictures;

[0051] When training the data, first randomly read four pictures, and then sequentially perform flipping (flip the original picture left and right), scaling (scale the size of the original picture), color gamut transformation (change the brightness, saturation, and hue of the original picture), and random cropping operations on the read pictures. It should be noted that picture flipping is to flip the original picture left and right, picture scaling is to scale the size of the original picture, and picture color gamut transformation is to change the brightness, saturation, and hue of the original picture.

[0052] S13: Arrange the processed photos with the first picture placed in the upper left, the second picture placed in the lower left, the third picture placed in the lower right, and the fourth picture placed in the upper right. Use a matrix method to intercept the fixed areas of the four pictures and splice them into a new picture containing a series of bounding boxes and label information;

[0053] After processing the four acquired images, place the first image in the upper left, the second image in the lower left, the third image in the lower right, and the fourth image in the upper right. Then, use a matrix method to intercept the fixed areas of the four images. Finally, splice the four placed images into a new image containing a series of bounding boxes and label information.

[0054] S14: Divide the preprocessed dataset into a training set, a validation set, and a test set according to a given ratio, and the ratio of the training set, the validation set, and the test set is 7:2:1;

[0055] Use the preprocessed new image as the dataset and divide it into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0056] Specifically, S20: Use the C2f_DCNV2 module to improve the backbone network of the model, extract features from the data and fuse the network. The CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important cross-channel and spatial features while maintaining low computational and parameter costs;

[0057] In this process, first use the C2f_DCNV2 module to improve the backbone network in the model, and then extract the features in the data; then fuse the improved features with the CoordAtt attention mechanism, which not only strengthens the multi-scale feature fusion process, but also improves the feature expression ability while maintaining the parameter cost; finally, use EIOU to optimize the loss function.

[0058] In addition, add the C2f_DCNV2 module in step S20. The C2f_DCNV2 module enhances the receptive field and feature extraction ability by integrating the deformable convolution DCNv2 in the C2f module of the backbone network, and improves the performance of the model in object detection and image segmentation tasks;

[0059] Add the module C2f_DCNV2 in front of the SPPF layer in the backbone network of the yolov8 network, which brings stable performance improvement. It should be noted that the C2f_DCNV2 module can enhance the receptive field and feature extraction ability by integrating the deformable convolution DCNv2 in the C2f module of the backbone network, thereby enhancing the model's feature extraction ability for multi-scale targets.

[0060] Furthermore, S20 also includes the following steps:

[0061] S21: Input the data into the improved backbone network, and perform feature extraction through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module in sequence;

[0062] Input the processed image data into the improved backbone network, and the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module can perform feature extraction on the image data in sequence.

[0063] S22: The CoordAtt attention mechanism embeds spatial coordinate information (in the height and width directions) into the channel attention mechanism, endowing the model with the ability to focus on important features at different positions. Although traditional attention mechanisms (such as the SE module) can weight global information, they often ignore the spatial structure of the feature map. Therefore, CoordinateAttention makes up for this defect by introducing coordinate information;

[0064] The CoordAtt attention mechanism combines channel attention and coordinate (spatial) information through input feature decomposition and global pooling, information aggregation and encoding, generating direction attention weights, and weight application and feature recalibration, achieving more efficient capture of spatial information, thereby improving the network's recognition effect on target positions and details.

[0065] S23: Use the EIOU regression loss function to solve the problems existing in CIOU. Based on the penalty term of CIOU, split the influence factors of the aspect ratio of the candidate box and the actual target area, and calculate the length and width values of the candidate box and the actual target area respectively. The expression formula of the EIOU regression loss function is as follows:

[0066]

[0067]

[0068] The EIOU regression loss function can solve the problems existing in CIOU. After splitting the influence factors of the aspect ratio of the candidate box and the actual target area based on the penalty term of CIOU, calculate the length and width values of the candidate box and the actual target area respectively. Among them, the expression of the EIOU regression loss function is as follows:

[0069]

[0070] It should be noted that c ωand c h are the lengths and widths of the closed candidate box and the minimum bounding rectangle of the actual region, and the formula contains three parts: the overlap loss L between the candidate box and the actual region IOU , the center distance loss L between the candidate box and the actual region dis , the length-width loss L between the candidate box and the actual region asp , loU is the anchor box loss function, and its result is obtained by dividing the overlapping part of the two regions of the predicted box and the ground truth box by the union part of the two regions; b, b gt are the center points of the predicted box and the ground truth box respectively; ρ 2 (b, b gt ) is the Euclidean distance between the center points of the predicted box and the ground truth box; c is the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box; α is a balancing parameter and does not participate in the gradient calculation; υ is a parameter used to measure the aspect ratio consistency; w and h are the width and height of the predicted box respectively; w gt , h gt are the width and height of the ground truth box respectively; The EIOU loss function calculates the length and width of the target box separately, can solve the problem that the GIOU loss function has large errors in the horizontal and vertical directions, and can improve the convergence speed and regression accuracy.

[0071] In addition, the feature extraction method in the above S21 is as follows: Given an intermediate feature map F ∈ R C×H×W as the input, first perform global max pooling and average pooling on the input by channel, send the two one-dimensional vectors after pooling into the fully connected layer for calculation and then add them to generate a one-dimensional channel attention M C ∈ R C×1×1 , then multiply the channel attention by the input elements to obtain the feature map F' adjusted by the channel attention. Secondly, perform global max pooling and average pooling on F' by space, splice the two two-dimensional vectors generated by pooling and then perform a convolution operation to finally generate a two-dimensional spatial attention M S ∈ R 1×H×W , and then multiply the spatial attention by F' element by element.

[0072] Regarding the specific implementation of the C2f_DCNV2 module in S21, it includes the following key steps:

[0073] S211: Feature transformation, perform feature transformation on the input data through two convolutional layers (cv1 and cv2) and extract features at different levels and degrees of abstraction;

[0074] By using the two convolutional layers cv1 and cv2, the input data can be feature-transformed, so that features at different levels of abstraction can be extracted.

[0075] S212: Branch processing. The input data is divided into two branches for processing. One branch is directly passed to the output, and the other branch is processed through multiple Bottleneck modules to increase the network's non - linear ability and representation ability.

[0076] First, the input data is divided into two branches and processed separately. One branch is directly passed to the output, and the other branch is processed through multiple Bottleneck modules, thereby increasing the network's non - linear ability and representation ability.

[0077] S213: Feature fusion. The features of different branches are concatenated in the channel dimension to enrich the feature expression ability.

[0078] The above - mentioned different branch features are concatenated in the channel dimension. After concatenation, the feature expression ability can be enriched.

[0079] Furthermore, the above - mentioned S22 specifically includes the following steps:

[0080] S221: Input feature decomposition and global pooling, which is used to process the input feature map and perform global pooling in both the height and width directions respectively.

[0081] The feature is input and decomposed as well as globally pooled. Then the input feature map is processed, and then global pooling is performed in both the height and width directions respectively. These two steps respectively aggregate the global information in the width and height directions, obtaining two one - dimensional feature vectors, which respectively represent the spatial information in the height and width directions.

[0082] S222: Information aggregation and encoding, which is used to perform feature encoding on the feature vectors in the height and width directions through a shared convolutional layer.

[0083] The feature vectors in the height and width directions are encoded through a shared convolutional layer. The role of this shared convolutional layer is to uniformly encode the feature vectors in the height and width directions to extract more abstract features.

[0084] S223: Generate direction attention weights, which is used to perform convolutions on the aggregated feature vectors in the height and width directions respectively to generate two different - direction attention weights.

[0085] The aggregated feature vectors are convolved in both the height and width directions respectively, thereby generating two different - direction attention weights. These two convolution operations ensure that the attention information in the height and width directions is not confused and can maintain the independence of spatial information.

[0086] S224: Weight application and feature recalibration are used to apply the generated attention weights in the height and width directions back to the original feature map respectively and recalibrate the features in each direction;

[0087] First, apply the generated attention weights in the height and width directions back to the original feature map respectively, and then recalibrate the features in each direction.

[0088] Specifically, S30: Train the improved YOLOv8 algorithm;

[0089] After the YOLOv8 algorithm is improved, train it again.

[0090] Furthermore, S30 also includes the following steps:

[0091] S31: Input the paths of the training set and validation set and the path of the initial training weight file into the parameters required for network model training for training to obtain the improved yolov8 model after training;

[0092] First, input the paths of the training set and validation set into the parameters required for network model training, then input the path of the initial training weight file into the parameters required for network model training, and then perform training, so as to obtain the improved yolov8 model after training.

[0093] Specifically, S40: Perform image preprocessing on the detected image and output the object detection result;

[0094] First, detect the pavement crack image, then perform preprocessing, and finally output the current detection result after recognition.

[0095] Furthermore, S40 includes the following steps:

[0096] S41: Output the object recognition result, including 4 results of longitudinal crack, transverse crack, reticular crack and pothole, written in Python language, verified the correctness of the code on a PC with PyCharm and connected to a camera for image analysis to conduct detection experiments to verify the effectiveness of the method;

[0097] Finally, output the object recognition result, and there are 4 object recognition results, namely longitudinal crack, transverse crack, reticular crack and pothole; Therefore, in order to verify the effectiveness of the method, detection experiments can be conducted on a PC and an embedded device, written in Python language, first verify the correctness of the code on a PC with PyCharm, if the code is correct, finally connect to a camera for image analysis to complete the judgment and positioning of the picture category.

[0098] The above embodiments are only used to illustrate the present invention and do not limit the technical solutions described in the present invention. Although the present specification has described the present invention in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation manners. Therefore, any modification or equivalent replacement of the present invention; all technical solutions and their improvements that do not depart from the spirit and scope of the invention are covered by the scope of the claims of the present invention.

Claims

1. A pavement crack detection method based on improved YOLOV8, characterized in that, It includes the following steps: Obtain the data of pavement cracks to be processed, preprocess the data, and enhance the preprocessed data during the training phase; Adopt the C2f_DCNV2 module to improve the backbone network of the model, extract and fuse features from the data. The CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially while maintaining low computational and parameter costs; Train the improved YOLOv8 algorithm; Preprocess the detection image and output the object detection result.

2. The pavement crack detection method based on improved YOLOV8 according to claim 1, characterized in that, The steps of obtaining the data of pavement cracks to be processed, preprocessing the data, and enhancing the preprocessed data during the training phase include: Sort out the obtained pavement crack detection data; Randomly read four pictures from the data during the training phase, and perform flipping, scaling, color gamut transformation, and random cropping operations on the four obtained pictures; Arrange the processed photos with the first picture in the upper left, the second picture in the lower left, the third picture in the lower right, and the fourth picture in the upper right. Use a matrix method to intercept the fixed areas of the four pictures and splice them into a new picture containing a series of bounding boxes and label information; Divide the preprocessed dataset into a training set, a validation set, and a test set according to a given ratio, and the ratio of the training set, the validation set, and the test set is 7:2:

1.

3. The pavement crack detection method based on improved YOLOV8 according to claim 2, characterized in that, In the step of adopting the C2f_DCNV2 module to improve the backbone network of the model, extract and fuse features from the data, and the CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially while maintaining low computational and parameter costs, add the C2f_DCNV2 module. The C2f_DCNV2 module enhances the receptive field and feature extraction ability by integrating the deformable convolution DCNv2 into the C2f module of the backbone network to improve the performance of the model in object detection and image segmentation tasks.

4. The pavement crack detection method based on improved YOLOV8 according to claim 3, characterized in that, The steps of adopting the C2f_DCNV2 module to improve the backbone network of the model, extract and fuse features from the data, and the CoordAtt attention mechanism enhances the ability of the convolutional neural network (CNN) to capture important features across channels and spatially while maintaining low computational and parameter costs include: Input the data into the improved backbone network, and perform feature extraction through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module in sequence; The CoordAtt attention mechanism embeds spatial coordinate information (in the height and width directions) into the channel attention mechanism, giving the model attention to important features at different positions. Among them, CoordAtt is used to make up for the defect caused by the traditional attention mechanism ignoring the spatial structure of the feature map; Use the EIOU regression loss function to solve the problems existing in CIOU. Based on the penalty term of CIOU, the influence factors of the aspect ratio of the candidate box and the actual target area are split, and the aspect ratio values of the candidate box and the actual target area are calculated separately. The expression formula of the EIOU regression loss function is as follows:

5. A road surface crack detection method based on improved YOLOV8 according to claim 4, characterized in that, The c ω and c h are the lengths and widths of the closed candidate box and the minimum bounding rectangle of the actual area. The formula contains three parts: the overlap loss L IOU between the candidate box and the actual area, the center distance loss L dis between the candidate box and the actual area, and the length-width loss L asp between the candidate box and the actual area. loU is the anchor box loss function, and its result is obtained by dividing the overlapping part of the predicted box and the ground truth box by the union part of the two areas; b, b gt are the center points of the predicted box and the ground truth box respectively; ρ 2 (b, b gt ) is the Euclidean distance between the center points of the predicted box and the ground truth box; c is the diagonal distance of the smallest closed area that can contain both the predicted box and the ground truth box; α is a balance parameter that does not participate in the gradient calculation; υ is a parameter used to measure the aspect ratio consistency; w, h are the width and height of the predicted box respectively; w gt , h gt are the width and height of the ground truth box respectively. The EIOU loss function calculates the length and width of the target box separately, solves the problem that the GIOU loss function has large errors in the horizontal and vertical directions, and can improve the convergence speed and regression accuracy.

6. The pavement crack detection method based on improved YOLOV8 according to claim 4, characterized in that, The way of feature extraction in the step of inputting the data into the improved backbone network and successively passing through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module is as follows: Given an intermediate feature map \(F\in\mathbb{R}\) C×H×W as the input, first perform global max pooling and average pooling on the input by channel, send the two resulting one-dimensional vectors after pooling into a fully connected layer for operation and then add them together to generate a one-dimensional channel attention \(M\) C \(\in\mathbb{R}\) C ×1×1 , then multiply the channel attention with the input elements to obtain the feature map \(F'\) adjusted by channel attention. Secondly, perform global max pooling and average pooling on \(F'\) by space, concatenate the two resulting two-dimensional vectors after pooling and then perform a convolution operation to finally generate a two-dimensional spatial attention \(M\) S \(\in\mathbb{R}\) 1×H×W , and then multiply the spatial attention with \(F'\) element-wise.

7. A road surface crack detection method based on improved YOLOV8 according to claim 6, characterized in that, The implementation process of the C2f_DCNV2 module in the step of inputting the data into the improved backbone network and successively passing through the Conv module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, Conv module, C2f_DCNV2 module, and SPPF module for feature extraction includes the following steps: Feature transformation, using two convolutional layers (cv1 and cv2) to perform feature transformation on the input data and extract features at different levels and abstraction degrees; Branch processing, dividing the input data into two branches for processing. One branch is directly passed to the output, and the other branch is processed through multiple Bottleneck modules to increase the non-linearity ability and representation ability of the network; Feature fusion, splicing the features of different branches in the channel dimension to enrich the expression ability of the features.

8. A pavement crack detection method based on improved YOLOV8 according to claim 4, characterized in that, The steps in which the CoordAtt attention mechanism embeds spatial coordinate information (in the height and width directions) into the channel attention mechanism, enabling the model to focus on important features at different positions. The steps for CoordAtt to make up for the defects caused by the traditional attention mechanism ignoring the spatial structure of the feature map include: Input feature decomposition and global pooling, used to process the input feature map and perform global pooling in the height and width directions respectively; Information aggregation and encoding, used to perform feature encoding on the feature vectors in the height and width directions through a shared convolutional layer; Generating directional attention weights, used to perform convolutions on the aggregated feature vectors in the height and width directions respectively to generate two different directional attention weights; Weight application and feature recalibration, used to apply the generated attention weights in the height and width directions back to the original feature map respectively and recalibrate the features in each direction.

9. A pavement crack detection method based on improved YOLOV8 according to claim 1, characterized in that, The steps for training the improved YOLOv8 algorithm include: Input the paths of the training set and validation set and the path of the initial training weight file into the parameters required for network model training to obtain the trained improved yolov8 model.

10. A pavement crack detection method based on improved YOLOV8 according to claim 1, characterized in that, The steps for preprocessing the detection image and outputting the object detection result include: Output the target recognition results, including four types of results: longitudinal cracks, transverse cracks, reticular cracks, and potholes. It is written in Python language, and the correctness of the code is verified on a PC using PyCharm. The camera is connected for image analysis to conduct detection experiments to verify the effectiveness of the method.

Citation Information

Cited By

  • Ground penetrating radar image crack identification method and system based on frequency domain enhanced YOLOv11

    CN121837191A