Pavement crack detection method based on SPTYolo
By improving the Yolov5 network, SPD-Conv, PConv and CT3 modules were introduced, combined with loss function optimization, the performance degradation caused by factors such as lighting and jitter in actual applications of the pavement crack detection model is solved, and efficient pavement crack detection is achieved.
Patent Information
- Application Number
- CN202311203605.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-08
AI Technical Summary
The existing road crack detection model is affected by factors such as light, vehicle jitter, and stains in actual engineering applications, resulting in a decline in performance. At the same time, the calculation quantity and detection speed have become bottlenecks, and the existing technology has not been effectively solved.
The pavement crack detection method based on SPTYolo, by improving the Yolov5 network, SPD-Conv, PConv and CT3 modules are introduced, combined with loss function optimization, the model's ability to extract fine-grained information, reduce the calculation amount and improve the detection speed.
The model's detection performance for jitter fuzzy cracks and small target cracks is improved, and the detection capability of large-span cracks is significantly enhanced, while reducing the calculation amount and improving the detection speed.
Smart Images

Figure CN120278936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and particularly to a road surface crack detection method based on SPTYolo. Background Art
[0002] In the practical engineering application of road surface crack detection, since the positions and sizes of cracks in pictures are always random, there are many random particle textures in the background, and the ratio of the area of the background to the area of the cracks in the image is always large. Some conventional object detection models do not perform well in this task. As a typical single-stage object detection network, the Yolo series of networks is very suitable for the needs of asphalt road surface crack detection due to its performance in detection scale, accuracy and speed. However, in practical engineering applications, the quality of asphalt road surface images is also affected by various factors such as lighting, vehicle jitter, stains, etc. during the acquisition process, resulting in common problems such as reflection, blurring, and shadows in the collected asphalt road surface crack images. These factors interact with the characteristics of asphalt road surface crack digital images, greatly affecting the performance of the model. In addition to the detection performance of the model, the computational amount and detection speed of the model are also important factors to be considered in engineering applications. Existing road surface crack detection models often only focus on improving the accuracy in experiments, while ignoring the hardware conditions in practical engineering applications.
[0003] Therefore, to solve the above technical problems, it is urgent to propose a new technical means. Summary of the Invention
[0004] In view of this, based on the Yolov5 series of networks, the present invention proposes a road surface crack detection method based on SPTYolo, aiming to reduce the influence of factors such as shadows, reflections, oil stains, jitter blurring, markings, repairs, etc. on the model performance during the acquisition process of road surface crack images, improve indicators such as the F1 value and mAP@0.5 of the model, and at the same time achieve an improvement in the detection speed FPS indicator and a decrease in the computational amount FOLPs of the model.
[0005] A road surface crack detection method based on SPTYolo provided by the present invention includes the following steps:
[0006] S1. Obtain a road surface crack image sample data set;
[0007] S2. Construct an SPTYolo detection model;
[0008] S3. Input the road surface crack image sample data set into the SPTYolo detection model for training;
[0009] S4. Determine whether the SPTYolo detection model is trained. If so, proceed to step S5. If not, update the parameters in the SPTYolo detection model and return to step S3;
[0010] S5. Input the image to be measured into the trained SPTYolo detection model, and output the target detection result image.
[0011] Furthermore, in step S2, the SPTYolo detection model is a model based on Yolov5, which is composed of a Backbone network, a Neck network, and a Head network. Among them, the output end of the Backbone network is connected to the input end of the Neck network, and the output end of the Neck network is connected to the input end of the Head network;
[0012] The Backbone network is composed of 5 SPD-Conv modules, 4 Conv modules, 4 PConv modules, and 1 CT3 module; the output end of the first SPD-Conv module is connected to the input end of the first Conv module, the output end of the first Conv module is connected to the input end of the second SPD-Conv module, the output end of the second SPD-Conv module is connected to the input end of the first PConv module, the output end of the first PConv module is connected to the input end of the second Conv module, the output end of the second Conv module is connected to the input end of the third SPD-Conv module, the output end of the third SPD-Conv module is connected to the input end of the second PConv module, the output end of the second PConv module is connected to the input end of the third Conv module, the output end of the third Conv module is connected to the input end of the fourth SPD-Conv module, the output end of the fourth SPD-Conv module is connected to the input end of the third PConv module, the output end of the third PConv module is connected to the input end of the fourth Conv module, the output end of the fourth Conv module is connected to the input end of the fifth SPD-Conv module, the output end of the fifth SPD-Conv module is connected to the input end of the fourth PConv module, and the output end of the fourth PConv module is connected to the input end of the CT3 module;
[0013] The Neck network consists of 6 Conv modules, 2 upsampling modules, 2 SPD modules, and 4 C3 modules; the output end of the first Conv module is connected to the input end of the first upsampling module, the output of the first upsampling module is concatenated with the output of the third PConv module in the Backbone network and input into the first C3 module, the output end of the first C3 module is connected to the input end of the second Conv module, the output end of the second Conv module is connected to the input end of the second upsampling module, the output of the second upsampling module is concatenated with the output of the second PConv module in the Backbone network and input into the second C3 module, the output end of the second C3 module is connected to the input end of the third Conv module, the output end of the third Conv module is connected to the input end of the first SPD module, the output of the first SPD module is concatenated with the output of the second Conv module and input into the fourth Conv module, the output end of the fourth Conv module is connected to the input end of the third C3 module, the output end of the third C3 module is connected to the input end of the fifth Conv module, the output end of the fifth Conv module is connected to the input end of the second SPD module, the output of the second SPD module is concatenated with the output of the first Conv module and input into the sixth Conv module, and the output end of the sixth Conv module is connected to the input end of the fourth C3 module;
[0014] The Head network adopts the Head network of the Yolov5 model.
[0015] Furthermore, in step S3, the parameters of the model are updated according to the loss calculated by the loss function, and the SDG optimizer is used to iteratively update the parameters;
[0016] The formula of the loss function is as follows:
[0017] LOSS = L IoU + L obj + L cls
[0018]
[0019] L obj = BCE obj (p o , p IoU )
[0020] L cls = BCE clsj (c o , c o )
[0021] Among them, L IoURepresents the regression loss of the prediction box, IoU represents the overlapping degree of the areas of the prediction box and the ground truth box, θ represents the angular loss, d represents the diagonal distance of the overlapping region between the prediction box and the ground truth box, y represents the ordinate of the center point of the prediction box, y gt represents the ordinate of the center point of the ground truth box, π represents pi, w represents the width of the prediction box, w gt represents the width of the ground truth box, h represents the height of the prediction box, h gt the height of the ground truth box;
[0022] L obj represents the object confidence loss, BCE represents the function for calculating the cross - entropy loss, p o represents the object confidence score in the prediction box, p IoU represents the IOU value between the prediction box and the corresponding ground truth box;
[0023] L cls represents the classification loss, c o is the predicted class score, c o is the ground truth class score.
[0024] Furthermore, in step S4, when MAP iterates 100 rounds or MAP no longer improves, the training is completed.
[0025] Advantages of the present invention: By enhancing the model's ability to extract fine - grained information, the present invention improves the model's detection performance for jitter - blurred cracks and small - target cracks; skillfully avoids redundant information in different channels of the same feature layer, performs convolution operations on some channels, and reduces the computational amount of the model; by enhancing the ability to capture the dependency relationship between image feature blocks, significantly enhances the model's detection performance for horizontal cracks with the characteristic of large span. The test results show that the present invention improves the Pr, Re, mAP, and F1 metrics while reducing the computational amount and improving the detection speed. Brief Description of the Drawings
[0026] The present invention will be further described below in conjunction with the drawings and embodiments:
[0027] Figure 1 is the flowchart of the present invention;
[0028] Figure 2 is the schematic diagram of the SPTYolo model of the present invention;
[0029] Figure 3 is the schematic diagram of the CT3 module of the present invention. Detailed Embodiments
[0030] The present invention will be further described below in conjunction with the accompanying drawings of the specification:
[0031] A pavement crack detection method based on SPTYolo provided by the present invention includes the following steps:
[0032] S1. Obtain a pavement crack image sample data set;
[0033] S2. Construct an SPTYolo detection model;
[0034] S3. Input the pavement crack image sample data set into the SPTYolo detection model for training;
[0035] S4. Determine whether the SPTYolo detection model is trained. If so, go to step S5. If not, update the parameters in the SPTYolo detection model and return to step S3;
[0036] S5. Input the image to be tested into the trained SPTYolo detection model and output the target detection result map. Through the above method, the ability of the model to extract fine-grained information can be enhanced, the detection performance of the model for jitter-blurred cracks and small target cracks, and the detection performance of the model for horizontal cracks with the characteristic of large span can be improved; and the calculation amount of the model is also reduced, and the calculation speed of the model is increased.
[0037] The SPD-Conv (Space-to-Depth Conv) module, PConv (Partial Convolution) module, and MHSA (Mult-Head Self-Attention) module mentioned in the present invention all adopt existing modules.
[0038] In this embodiment, in step S1, a road crack detection vehicle is used to collect pavement crack images of the highway pavement. More than 40,000 pavement images are collected in total, and by means of manual selection, 2,315 crack pictures with influencing factors such as shadow, reflection, oil stain, repair, marking, and jitter blur under different lighting conditions are selected from more than 40,000 pavement images as the data set of this embodiment.
[0039] In this embodiment, in step S2, the SPTYolo detection model is improved based on the Yolov5 model and consists of a Backbone network, a Neck network, and a Head network. Among them, the output end of the Backbone network is connected to the input end of the Neck network, and the output end of the Neck network is connected to the input end of the Head network, as Figure 2 shown;
[0040] The Backbone network consists of 5 SPD-Conv modules, 4 Conv modules, 4 PConv modules and 1 CT3 module; the output end of the first SPD-Conv module is connected to the input end of the first Conv module, the output end of the first Conv module is connected to the input end of the second SPD-Conv module, the output end of the second SPD-Conv module is connected to the input end of the first PConv module, the output end of the first PConv module is connected to the input end of the second Conv module, the output end of the second Conv module is connected to the input end of the third SPD-Conv module, the output end of the third SPD-Conv module is connected to the input end of the second PConv module, the output end of the second PConv module is connected to the input end of the third Conv module, the output end of the third Conv module is connected to the input end of the fourth SPD-Conv module, the output end of the fourth SPD-Conv module is connected to the input end of the third PConv module, the output end of the third PConv module is connected to the input end of the fourth Conv module, the output end of the fourth Conv module is connected to the input end of the fifth SPD-Conv module, the output end of the fifth SPD-Conv module is connected to the input end of the fourth PConv module, and the output end of the fourth PConv module is connected to the input end of the CT3 module;
[0041] In the present invention, first, an SPD-Conv module is added to the Backbone network in the original Yolov5 network. The SPD-Conv module splits the feature map at equal intervals in the spatial direction, then splices it in the channel direction, and finally adjusts the channels using a convolution with a convolution kernel size of 1*1 and a stride of 1, which can avoid information loss, reduce the computational amount of the model, and ensure the lightweight of the model;
[0042] Then, the C3 module in the Backbone network of the original Yolov5 network is replaced with a PConv module. In the PConv module, the calculation formula for the floating-point operations per second FLOPs1 is:
[0043]
[0044] And the calculation formula for the floating-point operations per second FLOPs2 of an ordinary Conv module is:
[0045] FLOPs2 = h × w × k 2 × c 2
[0046] where h represents the height of the input feature map, w represents the width of the input feature map, k represents the size of the convolution kernel, and c pn represents the number of channels for the convolutional operation of the input feature map parameters, and c represents the number of channels of the input feature map and the output feature map;
[0047] In this embodiment is 4. At this time, the FLOPs1 of PConv is only The PConv module can cleverly avoid the redundant information of different channels in the same feature layer, perform convolutional operations on some channels, and reduce the computational complexity of the model;
[0048] Finally, the CT3 module is added to the last layer of the Backbone, which can expand the receptive field of the Transformer and obtain the most general semantic features of the input asphalt pavement crack image. The CT3 module replaces the Bottleneck part of the C3 structure with a Transformer module with a multi-head attention mechanism MHSA (Mult-Head Self-Attention), as Figure 3 shown; The Transformer module includes two linear layers and an MHSA module. The output of the first linear layer is fused with the input of the first linear layer and input into the MHSA module. The output of the MHSA module and the input of the MHSA module are fused and input into the second linear layer. The output of the second linear layer and the input of the second linear layer are fused and output. Among them, the linear layer uses the Linear function;
[0049] The Neck network consists of 6 Conv modules, 2 upsampling modules, 2 SPD modules and 4 C3 modules; the output end of the first Conv module is connected to the input end of the first upsampling module, the output of the first upsampling module is concatenated with the output of the third PConv module in the Backbone network and input into the first C3 module, the output end of the first C3 module is connected to the input end of the second Conv module, the output end of the second Conv module is connected to the input end of the second upsampling module, the output of the second upsampling module is concatenated with the output of the second PConv module in the Backbone network and input into the second C3 module, the output end of the second C3 module is connected to the input end of the third Conv module, the output end of the third Conv module is connected to the input end of the first SPD module, the output of the first SPD module is concatenated with the output of the second Conv module and input into the fourth Conv module, the output end of the fourth Conv module is connected to the input end of the third C3 module, the output end of the third C3 module is connected to the input end of the fifth Conv module, the output end of the fifth Conv module is connected to the input end of the second SPD module, the output of the second SPD module is concatenated with the output of the first Conv module and input into the sixth Conv module, and the output end of the sixth Conv module is connected to the input end of the fourth C3 module;
[0050] The Head network in this application adopts the Head network of the Yolov5 network. The Head network of the Yolov5 network is a prior art and will not be elaborated here.
[0051] In this embodiment, in step S3, the road surface crack image sample data set is input into the SPTYolo detection model for training. The momentum is set to 0.937, the initial learning rate is 0.01, the cosine annealing algorithm is used to gradually decay the learning rate to 0.002, and the parameters of the model are updated according to the loss calculated by the loss function. The SDG optimizer is used to iteratively update the parameters;
[0052] The formula of the loss function is as follows:
[0053] LOSS = L IoU + L obj + L cls
[0054]
[0055] L obj = BCE obj (p o , p IoU )
[0056] L cls = BCE clsj(c o , c o )
[0057] Among them, L IoU represents the regression loss of the prediction box, IoU represents the overlapping degree of the areas of the prediction box and the ground truth box, θ represents the angle loss, d represents the diagonal distance of the overlapping region between the prediction box and the ground truth box, y represents the vertical coordinate of the center point of the prediction box, y gt represents the vertical coordinate of the center point of the ground truth box, π represents pi, w represents the width of the prediction box, w gt represents the width of the ground truth box, h represents the height of the prediction box, h gt the height of the ground truth box;
[0058] L obj represents the object confidence loss, BCE represents the function for calculating the cross-entropy loss, p o represents the object confidence score in the prediction box, p IoU represents the IOU value between the prediction box and the corresponding ground truth box;
[0059] L cls represents the classification loss, c o is the predicted class score, c o is the ground truth class score.
[0060] In this embodiment, in step S4, it is judged that the SPTYolo detection model MAP iterates 100 rounds or MAP no longer improves. If so, go to step S5. If not, update the parameters in the SPTYolo detection model and return to step S3. Here, MAP represents the average value of APs for all classes, and AP represents the accuracy of the predicted class. Among them, both MAP and AP are existing technologies and will not be elaborated here.
[0061] In this embodiment, in step S5, the image to be detected is input into the trained SPTYolo detection model, and the object detection result map is output.
[0062] In this embodiment, the road surface crack image sample dataset is used to test the performance of the SPTYolo model. Taking the classical models FasterRCNN-Resnet50 and SSD-VGG in the field of asphalt pavement crack detection and Yolov7-tiny, which is more advanced and has similar parameters compared with the Yolov5s version, as the control group, relevant indicators such as Pr (precision rate), Re (recall rate), mAP@0.5 (average value of 0.5-class AP), F1 (weighted average of Pr and Re), GFLOPs (number of floating-point operations per second in billions), and FPS (number of detections per second) are compared. The test results of each model for relevant indicators are shown in Table 1:
[0063]
[0064] Table 1 shows the relevant indicators of the test results of each model
[0065] According to Table 1, it can be concluded that except for the slightly lower performance of the SPTYolo model in Re compared with the Yolov7-tiny model, the SPTYolo model has a significant improvement in Pr, MAP@0.5%, F1, GFLOPs and FPS indicators compared with other models.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A road surface crack detection method based on SPTYolo, characterized in that: It includes the following steps: S1. Obtain a road surface crack image sample dataset; S2. Build an SPTYolo detection model; S3. Input the road surface crack image sample dataset into the SPTYolo detection model for training; S4. Determine whether the SPTYolo detection model is trained. If so, go to step S5. If not, update the parameters in the SPTYolo detection model and return to step S3; S5. Input the image to be tested into the trained SPTYolo detection model and output the target detection result image.
2. The pavement crack detection method based on SPTYolo according to claim 1, wherein: In step S2, the SPTYolo detection model is a model based on Yolov5, which consists of a Backbone network, a Neck network, and a Head network. Among them, the output end of the Backbone network is connected to the input end of the Neck network, and the output end of the Neck network is connected to the input end of the Head network; The Backbone network consists of 5 SPD-Conv modules, 4 Conv modules, 4 PConv modules, and 1 CT3 module; the output end of the first SPD-Conv module is connected to the input end of the first Conv module, the output end of the first Conv module is connected to the input end of the second SPD-Conv module, the output end of the second SPD-Conv module is connected to the input end of the first PConv module, the output end of the first PConv module is connected to the input end of the second Conv module, the output end of the second Conv module is connected to the input end of the third SPD-Conv module, the output end of the third SPD-Conv module is connected to the input end of the second PConv module, the output end of the second PConv module is connected to the input end of the third Conv module, the output end of the third Conv module is connected to the input end of the fourth SPD-Conv module, the output end of the fourth SPD-Conv module is connected to the input end of the third PConv module, the output end of the third PConv module is connected to the input end of the fourth Conv module, the output end of the fourth Conv module is connected to the input end of the fifth SPD-Conv module, the output end of the fifth SPD-Conv module is connected to the input end of the fourth PConv module, and the output end of the fourth PConv module is connected to the input end of the CT3 module; The Neck network consists of 6 Conv modules, 2 upsampling modules, 2 SPD modules, and 4 C3 modules. The output end of the first Conv module is connected to the input end of the first upsampling module. The output of the first upsampling module is concatenated with the output of the third PConv module in the Backbone network and input into the first C3 module. The output end of the first C3 module is connected to the input end of the second Conv module. The output end of the second Conv module is connected to the input end of the second upsampling module. The output of the second upsampling module is concatenated with the output of the second PConv module in the Backbone network and input into the second C3 module. The output end of the second C3 module is connected to the input end of the third Conv module. The output end of the third Conv module is connected to the input end of the first SPD module. The output of the first SPD module is concatenated with the output of the second Conv module and input into the fourth Conv module. The output end of the fourth Conv module is connected to the input end of the third C3 module. The output end of the third C3 module is connected to the input end of the fifth Conv module. The output end of the fifth Conv module is connected to the input end of the second SPD module. The output of the second SPD module is concatenated with the output of the first Conv module and input into the sixth Conv module. The output end of the sixth Conv module is connected to the input end of the fourth C3 module; The Head network adopts the Head network of the Yolov5 model.
3. The pavement crack detection method based on SPTYolo according to claim 1, characterized in that: In step S3, the parameters of the model are updated according to the loss calculated by the loss function, and the SDG optimizer is used to iteratively update the parameters; The formula of the loss function is as follows: LOSS = L IoU + L obj + L cls L obj = BCE obj (p o , p IoU ) L cls = BCE clsj (c o , c o ) Among them, L IoU represents the regression loss of the prediction box, IoU represents the overlapping degree of the areas of the prediction box and the ground truth box, θ represents the angular loss, d represents the diagonal distance of the overlapping region between the prediction box and the ground truth box, y represents the ordinate of the center point of the prediction box, y gt represents the ordinate of the center point of the ground truth box, π represents pi, w represents the width of the prediction box, w gt represents the width of the ground truth box, h represents the height of the prediction box, h gt the height of the ground truth box; L obj represents the target confidence loss, BCE represents the function for calculating the cross-entropy loss, and p o represents the target confidence score in the predicted bounding box, and p IoU represents the IOU value between the predicted bounding box and the corresponding ground truth bounding box; L cls represents the classification loss, c o is the predicted class score, c o is the true class score.
4. The pavement crack detection method based on SPTYolo according to claim 1, characterized in that: In step S4, when MAP iterates 100 rounds or MAP no longer improves, the training is completed.