Road disease detection method based on computer vision
By introducing the ISSP module in the Backbone layer of the Yolo v5 network, the model's ability to distinguish size and targets is improved, the problem of low detection accuracy of existing models is solved, and more efficient and accurate road disease detection is achieved.
Patent Information
- Application Number
- CN202510102676.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing road disease detection model based on Yolo v5 network has difficulties in distinguishing large targets and small targets, resulting in low detection accuracy.
The Backbone layer of Yolo v5 network was improved, and the ISSP module was introduced. This module enhances the receptive field of the output feature map through feature segmentation fusion module, multi-layer convolution operation and spatial pyramid pooling module, thereby improving the model's ability to distinguish size and targets.
Through the improved Yolo v5 network model, the accuracy and detection effect of road disease detection are significantly improved, and the large and small targets can be more effectively distinguished, improving detection efficiency and accuracy.
Smart Images

Figure CN120013908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a road damage detection method, and in particular to a road damage detection method based on computer vision. Background Art
[0002] Road disease detection is crucial in road maintenance. The main method of traditional manual detection of road diseases is for personnel to observe the road surface condition, and then manually record the location information, degree, type, etc. of the disease, and then calculate the damage indoors. This detection method requires huge workload for inspectors and is prone to errors. In addition, manual detection has disadvantages such as high cost, low accuracy, and traffic impact. With the maturity of deep learning technology, related technologies have also begun to be applied in the field of road disease detection.
[0003] Deep learning technology has been applied to almost every field of computer vision, such as target detection, image segmentation, super-resolution reconstruction and face recognition, and has broad application prospects in products such as image search, autonomous driving, user behavior analysis, text recognition, virtual reality and lidar. Computer vision based on deep learning technology can also have a profound impact on the field of road disease detection, realizing automatic recognition and automatic calculation in the process of road disease detection, greatly improving the accuracy of detection results.
[0004] Existing computer vision-based detection generally builds a detection model based on the Yolo network, and realizes the detection function after training it with a large amount of data. Among them, the Yolo v5 network is widely used in artificial intelligence detection due to its fast and lightweight characteristics. In the existing Yolo v5 network, the output end of its Backbone layer (backbone network layer) is provided with an SPP module (spatial pyramid pooling module), which uses a number of different-sized maximum pooling convolution kernels to input feature map pools, and then splices the outputs of each maximum pooling through the Concat module to convert feature maps of any size into feature vectors of fixed size, so that it can process input images of different sizes and avoid the loss of accuracy and computational redundancy caused by image scaling. However, the existing SPP module only performs a single convolution operation before pooling and after splicing, which limits the receptive field of its output feature map, making it difficult for the model to distinguish between large and small targets, thereby reducing the detection accuracy of the model. To this end, the existing Yolo v5 network needs to be improved to improve the model's ability to distinguish between large and small targets, thereby improving its detection accuracy. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a road damage detection method based on computer vision which is based on an improved Yolo v5 network model to achieve efficient and automated detection with high detection accuracy.
[0006] The road damage detection method based on computer vision of the present invention comprises the following steps: Collect road pictures through patrol vehicles equipped with cameras; Input road images into the detection model based on the improved Yolo v5 network; The improved Yolo v5 network includes a Backbone layer, a Bottleneck layer and a Head layer; The Backbone layer is provided with an ISSP module, the ISSP module includes a feature segmentation fusion module, the feature segmentation fusion module includes an input convolution layer, an SPP module and an output convolution layer connected in sequence, the input convolution layer includes at least two Conv modules connected in sequence, the SPP module includes a first Concat module and at least three maximum pooling modules arranged in parallel between the output end of the input convolution layer and the input end of the first Concat module, the input end of the first Concat module is also connected to the output end of the input convolution layer, the output end of the first Concat module is connected to the input end of the output convolution layer, and the output convolution layer includes at least two Conv modules connected in sequence; The ISPP module also includes a second Concat module. The output end of the feature segmentation and fusion module is connected to the input end of the second Concat module. The input end and output end of the second Concat module are also connected to Conv modules respectively. The input end of the Conv module located at the input end of the second Concat module is connected to the input end of the input convolution layer.
[0007] The advantage of the road disease detection method based on computer vision is that it realizes automatic detection of road images to be inspected through a detection model based on an improved Yolo v5 network, and its detection efficiency and accuracy are greatly improved compared with manual detection methods. At the same time, compared with the detection method based on the traditional Yolo v5 detection model, the improved Yolo v5 network adopted in this method includes an ISSP module. Compared with the existing SSP module, an input convolution layer and an output convolution layer are respectively set at the input end and the output end. The input convolution layer and the output convolution layer both include at least two Conv modules, which enables the ISPP module to perform multiple convolution operations on the input feature map, thereby improving the receptive field of the output feature map, so that it can effectively distinguish large targets from small targets, thereby improving the detection effect and accuracy of the model.
[0008] Furthermore, in the computer vision-based road damage detection method of the present invention, the input convolution layer includes three Conv modules, and the output convolution layer includes two Conv modules.
[0009] Furthermore, the road damage detection method based on computer vision of the present invention further includes the step of training the detection model before detecting the road image through the detection model. The method of training the detection model includes the following steps: Collect more than 10,000 frames of road disease images through patrol vehicles equipped with cameras, and annotate the images; Change the size of the road damage image to 608*608, and input the image data into the detection model for training; Filter the candidate box set extracted by the detection model and output the result set; Determine whether the output meets the preset accuracy requirements based on the result set. If so, keep the model. Otherwise, repeat the above steps until the result set meets the preset accuracy requirements. The method for screening the candidate box set extracted by the detection model and outputting the result set comprises the following steps: Step 1: Eliminate candidate boxes whose loss function is greater than a predetermined value; Step 2: Get the candidate box with the highest confidence, record it as candidate box M and add it to the result set, and remove it from the candidate box set; Step 3: Get the confidence optimization value of each candidate box in the candidate box set, and remove all candidate boxes with confidence optimization values of 0 in the candidate box set. If the candidate box set is not empty after removal, return to step 2, otherwise output the result set; The confidence optimization value s of the candidate box is obtained by the following formula; ; Among them, M represents the candidate box M with the highest confidence in step 2, b i represents the i-th candidate box in the candidate box, is the preset threshold, IoU is the candidate box b i The intersection-over-union ratio with the candidate box M, s i is the candidate box b i Confidence level; R NIou is the penalty factor, and its value is obtained by the following formula; ; Among them, b represents the candidate box b i The center point, b gt represents the center point of the candidate box M, and ρ represents the center point between point b and point b gt The distance between them, c is the distance containing the candidate box b i And the diagonal length of the minimum rectangle of the candidate box M.
[0010] In this training method, the candidate box screening method is based on R NIoUThe NMS algorithm is improved, which not only considers the overlapping area of the candidate boxes, but also takes into account the distance between the center points of the two boxes by introducing a penalty factor, so that in the case of non-overlap, the candidate box will also move toward the target box, thereby improving the suppression effect on duplicate targets. For targets that are too close, due to the introduction of the penalty factor, the intersection-over-union ratio is reduced, thereby improving the detection accuracy.
[0011] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it in accordance with the contents of the specification, the embodiments of the present invention are described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is the improved Yolo v5 network structure diagram.
[0013] Figure 2 This is the network structure diagram of the ISPP module.
[0014] In the figure, the input convolution layer is 1, the SPP module is 2, the output convolution layer is 3, the first Concat module is 4, the maximum pooling module is 5, and the second Concat module is 6. DETAILED DESCRIPTION
[0015] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0016] Example 1: See Figures 1 to 2 ,The road disease detection method based on computer vision of this embodiment includes the following steps; Collect road pictures through patrol vehicles equipped with cameras; Input road images into the detection model based on the improved Yolo v5 network; The improved Yolo v5 network includes a Backbone layer, a Bottleneck layer and a Head layer; The Backbone layer is provided with an ISSP module, which includes a feature segmentation and fusion module, which includes an input convolution layer 1, an SPP module 2 and an output convolution layer 3 connected in sequence, the input convolution layer includes at least two Conv modules connected in sequence, the SPP module includes a first Concat module 4 and at least three maximum pooling modules 5 arranged in parallel between the output end of the input convolution layer and the input end of the first Concat module, the input end of the first Concat module is also connected to the output end of the input convolution layer, the output end of the first Concat module is connected to the input end of the output convolution layer, and the output convolution layer includes at least two Conv modules connected in sequence; The ISPP module also includes a second Concat module. The output end of the feature segmentation and fusion module is connected to the input end of the second Concat module. The input end and output end of the second Concat module are also connected to Conv modules respectively. The input end of the Conv module located at the input end of the second Concat module is connected to the input end of the input convolution layer.
[0017] The road disease detection method based on computer vision realizes automatic detection of road images to be inspected through a detection model based on an improved Yolo v5 network, and its detection efficiency and accuracy are greatly improved compared with the manual detection method. At the same time, compared with the detection method based on the traditional Yolo v5 detection model, the improved Yolo v5 network adopted in this method includes an ISSP module. Compared with the existing SSP module, an input convolution layer and an output convolution layer are respectively set at the input end and the output end. The input convolution layer and the output convolution layer both include at least two Conv modules, which enables the ISPP module to perform multiple convolution operations on the input feature map, thereby improving the receptive field of the output feature map, so that it can effectively distinguish large targets from small targets, thereby improving the detection effect and accuracy of the model.
[0018] The Backbone layer is used to extract image features, the Bottleneck layer is used to further extract features and enhance images, and the Head layer is used for target detection.
[0019] The input end of each maximum pooling module is respectively connected to the Conv module at the tail end of the input convolutional layer, and the output end is respectively connected to the input end of the first Concat module.
[0020] During operation, the feature map is input into the second Concat module through the Conv module at the second Concat input end and the feature segmentation and fusion module respectively, and then is spliced by the second Concat module and output through the Conv module at its output end.
[0021] The Conv module is a convolution module, which is used to perform convolution operations on the input feature map, the maximum pooling module is used to perform maximum pooling operations on the feature map, and the Concat module is used to perform splicing operations on several feature maps.
[0022] The Backbone layer includes a Focus module, a CBL module, a CSP1_1 module, a CBL module, a CSP1_3 module, a CBL module, an ISSP module and a CSP2_1 module connected in sequence. The Focus module, the CBL module, each CSP module, the Conv module, the Concat module, the Bottleneck layer and the Head layer are all existing functional modules of Yolo v5. Their specific network structure and functions will not be repeated here.
[0023] Preferably, the input convolution layer includes three Conv modules, and the output convolution layer includes two Conv modules.
[0024] Preferably, before detecting the road image by using the detection model, the method further includes the step of training the detection model, and the method of training the detection model includes the following steps: Collect more than 10,000 frames of road disease images through patrol vehicles equipped with cameras, and annotate the images; When annotating, the image labels include the following nine types: Corrugation, Crocodile_Cracks, Longitudinal_Cracks, Pothole, Slippage, Transverse_Cracks, Patching, Looseness, Block_Cracks; Change the size of the road damage image to 608*608, and input the image data into the detection model for training; Filter the candidate box set extracted by the detection model and output the result set; Determine whether the output meets the preset accuracy requirements based on the result set. If so, keep the model. Otherwise, repeat the above steps until the result set meets the preset accuracy requirements. The method for screening the candidate box set extracted by the detection model and outputting the result set comprises the following steps: Step 1: Eliminate candidate boxes whose loss function is greater than a predetermined value; Step 2: Get the candidate box with the highest confidence, record it as candidate box M and add it to the result set, and remove it from the candidate box set; Step 3: Get the confidence optimization value of each candidate box in the candidate box set, and remove all candidate boxes with confidence optimization values of 0 in the candidate box set. If the candidate box set is not empty after removal, return to step 2, otherwise output the result set; The confidence optimization value s of the candidate box is obtained by the following formula; ; Among them, M represents the candidate box M with the highest confidence in step 2, b i represents the i-th candidate box in the candidate box set, is the preset threshold, IoU is the candidate box b i The intersection-over-union ratio with the candidate box M, s i is the candidate box b i Confidence level; R NIou is the penalty factor, and its value is obtained by the following formula; ; Among them, b represents the candidate box b i The center point, b gt represents the center point of the candidate box M, and ρ represents the center point between point b and point b gt The distance between them, c is the distance containing the candidate box b i And the diagonal length of the minimum rectangle of the candidate box M.
[0025] In the existing detection model based on Yolo v5, the Head layer performs multi-scale target detection on the feature map extracted by the backbone network to extract candidate boxes, and then screens out duplicate targets through the NMS (non-maximum suppression) algorithm. The NMS algorithm mainly screens candidate boxes through IoU (intersection over union). In this screening method, when the target is too close, the NMS algorithm will directly delete the detection box that exceeds the set threshold, resulting in a decrease in detection accuracy. When the target is far away, it will directly identify non-duplicate targets.
[0026] In this training method, the candidate box screening method is based on R NIoU The NMS algorithm is improved, which not only considers the overlapping area of the candidate boxes, but also takes into account the distance between the center points of the two boxes by introducing a penalty factor, so that in the case of non-overlap, the candidate box will also move toward the target box, thereby improving the suppression effect on duplicate targets. For targets that are too close, due to the introduction of the penalty factor, the intersection-over-union ratio is reduced, thereby improving the detection accuracy.
[0027] The above is only a preferred embodiment of the present invention, which is used to assist those skilled in the art to implement the corresponding technical solution, and is not used to limit the scope of protection of the present invention. The scope of protection of the present invention is defined by the attached claims. It should be pointed out that for those of ordinary skill in the art, on the basis of the technical solution of the present invention, several equivalent improvements and variations can be made, and these improvements and variations should also be regarded as the scope of protection of the present invention. At the same time, it should be understood that although this specification is described in accordance with the above-mentioned embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions of each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A road damage detection method based on computer vision, comprising the following steps; Collect road pictures through patrol vehicles equipped with cameras; Input road images into the detection model based on the improved Yolo v5 network; The improved Yolo v5 network includes a Backbone layer, a Bottleneck layer and a Head layer; The invention is characterized in that: the backbone layer is provided with an ISSP module, the ISSP module includes a feature segmentation fusion module, the feature segmentation fusion module includes an input convolution layer (1), an SPP module (2) and an output convolution layer (3) connected in sequence, the input convolution layer includes at least two Conv modules connected in sequence, the SPP module includes a first Concat module (4) and at least three maximum pooling modules (5) arranged in parallel between the output end of the input convolution layer and the input end of the first Concat module, the input end of the first Concat module is also connected to the output end of the input convolution layer, the output end of the first Concat module is connected to the input end of the output convolution layer, and the output convolution layer includes at least two Conv modules connected in sequence; The ISPP module also includes a second Concat module (6), the output end of the feature segmentation and fusion module is connected to the input end of the second Concat module, the input end and output end of the second Concat module are also respectively connected to a Conv module, and the input end of the Conv module located at the input end of the second Concat module is connected to the input end of the input convolution layer.
2. The road damage detection method based on computer vision according to claim 1, characterized in that: The input convolution layer includes three Conv modules, and the output convolution layer includes two Conv modules.
3. The road damage detection method based on computer vision according to claim 1, characterized in that: Before detecting the road image by using the detection model, the step of training the detection model is also included. The method for training the detection model includes the following steps: Collect more than 10,000 frames of road disease images through patrol vehicles equipped with cameras, and annotate the images; Change the size of the road damage image to 608*608, and input the image data into the detection model for training; Filter the candidate box set extracted by the detection model and output the result set; Determine whether the output meets the preset accuracy requirements based on the result set. If so, keep the model. Otherwise, repeat the above steps until the result set meets the preset accuracy requirements. The method for screening the candidate box set extracted by the detection model and outputting the result set comprises the following steps: Step 1: Eliminate candidate boxes whose loss function is greater than a predetermined value; Step 2: Get the candidate box with the highest confidence, record it as candidate box M and add it to the result set, and remove it from the candidate box set; Step 3: Get the confidence optimization value of each candidate box in the candidate box set, and remove all candidate boxes with confidence optimization values of 0 in the candidate box set. If the candidate box set is not empty after removal, return to step 2, otherwise output the result set; The confidence optimization value s of the candidate box is obtained by the following formula; ; Among them, M represents the candidate box M with the highest confidence in step 2, b i represents the i-th candidate box in the candidate box set, is the preset threshold, IoU is the candidate box b i The intersection-over-union ratio with the candidate box M, s i is the candidate box b i Confidence level; R NIou is the penalty factor, and its value is obtained by the following formula; ; Among them, b represents the candidate box b i The center point, b gt represents the center point of the candidate box M, and ρ represents the distance between point b and point b. gt The distance between them, c is the distance containing the candidate box b i And the diagonal length of the minimum rectangle of the candidate box M.