Road defect detection method and device, computer equipment and storage medium

By introducing multi-scale feature fusion, attention mechanism, weighted bidirectional feature pyramid network and MPDIoU loss function in the YOLOv8 model, the problems of low road defect detection efficiency and unstable detection performance in the prior art are solved, and a more efficient, accurate and robust road defect detection effect is achieved.

CN120107231AInactive Publication Date: 2025-06-06ANHUI UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510311501.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as low efficiency, high cost, and easy to miss inspection in road defect detection, and traditional YOLO models are difficult to maintain stable detection performance in complex environments.

Method used

By introducing multi-scale feature fusion and attention mechanisms, the feature extraction capability of the YOLOv8 model is enhanced; combining the weighted bidirectional feature pyramid network and neck network to achieve efficient bidirectional cross-scale connection; replacing the original loss function as the MPDIoU loss function, and introducing an improved non-maximum suppression algorithm Soft-NMS.

Benefits of technology

The feature extraction accuracy and detection accuracy of the YOLOv8 model in complex backgrounds are improved, the stability and adaptability of the model are enhanced, the target loss problem caused by hard thresholds is reduced, and the accuracy and robustness of road defect detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107231A_ABST
    Figure CN120107231A_ABST
Patent Text Reader

Abstract

The invention discloses a road defect detection method and device, computer equipment and a storage medium, and belongs to the technical field of road defect detection. The method is improved based on an existing YOLOv8 model, and comprises the following steps: adding an attention mechanism module in a backbone network of the YOLOv8 model for enhancing feature representation; a weighted bidirectional feature pyramid network is combined with a neck network, so that more efficient bidirectional cross-scale connection is realized; a loss function in a head network is replaced with an MPDIOU loss function, and the problem that an existing loss function cannot be optimized when a prediction bounding box and a true value bounding box have the same length-width ratio but the height value and the width value are completely different is solved. The improved non-maximum suppression algorithm Soft-NMS is introduced, so that the method is more suitable for a scene with dense targets, and the accuracy and robustness of road defect detection are improved. According to the invention, by improving the YOLOv8 model, the operation speed is improved, and the detection precision is ensured at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road defect detection, and more specifically, to a road defect detection method, device, computer equipment and storage medium. Background Art

[0002] With the acceleration of urbanization and the increase in road traffic demand, the maintenance and management of road infrastructure has become increasingly important. Road defects such as cracks and potholes not only affect the comfort and safety of vehicle driving, but may also cause traffic accidents and casualties. Therefore, timely and effective detection and repair of road defects are crucial to ensuring public safety and the long-term use of infrastructure.

[0003] Traditional road defect detection methods mainly rely on manual inspections and simple mechanical inspections. These methods have many shortcomings. For example, manual inspections are inefficient and costly, and are subject to subjective judgment by personnel, making them prone to missed inspections and false detections. Although mechanical inspections increase speed, the equipment is expensive and has high environmental requirements, making them difficult to be widely used in complex urban roads.

[0004] With the continuous development of computer vision technology, deep learning algorithms have made significant progress in the field of target detection. As an advanced target detection algorithm, Yolov8 (You Only Look Once) has become the first choice in many fields due to its fast and accurate characteristics. However, road defects such as cracks and potholes have complex and diverse forms, and traditional YOLO models may not achieve the expected results in feature extraction and classification. In addition, the road environment is complex and changeable, and external factors such as light and weather will interfere with the detection results, further increasing the difficulty of detection.

[0005] After searching, the Chinese patent application, application number 202411517813.8, published on January 24, 2025, discloses a road defect detection method based on an improved YOLOv8s model. The method includes: collecting road images with different types of crack defects, inputting the collected road images into the sample model to generate defect samples; dividing the defect samples into a training set and a test set, and training the improved YOLOv8s model; collecting the road image to be detected and inputting the trained YOLOv8s model, the feature extraction network in the YOLOv8s model extracts the features of the road image to be detected; the feature fusion network in the YOLOv8s model fuses the features of the road image to be detected; the target detection network in the YOLOv8s model performs defect detection; the output network in the YOLOv8s model outputs the detection results to complete the road defect detection. This method improves the operation speed by improving the YOLOv8s model while ensuring the accuracy of detection, but this method does not take into account the complex and diverse forms of road defects, and its improved YOLOv8s model is difficult to maintain stable performance in complex environments. Summary of the invention

[0006] 1. Technical problems to be solved

[0007] In view of the shortcomings of the prior art, the present invention provides a road defect detection method, device, computer equipment and storage medium, which enhances the perception of defects of different scales by introducing multi-scale feature fusion, and adopts the attention mechanism to improve the feature extraction accuracy of the YOLOv8 model under complex backgrounds. In addition, by optimizing the loss function to make it more suitable for road defect detection tasks, the detection accuracy and stability of the YOLOv8 model are further improved.

[0008] 2. Technical solution

[0009] The purpose of the present invention is achieved through the following technical solutions.

[0010] A road defect detection method comprises the following steps:

[0011] Collect a road defect image dataset, perform data enhancement on the road defect image dataset, and divide the enhanced road defect image dataset into a training set, a validation set, and a test set;

[0012] An improved YOLOv8 model is provided, wherein the YOLOv8 model includes a backbone network, a neck network and a head network; wherein an attention mechanism module is added to the backbone network, a weighted bidirectional feature pyramid network is combined with the neck network, a loss function in the head network is replaced with an MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced;

[0013] Train the YOLOv8 model based on the training set, evaluate the YOLOv8 model during the training process based on the validation set, and obtain the trained YOLOv8 model;

[0014] The test set is input into the trained YOLOv8 model for road defect detection to obtain the detection results.

[0015] As a further improvement of the present invention, the backbone network includes a Conv convolutional layer, a C2f module, an attention mechanism module and an SPPF module; an attention mechanism module is added before the SPPF module.

[0016] As a further improvement of the present invention, the attention mechanism module includes a channel attention module and a spatial attention module; the backbone network performs feature extraction on the road defect image, and the channel attention module and the spatial attention module enhance feature representation to obtain a feature map.

[0017] As a further improvement of the present invention, the neck network includes a Conv convolutional layer, a C2f module, an Upsample module and a weighted bidirectional feature pyramid network; the weighted bidirectional feature pyramid network transmits information on the feature map through a path from the bottom layer to the high layer and from the high layer to the bottom layer, and fuses the feature map to obtain a fused feature map.

[0018] As a further improvement of the present invention, the weighted bidirectional feature pyramid network fuses the feature graphs, and the calculation formula is:

[0019]

[0020] Where m represents a natural number. represents the intermediate features of the mth layer, represents the input features of the input node of the mth layer, represents the input features of the input node of the m+1th layer, represents the output features of the mth layer, represents the output features of the output node of the m-1th layer, Conv represents the depth-separable convolutional layer, Resize represents the upsampling or downsampling operation, and w 1 、w 2 、w 3 、w 1 ′、w 2 ′ and w 3 ′ represent different weight parameters, and ε represents the bias term.

[0021] As a further improvement of the present invention, the calculation formula of the MPDIoU loss function is:

[0022]

[0023] Among them, A represents the real box, B represents the predicted box, w represents the width of the fused feature map, h represents the height of the fused feature map, and d 1 Indicates the distance between the upper left corners of the two boxes, d 2 represents the distance between the lower right corners of the two boxes, ∩ represents the intersection of the real box and the predicted box, and ∪ represents the union of the real box and the predicted box.

[0024] As a further improvement of the present invention, the detection frame output by the head network is processed by an improved non-maximum suppression algorithm Soft-NMS, and the specific steps include:

[0025] (1) Sorting candidate boxes: sort all candidate boxes in descending order according to their confidence;

[0026] (2) Traversing candidate boxes: starting from the candidate box with the highest confidence, process each candidate box one by one;

[0027] (3) Calculate overlap: Calculate the overlap between the current candidate box and other candidate boxes;

[0028] (4) Decay confidence: For each candidate box that overlaps with the current candidate box, adjust the confidence of the candidate box according to the IoU value. The confidence decay formula is:

[0029]

[0030] Among them, score i represents the confidence of candidate box i, IoU(i,j) represents the interaction ratio between candidate box i and candidate box j, and σ represents a preset attenuation factor;

[0031] (5) Continue to process the next candidate box: After processing the current candidate box, continue to process the next candidate box until all candidate boxes are processed;

[0032] (6) Return result: The candidate box retained after being processed by the improved non-maximum suppression algorithm Soft-NMS is output as the detection box.

[0033] A road defect detection device, comprising:

[0034] A data processing module collects a road defect image dataset, performs data enhancement on the road defect image dataset, and divides the enhanced road defect image dataset into a training set, a validation set, and a test set;

[0035] A model improvement module improves the YOLOv8 model, wherein the YOLOv8 model includes a backbone network, a neck network, and a head network; wherein an attention mechanism module is added to the backbone network, a weighted bidirectional feature pyramid network is combined with the neck network, the loss function in the head network is replaced with an MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced;

[0036] Model training module: trains the YOLOv8 model based on the training set, evaluates the YOLOv8 model during training based on the validation set, and obtains the trained YOLOv8 model;

[0037] The defect detection module inputs the test set into the trained YOLOv8 model to perform road defect detection and obtain the detection results.

[0038] A computer device comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements any of the above-mentioned methods when executing the computer program.

[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above methods is executed.

[0040] 3. Beneficial effects

[0041] Compared with the prior art, the advantages of the present invention are:

[0042] (1) A road defect detection method, apparatus, computer device and storage medium of the present invention add an attention mechanism module on the basis of the original YOLOv8 model, enhance feature representation through the attention mechanism module, thereby improving the performance and generalization ability of the YOLOv8 model. At the same time, the weighted bidirectional feature pyramid network is integrated with the neck network to achieve more efficient bidirectional cross-scale connection. In addition, the original loss function of the YOLOv8 model is replaced by the MPDIoU loss function, which can effectively solve the problem that the existing loss function cannot be optimized when the predicted bounding box and the true bounding box have the same aspect ratio but completely different height and width values. Finally, by introducing the improved non-maximum suppression algorithm Soft-NMS, compared with the traditional non-maximum suppression algorithm, the target loss problem caused by the hard threshold is reduced, which is more suitable for scenes with dense targets and improves the accuracy and robustness of road defect detection.

[0043] (2) A road defect detection method, device, computer equipment and storage medium of the present invention, by improving the YOLOv8 model, can achieve better results for road defect detection by using only lightweight detectors. The use of detectors with simpler structures in the road defect detection process can greatly reduce the resources required for calculation, improve the detection speed of the YOLOv8 model and reduce the video memory occupancy required in the detection process, avoid unnecessary waste, effectively utilize limited computing resources, and thus realize the practical application of the YOLOv8 model in road defect detection, which has strong practicality and wide applicability.

[0044] (3) The road defect detection method, device, computer equipment and storage medium of the present invention can be applied to the real-time monitoring and maintenance of urban roads and highways. By efficiently and accurately detecting defects on the road surface (e.g., cracks and potholes), the timeliness and accuracy of road maintenance can be improved, and reliable technical support can be provided to traffic management departments to ensure driving safety and long-term maintenance of infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flow chart of a method according to an embodiment of the present invention;

[0046] Figure 2 A schematic diagram of the structure of the improved YOLOv8 model according to an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of the structure of the attention mechanism module in an embodiment of the present invention;

[0048] Figure 4 A schematic diagram of a weighted bidirectional feature pyramid network structure according to an embodiment of the present invention;

[0049] Figure 5 This is a diagram of the detection results of an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0051] Example

[0052] like Figure 1As shown, a road defect detection method provided by this embodiment includes: collecting a road defect image dataset, performing data enhancement on the road defect image dataset, and dividing the data enhanced road defect image dataset into a training set, a validation set, and a test set; improving the YOLOv8 model, the YOLOv8 model includes a backbone network, a neck network, and a head network, wherein an attention mechanism module is added to the backbone network, the weighted bidirectional feature pyramid network is combined with the neck network, the loss function in the head network is replaced by the MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced; the YOLOv8 model is trained based on the training set, and the YOLOv8 model in the training process is evaluated based on the validation set to obtain a trained YOLOv8 model; the test set is input into the trained YOLOv8 model to perform road defect detection and obtain a detection result.

[0053] Specifically in this embodiment, a common road defect image dataset is collected, and the existing mosaic data enhancement algorithm is used to enhance the road defect image dataset. The data-enhanced road defect image dataset is annotated using the existing Labelimg tool and converted into the YOLO algorithm training format. In this embodiment, the areas of road defects are marked as D00, D10, D20, D40, D43, D44, D50 and Repair, representing longitudinal cracks, transverse cracks, cracks, potholes, blurred intersections, blurred white lines, manhole covers and repairs, respectively. Finally, the data-enhanced road defect image dataset is randomly divided into a training set, a validation set and a test set in a ratio of 8:1:1 using an existing python script.

[0054] Furthermore, if Figure 2 As shown, the improved YOLOv8 model. In the prior art, the YOLOv8 model includes a backbone network, a neck network and a head network. The backbone network is used for feature extraction, the neck network is used for feature fusion, and the head network is used for target detection. It is worth noting that in this embodiment, an attention mechanism module (Convolutional Block Attetion Module, CBAM) is added to the backbone network, the weighted bidirectional feature pyramid network (BiFPN_Add2) is combined with the neck network, the loss function in the head network is replaced by the MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced.

[0055] Specifically, in the backbone network of the YOLOv8 model, an attention mechanism module is added to enhance the convolutional neural network. Figure 3As shown, in this embodiment, the attention mechanism module includes a channel attention module (Channel Attetion Module, CAM) and a spatial attention module (Spatial Attetion Module, SAM). In this embodiment, an attention mechanism module is added before the SPPF module. The backbone network includes a Conv convolution layer, a C2f module, an attention mechanism module and an SPPF module. Furthermore, the structure of the backbone network is a Conv convolution layer, a Conv convolution layer, a C2f module, a Conv convolution layer, a C2f module, a Conv convolution layer, a C2f module, a Conv convolution layer, a C2f module, a Conv convolution layer, a C2f module, an attention mechanism module and an SPPF module connected in sequence. Thus, the road defect image data set is input into the backbone network of the YOLOv8 model, the backbone network performs feature extraction on the road defect image, and the channel attention module and the spatial attention module enhance the feature representation, thereby improving the performance and generalization ability of the YOLOv8 model, and finally obtaining a feature map.

[0056] Specifically, the backbone network is used to extract features from road defect images. The backbone network starts with a Conv convolution layer, using a 3×3 convolution kernel with 64 output channels and a stride of 2. This layer downsamples the input image to 1 / 2 of the original. Next, it enters a Conv convolution layer, using a 3×3 convolution kernel with 128 output channels and a stride of 2. This layer downsamples the input image to 1 / 4 of the original. Then, the backbone network repeats the C2f module 3 times in a row, and each C2f module uses 128 output channels. Next, the backbone network enters a Conv convolution layer, using a 3×3 convolution kernel with 256 output channels and a stride of 2. This layer downsamples the input image to 1 / 8 of the original. Then, the backbone network repeats the C2f module 6 times in a row, and each C2f module uses 256 output channels. Next, the backbone network enters a Conv convolution layer, using a 3×3 convolution kernel with 512 output channels and a step size of 2. This layer downsamples the input image to 1 / 16 of the original. Then, the backbone network repeats the C2f module for 6 consecutive times, and each C2f module uses 512 output channels. Next, the backbone network enters a Conv convolution layer, using a 3×3 convolution kernel with 1024 output channels and a step size of 2. This layer downsamples the input image to 1 / 32 of the original. Then, the backbone network repeats the C2f module for 3 consecutive times, and each C2f module uses 1024 output channels. Next, the backbone network uses the attention mechanism module to perform attention mechanism processing on the feature map of 1024 channels, and finally uses the SPPF module to perform spatial pyramid pooling processing on the feature map of 1024 channels to obtain the final feature map.

[0057] In this embodiment, the attention mechanism module first weights the features of each channel through the channel attention module, strengthens the features of important channels, and suppresses the features of unimportant channels. The channel attention module obtains the global information of each channel by performing global average pooling and maximum pooling operations on the input feature map, and then generates the weight of each channel through a multi-layer perceptron, and finally performs weighted processing on each channel. Next, the attention mechanism module uses the spatial attention module to assign a weight to each spatial position and pay attention to the key areas in the feature map. The spatial attention module generates global feature information for each spatial position by performing average pooling and maximum pooling operations on each channel, and then uses convolution operations to generate spatial attention weights. As a result, the channel attention module and the spatial attention module work together to enable the backbone network to focus on important features, thereby improving the performance of target detection tasks.

[0058] In the neck network of the YOLOv8 model, the weighted bidirectional feature pyramid network is combined with the neck network, so that the neck network includes a Conv convolution layer, a C2f module, an Upsample module and a weighted bidirectional feature pyramid network. In this embodiment, the Upsample module performs an upsampling operation. Figure 4 As shown in FIG. 1 , the weighted bidirectional feature pyramid network transmits information of the feature map through the path from the bottom layer to the high layer and from the high layer to the bottom layer, fuses the feature map, and outputs the fused feature map.

[0059] Specifically, the neck network receives the feature map from the backbone network and processes it to generate the prediction result of the target detection. First, the number of channels is adjusted using a 1×1 convolution kernel with 512 output channels. Next, the resolution of the feature map is increased by 2 times through the nearest neighbor upsampling operation. Then, the feature map (P4) output by the backbone network is fused with the current feature map using a weighted bidirectional feature pyramid network, using 128 channels. Then, the neck network repeats the C2f module 3 times in a row, each C2f module uses 512 output channels. Next, the number of channels is adjusted using a 1×1 convolution kernel with 256 output channels. The neck network increases the resolution of the feature map by 2 times through the nearest neighbor upsampling operation. Then, the feature map (P3) output by the backbone network is fused with the current feature map using a weighted bidirectional feature pyramid network, using 64 channels. The neck network repeats the C2f module 3 times in a row, each C2f module uses 256 output channels. Then, a 3×3 convolution kernel with 512 output channels and a step size of 2 is used to downsample the feature map. Next, a weighted bidirectional feature pyramid network is used to fuse the current feature map with the feature map (P4) output by the backbone network, using 128 channels. Then, the C2f module is repeated 3 times in a row, and each P4 module uses 512 output channels. Then, a 3×3 convolution kernel with 512 output channels and a step size of 2 is used to downsample the feature map. Then, a weighted bidirectional feature pyramid network is used to fuse the current feature map with the feature map (P5) output by the backbone network, using 128 channels. Then, the C2f module is repeated 3 times in a row, and each C2f module uses 1024 output channels. In this embodiment, the weighted bidirectional feature pyramid network fuses the feature maps, and the calculation formula is:

[0060]

[0061] Where m represents a natural number. represents the intermediate features of the mth layer, represents the input features of the input node of the mth layer, represents the input features of the input node of the m+1th layer, represents the output features of the mth layer, represents the output features of the output node of the m-1th layer, Conv represents the depth-separable convolutional layer, Resize represents the upsampling or downsampling operation, and w 1 、w 2 、w 3 、w 1 ′、w 2 ′ and w 3′ respectively represent different weight parameters, and ε represents a bias term. In this embodiment, the value range of m is [1,6]. Therefore, in this embodiment, the weighted bidirectional feature pyramid network is combined with the neck network to achieve a more efficient bidirectional cross-scale connection, and the feature graph is fused. It is worth noting that in this embodiment, on the basis of the existing functional pyramid network (FPN), the edge of context information is added, and each edge is multiplied by a corresponding weight to more effectively perform feature fusion.

[0062] Further, in the head network of the YOLOv8 model, the head network detects the fused feature map to obtain the detection result. In this embodiment, the Detect module is used to perform target detection, receive the feature map (P3, P4 and P5) output by the neck network, and output the detection result. In this embodiment, in the head network, the original loss function is replaced with the MPDIoU loss function. Therefore, in this embodiment, the head network adopts a decoupled head structure while combining the MPDIoU loss function. It is worth noting that the MPDIoU loss function is a loss function for boundary regression, which can effectively solve the problem that the existing loss function cannot be optimized when the predicted bounding box and the true bounding box have the same aspect ratio but the height and width values ​​are completely different. In this embodiment, the calculation formula of the MPDIoU loss function is:

[0063]

[0064] Among them, A represents the real box, B represents the predicted box, w represents the width of the fused feature map, h represents the height of the fused feature map, and d 1 Indicates the distance between the upper left corners of the two boxes, d 2 represents the distance between the lower right corners of the two boxes, ∩ represents the intersection of the real box and the predicted box, and ∪ represents the union of the real box and the predicted box.

[0065] In the YOLOv8 model, an improved non-maximum suppression algorithm Soft-NMS is introduced to process the target detection frame. The traditional non-maximum suppression algorithm (NMS) ensures the uniqueness of the detection result by discarding the bounding box with low overlap, while the improved non-maximum suppression algorithm Soft-NMS in this embodiment improves the adaptability to dense occlusion scenes by introducing a softened weight in the overlap calculation. Compared with the traditional non-maximum suppression algorithm, it reduces the problem of target loss caused by hard thresholds, is more suitable for scenes with dense targets, and improves the accuracy and robustness of detection. In this embodiment, the specific steps of the improved non-maximum suppression algorithm Soft-NMS for processing the detection frame include: (1) sorting candidate frames: sorting all candidate frames in descending order according to confidence; (2) traversing candidate frames: starting from the candidate frame with the highest confidence, processing each candidate frame one by one; (3) calculating overlap: calculating the overlap between the current candidate frame and other candidate frames; (4) attenuating confidence: for each candidate frame that overlaps with the current candidate frame, adjusting the confidence of the candidate frame according to the IoU value; (5) continuing to process the next candidate frame: after processing the current candidate frame, continue to process the next candidate frame until all candidate frames are processed; (6) returning the result: outputting the candidate frame retained after processing by the improved non-maximum suppression algorithm Soft-NMS as the detection frame. In this embodiment, the confidence attenuation formula in step (4) is:

[0066]

[0067] Among them, score i represents the confidence of box i, IoU(i,j) represents the interaction ratio of box i and box j, σ represents a preset attenuation factor, usually taking a small constant value (such as 0.5). If the attenuated confidence is less than a certain threshold, the box will be discarded.

[0068] In this embodiment, by introducing the improved non-maximum suppression algorithm Soft-NMS, it is possible to better handle the situation of overlapping targets, better retain the detection frame close to the real target, and improve the accuracy and robustness of target detection.

[0069] After obtaining the improved YOLOv8 model, the YOLOv8 model is trained based on the training set. In this embodiment, the existing GeForce RTX 4060gpu is used to train the YOLOv8 model, the hyperparameters are set to 300 epochs, the batch is set to 16, the input road defect image size is set to 640 pixels × 640 pixels, and the improved YOLOv8 model is used for training. The YOLOv8 model in the training process is evaluated based on the validation set to obtain a trained YOLOv8 model. The test set is input into the trained YOLOv8 model for road defect detection to obtain the detection results. Figure 5 As shown in the figure, it is the result of road defect detection. It can be seen that the improved YOLOv8 model has a stronger feature extraction ability for road defect images and shows good accuracy and robustness in the detection task. At the same time, the detection results also show that the improved YOLOv8 model has certain accuracy and reliability in road defect detection.

[0070] This embodiment also provides a road defect detection device, including a data processing module, a model improvement module, a model training module and a defect detection module. The data processing module collects a road defect image data set, performs data enhancement on the road defect image data set, and divides the road defect image data set after data enhancement into a training set, a validation set and a test set. The model improvement module improves the YOLOv8 model, which includes a backbone network, a neck network and a head network; wherein an attention mechanism module is added to the backbone network, a weighted bidirectional feature pyramid network is combined with the neck network, the loss function in the head network is replaced with an MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced. The model training module trains the YOLOv8 model based on the training set, and evaluates the YOLOv8 model in the training process based on the validation set to obtain a trained YOLOv8 model. The defect detection module inputs the test set into the trained YOLOv8 model to perform road defect detection and obtain a detection result. A road defect detection device provided in this embodiment can implement any of the road defect detection methods, and the specific working process of a road defect detection device can refer to the corresponding process in the road defect detection method embodiment. The method and device provided in this embodiment can be implemented in other ways. For example, the device embodiment described above is only schematic; for example, the division of a module is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual connection or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.

[0071] This embodiment also provides a computer device. A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the road defect detection method when executing the computer program.

[0072] This embodiment also provides a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a road defect detection method described in this embodiment is executed. The computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device or device; the program code contained in the computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0073] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited thereto, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all belong to the protection scope of the present invention. In addition, the word "including" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate the name, and do not indicate any specific order.

Claims

1. A road defect detection method, comprising the following steps: Collect a road defect image dataset, perform data enhancement on the road defect image dataset, and divide the enhanced road defect image dataset into a training set, a validation set, and a test set; An improved YOLOv8 model is provided, wherein the YOLOv8 model includes a backbone network, a neck network and a head network; wherein an attention mechanism module is added to the backbone network, a weighted bidirectional feature pyramid network is combined with the neck network, a loss function in the head network is replaced with an MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced; Train the YOLOv8 model based on the training set, evaluate the YOLOv8 model during training based on the validation set, and obtain the trained YOLOv8 model; The test set is input into the trained YOLOv8 model for road defect detection to obtain the detection results.

2. A road defect detection method according to claim 1, characterized in that: The backbone network includes a Conv convolutional layer, a C2f module, an attention mechanism module and an SPPF module; an attention mechanism module is added before the SPPF module.

3. A road defect detection method according to claim 2, characterized in that: The attention mechanism module includes a channel attention module and a spatial attention module; the backbone network extracts features from road defect images, and the channel attention module and the spatial attention module enhance feature representation to obtain a feature map.

4. A road defect detection method according to claim 3, characterized in that: The neck network includes a Conv convolution layer, a C2f module, an Upsample module and a weighted bidirectional feature pyramid network; the weighted bidirectional feature pyramid network transmits information on feature maps through paths from the bottom layer to the high layer and from the high layer to the bottom layer, and fuses the feature maps to obtain a fused feature map.

5. A road defect detection method according to claim 4, characterized in that: The weighted bidirectional feature pyramid network fuses the feature maps, and the calculation formula is: Where m represents a natural number. represents the intermediate features of the mth layer, represents the input features of the input node of the mth layer, represents the input features of the input node of the m+1th layer, represents the output features of the mth layer, represents the output features of the output node of the m-1th layer, Conv represents the depth-wise separable convolutional layer, Resize represents the upsampling or downsampling operation, w1, w2, w3, w1′, w2′ and w3′ represent different weight parameters respectively, and ε represents the bias term.

6. A road defect detection method according to claim 1, characterized in that: The calculation formula of the MPDIoU loss function is: Among them, A represents the real box, B represents the predicted box, w represents the width of the fused feature map, h represents the height of the fused feature map, d1 represents the distance between the upper left corners of the two boxes, d2 represents the distance between the lower right corners of the two boxes, ∩ represents the intersection of the real box and the predicted box, and ∪ represents the union of the real box and the predicted box.

7. A road defect detection method according to claim 1, characterized in that: The detection frame output by the head network is processed by the improved non-maximum suppression algorithm Soft-NMS. The specific steps include: (1) Sorting candidate boxes: sort all candidate boxes in descending order according to their confidence; (2) Traversing candidate boxes: starting from the candidate box with the highest confidence, process each candidate box one by one; (3) Calculate overlap: Calculate the overlap between the current candidate box and other candidate boxes; (4) Decay confidence: For each candidate box that overlaps with the current candidate box, adjust the confidence of the candidate box according to the IoU value. The confidence decay formula is: Among them, score i represents the confidence of candidate box i, IoU(i,j) represents the interaction ratio between candidate box i and candidate box j, and σ represents a preset attenuation factor; (5) Continue to process the next candidate box: After processing the current candidate box, continue to process the next candidate box until all candidate boxes are processed; (6) Return result: The candidate box retained after being processed by the improved non-maximum suppression algorithm Soft-NMS is output as the detection box.

8. A road defect detection device, characterized in that: include: A data processing module collects a road defect image dataset, performs data enhancement on the road defect image dataset, and divides the enhanced road defect image dataset into a training set, a validation set, and a test set; A model improvement module improves the YOLOv8 model, wherein the YOLOv8 model includes a backbone network, a neck network, and a head network; wherein an attention mechanism module is added to the backbone network, a weighted bidirectional feature pyramid network is combined with the neck network, the loss function in the head network is replaced with an MPDIoU loss function, and an improved non-maximum suppression algorithm Soft-NMS is introduced; Model training module: trains the YOLOv8 model based on the training set, evaluates the YOLOv8 model during training based on the validation set, and obtains the trained YOLOv8 model; The defect detection module inputs the test set into the trained YOLOv8 model to perform road defect detection and obtain the detection results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Road defect detection method based on improved YOLOv8s model

    CN119360212A

Cited By

  • A warehouse damage detection method embedding MDA and WBFF mechanisms

    CN122675806A