Road hidden disease detection method, device, storage medium and program product
By generating a hybrid dataset of virtual and real images using a diffusion model, a lightweight YOLO model is constructed. This optimizes ground-penetrating radar (GPR) defect detection, solves the problem of model training relying on large amounts of data, and achieves efficient detection of hidden road defects.
Patent Information
- Application Number
- CN202411769140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The training process of ground-penetrating radar target detection models relies heavily on a large amount of high-quality data. On-site data collection and verification are labor-intensive and time-consuming, resulting in poor economic costs and engineering efficiency.
Virtual images are generated by a diffusion model and mixed with real images to construct a lightweight YOLO model hybrid dataset, reducing the amount of real image acquisition. Combined with the C2fGhost module and P6 detection layer to optimize the model, an efficient disease detection system is generated.
This reduces the manpower and material resources required for real image acquisition, improves the robustness of the model and detection efficiency, and enables automated detection of hidden road defects.
Smart Images

Figure CN119723331B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to road engineering, geophysical exploration, and computer science and technology, and in particular to a method, device, storage medium, and program product for detecting hidden road defects. Background Technology
[0002] Under the combined effects of traffic loads and complex environments, road structures are prone to developing hidden defects such as cracks and voids, severely impacting road structural stability and posing traffic safety hazards. Currently, some non-destructive testing equipment is being used in hazard screening processes. Among them, ground-penetrating radar (GPR) is an effective and user-friendly technology that, compared to traditional destructive testing methods, has the advantages of minimal impact on road structures, high testing efficiency, and low labor costs. GPR has wide applications in road infrastructure operation and maintenance. It extracts the characteristics and specific locations of hidden defects in road structures based on the propagation patterns of transmitted and received electromagnetic waves in different media, presenting the raw data in the form of two-dimensional images.
[0003] In recent years, further developments in digital signal and image processing technologies have made it possible to detect objects in images. In particular, in the field of deep learning, which eliminates the need for manual extraction of target features, various convolutional neural network target detection models, represented by YOLOv8, have made rapid progress. Their combined application with vehicle-mounted ground-penetrating radar has gradually enabled the automated detection of hidden road defects.
[0004] However, the practical engineering application of ground-penetrating radar still faces the following serious challenges:
[0005] The training process of a series of target detection models, represented by YOLOv8, heavily relies on a large amount of high-quality data, while the collection and verification of ground-penetrating radar field data is still labor-intensive and time-consuming, and its economic cost and engineering efficiency are unsatisfactory. Summary of the Invention
[0006] This disclosure is made in view of the above-mentioned problems. This disclosure provides a method, apparatus, storage medium, and program product for detecting hidden road defects.
[0007] According to the first aspect of this disclosure, a method for detecting hidden road defects is provided, comprising:
[0008] Acquire the first number of images of hidden road defects in the target area using image acquisition equipment;
[0009] A virtual image for each of the hidden road defects images is generated using a diffusion model, and the labels of the virtual images are the same as the labels of each of the hidden road defects images.
[0010] Based on the virtual image, the road hidden defects image, the label of the virtual image, and the label of the road hidden defects image, a second number of mixed datasets are generated; each mixed dataset includes a training set, a validation set, and a test set, wherein the training set includes images obtained from the road hidden defects image and images obtained from the virtual image, and the images in the validation set and the test set are both obtained from the road hidden defects image;
[0011] In the process of processing each of the hybrid datasets using a pre-constructed lightweight YOLO model, the performance data of the lightweight YOLO model on the test set included in each of the hybrid datasets is obtained;
[0012] Based on the performance data, the optimal mixed dataset is obtained from the second number of mixed datasets;
[0013] The lightweight YOLO model and the optimal hybrid dataset are used to detect road defects in the area to be detected.
[0014] Furthermore, according to the road hidden defect detection method of the first aspect of this disclosure, a virtual image of each of the road hidden defect images is generated by a diffusion model, including:
[0015] A preset implicit diffusion method is used to diffuse each of the hidden road defects images in the latent space to obtain a virtual image of each of the hidden road defects images.
[0016] Furthermore, according to the road hidden defect detection method of the first aspect of this disclosure, a second number of mixed datasets are generated based on the virtual image, the road hidden defect image, the label of the virtual image, and the label of the road hidden defect image, including:
[0017] The hidden road defects images are divided into a second number of image sets. Each image set includes a first training image, a verification image, and a test image with labels. The sum of the number of the first training image, the verification image, and the test image is the first number. The number of the verification image and the test image are the same. The number of the first training image, the verification image, or the test image in different image sets is different.
[0018] Based on the preset mixed image ratio, and considering the number of first training images, verification images, and test images in each image set, the target number of second training images in each image set is calculated.
[0019] The second number of mixed datasets is obtained based on each of the image sets and the target number of second training images obtained from the labeled virtual images.
[0020] Furthermore, according to the road hidden defects detection method of the first aspect of this disclosure, before obtaining the second number of mixed datasets based on each of the image sets and the target number of second training images obtained from the labeled virtual images, the method further includes:
[0021] The pixel dimensions of the first and second training images belonging to each of the image sets are cropped to a preset size.
[0022] Furthermore, according to the method for detecting hidden road defects according to the first aspect of this disclosure, before acquiring a first number of images of hidden road defects collected by an image acquisition device, the method further includes:
[0023] Adjust the device parameters of the image acquisition device so that the road hidden defects image meets the requirements of zero-point correction, automatic gain control, and background removal.
[0024] Furthermore, according to the road hidden defects detection method of the first aspect of this disclosure, based on the performance data, an optimal mixed dataset is obtained from the second number of mixed datasets, including:
[0025] An average accuracy curve is generated based on the average accuracy of the lightweight YOLO model's performance data across the second number of mixed datasets; and a total time cost curve is generated based on the total time cost of the lightweight YOLO model's performance data across the mixed datasets; wherein the horizontal axis of both the average accuracy curve and the total time cost curve represents the ratio between road hidden defect images and virtual images in the training set included in the mixed datasets.
[0026] Obtain all x-axis values from the average accuracy curve that make the average accuracy of the average accuracy curve greater than a preset accuracy threshold;
[0027] From all the horizontal axis values, obtain the target horizontal axis value that makes the total time cost of the total time cost curve less than or equal to the preset time;
[0028] The mixed dataset corresponding to the target x-coordinate value is taken as the optimal mixed dataset.
[0029] According to a second aspect of this disclosure, a road hidden defect detection device is provided, comprising:
[0030] The first acquisition module is used to acquire a first number of images of hidden road defects collected in the target area by an image acquisition device.
[0031] The first acquisition module is used to generate a virtual image for each of the hidden road defects images through a diffusion model, wherein the label of the virtual image is the same as the label of each of the hidden road defects images;
[0032] A generation module is used to generate a second number of mixed datasets based on the virtual image, the road hidden disease image, the label of the virtual image, and the label of the road hidden disease image; each mixed dataset includes a training set, a validation set, and a test set, wherein the training set includes images obtained from the road hidden disease image and images obtained from the virtual image, and the images in the validation set and the test set are both obtained from the road hidden disease image;
[0033] The second acquisition module is used to obtain the performance data of the lightweight YOLO model on the test set included in each of the hybrid datasets during the process of processing each of the hybrid datasets using a pre-constructed lightweight YOLO model.
[0034] The second acquisition module is used to acquire the optimal mixed dataset from the second number of mixed datasets based on the performance data;
[0035] The detection module is used to detect road defects in the area to be detected using the lightweight YOLO model and the optimal hybrid dataset.
[0036] According to a third aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method described in the first aspect.
[0037] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program / instructions stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.
[0038] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0039] As will be described in detail below, the road hidden defect detection method according to embodiments of this disclosure obtains virtual images of actually collected road hidden defect images through the diffusion principle. A hybrid dataset is then generated using the virtual image and the road hidden defect image, and processed by a YOLO model. The images in the hybrid dataset processed by the YOLO model are not entirely actual collected images; only a portion of the actual images need to be collected to obtain the hybrid dataset processed by the YOLO model. Compared to existing technologies where the images in the dataset processed by the YOLO model are entirely actual collected images, the solution of this application can significantly reduce the manpower and material resources consumed in the actual image acquisition process.
[0040] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description
[0041] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0042] Figure 1 This is a flowchart illustrating a method for detecting hidden road defects according to an embodiment of this disclosure.
[0043] Figure 2 This is a schematic diagram illustrating zero-point correction, background removal, and automatic gain control according to embodiments of the present disclosure.
[0044] Figure 3 This is a comparison diagram showing the effect of applying a virtual image and a road hidden defect image according to an embodiment of this disclosure.
[0045] Figure 4 This is a schematic diagram illustrating the application of the C2fGhost module according to an embodiment of this disclosure.
[0046] Figure 5 This diagram illustrates the network structure of a lightweight YOLO model with an added P6 layer, according to an embodiment of this disclosure.
[0047] Figure 6 This is a schematic diagram illustrating the performance of the YOLO model applied according to an embodiment of this disclosure on a test set.
[0048] Figure 7 This is a structural diagram illustrating a road hidden defect detection device according to an embodiment of the present disclosure.
[0049] Figure 8 This is a hardware block diagram illustrating an electronic device according to an embodiment of the present disclosure.
[0050] Figure 9 This is a schematic diagram illustrating a computer-readable storage medium according to an embodiment of the present disclosure. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0052] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0053] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0054] In related technologies, the model training process of the scheme combining digital signal and image processing technology with vehicle-mounted ground-penetrating radar for detecting hidden road defects heavily relies on a large amount of high-quality data. However, the collection and verification of ground-penetrating radar field data is still labor-intensive and time-consuming, and its economic cost and engineering efficiency are unsatisfactory.
[0055] To alleviate the technical problems existing in related technologies, this disclosure provides a method, device, storage medium, and program product for detecting hidden road defects. This method uses the diffusion principle to obtain virtual images of actually acquired road defect images, and generates a hybrid dataset using the virtual image and the road defect images. The hybrid dataset is then processed by a YOLO model. The images in the hybrid dataset processed by the YOLO model are not entirely actual acquired images; only a portion of the actual images need to be acquired to obtain the hybrid dataset processed by the YOLO model. Compared to existing technologies where the images in the dataset processed by the YOLO model are entirely actual acquired images, the solution of this application can significantly reduce the manpower and material resources consumed in the actual image acquisition process.
[0056] To facilitate understanding of this embodiment, a detailed description of the road hidden defects detection method disclosed in this disclosure embodiment will be provided first. The execution subject of the road hidden defects detection method provided in this disclosure embodiment is generally an electronic device with a certain computing power, such as a terminal device, a server, or other processing device. In some possible implementations, the road hidden defects detection method can be implemented by a processor calling computer-readable instructions stored in memory.
[0057] See Figure 1 The diagram shows a flowchart of a method for detecting hidden road defects provided in this embodiment of the present disclosure. The method includes the following steps:
[0058] Step 101: Obtain the first number of hidden road defects images collected in the target area using an image acquisition device.
[0059] In this embodiment, the image acquisition equipment includes, but is not limited to, ground-penetrating radar equipment. To meet the requirements of data balance and diversity as much as possible, the first number of road hidden defects images may include images of multiple types of road hidden defects, including but not limited to cracks and cavities. The first number can be set manually based on experience or according to actual needs. For example, the first number can be set to 500 images, which may include 250 images of cavities and 250 images of cracks.
[0060] In practical applications, when image acquisition equipment collects images of hidden road defects in a target area, the equipment first collects images of suspected defects. Then, professionals use an endoscope to verify whether there are defects inside the road. If defects are confirmed, the image is used as the image of hidden road defects.
[0061] In practice, professionals can also cross-validate the selected images to preserve images with cleaned disease features.
[0062] In some embodiments, in order to improve the precision and accuracy of the images acquired by the image acquisition device, before acquiring the first number of road hidden defects images acquired by the image acquisition device, the device parameters of the image acquisition device may be adjusted so that the road hidden defects images meet the requirements of zero-point correction, automatic gain control and background removal.
[0063] Zero-point correction refers to eliminating the error caused by the air layer between the ground-penetrating radar antenna and the ground by determining the ground reflection position, thus unifying the starting position of multiple reflected waves to the same zero point. Automatic gain refers to amplifying the amplitude of reflected signals deep within the road structure layer to compensate for energy attenuation during electromagnetic wave propagation and enhance the defect characteristics of deep targets. Background elimination refers to calculating the average amplitude of all single-channel signal waves, subtracting this average from the original signal amplitude, eliminating various influences such as direct waves, background noise, and DC bias, and improving the signal-to-noise ratio of the target area. For ease of understanding, examples are given below. Figure 2 The diagram shows zero-point correction, background elimination, and automatic gain control.
[0064] Step 102: Generate a virtual image for each hidden road defect image using a diffusion model. The labels of the virtual images are the same as those of each hidden road defect image.
[0065] In this embodiment, the labels for road hidden defects images can be assigned by professionals who log in to a specific website and, under cross-validation by no fewer than two people, label the defects in the road hidden defects images. The cross-validation process requires an error rate of no more than 5%.
[0066] In this embodiment, each image of a hidden road defect can have one or more virtual images, such as two.
[0067] In some alternative embodiments, generating a virtual image of each hidden road defect image using a diffusion model may include:
[0068] A preset implicit diffusion method is used to diffuse each road hidden defect image in the latent space to obtain a virtual image of each road hidden defect image.
[0069] To improve the efficiency of virtual data (i.e., virtual images) generation, a stable diffusion model is adopted. Although the stable diffusion model supports both text-to-image and image-to-image generation functions, in order to limit the input of subjective information, this embodiment only deploys the model's img2img module, i.e., the image-to-image generation module.
[0070] Specifically, the stable diffusion model employs the implicit diffusion principle and the latent space acceleration method:
[0071] The implicit diffusion process is divided into forward diffusion and reverse diffusion. Forward diffusion retains some information from the input real image, and then adds noise to create a completely noisy image. Afterward, according to the Markov chain principle, reverse diffusion is performed to denoise, resulting in the generated image. The generated image retains some original image information and produces appropriate texture changes, enhancing data diversity while doubling the amount of data. Finally, the generated images are sequentially numbered to form a generated image database. In this process, the hyperparameter denoising strength determines the amount of original image information retained in the generated image to the greatest extent, with a range controlled within [0-1], where 0 represents retaining all original image information and 1 represents losing all original image information. Extensive testing has shown that the dynamic adjustment range of this parameter in this disclosure should not exceed [0.2-0.4], and 0.23 is recommended as the control parameter. Technicians can fine-tune it based on their own real dataset, ensuring that the similarity between the generated image and the real image reaches a certain level. Figure 3 Furthermore, for each input image, multiple images can be generated through repeated diffusion, meaning there is theoretically no limit to the number of virtual images generated. However, to ensure the robustness of the mixed dataset, it is recommended to generate only one corresponding virtual image for each input image, and to consider generating two images if the amount of real data is insufficient.
[0072] Latent space acceleration involves introducing the forward and backward diffusion image generation process into the latent space. Specifically, the input image is first processed by vector transformation to compress the image size within the latent space, followed by a stable diffusion process. Finally, the latent space vector is inversely transformed to output a two-dimensional generated image. This process can improve the diffusion speed by at least ten times. Furthermore, the hyperparameter batch size has a slight impact on the diffusion speed; a setting of 8 is recommended. Testing showed that, under the server configuration and environment shown in Table 1, the speed of generating a single image is approximately 1.72 seconds, meeting practical engineering efficiency requirements.
[0073] Please refer to Figure 3 , Figure 3 This is a comparison diagram of the effects of the virtual image and the hidden road defects image given in this embodiment.
[0074] Step 103: Based on the virtual image, the road hidden disease image, the label of the virtual image, and the label of the road hidden disease image, generate a second number of mixed datasets; each mixed dataset includes a training set, a validation set, and a test set. The training set includes images obtained from the road hidden disease image and images obtained from the virtual image. The images in the validation set and the test set are both obtained from the road hidden disease image.
[0075] In some embodiments, a dynamic random sampling algorithm for small sample cases can be written using Python's standard library `random`. This dynamic random sampling algorithm generates a mixed dataset, which can achieve the following objectives:
[0076] The size of the training, validation, and testing datasets for the mixed datasets is limited to an 8:1:1 ratio, with the option to fine-tune the ratio based on actual engineering needs.
[0077] The ratio of crack and cavity disease data in the training, validation and test sets of the mixed dataset is restricted and strictly controlled to 1:1 to meet the data balance requirements.
[0078] Restrict the data types in the validation and test sets (containing only real data) to ensure that only real data is used to validate and optimize the detection performance during model training;
[0079] The ratio of real to generated images in the training set is dynamically adjusted, gradually increasing from 1:0 to 1:5 (as the total size of the training set increases), thus establishing 11 mixed datasets.
[0080] The total number of real data images in the mixed dataset was limited to approximately 500 to validate the small sample size.
[0081] It should be understood that the first number of hidden road defects images are acquired in real time by the image acquisition device, so this first number of hidden road defects images are the real data in this embodiment, while the virtual images are images virtualized by the diffusion model, so the virtual images are not real data.
[0082] In achieving the above objectives, in some embodiments, step 103 may specifically include the following steps:
[0083] The images of hidden road defects are divided into a second number of image sets. Each image set includes a first training image, a validation image, and a test image with labels. The sum of the number of the first training image, the validation image, and the test image is the first number. The number of validation images and test images is the same. The number of the first training image, the validation image, or the test image in different image sets is different.
[0084] Based on the preset image mixing ratio, and considering the number of first training images, validation images, and test images in each image set, the target number of second training images in each image set is calculated.
[0085] A second number of mixed datasets is obtained based on each image set and a second number of training images obtained from the target number of labeled virtual images.
[0086] In this embodiment, the second quantity can be set manually based on experience or according to actual needs. For example, the second quantity can be set to 11.
[0087] As an example, taking the initial set of 500 images of hidden road defects as an example, the final 11 mixed datasets are shown in Table 1:
[0088] Table 1
[0089]
[0090] In some embodiments, in order to improve the processing speed of subsequent YOLO models, before obtaining a second number of mixed datasets based on each image set and a second number of training images of a target number obtained from virtual images, the pixel dimensions of the first and second training images belonging to each image set may be cropped to a preset size.
[0091] The preset size here can be set manually based on experience or actual needs, such as setting the preset size to 1180×1180 or 1848×1848.
[0092] Step 104: During the process of processing each mixed dataset using a pre-constructed lightweight YOLO model, obtain the performance data of the lightweight YOLO model on the test set included in each mixed dataset.
[0093] In this embodiment, the C2f module in the lightweight YOLO model is replaced with the C2fGhost module, which has a smaller number of model parameters. Please refer to... Figure 4 , Figure 4 This is a schematic diagram of the C2fGhost module. The working principle of the C2fGhost module is as follows:
[0094] The features output from the previous layer of the C2fGhost module are used as input features into the C2fGhost module. After one regular convolution (GhostConv) process, the convolutional feature map is split into two branches. One branch passes through two consecutive GhostConv modules for feature extraction. The feature map obtained after passing through the two GhostConv modules is concatenated with the feature map that has not passed through this module. Finally, after a regular convolution, the final feature map is output. The GhostConv module can significantly reduce the total number of model parameters and reduce computational complexity.
[0095] In this embodiment, a P6 detection layer with a lower output feature map resolution is added to the original detection layers of the YOLO model. The improved model includes P3, P4, P5, and P6 detection layers, which can better identify multi-scale features of the image. Specifically, the shallow network detection layers (P3, P4, P5) mainly identify image detail features, which is beneficial for small target detection. The deep network detection layer (P6) is located deeper in the YOLO network structure, which is more beneficial for identifying large target lesions. In addition, the lower resolution of the P6 output feature map can save computing power and improve the model training speed. The corresponding improved lightweight YOLO model network structure is added as follows: Figure 5 As shown.
[0096] To verify the performance of the YOLO model in this embodiment, the following comparative verification implementation scheme is also provided in some embodiments.
[0097] Based on YOLO's performance metrics, average accuracy (map) and total time cost (TC), the relevant metrics are defined as follows:
[0098] "Positive" and "Negative" are the predicted sample labels, where "Positive" represents the disease in the image to be detected (positive sample), and "Negative" represents the background in the image to be detected (negative sample). "True" and "False" represent the prediction results, where "True" indicates a correct prediction and "False" indicates an incorrect prediction. Therefore, FP indicates that a negative sample was incorrectly predicted as a positive sample, FN indicates that a positive sample was incorrectly predicted as a negative sample, TN indicates that a negative sample was correctly predicted as a negative sample, and TP indicates that the classification was correctly predicted.
[0099] Precision, also known as accuracy, assesses the accuracy of disease detection; that is, the proportion of data that are actually positive out of all data predicted as positive.
[0100]
[0101] Recall rate, also known as completeness, assesses whether disease detection is comprehensive, that is, the proportion of data that are actually positive samples that are correctly predicted as positive samples.
[0102]
[0103] In average precision mapping, precision and recall are typically not optimized simultaneously, leading to a neglect of the overall results. Therefore, AP is introduced. i A map is used to balance the calculation results of the two. For a disease type, a Precision-Recall (PR) curve is obtained by plotting Precision on the vertical axis and Recall on the horizontal axis. The area under the PR curve represents the Aptitude (AP). iThe value of . Furthermore, map represents the average AP of all classes across the entire test set. i .
[0104]
[0105] Where i represents the category (crack or cavity); N i This represents the total number of a certain type of disease in the test set; k is the subscript symbol that distinguishes a certain type of disease.
[0106] Furthermore, this disclosure defines the Total Time Cost (TC) as a metric to quantitatively describe the training cost, and its calculation method is as follows:
[0107] TC=TT+TG (4)
[0108] Where TT represents the training time for improving the YOLO model during training; TG represents the time required to complete data augmentation, which is the product of the total number of images generated and the efficiency of generating a single image (1.72 seconds / image).
[0109] The improved and baseline models were trained, validated, and tested on a multi-gradient mixed dataset using the controlled variable method. The optimal lightweight model was obtained by combining its performance metrics on 11 test sets. The results are shown in Table 2.
[0110] Table 2
[0111]
[0112]
[0113] Based on Table 2, the improved YOLO model achieved the highest recognition accuracy and the fastest training speed after introducing the C2fGhost module and the P6 detection layer. Compared with the baseline model, the average accuracy map was improved by 1.54%, while the total time cost (TC) was reduced by 7.2% on average. Therefore, YOLOv8-Ghost-P6 is determined to be the optimal lightweight model.
[0114] Step 105: Based on the performance data, obtain the optimal mixed dataset from the second number of mixed datasets.
[0115] In some embodiments, step 105 may include the following steps:
[0116] Based on the average accuracy of the lightweight YOLO model in the second number of mixed datasets, an average accuracy curve is generated; and based on the total time cost of the lightweight YOLO model in the mixed datasets, a total time cost curve is generated; wherein, the x-axis of the average accuracy curve and the total time cost curve are both the ratio between the road hidden disease images and virtual images in the training set included in the mixed datasets.
[0117] Obtain all x-axis values from the average accuracy curve that make the average accuracy of the average accuracy curve greater than the preset accuracy threshold;
[0118] From all x-axis values, obtain the target x-axis value that makes the total time cost of the total time cost curve less than or equal to the preset time;
[0119] The mixed dataset corresponding to the target x-coordinate value is taken as the optimal mixed dataset.
[0120] In this embodiment, both the accuracy threshold and the preset time can be preset by the user according to their own needs and actual requirements. For example, the accuracy threshold can be set to 90%, and the preset time can be set to 20 minutes.
[0121] Please see Figure 6 , Figure 6 Taking the second number of mixed datasets as an example, Table 1 shows the performance of the YOLO model on the test set (average accuracy (map) and total time cost (TC)) using the 11 multi-gradient mixed datasets shown in Table 1 for training, validation and testing.
[0122] analyze Figure 6 On the left, as the proportion of generated data in the training set increases and the total size of the training set continues to grow, the map value first rises rapidly and then slowly until it reaches nearly 0.99. When the ratio of generated data to real data reaches 5:1, the mixed dataset brings a significant improvement of nearly 0.13 compared to the original small sample data. This shows that the method of using a stable diffusion model to generate data to overcome the small sample problem in actual engineering has achieved significant results.
[0123] analyze Figure 6 As the proportion of generated data in the training set increases and the total size of the training set continues to grow, the total time cost (TC) gradually increases. Moreover, the time cost (TC) required to run a stable diffusion model gradually becomes dominant. When the ratio of generated data to real data reaches 5:1, the TC reaches nearly 5 times that under the original small sample condition.
[0124] Based on the above analysis, in order to balance the identification accuracy and total time cost of diseases in the case of small samples, it is recommended that the optimal hybrid dataset should meet the following requirements:
[0125] The ratio of generated data to real data in the training set should be controlled at approximately 1:1 to 3:1. At this ratio, a small-sample ground-penetrating radar system for detecting underground defects based on stable diffusion and lightweight YOLO can meet the accuracy and efficiency requirements of engineering detection. (Combined with...) Figure 6 It can be observed that when the ratio of generated data to real data in the training set is controlled at around 1:1 to 3:1, the average accuracy (map) is greater than 90%, and the total time cost (TC) is greater than 20 minutes.
[0126] Step 106: Use the YOLO model and the optimal hybrid dataset to detect road defects in the area to be detected.
[0127] When detecting road defects in the area to be detected, images of hidden road defects can be collected from the area to be detected according to the optimal ratio of data in the training set, test set, and validation set in the mixed dataset. These images can then be diffused to form a mixed dataset for processing the YOLO model.
[0128] In the solution provided in this embodiment, a virtual image of the actual collected road hidden defects can be obtained through the diffusion principle. A hybrid dataset is then generated using the virtual image and the road hidden defects image, and processed by the YOLO model. The images in the hybrid dataset processed by the YOLO model are not entirely actual collected images; only a portion of the actual images need to be collected to obtain the hybrid dataset processed by the YOLO model. Compared to existing technologies where the images in the dataset processed by the YOLO model are entirely actual collected images, the solution of this application embodiment can significantly reduce the manpower and material resources consumed in the actual image acquisition process.
[0129] Based on this, the solution of this embodiment also has the following advantages:
[0130] Preprocessing of real-collected road hidden defects images using zero-point correction, automatic gain control, and background removal enhances defect features, eliminates errors and signal interference, and provides a high-quality data foundation for subsequent image generation and improved YOLO model training.
[0131] This disclosure specifies a technical method for data augmentation using a stable diffusion model, which overcomes the time-consuming and labor-intensive nature of collecting and verifying original real-world disease data from vehicle-mounted ground-penetrating radar. It allows for the creation of a large-scale mixed dataset with only a small sample of real-world disease data, thus overcoming the small sample problem and improving the robustness and stability of the model.
[0132] This disclosure presents a lightweight improvement to the YOLO model by introducing the C2fGhost module to replace the C2f module, thereby reducing the number of model parameters. A P6 detection layer is added to the original detection layer of the YOLO model to expand the receptive field of the model, improve the network's ability to perceive global features, reduce the resolution of the output feature map, and reduce the time required to train the model, providing a new model for its rapid deployment and application in practical engineering.
[0133] This disclosure introduces average accuracy (map) and total time cost (TC) performance metrics to achieve a balance between detection accuracy and efficiency, thereby establishing a complete optimal lightweight model and hybrid dataset screening system.
[0134] This disclosure also provides a road hidden defect detection device, which is used to perform the road hidden defect detection method provided in any of the above embodiments. Figure 7 As shown, the device includes:
[0135] The first acquisition module 71 is used to acquire a first number of images of hidden road defects collected in the target area by the image acquisition device.
[0136] The first obtaining module 72 is used to generate a virtual image for each of the hidden road defects images through a diffusion model, wherein the label of the virtual image is the same as the label of each of the hidden road defects images;
[0137] The generation module 73 is used to generate a second number of mixed datasets based on the virtual image, the road hidden disease image, the label of the virtual image, and the label of the road hidden disease image; each mixed dataset includes a training set, a validation set, and a test set, wherein the training set includes images obtained from the road hidden disease image and images obtained from the virtual image, and the images in the validation set and the test set are both obtained from the road hidden disease image;
[0138] The second acquisition module 74 is used to obtain the performance data of the lightweight YOLO model on the test set included in each of the mixed datasets during the process of processing each of the mixed datasets using a pre-constructed lightweight YOLO model.
[0139] The second acquisition module 75 is used to acquire the optimal mixed dataset from the second number of mixed datasets based on the performance data.
[0140] The detection module 76 is used to detect road defects in the area to be detected using the lightweight YOLO model and the optimal hybrid dataset.
[0141] In some embodiments, the first obtaining module 72 is used for:
[0142] A preset implicit diffusion method is used to diffuse each of the hidden road defects images in the latent space to obtain a virtual image of each of the hidden road defects images.
[0143] In some embodiments, the generation module 73 is used for:
[0144] The hidden road defects images are divided into a second number of image sets. Each image set includes a first training image, a verification image, and a test image with labels. The sum of the number of the first training image, the verification image, and the test image is the first number. The number of the verification image and the test image are the same. The number of the first training image, the verification image, or the test image in different image sets is different.
[0145] Based on the preset mixed image ratio, and considering the number of first training images, verification images, and test images in each image set, the target number of second training images in each image set is calculated.
[0146] The second number of mixed datasets is obtained based on each of the image sets and the target number of second training images obtained from the labeled virtual images.
[0147] In some embodiments, the device is further used for:
[0148] Before obtaining the second number of mixed datasets based on each of the image sets and the target number of second training images obtained from the labeled virtual images, the pixel dimensions of the first and second training images belonging to each of the image sets are cropped to a preset size.
[0149] In some embodiments, the device is further used for:
[0150] Before acquiring the first number of road hidden defects images through the image acquisition device, the device parameters of the image acquisition device are adjusted so that the road hidden defects images meet the requirements of zero-point correction, automatic gain control, and background removal.
[0151] In some embodiments, the second acquisition module 75 is used for:
[0152] An average accuracy curve is generated based on the average accuracy of the lightweight YOLO model's performance data across the second number of mixed datasets; and a total time cost curve is generated based on the total time cost of the lightweight YOLO model's performance data across the mixed datasets; wherein the horizontal axis of both the average accuracy curve and the total time cost curve represents the ratio between road hidden defect images and virtual images in the training set included in the mixed datasets.
[0153] Obtain all abscissa values from the average accuracy curve that make the average accuracy of the average accuracy curve greater than a preset accuracy threshold;
[0154] From all the horizontal axis values, obtain the target horizontal axis value that makes the total time cost of the total time cost curve less than or equal to the preset time;
[0155] The mixed dataset corresponding to the target x-coordinate value is taken as the optimal mixed dataset.
[0156] The road hidden defects detection device and the road hidden defects detection method provided in this disclosure are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.
[0157] This disclosure also provides an electronic device for performing the above-described method for detecting hidden road defects. Please refer to... Figure 8 It illustrates a schematic diagram of an electronic device provided by some embodiments of this disclosure. For example... Figure 8 As shown, the electronic device 8 includes: a processor 800, a memory 801, a bus 802, and a communication interface 803. The processor 800, the communication interface 803, and the memory 801 are connected via the bus 802. The memory 801 stores a computer program that can run on the processor 800. When the processor 800 runs the computer program, it executes the road hidden defects detection method provided in any of the foregoing embodiments of this disclosure.
[0158] The memory 801 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this device network element and at least one other network element is achieved through at least one communication interface 803 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0159] Bus 802 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 801 is used to store programs. After receiving an execution instruction, the processor 800 executes the program. The road hidden defects detection method disclosed in any of the foregoing embodiments of this disclosure can be applied to the processor 800, or implemented by the processor 800.
[0160] The processor 800 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 800 or by instructions in software form. The processor 800 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 801. Processor 800 reads the information in memory 801 and, in conjunction with its hardware, completes the steps of the above method.
[0161] The electronic equipment provided in this disclosure and the road hidden defects detection method provided in this disclosure are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.
[0162] This disclosure also provides a computer-readable storage medium corresponding to the road hidden defect detection method provided in the foregoing embodiments. Please refer to... Figure 9 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the road hidden defects detection method provided in any of the foregoing embodiments.
[0163] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.
[0164] The computer-readable storage medium provided in the above embodiments of this disclosure and the road hidden defects detection method provided in the embodiments of this disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0165] It should be noted that:
[0166] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this disclosure may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0167] Similarly, it should be understood that, in order to simplify this disclosure and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this disclosure, various features of this disclosure are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting a schematic diagram in which the claimed disclosure requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this disclosure.
[0168] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this disclosure and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0169] The above description is merely a preferred embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for detecting hidden road defects, characterized in that, include: Acquire a first number of images of hidden road defects in the target area using an image acquisition device; A virtual image for each of the hidden road defects images is generated using a diffusion model, and the labels of the virtual images are the same as the labels of each of the hidden road defects images. Based on the virtual image, the road hidden defects image, the label of the virtual image, and the label of the road hidden defects image, a second number of mixed datasets are generated; each mixed dataset includes a training set, a validation set, and a test set, wherein the training set includes images obtained from the road hidden defects image and images obtained from the virtual image, and the images in the validation set and the test set are both obtained from the road hidden defects image; In the process of processing each of the hybrid datasets using a pre-constructed lightweight YOLO model, the performance data of the lightweight YOLO model on the test set included in each of the hybrid datasets is obtained; Based on the performance data, the optimal mixed dataset is obtained from the second number of mixed datasets; The lightweight YOLO model and the optimal hybrid dataset are used to detect road defects in the area to be detected. Specifically, based on the virtual image, the road hidden defect image, the label of the virtual image, and the label of the road hidden defect image, a second number of mixed datasets are generated, including: The hidden road defects images are divided into a second number of image sets. Each image set includes a first training image, a verification image, and a test image with labels. The sum of the number of the first training image, the verification image, and the test image is the first number. The number of the verification image and the test image are the same. The number of the first training image, the verification image, or the test image in different image sets is different. Based on the preset mixed image ratio, and considering the number of first training images, verification images, and test images in each image set, the target number of second training images in each image set is calculated. The second number of mixed datasets is obtained based on each of the image sets and the target number of second training images obtained from the labeled virtual images.
2. The method according to claim 1, characterized in that, Virtual images of each of the hidden road defects are generated using a diffusion model, including: A preset implicit diffusion method is used to diffuse each of the hidden road defects images in the latent space to obtain a virtual image of each of the hidden road defects images.
3. The method according to claim 1, characterized in that, Before obtaining the second number of hybrid datasets based on each of the image sets and the target number of second training images obtained from the labeled virtual images, the method further includes: The pixel dimensions of the first and second training images belonging to each of the image sets are cropped to a preset size.
4. The method according to claim 1, characterized in that, Before acquiring the first number of images of hidden road defects captured by the image acquisition device, the process also includes: Adjust the device parameters of the image acquisition device so that the road hidden defects image meets the requirements of zero-point correction, automatic gain control and background removal.
5. The method according to claim 1, characterized in that, Based on the performance data, the optimal mixed dataset is obtained from the second number of mixed datasets, including: An average accuracy curve is generated based on the average accuracy of the lightweight YOLO model's performance data across the second number of mixed datasets; and a total time cost curve is generated based on the total time cost of the lightweight YOLO model's performance data across the mixed datasets; wherein the horizontal axis of both the average accuracy curve and the total time cost curve represents the ratio between road hidden defect images and virtual images in the training set included in the mixed datasets. Obtain all abscissa values from the average accuracy curve that make the average accuracy of the average accuracy curve greater than a preset accuracy threshold; From all the horizontal axis values, obtain the target horizontal axis value that makes the total time cost of the total time cost curve less than or equal to the preset time; The mixed dataset corresponding to the target x-coordinate value is taken as the optimal mixed dataset.
6. A device for detecting hidden road defects, characterized in that, include: The first acquisition module is used to acquire a first number of images of hidden road defects collected in the target area by an image acquisition device. The first acquisition module is used to generate a virtual image for each of the hidden road defects images through a diffusion model, wherein the label of the virtual image is the same as the label of each of the hidden road defects images; A generation module is used to generate a second number of mixed datasets based on the virtual image, the road hidden disease image, the label of the virtual image, and the label of the road hidden disease image; each mixed dataset includes a training set, a validation set, and a test set, wherein the training set includes images obtained from the road hidden disease image and images obtained from the virtual image, and the images in the validation set and the test set are both obtained from the road hidden disease image; The second acquisition module is used to obtain the performance data of the lightweight YOLO model on the test set included in each of the hybrid datasets during the process of processing each of the hybrid datasets using a pre-constructed lightweight YOLO model. The second acquisition module is used to acquire the optimal mixed dataset from the second number of mixed datasets based on the performance data; The detection module is used to detect road defects in the area to be detected using the lightweight YOLO model and the optimal hybrid dataset. The generation module generates a second set of mixed datasets based on the virtual image, the road hidden defects image, the label of the virtual image, and the label of the road hidden defects image, including: The hidden road defects images are divided into a second number of image sets. Each image set includes a first training image, a verification image, and a test image with labels. The sum of the number of the first training image, the verification image, and the test image is the first number. The number of the verification image and the test image are the same. The number of the first training image, the verification image, or the test image in different image sets is different. Based on the preset mixed image ratio, and considering the number of first training images, verification images, and test images in each image set, the target number of second training images in each image set is calculated. The second number of mixed datasets is obtained based on each of the image sets and the target number of second training images obtained from the labeled virtual images.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-5.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Multi-dimensional pavement disease data set construction method
CN115329109A
DDPM-YOLO-based side-scan sonar image confrontation enhancement generation method
CN118397442A