Lightweight Detection Method for Fracture Penetration Degree Optimized by Ground Penetrating Radar and Swin Transformer

By combining ground penetrating radar and Swin Transformer optimized lightweight detection methods, the accuracy and real-time problems in asphalt pavement crack detection are solved, and efficient and accurate crack penetration degree detection is achieved. It is suitable for large-area rapid scanning and key area inspection, reducing calculation complexity and resource requirements.

CN119669680BActive Publication Date: 2025-07-11CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411722449.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-07-11
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, low real-time, high computational complexity and lack of time-frequency analysis capabilities in asphalt pavement crack detection, especially in complex environments, it is difficult to effectively identify the degree of throughput of cracks.

Method used

The lightweight crack penetration degree detection method based on ground penetrating radar and Swin Transformer optimization is adopted. By collecting and preprocessing radar data, a multi-feature fusion data set is constructed, and a two-stage and one-stage detection model based on Swin Transformer is designed. Combined with the FPN network and ROIAlign pooling method, the bounding box position is optimized using the Smooth L1 loss function to achieve efficient feature extraction and detection.

Benefits of technology

It improves the accuracy and efficiency of crack detection, can quickly identify cracks on large-area roads under low computing resources, reduce maintenance costs, extend the service life of the road, ensure traffic safety, and provide scientific basis for road design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669680B_ABST
    Figure CN119669680B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight method for detecting the degree of crack penetration optimized based on ground penetrating radar and Swin Transformer, including the following steps: Step 1, collect ground penetrating radar data and perform preprocessing and feature extraction; Step 2, construct a dataset with multi-feature fusion; Step 3, design a two-stage through-crack detection model optimized based on Swin Transformer; Step 4, design a one-stage through-crack detection model optimized based on Swin Transformer-YOLOv8; Step 5, train and predict the training set and test set data according to the model architectures in Step 3 and Step 4. The present invention improves the accuracy of crack penetration degree detection, provides a scientific basis for road maintenance; greatly improves the detection speed, meets the requirements of large-area rapid detection; reduces the computational complexity, relieves the performance pressure of equipment, and improves the detection efficiency; enhances the generalization ability of the model and can adapt to different road conditions; promotes the development of intelligent road detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of ground penetrating radar and deep learning object detection, and particularly to a lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer. Background Technique

[0002] With the rapid development of the economy and the acceleration of the urbanization process, the highway traffic volume has increased sharply. As the main urban road material, the performance and service life of asphalt pavement have received more and more attention. Under the long-term action of load stress and temperature stress, asphalt pavement is prone to diseases such as cracks. If these diseases are not detected and repaired in time, it will lead to further damage to the road surface structure and even cause serious safety accidents such as road collapse. Therefore, the early identification and evaluation of asphalt pavement cracks are particularly important.

[0003] Traditional asphalt pavement crack detection methods mainly rely on manual visual inspection. This method is inefficient, inaccurate, and has great subjectivity. With the development of ground penetrating radar (GPR) technology, as a non-destructive testing technology, it has been widely used in road crack detection due to its advantages such as simple operation, high speed, and high resolution. However, the analysis of ground penetrating radar data requires professional knowledge and experience, and its ability to process non-linear and non-stationary signals is limited.

[0004] Applying deep learning technology to the crack detection of ground penetrating radar images can significantly improve the automation level and accuracy of crack detection.

[0005] In the field of asphalt pavement crack detection, existing technical solutions mainly rely on traditional image processing techniques and emerging deep learning algorithms. Although these methods can provide useful results in some cases, they are usually difficult to adapt to the complexity and diversity of crack morphologies, especially in the case of crack intersections or branches. In addition, these methods also show limitations in processing non-linear and non-stationary signals because they are often based on linear assumptions and cannot effectively capture the dynamic characteristics of signals. However, existing deep learning methods still have limitations in processing non-linear and non-stationary signals.

[0006] The deficiencies of existing methods in crack detection are as follows: (1) The accuracy and robustness of existing methods are insufficient in complex environments, especially when the crack morphology is complex or there are intersections, branches, etc.; (2) Traditional ground penetrating radar analysis methods require manual operation and analysis, which limits the real-time performance and automation degree of detection; (3) Although some deep learning methods have improved the detection accuracy, they have high computational complexity, resulting in slow detection speed and are not suitable for large-scale applications; (4) Existing methods lack the ability to perform time-frequency analysis on ground penetrating radar signals and cannot fully utilize the time-frequency characteristics of signals to improve detection performance.

[0007] The present invention comprehensively applies ground penetrating radar equipment and deep learning technology, and designs two through-crack detection models, which are respectively suitable for refined detection in key areas and rapid general detection in large ranges. It can not only improve the efficiency and quality of road maintenance, reduce maintenance costs, but also extend the service life of roads by timely discovering and repairing cracks, reduce traffic accidents caused by road collapses, etc., and ensure the safety of life and property. Moreover, it can provide a scientific basis for road design and construction, optimize road structure design, and improve the stability and durability of roads; in the fields of environmental monitoring, disaster prevention, etc., the present invention also has broad application prospects. Summary of the Invention

[0008] In view of this, the present invention provides a lightweight method for detecting the degree of crack penetration optimized based on ground penetrating radar and Swin Transformer.

[0009] To solve the above technical problems, the present invention adopts the following technical solutions:

[0010] A lightweight method for detecting the degree of crack penetration optimized based on ground penetrating radar and Swin Transformer includes the following steps:

[0011] Step 1, collect ground penetrating radar data, and perform preprocessing and feature extraction

[0012] Use ground penetrating radar equipment to scan the asphalt pavement and collect radar signal data below the pavement; process the original radar data; extract key information helpful for crack detection from the preprocessed data;

[0013] Step 2, construct a multi-feature fusion data set;

[0014] Combine the instantaneous amplitude, instantaneous phase and instantaneous frequency feature images of the ground penetrating radar signal to form a multi-feature fusion data set; then expand and annotate the data set to complete the construction of the multi-feature fusion data set for ground penetrating radar through-crack.

[0015] Step 3, design a two-stage through-crack detection model optimized based on Swin Transformer

[0016] (3a) Optimize the overall framework of the network

[0017] Introduce Swin Transformer into the backbone network, introduce the FPN network after the backbone network, and improve the two-stage detection network by replacing ROI pooling with ROIAlign in combination with the advantages of Mask R-CNN;

[0018] (3b) Design of FPN multi-scale fusion module

[0019] Generate feature maps with different depths and scales through a top-down feature extraction network, dynamically adjust the resolution of the deepest feature map using the upsampling method, fuse the feature maps of two adjacent layers with different scales, and use the feature map containing the location information of the shallow network and the semantic information of the deep network after fusion for object detection;

[0020] (3c) Optimization of ROI pooling method

[0021] After calculating the values of four positions through bilinear interpolation, obtain the ROI output of the same size through Maxpooling;

[0022] (3d) Optimization of Smooth L1 loss function

[0023] The bounding box position loss function adopts the Smooth L1 function, and the Smooth L1 loss function combines the L1 loss function and the L2 loss function;

[0024] Step 4, design a one-stage through-crack detection model optimized based on Swin Transformer-YOLOv8

[0025] (4a) Integrate the Swin Transformer network

[0026] Integrate YOLOv8 and Swin Transformer, and introduce the SwinTransformer network in the Backbone part of YOLOv8;

[0027] (4b) Weighted bidirectional feature fusion network

[0028] First, continuously upsample the deep features and fuse them with the bottom features. After the upsampling is completed, then downsample the bottom features and fuse them with the deep features, and add two horizontal connection paths between the two feature extraction paths to fuse the feature maps generated at each stage in the backbone network with the feature maps to be detected;

[0029] (4c) Overall architecture of the Swin Transformer-YOLOv8 model

[0030] YOLOv8 incorporates the Swin Transformer and introduces the weighted bidirectional feature pyramid (BiFPN).

[0031] Step 5: Train and predict the training set and test set data according to the model architectures in Step 3 and Step 4.

[0032] Preferably, in Step 1, the ground penetrating radar data includes detailed information about the road surface structure, especially the location and depth of cracks.

[0033] The preprocessing includes zero-offset correction, FIR filtering, and signal gain adjustment processing.

[0034] The feature extraction includes amplitude information, frequency information, and phase information.

[0035] Preferably, in Step 2, the data set is augmented using methods such as random cropping, mirror flipping, and contrast enhancement, and finally, 8 data sets are labeled to complete the construction of the multi-feature fusion data set for ground penetrating radar through-crack.

[0036] Preferably, in Step 3, the L1 and L2 loss functions are as follows:

[0037]

[0038] where f(x i ) and y i represent the predicted value and the corresponding true value of the i-th sample respectively, and n is the number of samples.

[0039] Preferably, in Step 3, the formula for Smooth L1 is as follows:

[0040]

[0041] This function is a piecewise function. It is the L2 loss between [-1,1], which solves the problem of the breakpoint of L1 at 0. Outside the [-1,1] interval, it is the L1 loss, which solves the problem of gradient explosion of outliers. Therefore, it can limit the gradient from the following two aspects: one is that when the error between the predicted value and the true value is too large, the gradient value will not be too large; the other is that when the error between the predicted value and the true value is very small, the gradient value is small enough.

[0042] Preferably, the specific method adopted in Step 5 is:

[0043] Adopt the multi-feature data set of through-crack constructed in Step 2;

[0044] In step 3, the SGD optimizer is used as the basic iterator, the weight_decay weight decay factor is 0.0001, the momentum factor is 0.9, and single-GPU training is used. The batchsize is set to 4 in both the original network and the improved network, and the corresponding initial learning rates are set to 0.001. During training, the model first loads the initialization weight file of the Swin Transformer trained on the public dataset ImageNet, and the total number of training iterations is set to 100.

[0045] In step 4, the SGD optimizer is used as the basic iterator, the initial learning rate is set to 0.001, the weight_decay weight decay factor is 0.0005, and the momentum factor is 0.937. Single-GPU training is used, the batchsize is set to 4, and considering the comprehensive training speed and the original image size, the total number of training iterations is set to 100 rounds.

[0046] In both step 3 and step 4, accuracy, average precision, mean average precision, and P-R curve are used as evaluation metrics to measure the performance of the model.

[0047] In step 3, for the obtained datasets of the original signal, single feature, and multi-feature fusion, a two-stage network structure optimized based on Swin Transformer is used for training.

[0048] Preferably, the accuracy represents the proportion of correctly detected results among all model detection results, and the recall rate represents the proportion of correctly detected results among all true annotation boxes, as shown in the following formula:

[0049]

[0050]

[0051] Among them, TP represents correct model recognition, that is, the number of cracks correctly detected; FP represents incorrect model recognition, that is, the number of misdetected cracks; FN represents misjudging positive samples as negative samples, that is, the number of missed cracks. P and R can both measure the detection effect of the model in some scenarios, but these two indicators are inversely proportional. Therefore, AP is needed to comprehensively evaluate the model by combining P and R indicators.

[0052] Preferably, the average precision (AP) value is calculated from the area enclosed by the P-R curve and the coordinate axes. The larger the area of the P-R curve, the better the model performance. The formula is as follows:

[0053]

[0054] Preferably, the mean average precision (mAP) is the evaluation metric that best demonstrates the performance of the object detection model. The mAP calculates the average of the APs for all classes in the dataset, and the calculation formula is as follows:

[0055]

[0056] where n represents all classes in the dataset. The mAP mainly includes three evaluation criteria: mAP@0.5, mAP@0.75, and mAP@0.5:0.95. mAP@0.5 and mAP@0.75 refer to the mAP values when IoU = 0.5 and IoU = 0.75 respectively. mAP@"0.5:0.95" refers to the average of the mAP values corresponding to each IoU when IoU varies from 0.5 to 0.95 with a step of 0.05.

[0057] The present invention has achieved the following technical effects compared with the prior art:

[0058] (1) By combining the ground penetrating radar technology and the deep learning model, the present invention can accurately identify and locate the cracks in the asphalt pavement, especially the penetration degree of the cracks. This high-precision detection ability provides reliable data support for road maintenance and repair work. Specifically, the Hilbert-Huang transform and two-dimensional discrete Fourier transform technologies adopted by the present invention can analyze the crack characteristics from multiple spatial domains and extract more accurate crack feature images. These feature images are used as the input of the deep learning model, enabling the model to more accurately learn and identify the crack characteristics, thereby improving the accuracy of crack penetration degree detection. The application of this technology enables road managers to formulate more scientific and effective maintenance plans according to the actual conditions of the cracks, thereby extending the service life of the road and reducing the maintenance cost.

[0059] (2) The one-stage detection model Swin Transformer-YOLOv8 adopted by the present invention can achieve fast scanning and analysis of large areas of roads while maintaining high detection accuracy. Specifically, through the optimized network structure and algorithm, the model significantly improves the inference speed, enabling a large increase in the number of images that can be processed per second. This improvement greatly enhances the detection efficiency, making it possible to detect longer road sections within a limited time and providing strong guarantee for road safety management. For example, in practical applications, the detection speed of the present invention can reach 35.71 frames per second, which means that large areas of roads can be quickly covered and potential crack problems can be identified in a timely manner.

[0060] (3) When designing the model, the present invention fully considers the computational efficiency. By optimizing the algorithm and model structure, the computational complexity of the model is effectively reduced. Specifically, the Swin Transformer-YOLOv8 model adopted by the present invention combines the efficient feature extraction ability of Swin Transformer and the fast detection ability of YOLOv8 to achieve efficient detection at a relatively low computational cost. This design enables the detection model to run quickly even on devices with low performance, without the need for high-end computing resources. This design not only improves the accessibility and practicability of detection, but also reduces the energy consumption and cost of the device, making the detection technology of the present invention easier to deploy and use.

[0061] (4) When constructing the dataset, the present invention adopts a technical solution of multi-feature fusion, which not only enhances the recognition of crack features, but also improves the adaptability of the model to different road conditions. Specifically, by fusing features such as the instantaneous amplitude, instantaneous phase, and instantaneous frequency of ground-penetrating radar signals, the present invention constructs a multi-feature fusion dataset. This way of constructing the dataset enables the model to learn more comprehensive feature information during the training process, thereby improving the model's ability to recognize crack features under different road conditions. The application of this technology enables the model to maintain stable detection performance in a variety of environments. Whether it is new or old roads or different climate conditions, it can accurately identify cracks, ensuring the reliability of the detection results.

[0062] (5) The successful implementation of the present invention provides a new technical means for the field of road detection and promotes the development of intelligent road detection technology. Specifically, through the application of deep learning technology, the present invention not only improves the automation level of detection, but also lays a foundation for the intelligence and unmanned operation of future road detection technology. The application of this technology can not only improve the efficiency and accuracy of road detection, but also provide more scientific data support for road maintenance decision-making, with important scientific research value and application prospects. In addition, the technical solution of the present invention also provides a reference for other fields, such as non-destructive testing of infrastructure such as bridges and tunnels, further expanding the application scope of intelligent detection technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is the two-stage detection network diagram optimized based on Swin Transformer of the present invention;

[0064] Figure 2 It is the schematic diagram of ROI Align of the present invention;

[0065] Figure 3 It is the distribution diagram of the Smooth L1 curve of the present invention;

[0066] Figure 4 It is the fusion diagram of the backbone network module of the present invention;

[0067] Figure 5 It is the diagram of two FPN modules of the present invention;

[0068] Figure 6 It is the structure diagram of Swin Transformer - YOLOv8 of the present invention;

[0069] Figure 7 It is the loss function diagram of the single - feature dataset of the present invention;

[0070] Figure 8 It is the loss function diagram of the feature - fusion dataset of the present invention;

[0071] Figure 9 It is the diagram of the training process of the IA + IF one - stage network of the present invention;

[0072] Figure 10 It is the PR curve diagram of IA + IF of the present invention;

[0073] Figure 11 It is the experimental result diagram of the single - feature dataset of the present invention;

[0074] Figure 12 It is the experimental result diagram of the multi - feature fusion dataset of the present invention;

[0075] Figure 13 It is the comparison diagram of the detection effects between YOLOv8 and the optimized Swin Transformer - YOLOv8 of the present invention;

[0076] Figure 14 It is the comparison diagram of the detection effects between models of the present invention; Detailed implementation manners

[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0078] The present invention discloses a lightweight crack penetration degree detection method optimized based on ground - penetrating radar and Swin Transformer, including the following steps:

[0079] Step 1: Collect ground - penetrating radar data and perform pre - processing and feature extraction

[0080] Use a ground penetrating radar device to scan the asphalt pavement and collect radar signal data below the pavement; process the original radar data; extract key information from the preprocessed data that is helpful for crack detection;

[0081] The ground penetrating radar data includes detailed information about the pavement structure, especially the location and depth of cracks;

[0082] Preprocessing includes zero-offset correction, FIR filtering, and signal gain adjustment processing;

[0083] Feature extraction includes amplitude information, frequency information, and phase information;

[0084] Step 2, construct a multi-feature fusion dataset;

[0085] Combine the instantaneous amplitude, instantaneous phase, and instantaneous frequency feature images of the ground penetrating radar signal to form a multi-feature fusion dataset; use methods such as random cropping, mirror flipping, and contrast enhancement to expand the dataset, and finally label 8 datasets to complete the construction of the multi-feature fusion dataset for ground penetrating radar through-cracks;

[0086] Step 3, design a two-stage through-crack detection model optimized based on Swin Transformer

[0087] (3a) Optimize the overall framework of the network

[0088] Introduce Swin Transformer into the backbone network, introduce the FPN network after the backbone network, and combine the advantages of Mask R-CNN to replace ROI pooling with ROIAlign to improve the two-stage detection network; the improved overall network structure is as Figure 1 shown;

[0089] (3b) Design of the FPN multi-scale fusion module

[0090] Generate feature maps with different depths and scales through a top-down feature extraction network, use upsampling to dynamically adjust the resolution of the deepest feature map, fuse the feature maps of two adjacent layers with different scales, and use the feature map containing the position information of the shallow network and the semantic information of the deep network after fusion for object detection;

[0091] (3c) Optimization of the ROI pooling method

[0092] After calculating the values at four positions through bilinear interpolation, obtain the ROI output of the same size through Maxpooling; the schematic diagram of the ROIAlign calculation principle is shown in Figure 2;

[0093] (3d) Optimization of the Smooth L1 loss function

[0094] The bounding box position loss function uses the Smooth L1 function, which combines the L1 loss function and the L2 loss function;

[0095] The L1 and L2 loss functions are as follows:

[0096]

[0097] where f(x i ) and y i represent the predicted value and the corresponding true value of the i-th sample respectively, and n is the number of samples;

[0098] The Smooth L1 formula is as follows:

[0099]

[0100] This function is a piecewise function. It is the L2 loss between [-1, 1], which solves the problem of the breakpoint of L1 at 0. Outside the interval [-1, 1], it is the L1 loss, which solves the problem of gradient explosion of outliers. Therefore, it can limit the gradient from the following two aspects. One is that when the error between the predicted value and the true value is too large, the gradient value will not be too large; the other is that when the error between the predicted value and the true value is very small, the gradient value is small enough; the curve distribution is as Figure 3 shown;

[0101] Step 4, design a one-stage through-crack detection model optimized based on Swin Transformer-YOLOv8

[0102] (4a) Integrate the SwinTransformer network

[0103] Integrate YOLOv8 and SwinTransformer. In the Backbone part of YOLOv8, introduce the SwinTransformer network; as Figure 4 shown;

[0104] (4b) Weighted bidirectional feature fusion network

[0105] First, continuously upsample the deep features and fuse them with the bottom features. After the upsampling is completed, then downsample the bottom features and fuse them with the deep features, and add two horizontal connection paths between the two feature extraction paths to fuse the feature maps generated at each stage in the backbone network with the feature maps to be measured; as Figure 5 shown;

[0106] (4c) Overall architecture of the SwinTransformer-YOLOv8 model

[0107] YOLOv8 integrates Swin Transformer and introduces the weighted bidirectional feature pyramid (BiFPN); as Figure 6 shown;

[0108] Step 5: Train and predict the training set and test set data according to the model architectures in Steps 3 and 4:

[0109] Adopt the multi-feature dataset of through-crack constructed in Step 2;

[0110] In Step 3, the SGD optimizer is used as the basic iterator, the weight decay factor weight_decay is 0.0001, the momentum factor momentum is 0.9, and single-GPU training is used. The batchsize is set to 4 in both the original network and the improved network, and the corresponding initial learning rates are set to 0.001; during the training process, the model first loads the initialization weight file of Swin Transformer trained on the public dataset ImageNet, and the total number of training iterations is set to 100;

[0111] In Step 4, the SGD optimizer is used as the basic iterator, the initial learning rate is set to 0.001, the weight decay factor weight_decay is 0.0005, and the momentum factor momentum is 0.937; single-GPU training is used, the batchsize is set to 4, and considering the comprehensive training speed and the original image size, the total number of training iterations is set to 100 rounds;

[0112] In both Step 3 and Step 4, accuracy, average precision, mean average precision, and P-R curve are used as evaluation metrics to measure the performance of the model;

[0113] Accuracy represents the proportion of correctly detected results among all model detection results, and recall represents the proportion of correctly detected results among all true annotation boxes, as shown in the following formula:

[0114]

[0115] Among them, TP represents correct model recognition, that is, the number of cracks correctly detected; FP represents incorrect model recognition, that is, the number of cracks misdetected; FN represents judging positive samples as negative samples, that is, the number of cracks missed; P and R can both measure the detection effect of the model in some scenarios, but these two metrics are inversely proportional, so AP is needed to comprehensively evaluate the model by combining P and R metrics;

[0116] The average precision (AP) value is calculated from the area enclosed by the P-R curve and the coordinate axes. The larger the area of the P-R curve, the better the model performance. The formula is as follows:

[0117]

[0118] The mean average precision (mAP) is the evaluation metric that best demonstrates the performance of an object detection model. mAP calculates the average of the APs for all classes in the dataset, and the calculation formula is as follows:

[0119]

[0120] Where n represents all classes in the dataset. mAP mainly includes three evaluation criteria: mAP@0.5, mAP@0.75, and mAP@0.5:0.95. mAP@0.5 and mAP@0.75 refer to the mAP values when IoU = 0.5 and IoU = 0.75 respectively. mAP@"0.5:0.95" refers to the average of the mAP values corresponding to each IoU when IoU varies from 0.5 to 0.95 with a step of 0.05.

[0121] In step three, for the obtained datasets of the original signal, single feature, and multi - feature fusion, a two - stage network structure optimized based on Swin Transformer is used for training;

[0122] Example 1: A two - stage through - crack detection model optimized based on Swin Transformer

[0123] (1a) Comparison and analysis of experimental results on different datasets

[0124] For the obtained datasets of the original signal, single feature, and multi - feature fusion, the two - stage network structure optimized based on Swin Transformer proposed in step 3 of the technical solution is used for training to train the model's recognition and localization ability for through - cracks (t - c) and non - through - cracks (n - c), and to verify the influence of different datasets on the model performance. The single - feature datasets of ground - penetrating radar cracks include: instantaneous amplitude (IA), instantaneous phase (IP), and instantaneous frequency (IF), and the original signal (Origin) is added for comparison. The loss function during the training process is as Figure 7 shown. The loss function gradually stabilizes after 10,000 training iterations, and the model converges in the training on all four datasets.

[0125] There are 4 feature - fusion datasets, namely IA + IF, IA + IF, IF + IP, and IA + IP + IF. To intuitively and clearly compare the performance of the model on the fusion datasets, as Figure 8 shown, in addition to the loss - function curve of the feature - fusion datasets, the loss - function curve of the original - signal dataset is also included. During the training process of the feature - fusion datasets, the loss function drops rapidly in the first 400 iterations and then tends to stabilize. The loss of the model after convergence of IA + IF is the smallest.

[0126] After the loss function converges in the training process, the best.pt weight file generated last is used to load the weight file into the model, and the various indicators of the model are tested with the test set. The comprehensive and detailed results are shown in Table 1.

[0127] Table 1 Comparison results of different data sets

[0128]

[0129] Compare the model test results of the original signal and the single feature data set:

[0130] The two-stage network optimized based on Swin Transformer has better performance indicators in the original signal test set than the single feature test set, with an mAP accuracy of 0.838. The accuracy index of the IA dataset is slightly lower than that of the original dataset. This is because the IA dataset obtained from the echo signal is essentially a threshold processing of the abnormal area in the original signal. The part with smaller amplitude lacks other judgment information, which increases the difficulty of network feature extraction. IA reflects the rapid change of amplitude intensity at the abnormal position. When the electromagnetic wave encounters a through crack, due to the small dielectric constant of the air, the energy of the absorbed electromagnetic wave is low. Compared with other single feature datasets, it is difficult for the noise energy to cause significant amplitude changes in a short time.

[0131] In the IP dataset, the network's automatic identification ability is poor. It is possible that IP reflects the continuity of adjacent signals. When electromagnetic waves propagate to different media, the phase of the reflected wave will change, and factors such as scattered waves will cause more noise in the dataset. Therefore, the IP test set has the lowest accuracy.

[0132] The accuracy of the IF dataset is slightly higher than that of the IP dataset. Since high-frequency electromagnetic waves are more sensitive to structural layers with different dielectric constants, the differential calculation involving the phase change of the signal makes the energy of the high-frequency noise drop faster on the spectrum, which helps to suppress certain noise components.

[0133] Compare the model test results of the original signal and the multi-feature fusion data set:

[0134] On the fusion feature dataset, IA+IF has the highest average precision. Compared with the original signal dataset with better verification effect, the verification index mAP and R of the IA+IF dataset are improved by 3.7% and 2.7%, respectively. This may be due to the fact that the IA and IF datasets contain less noise and have complementary advantages. The IA dataset contains location and morphological features, and the IF dataset contains boundary texture information of structural stratification and through-cracks.

[0135] This model performs the worst on three feature fusion datasets. This may be because a direct weighting of 1:1:1 was performed on the three spectral images, which contains more spectral information. However, during the weighting process, no weighting ratio was assigned, thus amplifying the noise and affecting the training and validation results of the model.

[0136] (1b) Comparison and analysis of ablation experiment results within the model

[0137] Table 2 Results of ablation experiment on the performance of the model with single - feature datasets

[0138]

[0139] Table 3 Results of ablation experiment on the performance of the model with multi - feature fusion datasets

[0140]

[0141] In (1a), the influence of different datasets on the model performance was verified, and the improved module of the designed two - stage network structure optimized based on Swin Transformer was verified. Four improvements were mainly made to the two - stage network structure: deepening the feature extraction network and introducing the Swin Transformer deep learning network; adding the FPN structure for top - down feature map fusion; using ROIAlign to replace the original ROI pooling method; and finally, during the training process, using Smooth L1 as the bounding box loss function. The verification of the improvement of each module in the model for different datasets shows that the experimental results of the mAP@0.5 metric are shown in Table 2 and Table 3.

[0142] In this embodiment, a total of five groups of ablation experiments were designed, including the verification of 8 different datasets. Taking the Origin dataset as a reference, when the backbone network of the FasterRCNN network was replaced with Swin Transformer, the metrics on all datasets were significantly improved, and the mAP@0.5 increased by 8%. This shows that Swin Transformer can effectively improve the feature extraction ability of the model. Then, after adding the FPN structure, the metrics of the model were significantly improved, and the mAP@0.5 increased by approximately 2%. On this basis, using the ROIAlign pooling method, the network performance also had a good improvement of about 2%. Finally, when the bounding box loss function was replaced with Smooth L1, the accuracy increased slightly by less than 1%. Finally, compared with the original classic two - stage network Faster R - CNN, the two - stage network optimized based on Swin Transformer had an increase of approximately 13% in mAP@0.5. The experiment proves that for each dataset, the four improvement strategies introduced in technical solution step 3 all improve the accuracy of the model to varying degrees, and the combined use has a more obvious improvement effect, effectively improving the accuracy of through - crack identification.

[0143] (1c) Comparison and Analysis of Experimental Results between Models

[0144] Taking the original signal dataset as the comparison benchmark, the through-crack two-stage detection model optimized based on SwinTransformer will be compared with the following benchmark networks: Faster R-CNN, SSD, and YOLOv5 proposed in 2021. The reason for the comparison is that Faster R-CNN is a classic two-stage detection network, SSD and YOLOv5 are one-stage networks, and YOLOv5 also uses the FPN structure. Comparing with these networks can highlight the advantages and disadvantages of the two-stage network optimized based on Swin Transformer. Additionally, the floating-point operation count (FLOPs), the number of parameters, and the model file size after training are added. FLOPs refers to the number of floating-point operations per second of a computer and can measure the computational complexity of the model. The number of parameters is the sum of the parameters of all modules in the model and can measure the size of the model. For example, a 3×3 convolutional layer contains 9 parameters in the convolutional kernel and one bias parameter. Model Size is the space it occupies on the hard disk.

[0145] Table 4 Comparison of Detection Results between Models

[0146]

[0147] As can be seen from Table 4, the road through-crack detection model proposed in step 3 of the technical solution is superior to other algorithms in terms of the mAP@0.5 index. Compared with Faster R-CNN, the accuracy is about 13% higher; compared with YOLOv5 and SSD with one-stage network structures, the mAP@0.5 is about 11% and 15% higher respectively. The improved network in step 3 of the technical solution has a significant improvement in detection accuracy compared with the classic network, indicating the effectiveness of the improvement strategy introduced in step 3 of the technical solution.

[0148] Example 2: One-stage Through-Crack Detection Model Optimized Based on Swin Transformer-YOLOv8

[0149] (2a) Comparison and Analysis of Experimental Results on Different Datasets

[0150] To compare the performance of the multi-feature fusion dataset in the one-stage detection network, the original signal and seven other feature datasets were trained and tested. Through different comparative experiments, the role of different datasets in improving the model performance was evaluated. Among the 8 datasets, the IA+IF training set performed outstandingly, with stable and favorable trends in the metrics during its training process. For example, Figure 9 as shown, both its learning potential and generalization ability are superior to those of the original signal dataset, single-feature datasets, and the other three fusion datasets.

[0151] Corresponding to the training process, the PR curve of this dataset also shows relatively high precision and recall. As shown in Figure 10 it further verifies that IA+IF has good application effects on the designed one-stage detection model for through cracks.

[0152] In addition, to provide more comprehensive and detailed data support, the detection accuracy data of more datasets were also sorted out and presented in Tables 5 and 6. These tables detail the performance of each dataset under different evaluation metrics.

[0153] Table 5 Experimental results of single-feature dataset mAP@0.5 (%)

[0154]

[0155] Table 6 Experimental results of feature fusion dataset mAP@0.5 (%)

[0156]

[0157]

[0158] After a detailed analysis of the test results, it was found that for the one-stage object detection network, the combination of IA+IF in the multi-feature fusion dataset performed particularly well among the eight datasets, with a detection accuracy of 86.2%, slightly higher than that of the original signal dataset by 2.4%. This result further confirms the positive effect of the multi-feature fusion process on enhancing crack features, thus significantly improving the model's automatic recognition ability. However, the IA+IP+IF dataset did not bring about an improvement in performance in the one-stage detection network either, and it was also the fusion dataset with the worst performance. This shows that in the data fusion process, the influence of noise interference cannot be ignored, and excessive or improper fusion may introduce negative effects, thereby affecting the model's performance. Therefore, when performing feature fusion, it is necessary to select feature combinations to make full use of the complementarity between features and minimize noise interference as much as possible.

[0159] (2b) Comparison and analysis of ablation experiment results within the model

[0160] Meanwhile, to verify the improvement brought by the two improvement strategies of fusing YOLOv8 and Swin Transformer and the BiFPN module to the one-stage through-crack detection network model, three groups of ablation experiments were designed. Taking the detection accuracy as the index, the two improved parts of the model were evaluated and analyzed. The ablation experiment results of the detection accuracy are also shown in Tables 5 and 6. Figure 11 、 12 Better reflects the change trend of the detection accuracy.

[0161] From Tables 5, 6 and Figure 11 、 12 It can be seen that when YOLOv8 was not optimized, mAP@0.5 reached 71.6%. This shows the effectiveness of YOLOv8 as a basic model in the through-crack detection task. Then, when the Swin Transformer structure was introduced and fused with YOLOv8, it was observed that the accuracy of the optimized model increased on all datasets. Especially on the Origin dataset, the accuracy increased by about 8%. This improvement indicates that by fusing the complex network structure of Swin Transformer, the model's ability to automatically extract features has been enhanced. The introduction of Swin Transformer enhanced the model's perception ability of targets with different scales and deformations in the image, thus improving the detection accuracy.

[0162] Furthermore, based on the Swin Transformer-YOLOv8 model, the BiFPN structure was introduced in the Neck part, which increased the accuracy by about 4%. And the effectiveness of the BiFPN structure in improving the model accuracy was verified.

[0163] Figure 13 The detection effects of YOLOv8 and Swin Transformer-YOLOv8 that integrates the two improvement strategies were compared. It can be seen from the figure that for some cracks with unclear features, the optimized Swin Transformer-YOLOv8 can detect them while the original YOLOv8 cannot. This is because the BiFPN structure introduced in the Neck part performs weighted fusion of feature maps from different layers during feature fusion, so that the network can extract some unclear features in the original image when extracting features.

[0164] (2c) Comparison and analysis of experimental results among models

[0165] Select some one-stage and two-stage object detection models to reproduce on the ground penetrating radar image multi-feature fusion dataset, such as YOLOv7, Faster R-CNN, and the two-stage network optimized based on Swin Transformer proposed in step 3 of the technical solution. In addition to the metrics proposed above, FPS is also introduced, which is a metric that measures the detection speed of the model, that is, the number of images the model can detect in one second (unit: frame / s). The comparison results on the through-crack original signal dataset are shown in Table 7.

[0166] Table 7 Comparison of Model Detection Results

[0167]

[0168] After comparing with the two-stage detection model and other networks proposed in step 3 of the technical solution, the one-stage detection network of YOLOv8 fused with Swin Transformer introduced in step 4 shows excellent performance in multiple key metrics. Specifically, the three metrics that measure the model complexity and computational burden, namely Model Size, Params, and FLOPs, all rank first among the compared models, which fully demonstrates the remarkable effectiveness of the Swin Transformer-YOLOv8 optimized network in alleviating the device storage pressure and reducing the computational complexity.

[0169] From the perspective of detection accuracy, although it is slightly lower by about 1% compared with the two-stage detection model optimized by SwinTransformer in step 3 of the technical solution, in practical applications, such an accuracy gap is usually acceptable. More importantly, the FPS (frames per second) of this one-stage detection network has increased from 10.37 frames / s to 35.71 frames / s. Since the ground penetrating radar test images used correspond to 12 cm in the actual road, it can be calculated that the detection speed of this model can reach 4.2 m / s, that is, the detection speed is 15.42 km / h, an increase of about 10 km / h, and the detection speed has been significantly improved.

[0170] In addition, although the FPS of the Swin Transformer-YOLOv8 optimized network is about 1 frame / s slower than that of the YOLOv7 network, it has achieved a 2.6% improvement in accuracy. The trade-off between accuracy and speed is of great significance in practical applications, especially in scenarios with certain requirements for both accuracy and real-time performance.

[0171] Generally speaking, the one-stage detection network optimized by Swin Transformer-YOLOv8 maintains a high detection accuracy while significantly optimizing the model complexity and computational amount, making it more efficient and stable in processing actual detection tasks.

[0172] (2d) Model applicability analysis

[0173] The two-stage through-crack detection optimized based on Swin Transformer designed in step 3 of the technical solution has good accuracy in terms of precision. However, the FPS of this model for detecting and identifying through-crack images is only 10.7 frames per second, that is, it can only detect 10 images per second on average, and the detection rate can only reach 4.6 km / h. Therefore, this two-stage through-crack detection optimized based on Swin Transformer has high accuracy but low detection rate, and is suitable for the scenario of key area detection.

[0174] The one-stage through-crack detection model optimized by SwinTransformer-YOLOv8 designed in step 4 of the technical solution has a detection rate of 37 images per second, with an FPS 25.34 frames higher, and a detection speed of 15.42 km / h on the premise that the mAP@0.5 is only about 1% lower than the detection model in step 3 of the technical solution. Therefore, the one-stage through-crack detection model optimized by Swin Transformer-YOLOv8 is suitable for some general detection scenarios, that is, scenarios with a wide detection range, short detection time, and real-time detection while ensuring a certain detection accuracy.

[0175] The two-stage detection network based on SwinTransformer of the present invention introduces three optimization strategies. First, combined with the FPN structure, it makes full use of the position information of the underlying feature maps to improve the feature utilization rate of the network. Secondly, in order to avoid losing the features of some tiny cracks during the use of nearest neighbor interpolation in ROIpooling, the pooling method in the network is improved to ROIAlign using bilinear interpolation to improve the feature extraction ability for small cracks. Finally, in order to prevent the problem of gradient explosion during the training process, the bounding box loss function adopts Smooth L1.

[0176] Moreover, in order to improve the detection speed, the present invention optimizes Swin Transformer-YOLOv8 to meet the requirements of rapid detection of road through cracks. The backbone network of YOLOv8 integrates Swin Transformer, and then introduces BiFPN to adjust the contribution degrees of different feature layers. The training and test results of the optimized one-stage network of Swin Transformer-YOLOv8 on different feature datasets are compared, further proving that feature fusion can enhance crack features, and it has the best performance on IA+IF, reaching 86.2%. Then, through the ablation experiment of the model structure, it is proved that the detection accuracy of the optimized Swin Transformer-YOLOv8 network has been effectively improved, reaching 83.9% on the original signal dataset. And compared with the optimized network of the two-stage detection network based on Swin Transformer and other networks, it is found that its various indicators meet the design concept. While maintaining a certain accuracy, the detection speed FPS reaches 35.71 frames / s, laying a foundation for real-time detection.

[0177] The above are only preferred embodiments of the present invention, and do not impose any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer, characterized in that It includes the following steps: Step 1: Collect ground penetrating radar data, and perform preprocessing and feature extraction Use a ground penetrating radar device to scan the asphalt pavement and collect radar signal data below the pavement; process the original radar data; extract key information helpful for crack detection from the preprocessed data; Step 2: Construct a multi-feature fusion dataset; Combine the instantaneous amplitude, instantaneous phase, and instantaneous frequency feature images of the ground penetrating radar signal to form a multi-feature fusion dataset; then expand and annotate the dataset to complete the construction of the multi-feature fusion dataset for ground penetrating radar through-crack detection; Step 3: Design a two-stage through-crack detection model optimized based on SwinTransformer (3a) Optimize the overall framework of the network Introduce Swin Transformer into the backbone network, introduce the FPN network after the backbone network, and improve the two-stage detection network by replacing ROI pooling with ROIAlign by combining the advantages of Mask R-CNN; (3b) Design of the FPN multi-scale fusion module Generate feature maps with different depths and scales through a top-down feature extraction network, dynamically adjust the resolution of the deepest feature map using upsampling, fuse the feature maps of two adjacent layers with different scales, and use the feature map containing the location information of the shallow network and the semantic information of the deep network after fusion for target detection; (3c) Optimization of the ROI pooling method After calculating the values of four positions through bilinear interpolation, obtain the ROI output of the same size through Maxpooling; (3d) Optimization of the Smooth L1 loss function The bounding box position loss function uses the Smooth L1 function, and the Smooth L1 loss function combines the L1 loss function and the L2 loss function; Step 4: Design a one-stage through-crack detection model optimized based on SwinTransformer-YOLOv8 (4a) Integrate the Swin Transformer network Integrate YOLOv8 and Swin Transformer, and introduce the SwinTransformer network in the Backbone part of YOLOv8; (4b) Weighted bidirectional feature fusion network First, continuously upsample the deep features and fuse them with the bottom features. After the upsampling is completed, then downsample the bottom features and fuse them with the deep features, and add two horizontal connection paths between the two feature extraction paths to fuse the feature maps generated at each stage in the backbone network with the feature maps to be detected; (4c) Overall architecture of the Swin Transformer-YOLOv8 model YOLOv8 integrates Swin Transformer and introduces a weighted bidirectional feature pyramid (BiFPN); Step 5: Train and predict the training set and test set data according to the model architectures in Step 3 and Step 4.

2. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 1, wherein In the said Step 1, the ground penetrating radar data includes the position and depth of the cracks; The preprocessing includes zero-offset correction, FIR filtering, and signal gain adjustment processing; Feature extraction includes amplitude information, frequency information, and phase information.

3. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 1, characterized in that, In step 2, the methods of random cropping, mirror flipping, and contrast enhancement are used to augment the dataset. Finally, 8 datasets are labeled to complete the construction of the multi-feature fusion dataset for the ground-penetrating radar through cracks.

4. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 1, characterized in that In step 3, the L1 and L2 loss functions are as follows: where f(x i ) and y i represent the predicted value and the corresponding true value of the i-th sample respectively, and n is the number of samples.

5. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 1, characterized in that, In step 3, the formula for Smooth L1 is as follows: This function is a piecewise function. It is the L2 loss between [-1, 1], which solves the problem of the L1 having a breakpoint at 0. Outside the [-1, 1] interval, it is the L1 loss, which solves the problem of gradient explosion for outliers. Therefore, it can limit the gradient from the following two aspects: one is that when the error between the predicted value and the true value is too large, the gradient value will not be too large; the other is that when the error between the predicted value and the true value is very small, the gradient value is small enough.

6. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 1, characterized in that The specific method adopted in step 5 is as follows: Adopt the multi-feature dataset of through cracks constructed in step 2; In step 3, the SGD optimizer is used as the basic iterator, the weight_decay weight decay factor is 0.0001, the momentum factor is 0.9, and a single GPU is used for training. The batchsize is set to 4 in both the original network and the improved network, and the corresponding initial learning rates are set to 0.001 respectively; during the training process, the model first loads the initialization weight file of the Swin Transformer trained on the public dataset ImageNet, and the total number of training iterations is set to 100; In step 4, the SGD optimizer is used as the basic iterator, the initial learning rate is set to 0.001, the weight_decay weight decay factor is 0.0005, and the momentum factor is 0.937; a single GPU is used for training, the batchsize is set to 4, and considering the comprehensive training speed and the original image size, the total number of training iterations is set to 100 rounds; In step 3 and step 4, the accuracy, average precision, mean average precision, and P-R curve are used as evaluation indicators to measure the performance of the model; In step 3, for the obtained original signal, single-feature, and multi-feature fusion datasets, a two-stage network structure optimized based on Swin Transformer is used for training.

7. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 6, wherein The accuracy rate represents the proportion of correctly detected results among all model detection results, and the recall rate represents the proportion of correctly detected results among all true annotation boxes, as shown in the following formula: Where TP represents the correct recognition by the model, that is, the number of cracks correctly detected; FP represents the incorrect recognition by the model, that is, the number of cracks misdetected; FN represents the positive sample being judged as a negative sample, that is, the number of cracks missed; P and R can both measure the detection effect of the model in some scenarios, but these two indicators are inversely proportional. Therefore, AP is needed to comprehensively evaluate the model by combining the P and R indicators.

8. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 6, wherein The average precision (AP) value is calculated from the area enclosed by the P-R curve and the coordinate axes. The larger the area of the P-R curve, the better the model performance. The formula is as follows:

9. The lightweight crack penetration degree detection method optimized based on ground penetrating radar and Swin Transformer according to claim 6, characterized in that, The mean average precision (mAP) is the evaluation metric that best demonstrates the performance of the object detection model. mAP calculates the average of the APs for all classes in the dataset, and the calculation formula is as follows: Where n represents all classes in the dataset, and mAP includes three evaluation criteria: mAP@0.5, mAP@0.75, and mAP@0.5:0.95; mAP@0.5 and mAP@0.75 refer to the mAP values when IoU = 0.5 and IoU = 0.75 respectively, and mAP@"0.5:0.95" refers to the average of the mAP values corresponding to each IoU when IoU varies from 0.5 to 0.95 with a step of 0.05.

Citation Information

Patent Citations

  • Mountain crack detection method based on improved self-attention mechanism and transfer learning

    CN114022770A

  • Road surface crack small target identification method based on lightweight algorithm

    CN118781325A