Road dielectric constant inversion and disease discrimination method based on neural network

By constructing the GPR-TransUNet neural network model, the problems of low computational efficiency and limited accuracy of existing GPR dielectric constant inversion technology are solved, realizing end-to-end inversion of high-resolution dielectric constant and accurate detection of multi-scale defects.

CN121482553BActive Publication Date: 2026-04-10CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing GPR dielectric constant inversion technology suffers from low computational efficiency, limited inversion accuracy, excessive manual intervention, and poor adaptability to complex defects, making it difficult to meet the demand for efficient and accurate detection of hidden road defects.

Method used

A GPR-TransUNet neural network model was constructed. By extending skip connections and introducing a gated attention mechanism, a multi-scale feature fusion module, an asymmetric depth-space attention module, and a deep supervised decoding head were designed. Combined with an adaptive region-weighted loss function and an AdamW optimizer, high-resolution dielectric constant inversion and disease identification were achieved.

Benefits of technology

It achieves end-to-end inversion of high-resolution dielectric constant, accurately captures multi-scale disease characteristics, improves detection efficiency and accuracy, and the simulation verification results are highly consistent with borehole endoscopy detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482553B_ABST
    Figure CN121482553B_ABST
Patent Text Reader

Abstract

The application relates to a neural network-based road dielectric constant inversion and disease identification method. The method comprises the following steps: based on the baseline expansion of TransUNet, the number of jump connections is increased, and a gating attention mechanism is introduced; a multi-scale feature fusion module, an asymmetric depth-space attention module and a deep supervision decoding head are added; an AdamW optimizer and a Warmup-cosine annealing learning rate scheduling strategy are used; and a self-adaptive regional weighted loss function is constructed to amplify the dielectric constant mutation boundary loss weight. After the model is trained through a road GPR simulation data set and is adjusted to convergence through a verification set, a GPR-B-scan image to be detected is input, a dielectric constant distribution is output, and according to the dielectric constant characteristics of different diseases, precise identification of hidden diseases such as cavities, emptying, loose bodies and water-rich bodies can be realized. The method can realize high-resolution dielectric constant end-to-end inversion and precise capture of multi-scale disease characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image detection, in particular to a road dielectric constant inversion and disease discrimination method based on a neural network. BACKGROUND

[0002] As a national key infrastructure, the structural safety and service life of the road are directly related to the traffic efficiency and public safety. In the long-term service process, the internal structure of the road is prone to hidden diseases such as cavities, voids, loose bodies and water-rich bodies due to repeated load action, environmental erosion and other factors. Such diseases are difficult to be found by visual inspection in the early stage, but they can significantly reduce the bearing capacity and durability of the road. If not treated in time, they can easily cause serious problems such as road collapse and cracking, so there is an urgent need for efficient and accurate detection technology for internal diseases of the road.

[0003] Ground penetrating radar (GPR) has become the mainstream technology for detecting the internal structure of the road due to its advantages of non-destructive, high resolution and fast imaging. Its principle is to generate a B-scan spectrum reflecting the internal structure of the road by transmitting high-frequency electromagnetic waves and receiving the reflected signals of the underground medium. However, the original GPR spectrum is only the strength profile of the electromagnetic wave reflection signal, which cannot directly represent the physical properties of the material. It needs to be inverted into the relative dielectric constant, which is the most sensitive parameter to indicate the health of the road. The dielectric constant of the road material is usually 3-10, the dielectric constant of water is as high as 81, and the dielectric constant of air is close to 1. The distribution of dielectric constant can directly reflect the distribution of internal diseases of the road.

[0004] Current dielectric constant inversion methods can be mainly divided into two categories: one is the physically driven method, which realizes inversion by iteratively matching the observed and simulated waveforms. However, it has the defects of high computational complexity, sensitivity to initial model and susceptibility to noise interference. Especially, three-dimensional FWI can improve the spatial coverage, but it is too time-consuming to meet the efficient detection needs of engineering. The second is the data-driven method. Traditional machine learning relies on manual feature design, and its prediction ability for complex underground structure dielectric constant is limited, making it difficult to achieve fine inversion. Although existing deep learning methods can extract deep features, CNN lacks the ability to model the global context, which can easily lead to "one picture with multiple interpretations" of GPR spectrum. Traditional skip connection can easily introduce redundant noise, and the multi-scale disease feature adaptation is insufficient, the boundary detail prediction is fuzzy, which cannot meet the engineering needs of accurate discrimination of road hidden diseases.

[0005] In summary, the existing GPR dielectric constant inversion technology has the bottlenecks of low computational efficiency, limited inversion accuracy, high manual intervention and poor adaptability to complex diseases, which makes it difficult to efficiently support road maintenance decisions. SUMMARY

[0006] Therefore, it is necessary to provide a neural network-based road dielectric constant inversion and disease identification method capable of realizing high-resolution end-to-end dielectric constant inversion and accurately capturing multi-scale disease characteristics to solve the above technical problems.

[0007] A neural network-based road dielectric constant inversion and disease identification method, the method comprising:

[0008] A road GPR simulation dataset is constructed; the number of skip connections in a TransUNet baseline model is expanded, a gating attention mechanism is introduced at each skip connection, and a multi-scale feature fusion module, an asymmetric depth-space attention module, and a deep supervision decoding head are designed to construct a GPR-TransUNet neural network model;

[0009] AdamW is used as an optimizer, a learning rate scheduling mode is set in combination with a Warmup-cosine annealing strategy, an adaptive regional weighted loss function is constructed, and the adaptive regional weighted loss function amplifies the loss weight of the dielectric constant mutation boundary through a dynamic weighting mechanism;

[0010] The training set in the road GPR simulation dataset is input into the GPR-TransUNet neural network model, the model is trained using the configured training strategy and the adaptive regional weighted loss function, the loss change and evaluation index fluctuation of the model in the training process are monitored through the validation set, the model parameters are adjusted until the model converges, and a trained GPR-TransUNet neural network model is obtained;

[0011] The GPR-B-scan image of the road to be detected is input into the trained GPR-TransUNet neural network model, and the dielectric constant distribution of the road to be detected is output; according to the dielectric constant characteristics corresponding to different diseases in the dielectric constant distribution, the hollow, void, loose body, and water-rich body diseases in the road to be detected are identified, and the road concealed disease identification is completed.

[0012] The above-mentioned neural network-based road dielectric constant inversion and disease discrimination method, first, a simulation data set is constructed through the gprMax platform, which covers the real structure characteristics of the road and provides rich disease feature samples for the model, avoids the generalization ability caused by single data, and lays a data foundation for subsequent accurate inversion. In the TransUNet baseline model, the number of skip connections is expanded, a gating attention mechanism is introduced at each skip connection, a multi-scale feature fusion module, an asymmetric depth-space attention module and a deep supervision decoding head are designed to form a GPR-TransUNet neural network model. The asymmetric depth-space attention module decouples the physical features of the vertical depth axis and the horizontal scanning axis of the GPR graph, accurately captures the correlation of the layered structure and the local anomaly; the multi-scale feature fusion module extracts features in parallel through multiple receptive fields, adapts to different scale diseases such as small cavities and large range cavities, and the ablation experiment proves that it has the greatest contribution to the accuracy improvement; the extended skip connection and the introduced gating attention mechanism filter the effective information in the low-level detailed features of the encoder and the redundant noise, solving the problem of interference introduced by traditional skip connection; the deep supervision decoding head is constrained by multiple layers of supervision, guiding the model to learn the dielectric constant mapping rule at each scale, avoiding the loss of boundary details in the decoding process, and the four cooperate to strengthen the modeling ability of the model to global semantics and local details, providing structural support for high-resolution inversion. The adaptive regional weighted loss function amplifies the loss weight of the dielectric constant mutation boundary by dynamically weighting, solves the problem of insufficient attention to disease boundaries in traditional loss, and improves the edge description accuracy; the AdamW optimizer combined with the Warmup-cosine annealing strategy avoids the local optimum caused by improper learning rate in the early training, and through the smooth speed reduction in the later period, it ensures the convergence stability and ensures that the model fully learns the data features. Finally, the trained model can directly map the GPR-B-scan image to be detected into a dielectric constant distribution, and realize accurate discrimination combined with the dielectric characteristics of different diseases. The inversion result in the simulation verification is highly consistent with the drilling endoscopy detection, which verifies the effectiveness of the present application in end-to-end inversion and multi-scale disease capture. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A flowchart of a neural network-based road dielectric constant inversion and disease discrimination method in one embodiment;

[0014] Figure 2 A schematic diagram of a GPR-TransUNet neural network model in one embodiment;

[0015] Figure 3 A network structure schematic diagram of a multi-scale feature fusion (MSF) module in one embodiment;

[0016] Figure 4Flowchart of the gating attention mechanism (GSF) in another embodiment;

[0017] Figure 5 Flowchart of the deep supervision mechanism (DS-Head) in an embodiment;

[0018] Figure 6 Structure diagram of the asymmetric deep-spatial attention mechanism (ASDS) in an embodiment;

[0019] Figure 7 GPR-TransUNet training loss and evaluation index change curve diagram in an embodiment; Figure 7 (a) is a loss change curve diagram, Figure 7 (b) is an index change diagram;

[0020] Figure 8 Main model dielectric inversion result diagram in an embodiment;

[0021] Figure 9 GPR-TransUNet inversion result diagram for different types of diseases in an embodiment;

[0022] Figure 10 GPR-TransUNet inversion result and endoscopic view diagram of the ground penetrating radar measured spectrum in an embodiment; Figure 10 (a) is a road section 1 diagram; Figure 10 (b) is a dielectric constant inversion result diagram of road section 1; Figure 10 (c) is an endoscopic view diagram of road section 1; Figure 10 (d) is a road section 2 diagram; Figure 10 (e) is a dielectric constant inversion result diagram of road section 2; Figure 10 (f) is an endoscopic view diagram of road section 2; Figure 10 (g) is a road section 3 diagram; Figure 10 (h) is a dielectric constant inversion result diagram of road section 3; Figure 10 (i) is an endoscopic view diagram of road section 3. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0024] In an embodiment, as shown in Figure 1 , a neural network-based road dielectric constant inversion and disease discrimination method is provided, comprising the following steps:

[0025] Step 102, construct a road GPR simulation dataset; extend the number of skip connections in the TransUNet baseline model, introduce a gated attention mechanism at each skip connection, and design a multi-scale feature fusion module, an asymmetric depth-space attention module, and a deep supervision decoding head to construct a GPR-TransUNet neural network model.

[0026] The road GPR simulation dataset is the basis for model training. The dataset constructed by the gprMax three-dimensional forward platform can simulate the homogeneous and heterogeneous structures of real roads, and contains four typical diseases of cavities, voids, loose bodies and water-rich bodies, ensuring the diversity and authenticity of the data.

[0027] TransUNet uses 3 skip connections for feature fusion. This application extends to 4 skip connections, and introduces a gated attention mechanism (GatedAttention) at each skip connection, allowing the model to adaptively select and enhance key features. Focus on more discriminative areas to improve the accuracy of the prediction details.

[0028] Because the features output by the Transformer have differences in spatial layout compared to convolutional features, this application designs an MSF module to reconstruct the global features of the Transformer into a convolution-friendly tensor representation, and then inputs the decoder after supplementing multi-scale local information.

[0029] As Figure 2 shown is the GPR-TransUNet neural network model structure. GPR-TransUNet improves the effectiveness of the training signal by using a hybrid loss function + deep supervision mechanism, optimizes the convergence of the training process by using Warmup + cosine annealing, and improves the feature expression capability by attention enhancement and multi-scale fusion. The improved network can more accurately restore the underground dielectric constant distribution. The multi-scale feature fusion (MSF) module network structure is shown in Figure 3 .

[0030] The road GPR simulation dataset is the basis for model training. The dataset constructed by the gprMax three-dimensional forward platform can simulate the homogeneous and heterogeneous structures of real roads, and contains four typical diseases of cavities, voids, loose bodies and water-rich bodies, ensuring the diversity and authenticity of the data.

[0031] The TransUNet baseline model combines the local feature extraction capability of CNN and the global context modeling advantage of Transformer. The expansion of the number of skip connections can enhance the integrity of feature transmission. The addition of the gating attention mechanism, multi-scale feature fusion module, asymmetric depth-spatial attention module, and deep supervision decoding head solves the problems of noise interference, poor multi-scale disease adaptation, insufficient feature extraction, and boundary detail loss in traditional models, and cooperatively improves the inversion accuracy of the model.

[0032] In step 104, AdamW is used as an optimizer, combined with a Warmup-cosine annealing strategy to set the learning rate scheduling method, and an adaptive regional weighted loss function is constructed. The adaptive regional weighted loss function amplifies the loss weight of the dielectric constant mutation boundary through a dynamic weighting mechanism.

[0033] The AdamW optimizer effectively reduces the risk of overfitting and improves the stability of parameter updates through weight decay. The Warmup-cosine annealing strategy linearly increases the learning rate in the early stage of training to avoid the model converging to a local optimum too early, and then reduces the learning rate through a cosine function to achieve smooth convergence. The adaptive regional weighted loss function dynamically amplifies the loss weight of the boundary region based on the characteristics of the dielectric constant mutation disease boundary pixels, which have a small proportion but are important in physical meaning, to solve the problem of blurred inversion results caused by the traditional loss function's lack of attention to the boundary, and to improve the edge depiction accuracy.

[0034] In step 106, the training set in the road GPR simulation dataset is input into the GPR-TransUNet neural network model, and the configured training strategy and adaptive regional weighted loss function are used for model training. The loss change and evaluation index fluctuation of the model during training are monitored through the validation set, and the model parameters are adjusted until the model converges, obtaining the trained GPR-TransUNet neural network model.

[0035] The training set provides sufficient learning samples for the model, enabling the model to master the mapping relationship between GPR-B-scan images and dielectric constant distribution. The validation set is used to monitor the training effect of the model in real time. By observing the loss change curve and evaluation index (MAE, RMSE, R²) fluctuation, it is determined whether the model is overfitting or underfitting, and then the model parameters (such as learning rate, iteration number, etc.) are adjusted to ensure that the model has good generalization ability and inversion accuracy after convergence.

[0036] Step 108, input the GPR-B-scan image of the road to be detected into the trained GPR-TransUNet neural network model, output the dielectric constant distribution of the road to be detected; according to the dielectric constant characteristics corresponding to different diseases in the dielectric constant distribution, identify the cavity, void, loose body and water-rich body diseases in the road to be detected, and complete the road hidden disease discrimination.

[0037] The trained model can realize end-to-end mapping from GPR original signal to dielectric constant distribution, without manual intervention, and improve the detection efficiency. The dielectric constant characteristics corresponding to different diseases have significant differences. Air (dielectric constant ≈ 1) corresponds to cavity and void diseases, water (dielectric constant ≈ 81) corresponds to water-rich body diseases, and road materials (dielectric constant 3-10) and uneven distribution correspond to loose body diseases. According to the characteristics, the disease type and distribution range can be accurately distinguished.

[0038] The above-mentioned road dielectric constant inversion and disease discrimination method based on neural network first constructs a simulation data set through the gprMax platform, which covers the real structure characteristics of the road and provides rich disease feature samples for the model, avoids the insufficient generalization ability caused by single data, and lays a data foundation for subsequent accurate inversion. In the TransUNet baseline model, the number of skip connections is expanded, a gating attention mechanism is introduced at each skip connection, a multi-scale feature fusion module, an asymmetric depth-space attention module and a deep supervision decoding head are designed to form a GPR-TransUNet neural network model. The asymmetric depth-space attention module decouples the physical features of the vertical depth axis and the horizontal scanning axis of the GPR graph, accurately capturing the correlation between the layered structure and the local anomaly. The multi-scale feature fusion module extracts features in parallel through multiple receptive fields, adapts to different scale diseases such as small cavities and large range cavities, and the ablation experiment proves that it contributes most to the accuracy improvement. The extended skip connection and the introduced gating attention mechanism filter the effective information in the low-level detailed features of the encoder and filter the redundant noise, solving the problem of interference introduced by traditional skip connection; the deep supervision decoding head constrains the model through multi-layer supervision to learn the dielectric constant mapping rule at each scale, avoiding the loss of boundary details in the decoding process. The four work together to strengthen the modeling ability of the model for global semantics and local details, providing structural support for high-resolution inversion. The adaptive regional weighted loss function dynamically weights the loss weight of the dielectric constant mutation boundary to solve the problem of insufficient attention to disease boundaries by traditional loss, improving the edge description accuracy; the AdamW optimizer combined with the Warmup-cosine annealing strategy avoids local optimization caused by improper learning rate in the early training stage, and through the smooth speed reduction in the later stage, it ensures the convergence stability and ensures that the model fully learns the data features. Finally, the trained model can directly map the GPR-B-scan image to be detected to the dielectric constant distribution, and realize accurate discrimination combined with the dielectric characteristics of different diseases. The inversion results in the simulation verification are highly consistent with the drilling endoscopy detection, which verifies the effectiveness of the application in end-to-end inversion and multi-scale disease capture.

[0039] In one of the embodiments, the process of constructing the road GPR simulation data set includes:

[0040] Based on the gprMax three-dimensional forward platform, homogeneous and heterogeneous road models are built, four typical diseases of cavities, cavities, loose bodies and water-rich bodies are randomly embedded in the two types of models, and multiple B-scan images of road GPR are generated; the B-scan images are divided into training set, validation set and test set according to the ratio of 8:1:1, and the real dielectric constant distribution of the corresponding B-scan image is extracted from the gprMax forward simulation result, which is mapped to a uint16 grayscale image consistent with the size of the B-scan image as label data, completing the construction of the data set.

[0041] Specifically, the forward simulation of road structure was carried out based on GprMax software, the road model simulation parameters are shown in Table 1, and the setting parameters of GprMax software are shown in Table 2. The ground penetrating radar image data set of various typical internal diseases of road was constructed.

[0042] The geometric and material parameters of the road structure were simplified according to the actual pavement engineering structure. The model is composed of 20 cm asphalt surface layer and 80 cm cement stabilized gravel base layer, with a total thickness of 1.0 m. A PLM layer is provided around the model to absorb the electromagnetic waves propagating to the boundary of the model and prevent them from reflecting back to the calculation area. To enhance the diversity of data, two types of basic models are designed. One is a homogeneous material model, i.e. the internal materials of asphalt layer and water stable layer are uniform and consistent, and the other is a non-homogeneous material model, in which different sizes of aggregates are distributed in the asphalt layer and water stable layer according to the gradation requirements to more realistically simulate the internal non-homogeneous structure of the road. On this basis, four basic diseases of cavity, void, loose body and water-rich body are randomly embedded in the homogeneous and non-homogeneous models respectively. The position, size and combination method are randomly set to enrich the diversity of disease performance.

[0043] Table 1

[0044]

[0045] Table 2

[0046]

[0047] A total of 2000 B-scan images were generated, including 1000 homogeneous and 1000 non-homogeneous models. According to the common planning principle of deep learning task, the data was randomly divided into training set, validation set and test set, with the ratio of 8:1:1. The training set was 1600, the validation set was 200, and the test set was 200.

[0048] The label data of this study was directly derived from the dielectric constant distribution generated by gprMax forward simulation. For each B-scan image, the corresponding true dielectric constant field was mapped to a uint16 grayscale image with the same input size, where the grayscale value of each pixel represented the dielectric constant at that position. During model training, uint16 was converted to float32, but the pixel value itself was not changed. The real value matrix of this pixel was used as the supervision signal, and the network output was the continuous dielectric constant distribution with the same size, realizing end-to-end pixel-level regression.

[0049] To facilitate the display of results and intuitive comparison, color mapping is used in the experiment to convert the predicted dielectric constant values into pseudo-color images, making different numerical intervals visually present as gradient color bands. This strategy not only preserves the continuous characteristics of the true dielectric constant, but also improves the legibility of the image, which helps to evaluate the prediction accuracy of the model in the transition between layers, disease boundaries and local details.

[0050] In one of the embodiments, the implementation process of the gated attention mechanism includes:

[0051] The high-level semantic features X output by the decoder in the GPR-TransUNet neural network model are spliced with the low-level detail features S output by the encoder in the channel dimension to obtain joint features Z; ;

[0052] The joint features Z are compressed in the channel by a first 1x1 convolutional layer, followed by batch normalization processing and ReLU activation function operations, and then a second 1x1 convolutional layer and a Sigmoid activation function operation are performed to generate gating weights consistent with the channel number of the low-level detail features;

[0053] The gating weights are multiplied by the low-level detail features element by element to obtain weighted screening features, which realizes the enhancement of effective features and the suppression of noise features in the skip connection;

[0054] The gating weights consistent with the channel number of the low-level detail features are generated as follows:

[0055] ;

[0056] wherein, represents the gating weight, is the weight parameter of the first 1x1 convolutional layer, is the weight parameter of the second 1x1 convolutional layer, is the Sigmoid function, is the ReLU function, is the batch normalization operation.

[0057] Specifically, a lightweight gating unit is embedded in the skip path, and the skip features are adaptively re-labeled by a learnable channel-level weight vector, realizing "useful information enhancement and noise component suppression", so as to improve the sensitivity and robustness of dielectric constant inversion to key scattering details. The structure diagram is as shown in Figure 5 .

[0058] The high-level semantic features X output by the decoder contain global context information, and the low-level detail features S output by the encoder contain local texture and edge information. The joint features Z formed by splicing the two types of features in the channel dimension have the advantages of both types of features. The first 1×1 convolutional layer is used to compress the channel number of the joint features Z, reduce the calculation amount and fuse the feature information. The batch normalization processing can accelerate the model training convergence, and the ReLU activation function enhances the non-linear expression ability of the features. The second 1×1 convolutional layer maps the features to a vector with the same channel number as the low-level detail features S. After the Sigmoid activation function, the gate weight G is obtained. The closer the value of G is to 1, the more important the feature at the corresponding position is, and the closer the value of G is to 0, the more redundant noise the feature at the position is. Element-wise multiplication of the gate weight G and the low-level detail features S can retain and enhance useful detail features and filter redundant noise, solving the problem of noise introduction caused by the semantic gap in traditional skip connection. The beneficial effect is that this mechanism realizes adaptive selection of skip connection features, improves the effectiveness of feature transmission, makes the model focus more on key features related to diseases, and thus improves the accuracy of dielectric constant inversion and the clarity of boundary details.

[0059] In one of the embodiments, the implementation process of the asymmetric depth-spatial attention module includes:

[0060] Performing linear transformation on the channel dimension of the module input feature map to generate a Value feature map with the same size as the input feature map through a 1×1 convolutional layer;

[0061] Respectively constructing a depth branch and a spatial branch to capture features in different dimensions, wherein the depth branch uses a k×1 size convolutional kernel to extract correlation features of the vertical axis of the GPR graph, and obtains a depth attention map;

[0062] The spatial branch uses a 1×k size convolutional kernel to extract correlation features of the horizontal axis of the GPR graph, and obtains a spatial attention map;

[0063] Adding the depth attention map and the spatial attention map, compressing to the range [0, 1] through the Sigmoid activation function to generate the final attention gate;

[0064] Element-wise multiplication of the attention gate and the Value feature map obtains the output feature map after re-calibration, which realizes accurate focusing on the vertical layered features and horizontal abnormal features in the GPR graph.

[0065] Specifically, Asymmetric Depth-Spatial Attention (ADSA) is a customized attention module for the physical characteristics of GPR (Ground Penetrating Radar) B-scan maps. Traditional convolutional neural networks usually use symmetric convolution kernels when processing images. This design assumes that the features of the image are isotropic in all spatial directions. However, GPR B-scan is physically anisotropic, with two dimensions having completely different physical meanings. The vertical axis represents the propagation time of the electromagnetic wave (depth). The features on this axis mainly reflect the layered structure (such as the road surface, base layer, and void) and signal attenuation. These features usually have strong continuity in the horizontal direction. The horizontal axis represents the scanning track of the radar. The features on this axis are mainly used to identify local anomalies.

[0066] Using a symmetric convolution kernel confuses these two completely different physical features. It forces the network to use the same tool to process both horizontal and vertical signals, which goes against the physical imaging principle of GPR and results in low feature extraction efficiency and weak ability to distinguish disease types.

[0067] The ADSA module solves the above problems by decoupling the two orthogonal physical axes of the GPR map, including the following steps:

[0068] Step 1: Linear Transformation

[0069] Linear transformation Convvalue This unit is used to perform a linear transformation on the input features Figure X of the channel dimension to generate a feature map V with the same size as the input (called the Value feature map).1 The convolution of 1 is to perform a linear reorganization and extraction of the content contained in the input x to generate a "to-be-processed" feature map V as follows:

[0070] ;

[0071] Step 2: Asymmetric Attention Mechanism

[0072] Depth branch, this branch is used to capture the correlation of the vertical axis (depth / time axis) in the GPR map. An asymmetric (k 1) convolution kernel is used, which can only move and perceive in the vertical direction and is not sensitive to changes in the horizontal direction. It is suitable for extracting the vertical disease electromagnetic wave reflection characteristics in the road surface and base layer of the GPR map. The feature extraction process of the depth branch is as follows:

[0073] ;

[0074] Spatial branch, this branch is used to capture the correlation of the horizontal axis (space / line axis) in the GPR map. An asymmetric (1 k) the convolution kernel, the receptive field of which can only move and perceive in the horizontal direction, and is insensitive to changes in the vertical direction. Suitable for identifying reflection patterns with certain lateral size caused by local abnormal bodies. The spatial branch feature extraction process is:

[0075] ;

[0076] Step 3: Attention gate fusion

[0077] This step is used to generate the final attention gate M. This fusion method enables M to activate both "vertical feature significant" and "lateral abnormal significant" regions. The fused attention map is compressed to the range [0, 1] using the Sigmoid activation, which converts it to a gating unit. The final attention gate generation process is:

[0078] ;

[0079] Step 4: Feature re-labeling

[0080] This step is the final application of the attention mechanism. The attention gate M generated in step 3 is multiplied element by element with the value V generated in step 1. The features in V corresponding to the high-value regions of M will be completely preserved, while the features in V corresponding to the low-value regions of M will be effectively suppressed. Finally, the output feature map Y is:

[0081] ;

[0082] This module decouples the vertical and horizontal dimensional physical features of the GPR profile, avoids the confusion of symmetric convolution kernels for different dimensional features, improves the relevance and effectiveness of feature extraction, and enables the model to more accurately capture road stratified structures and local disease abnormalities, providing support for high-resolution permittivity inversion.

[0083] In one embodiment, the construction process of the multi-scale feature fusion module includes:

[0084] Four parallel feature extraction branches are set up, each branch uses a 1x1, 3x3, 5x5, and 7x7 size convolution kernel for convolution operation, and padding is set in the convolution process to keep the spatial resolution of the output feature map consistent with the input feature map;

[0085] After the convolution operation in each branch, batch normalization processing and ReLU activation function operation are performed in sequence to enhance the feature nonlinear expression ability;

[0086] The output features of the four branches are spliced in the channel dimension, and a 1x1 convolutional layer is used to compress the channels and integrate the features, so that the number of output channels is restored to the number of input feature channels, and batch normalization and ReLU activation functions are performed again;

[0087] A learnable residual scaling factor is introduced to weight and integrate the integrated features and the module input features to obtain a multi-scale feature integration result.

[0088] Specifically, the multi-scale feature integration module designed in the application explicitly extracts multi-level features in the 1x1 to 7x7 receptive field range through a parallel multi-branch structure, and integrates them in an adaptive manner to enhance the model's multi-scale representation ability for complex underground structures, thereby improving the accuracy and robustness of the dielectric constant inversion. The structure diagram is shown in Figure 4 The module contains four parallel branches, each branch uses standard convolution operation with different kernel sizes (1x1, 3x3, 5x5, and 7x7), and introduces appropriate pendding while keeping the spatial resolution unchanged. After each convolution, batch normalization (Batch Normalization) and ReLU activation function are connected. The outputs of the four branches are spliced in the channel dimension. To further compress the channel dimension and enhance feature interaction, a 1x1 convolutional layer is introduced to integrate the spliced features, restore their channel number to C, and again pass through batch normalization and ReLU activation. Finally, the module introduces a learnable residual scaling factor to weight and integrate the fused features and the original input. This module extracts and adaptively integrates features through multiple receptive fields, allowing the model to adapt to different scales of diseases in the road and enhancing the model's multi-scale representation ability for complex underground structures. Ablation experiments confirm that it has the greatest contribution to improving the inversion accuracy.

[0089] In another embodiment, in a traditional encoder-decoder network, the supervision signal usually only acts on the output of the last layer. This single-point supervision mechanism easily leads to insufficient gradient transmission in the middle layers, which in turn affects the sufficiency of feature learning. At the same time, since deep features gradually lose detailed information during the upsampling process, the model often shows fuzzy and inaccurate defects when predicting object boundaries. Therefore, a deep supervision mechanism is introduced to constrain the decoding process at multiple scales. The specific structure is as Figure 6The output of each decoding layer is passed through a light segmentation head, and the base loss is calculated for each auxiliary output, respectively, and the weights are assigned as 0.5, 0.3, and 0.2, respectively. Finally, the final loss is obtained by adding the main output loss. By introducing supervision at different scale output layers, the model can simultaneously learn local details and global semantics, effectively constrain the low-level features containing rich spatial information, and promote the prediction ability of the network in the boundary and fine-grained regions.

[0090] In one embodiment, the calculation process of the adaptive region weighted loss function includes:

[0091] The spatial gradient information is extracted from the real permittivity distribution of the label data, and the gradient approximation value in the horizontal direction and the gradient approximation value in the vertical direction of the real permittivity distribution are calculated;

[0092] The square root of the sum of the squares of the gradient approximation value in the horizontal direction and the gradient approximation value in the vertical direction is calculated pixel by pixel, and the original size is restored by zero padding to synthesize the gradient amplitude map;

[0093] The gradient amplitude map is input into an auxiliary detection network composed of two convolution layers, ReLU nonlinear activation layers, and Sigmoid output layers to generate a dynamic key region weighted mask with the same size as the predicted permittivity map output by the GPR-TransUNet neural network model;

[0094] The base loss is calculated based on the global pixel-level error between the predicted permittivity map and the real permittivity distribution;

[0095] The key region loss is calculated based on the predicted permittivity map, the real permittivity distribution, and the dynamic key region weighted mask, and the adaptive region weighted loss function is obtained by weighting and summing the base loss and the key region loss through a preset coefficient.

[0096] In one embodiment, the base loss is calculated based on the global pixel-level error between the predicted permittivity map and the real permittivity distribution

[0097] ;

[0098] wherein, is the height of the feature map, is the width of the feature map, is the pixel value of the predicted permittivity map at position (i, j), is the pixel value of the real permittivity distribution at position (i, j).

[0099] In one embodiment, the key region loss is calculated based on the predicted permittivity map, the real permittivity distribution, and the dynamic key region weighted mask, including:

[0100] Based on the predicted dielectric constant map, the actual dielectric constant distribution, and the dynamic critical region weighted mask, the critical region loss is calculated as follows:

[0101] ;

[0102] in, For feature map height, The width of the feature map. To predict the dielectric constant plot in ( i, j The pixel value at position ) The true dielectric constant is distributed in ( i, j The pixel value at position ) Weighted mask for dynamic key regions in ( i, j The pixel value at the location.

[0103] Specifically, the core idea of ​​the adaptive region-weighted loss function is to utilize a trainable auxiliary network to automatically learn and amplify the loss weights for dielectric constant abrupt changes. Standard loss functions (MSE or L1) treat all pixels in an image equally. In the GPR dielectric constant inversion task, background pixels constitute the vast majority, while the boundary pixels of lesions, though fewer in number, are physically significant. Traditional loss functions cause the model to ignore the precision of these boundaries in order to reduce the overall average error, resulting in blurred boundaries, loss of detail, and inability to accurately delineate the extent of lesions. The adaptive region-weighted loss function addresses these issues through a dynamic weighting mechanism. It identifies which pixels represent key boundaries of dielectric constant abrupt changes and then applies a higher penalty to the prediction error of these pixels.

[0104] From the actual dielectric constant label image Spatial gradient information is extracted to locate regions where the dielectric constant changes abruptly. This step yields a gradient map. Used only as an auxiliary detection network The input is treated as a constant during training.

[0105] Calculate separately Gradient approximations in the horizontal and vertical directions and .

[0106] ;

[0107] ;

[0108] Calculated pixel by pixel and The square root of the sum of squares is used to synthesize the gradient magnitude map by restoring the original size by padding with zeros. and set the gradient back propagation to be blocked.

[0109] ;

[0110] The generated gradient map without gradient is input to the trainable auxiliary detection network , is a lightweight convolutional neural network composed of two convolutional layers, ReLU nonlinear activation layers and a Sigmoid output layer. The is forward propagated, and according to the saliency of the region, the dynamic key region weighted mask with the same size as the predicted permittivity map output by the main inversion network is automatically generated. . is a spatial attention map, and the value of each pixel on it is between [0, 1], which represents the probability or weight of the corresponding position being the disease boundary.

[0111] The basic loss is calculated and the global pixel-level error between The key region error of the disease boundary adjusted by the mask is calculated. This calculation is realized by simultaneously acting on and ,

[0112] ;

[0113] ;

[0114] This step is equivalent to applying a square weight to the pixel-wise square error between and This square weighting mechanism will amplify the error penalty in the area with high value, while suppressing the error impact in the area with low value.

[0115] The final total loss is obtained by weighted sum of and

[0116] .

[0117] ​​​The loss function solves the problem of insufficient attention to disease boundaries in traditional loss functions, and through a dynamic weighting mechanism, the model pays more attention to key boundary areas during training, improving the edge fidelity of permittivity inversion and providing clearer boundary basis for accurate disease identification.

[0118] In one of the embodiments, the learning rate is set as

[0119] ;

[0120] wherein, is the current round learning rate, is the minimum learning rate, is the initial learning rate, is the current training round, is the total training round.

[0121] In specific embodiments, the optimizer AdamW is used to reduce the risk of overfitting and improve the stability of parameter updating. The learning rate scheduling adopts a Warmup+cosine annealing strategy. In the early stage of training, the learning rate is gradually increased using linear Warmup to avoid the model converging to a local optimum too early. In the later stage, the learning rate is gradually reduced through cosine annealing, so as to realize a smoother convergence process. This combined scheduling strategy shows more stable training curves and better final accuracy in experiments. This learning rate scheduling strategy balances the global search ability and local refinement ability of the model, avoids the problems of gradient explosion caused by too high learning rate in the early stage of training or slow convergence caused by too low learning rate in the early stage of training, ensures that the model can stably converge to the optimal state, and improves the inversion accuracy.

[0122] In one of the embodiments, the GPR-B-scan image of the road to be detected is input into the trained GPR-TransUNet neural network model, and the permittivity distribution of the road to be detected is output, including:

[0123] After inputting the GPR-B-scan image of the road to be detected into the trained GPR-TransUNet, the asymmetric depth-spatial attention module is first used for the encoder. Through k×1 convolution, the layered features of the vertical depth axis of the GPR image are captured, and 1×k convolution is used to extract the local abnormal features of the horizontal scan axis. Then, the attention gate is generated by fusion, and the enhanced feature map in the encoding stage is obtained.

[0124] The enhanced feature map is optimized by the multi-scale feature fusion module. Different receptive field features are extracted in parallel through multi-size convolution kernels. After the outputs of each branch are processed by BN-ReLU, the residual scaling factor is introduced to weight the sum of the fused features and the original high-dimensional features after channel dimension splicing, and the feature map after feature fusion is obtained.

[0125] In the process of feature transmission to the decoder, the gating attention module splices the low-level detail features of the encoder with the high-level semantic features of the decoder, then compresses the dimensions through 1x1 convolution-BN-ReLU, and generates gating weights through Sigmoid; the gating weights are multiplied with the original features of the encoder pixel by pixel to filter out the detail features useful for dielectric inversion, and the filtered features are transmitted to the corresponding layer of the decoder;

[0126] According to the deep supervision decoding head, the output of each decoding layer is supervised and constrained, and an auxiliary prediction is generated through the light segmentation head. The decoder output is compressed in the channel and activated linearly to map the feature pixel value to the actual dielectric constant value, forming a dielectric constant distribution consistent with the input image size.

[0127] Specifically, the GPR-B-scan image of the road to be detected is first standardized and preprocessed. The mean and standard deviation of the training set image are used to normalize each pixel value, ensuring that the input feature distribution is consistent with the training phase. In the encoder stage, the preprocessed image is down-sampled through 3 layers of "convolution-BN-ReLU" convolution blocks, and the channel number is gradually increased from the initial 3 channels to 256 channels to capture local reflection features. The asymmetric deep-spatial attention module enhances the output of each convolution block to generate enhanced feature maps in the encoding stage. The multi-scale feature fusion module optimizes the 256-channel high-dimensional features output by the last layer of the encoder. Through 4 parallel branches (1x1, 3x3, 5x5, 7x7 convolution kernel), multi-receptive field features are extracted, and after concatenation, the channel number is compressed to 256 through 1x1 convolution. Combined with the residual scaling factor, the weighted fusion is obtained to get the feature fusion feature map. In the decoder stage, the fused feature map is up-sampled through 3 layers of deconvolution (step 2), and after each up-sampling, the encoder feature map (same size, same channel number) filtered by the gating attention module is concatenated. After "convolution-BN-ReLU" processing, it is gradually restored to the input image size; the deep supervision decoding head sets a light segmentation head (1x1 convolution+BN) at the output end of each decoding layer to generate an auxiliary prediction map. Although the inference stage does not rely on auxiliary output, the multi-layer supervision in the training stage has made the decoder have precise feature mapping capability. Finally, the 3-channel feature map output by the decoder is compressed to 1 channel through 1x1 convolution, and the pixel value is mapped to the dielectric constant value in the interval of 1~81 through the linear activation function, forming a dielectric constant distribution matrix consistent with the input image size. The beneficial effects are that this process realizes the end-to-end mapping from the GPR original signal to the dielectric constant distribution, without the need for manual intervention, improving the detection efficiency; the synergistic effect of each core module ensures the high resolution and high accuracy of the dielectric constant distribution, providing reliable physical parameter support for subsequent disease discrimination.

[0128] In a specific embodiment, a unified software and hardware experimental platform was constructed. The experiment relied on the Windows 11 operating system, the PyTorch deep learning framework, and the Python programming language to complete the algorithm improvement and development. The detailed configuration parameters of the software and hardware environment are shown in Table 3.

[0129] During the model training phase, the total number of training epochs was set to 200; the batch size was set to 4; the initial learning rate (lr0) was set to 0.001; and the optimizer was Warmup + cosine annealing, as shown in Table 3, to ensure that the construction from environment to policy is reproducible.

[0130] Table 3

[0131]

[0132] In the fields of statistics and machine learning, evaluating the predictive performance of regression models is a core step in the model validation process. To achieve objective and quantitative evaluation, a series of mathematical indicators are typically used to measure the consistency between model predictions and actual observations. The mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²) used in this study are three of the most representative indicators, each revealing the model's performance characteristics from different dimensions.

[0133] MAE is the arithmetic mean of the absolute errors of all individual observations. It quantifies the average magnitude of the model's predicted residuals. Its mathematical expression is:

[0134] ;

[0135] In the formula, n is the sample size. Representing the Actual value of each sample This represents the predicted value of the i-th sample.

[0136] The MAE (Modulation Error Analysis) is based on the L1 norm and directly reflects the expected magnitude of the prediction error. Its units are consistent with the dependent variable, making it highly intuitive and interpretable. The core statistical property of the MAE lies in its robustness. Because it linearly penalizes the error, it is not highly sensitive to outliers in the data. This means that individual extreme errors will not disproportionately affect the overall result, allowing it to more stably reflect the model's general performance on a majority of samples.

[0137] The root mean square error (RMSE) is the square root of the mean of the squares of the prediction errors. It represents the standard deviation of the prediction residuals. Its mathematical expression is:

[0138] ;

[0139] In the formula, n is the sample size. represent the actual value of the i-th sample represent the actual value of the i-th sample represent the predicted value of the i-th sample.

[0140] RMSE is based on the L2 norm and its dimension is restored to the same as the dependent variable by the square root operation for ease of interpretation. Unlike MAE, RMSE performs a quadratic penalty on the error. This feature makes it highly sensitive to large prediction errors. Specifically, larger residuals are amplified by squaring during the calculation process, thus having a significant impact on the final value of RMSE.

[0141] R² is a statistical measure of the goodness of fit of a model, which represents the proportion of the total variance of the dependent variable that can be explained by the model. Its mathematical expression is:

[0142] ;

[0143] In the formula, RSS (Residual Sum of Squares) is the sum of squares of residuals, and TSS (Total Sum of Squares) is the total sum of squares. In the formula, n is the sample size represent the actual value of the i-th sample represent the actual value of the i-th sample represent the predicted value of the i-th sample, y represents the arithmetic mean of all actual observations.

[0144] The core value of R² lies in providing a relative performance benchmark. It evaluates the magnitude of improvement in model prediction compared to a benchmark, rather than the absolute size of prediction error. Therefore, it is independent of the scale of the dependent variable, facilitating the comparison of explanatory power between different data sets or models.

[0145] To verify the independent and synergistic effects of the multi-scale feature fusion (MSF), the gated skip connection (GSF), the asymmetric depth-spatial attention (ASDS), and the deep supervision mechanism (DS-Head) proposed in this application in GPR-TransUNet, ablation experiments were conducted on each improved module. The basic TransUNet was used as a control, and MSF, GSF, ASDS, and DS-Head were introduced respectively under the same training configuration and data division, and their effects on the dielectric constant inversion accuracy were studied.

[0146] The ablation scheme is shown in Table 4. Model 1 is the baseline without model improvement, named Model 2-5 respectively adds MSF, GSF, ASDS and DSHead; Model 6-11 is the combination of MSF, GSF, ASDS and DSHead modules two by two; Model 12-15 is the combination of MSF, GSF, ASDS and DSHead modules three; Model 16 is the complete GPR-TransUNet of all four fusion. The model evaluation indexes are RMSE, MAE and determination coefficient R 2 .

[0147] Table 4

[0148]

[0149] From the ablation experiment results in Table 4, it can be seen that the original TransUNet (Model 1) is at the lowest level in RMSE, MAE and R² indexes, indicating that the baseline model has limitations in dealing with complex ground penetrating radar inversion tasks. After introducing the improved modules one by one, the model performance is improved to varying degrees.

[0150] In the single module introduction stage, the model (Model 2) with single MSF introduction performs the most outstanding, with RMSE reduced from 2.927 to 2.845 and R² increased to 0.652, indicating that multi-scale features can effectively enhance the adaptability to the spatial scale of underground structures. It is worth noting that the model (Model 4) with single ASDS introduction also shows significant advantages, with RMSE reduced to 2.860 and R² reaching 0.648, which is better than the effect of single GSF (Model 3) or DS-Head (Model 5). This result confirms that the asymmetric depth-space attention mechanism (ASDS) has an independent and key role in suppressing background noise and focusing on effective target features. In contrast, if there is a lack of strong feature extraction or attention guidance, relying only on gating or deep supervision mechanism (Model 3 and Model 5), the performance improvement is relatively limited.

[0151] Further analysis of the synergistic effect of double modules can find that the combination of ASDS and other modules produces a qualitative leap. In particular, Model 8 (MSF+ASDS), its RMSE is greatly reduced to 2.795, and R² is increased to 0.665, which is significantly better than Model 7 (MSF+GSF) (RMSE 2.848). This shows that the rich multi-scale features provided by MSF can more accurately locate and restore the dielectric constant anomaly area under the attention weighting of ASDS, and the two have strong complementarity.

[0152] In the multi-module deep fusion stage, the role of ASDS is further highlighted. Compared with the three-module combination Model 14 (MSF+GSF+DS, RMSE 2.840) without ASDS and the combination Model 12 (MSF+GSF+ASDS, RMSE 2.770) with ASDS, the latter performs significantly better, proving that ASDS is a key component to improve the upper limit of the model.

[0153] Finally, Model 16, which integrates the four modules simultaneously, achieves the global optimal state. Its RMSE drops to 2.750, MAE stabilizes at 0.893, and R² reaches 0.674. This result fully demonstrates that the multi-scale capture of MSF, the spatial attention focus of ASDS, the feature gating selection of GSF, and the deep gradient supervision of DS-Head form a closed-loop synergy in GPR-TransUNet. This synergy not only enhances the network's ability to characterize complex underground media, but also effectively solves the problems of gradient vanishing and feature redundancy in deep network training, thereby achieving high-precision inversion of the dielectric constant of ground penetrating radar.

[0154] Figure 7 The loss and evaluation index changes of the proposed GPR-TransUNet within 200 training cycles are shown. Figure 7 (a) shows that the training set and validation set losses quickly decrease in the initial stage and gradually tend to be stable after about 50 cycles; the convergence trend of the validation curve is basically the same as that of the training curve, indicating that the model has no obvious overfitting and the overall convergence is stable.

[0155] Figure 7 (b) shows that RMSE and MAE quickly decrease from the initial high values and enter a slow decline stage around 60 cycles, and R 2 then rapidly increases in the early stage and gradually approaches 0.5, and finally remains smooth fluctuation in the training process. The amplitude of each index curve significantly decreases and presents nearly smooth changes in the middle and late stages, which is related to the learning rate scheduling strategy of cosine annealing combined with Warmup. This strategy alleviates gradient shock through the Warmup stage in the early training, and then gradually reduces the learning rate using cosine annealing, making the parameter update step smooth and decreasing, which effectively suppresses the loss and index fluctuations in the later stage. RMSE represents the root mean square error. MAE represents the mean absolute error. R² represents the coefficient of determination. Epoch represents the training round.

[0156] The smooth convergence of loss and evaluation index not only verifies the training stability of GPR-TransUNet, but also shows that the learning rate strategy of cosine annealing and Warmup plays a positive role in ensuring global search and later refinement.

[0157] Table 5

[0158]

[0159] To verify the effectiveness of the proposed model, three powerful mainstream network models, U-Net++, DeepLabV3+ and ConvNeXt, were selected for comparison on the same ground penetrating radar permittivity inversion dataset. Table 5 lists the key evaluation indicators of each model. The results show that the proposed model achieves the best overall performance with an RMSE of 2.750, an MAE of 0.893, and an R² of 0.674. In contrast, U-Net++ has an RMSE of 3.337 and an MAE of 1.173, and an R² of 0.525, indicating that it has obvious shortcomings in continuous value regression and detail recovery. Although DeepLabV3+ and ConvNeXt have improved over U-Net++, their RMSE, MAE, and R² are still inferior to the proposed model, Figure 8 The permittivity inversion images of the radar image by U-Net++, DeepLabV3+, ConvNeXt and GPR-TransUNet models. Figure 8 The results show that the proposed improved network has stronger expression and generalization ability in capturing the spatial continuity and local details of permittivity.

[0160] To further verify the actual performance of the proposed GPR-TransUNet in the ground penetrating radar permittivity inversion task, this section gives the visualization comparison of typical test samples. The examples include complete profiles of asphalt layer and water-stable layer, covering typical disease structures such as void, loose body, water-rich body and void, to investigate the restoration ability of the model in complex medium environment. Figure 9The inversion results of different types of diseases by GPR-TransUNet. Overall, the GPR-TransUNet model can accurately recover the spatial distribution characteristics of dielectric constant under different disease scenarios. Comparing the inversion results of different typical diseases, the local details of the void, void and water-rich body area are almost consistent with the label, indicating that the model has good ability in the recovery of interface structure and dielectric characteristics. However, for the loose body area, the prediction result is slightly blurred compared to the real label, the main reason is that the spatial details are lost in the feature extraction process of the convolution operator, so that the texture cannot be faithful. The error map further reveals the spatial distribution characteristics of the prediction bias. Overall, the error is mainly concentrated in the disease boundary, reflecting that the dielectric constant gradient is larger in the interface transition area, and the model still has some difficulty in representing the fine edge. For the loose body area, the error presents a more dispersed overall distribution, rather than being limited to the boundary. This phenomenon shows that convolution feature extraction may not be able to fully capture the continuous and rapid change of dielectric information when facing the texture of the loose body structure, resulting in an increase in overall prediction bias. To verify the engineering applicability of the GPR-TransUNet model in real road disease detection, this study carried out radar measurement and drilling sampling on a national and provincial trunk highway in X province, Figure 10 respectively, are the B-scan image of the ground penetrating radar scan with a center frequency of 900 MHz, the dielectric constant inversion image of the GPR-TransUNet model, and the endoscopic image taken after drilling, Figure 10 (a) is a schematic diagram of section 1; Figure 10 (b) is a schematic diagram of the dielectric constant inversion result of section 1; Figure 10 (c) is a schematic diagram of endoscopic inspection of section 1; Figure 10 (d) is a schematic diagram of section 2; Figure 10 (e) is a schematic diagram of the dielectric constant inversion result of section 2; Figure 10 (f) is a schematic diagram of endoscopic inspection of section 2; Figure 10 (g) is a schematic diagram of section 3; Figure 10 (h) is a schematic diagram of the dielectric constant inversion result of section 3; Figure 10 (i) is a schematic diagram of endoscopic inspection of section 3.

[0161] The inversion results show that there are large areas of low dielectric constant anomaly in the three sections, with an average relative dielectric constant close to 1, showing a band or block distribution, which is highly consistent with the typical void characteristics. To further verify the reliability of the model, drilling sampling is carried out in the abnormal area, and direct observation is carried out through endoscopic imaging in the hole. The field endoscopic results confirm that there are obvious void and cavity structures in all abnormal areas, which are completely consistent with the model prediction.

[0162] The example shows that the GPR-TransUNet can not only accurately reconstruct the permittivity distribution under laboratory conditions, but also stably capture the electrical anomalies of large-scale diseases in real road environments. Compared with traditional manual interpretation relying only on radar maps, the model of the application can directly provide quantitative permittivity profiles, significantly improving the objectivity and accuracy of disease identification, and providing reliable technical support for road structure health monitoring and rapid decision-making.

[0163] It should be understood that, although Figure 1 The steps in the flowchart of the method can be displayed in sequence according to the arrows, but these steps are not necessarily executed in sequence according to the arrows. Unless otherwise stated in the application, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, Figure 1 At least part of the steps in the method can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of the sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or sub-steps or stages of other steps.

[0164] The technical features of the above embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0165] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the present application. It should be pointed out that, for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A neural network-based road permittivity inversion and disease discrimination method, characterized in that, The method comprises: constructing a road GPR simulation dataset; expanding the number of skip connections in the TransUNet baseline model, introducing a gated attention mechanism at each skip connection, and designing a multi-scale feature fusion module, an asymmetric depth-spatial attention module, and a deep supervision decoding head to form a GPR-TransUNet neural network model; using AdamW as an optimizer, setting the learning rate scheduling mode by combining the Warmup-cosine annealing strategy, and constructing an adaptive regional weighted loss function that amplifies the loss weight of the dielectric constant mutation boundary through a dynamic weighting mechanism; inputting the training set in the road GPR simulation dataset into the GPR-TransUNet neural network model, using the configured training strategy and the adaptive regional weighted loss function to train the model, monitoring the loss change and evaluation index fluctuation of the model during training through the validation set, adjusting the model parameters until the model converges, and obtaining the trained GPR-TransUNet neural network model; inputting the GPR-B-scan image of the road to be detected into the trained GPR-TransUNet neural network model, outputting the dielectric constant distribution of the road to be detected, identifying the cavity, void, loose body, and water-rich body diseases in the road to be detected according to the dielectric constant characteristics corresponding to different diseases in the dielectric constant distribution, and completing the road hidden disease discrimination.

2. The road permittivity inversion and disease discrimination method according to claim 1, characterized by, The process of constructing the road GPR simulation dataset comprises: Based on the gprMax three-dimensional forward platform, build two types of road models of homogeneous and heterogeneous, randomly embed four types of typical diseases of cavity, void, loose body and water-rich body in the two types of models respectively, generate multiple B-scan images of road GPR; divide the B-scan images into training set, validation set and test set according to the ratio of 8:1:1, and extract the true dielectric constant distribution of the corresponding B-scan image from the gprMax forward simulation result, map it to a uint16 gray image consistent with the size of the B-scan image as label data, and complete the dataset construction.

3. The road permittivity inversion and disease discrimination method according to claim 1, characterized by, The implementation process of the gated attention mechanism comprises: concatenate the high-level semantic features X output by the decoder and the low-level detail features S output by the encoder in the channel dimension of the GPR-TransUNet neural network model to obtain joint features Z; perform channel compression on the joint features Z through a first 1x1 convolutional layer, sequentially perform batch normalization processing and ReLU activation function operation, and then generate gating weights consistent with the channel number of the low-level detail features through a second 1x1 convolutional layer and a Sigmoid activation function operation; multiply the gating weights and the low-level detail features element by element to obtain weighted screening features, realize the enhancement of effective features and the suppression of noise features in the skip connection; The generation of gating weights consistent with the channel number of the low-level detail features is: wherein, denotes a gating weight, is a first 1 x 1 convolutional layer weight parameter, is a second 1 x 1 convolutional layer weight parameter, is a Sigmoid function, is a ReLU function, is a batch normalization operation, and Z is a joint feature.

4. The road permittivity inversion and disease discrimination method according to claim 1, characterized by, The implementation process of the asymmetric depth-spatial attention module comprises: The channel dimension linear transformation is performed on the module input feature map, and a Value feature map with the same size as the input feature map is generated through a 1*1 convolution layer; The depth branch and the spatial branch are respectively constructed to capture different dimensional features, wherein the depth branch adopts a k*1 size convolution kernel to extract the correlation features of the vertical axis of the GPR graph, and a depth attention map is obtained; The spatial branch adopts a 1*k size convolution kernel to extract the correlation features of the horizontal axis of the GPR graph, and a spatial attention map is obtained; The depth attention map and the spatial attention map are added, and are compressed to the range of [0, 1] through a Sigmoid activation function, to generate a final attention gate; The attention gate and the Value feature map are multiplied element by element, to obtain an output feature map after re-calibration, so as to accurately focus on the vertical layered features and the horizontal abnormal features in the GPR graph.

5. The road permittivity inversion and disease discrimination method according to claim 1, characterized by, The construction process of the multi-scale feature fusion module includes: Four parallel feature extraction branches are set, each branch adopts a 1*1, 3*3, 5*5 and 7*7 size convolution kernel for convolution operation, and padding is set in the convolution process to keep the spatial resolution of the output feature map of each branch consistent with the input feature map; After the convolution operation of each branch, batch normalization processing and ReLU activation function operation are sequentially performed to enhance the feature nonlinear expression ability; The output features of the four branches are spliced in the channel dimension, and the spliced features are subjected to channel compression and feature integration through a 1*1 convolution layer, so that the output channel number is restored to the channel number of the input feature map, and batch normalization processing and ReLU activation function operation are performed again; A learnable residual scaling factor is introduced to weight and fuse the integrated features and the module input features, to obtain a multi-scale feature fusion result.

6. The road permittivity inversion and disease discrimination method of claim 1, wherein, The calculation process of the adaptive region weighted loss function includes: The spatial gradient information is extracted from the real permittivity distribution of the label data, and the gradient approximation value of the real permittivity distribution in the horizontal direction and the gradient approximation value in the vertical direction are calculated; The square root of the sum of the squares of the gradient approximation value in the horizontal direction and the gradient approximation value in the vertical direction is calculated pixel by pixel, and is restored to the original size through zero padding to synthesize a gradient amplitude map; The gradient amplitude map is input into an auxiliary detection network composed of two convolution layers, a ReLU nonlinear activation layer and a Sigmoid output layer, to generate a dynamic key region weighting mask with the same size as the predicted permittivity map output by the GPR-TransUNet neural network model; The basic loss is calculated based on the global pixel-level error between the predicted permittivity map and the real permittivity distribution; The key region loss is calculated based on the predicted permittivity map, the real permittivity distribution and the dynamic key region weighting mask, and the adaptive region weighted loss function is obtained by weighting and summing the basic loss and the key region loss through a preset coefficient.

7. The road permittivity inversion and disease discrimination method according to claim 6, characterized by, The basic loss is calculated based on the global pixel-level error between the predicted permittivity map and the real permittivity distribution wherein, is the feature map height, is the feature map width, is the pixel value of the predicted permittivity map at position (i,j), is the pixel value of the real permittivity distribution at position (i,j).

8. The road permittivity inversion and disease discrimination method according to claim 6, characterized by, The key region loss is calculated based on the predicted permittivity map, the real permittivity distribution and the dynamic key region weighting mask, including: The key region loss is calculated based on the predicted permittivity map and the real permittivity distribution and a dynamic key region weighted mask as in, For feature map height, The width of the feature map. To predict the dielectric constant plot in ( i,j The pixel value at position ) The true dielectric constant is distributed in ( i,j The pixel value at position ) Weighted mask for dynamic key regions in ( i,j The pixel value at the location.

9. The road permittivity inversion and disease discrimination method of claim 1, wherein, The learning rate is set as wherein, is the current epoch learning rate, is the minimum learning rate, is the initial learning rate, is the current training epoch, is the total training epochs.

10. The road permittivity inversion and disease discrimination method of claim 1, wherein, The GPR-B-scan image of the road to be detected is input into the trained GPR-TransUNet neural network model, and the permittivity distribution of the road to be detected is output, including: After the GPR-B-scan image of the road to be detected is input into the trained GPR-TransUNet, the asymmetric depth-space attention module is first applied to the encoder, the layered features of the GPR image vertical depth axis are captured through k×1 convolution, the local abnormal feature of the horizontal scan axis is extracted through 1×k convolution, and then the attention gate is generated by fusion to obtain the enhanced feature map in the encoding stage; The enhanced feature map is optimized by the multi-scale feature fusion module, different receptive field features are extracted by parallel multi-size convolution kernels, the outputs of each branch are processed by BN-ReLU, then spliced in the channel dimension, and a residual scaling factor is introduced to weight the sum of the fused features and the original high-dimensional features, to obtain the feature map after feature fusion; In the process of feature transmission to the decoder, the gating attention module splices the low-level detail features of the encoder and the high-level semantic features of the decoder, then compresses the dimension through 1×1 convolution-BN-ReLU, and generates gating weights through Sigmoid; the gating weights are multiplied with the original features of the encoder pixel by pixel, and the useful detail features for dielectric inversion are screened out, and the screened features are transmitted to the corresponding layer of the decoder; According to the deep supervision decoding head, the output of each decoding layer is supervised and constrained, the auxiliary prediction is generated through the light segmentation head, the decoder output is compressed in the channel and linearly activated, the pixel value of the feature map is mapped to the actual permittivity value, and the permittivity distribution consistent with the input image size is formed.

Citation Information

Patent Citations

  • Road cavity detection method based on ground penetrating radar three-dimensional forward modeling and SAM2-Unet fusion

    CN120275956A

  • Road underground disease detection model based on medium inversion and deep learning

    CN120949348A