Ground penetrating radar positioning method, device and equipment based on time-frequency space projection and medium
By using a multi-head neural network based on time-frequency spatial projection and employing a learnable Gabor kernel and a depth-matched filter layer, the problems of mismatch and robustness in ground-penetrating radar positioning are solved, achieving high-precision location identification and offset estimation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-02
AI Technical Summary
Existing ground-penetrating radar (GPR) positioning technology suffers from numerous mismatches, overfitting, and insufficient robustness and practicality in global positioning tasks, mainly due to the sparsity of underground features and the variability of dielectric constant.
Feature extraction is performed using a multi-head neural network based on time-frequency spatial projection, including an echo direction encoding unit, a location identification unit, and an offset estimation unit. A multi-scale filter bank is constructed using a learnable Gabor kernel, and a direction selection attention mechanism and a depth-matched filter layer are combined. The network is trained through contrastive learning and supervised learning.
It effectively solves the problems of mismatch and overfitting in global positioning tasks of ground penetrating radar, improves robustness and practicality, and achieves high-precision location identification and offset estimation.
Smart Images

Figure CN122131265A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of radar positioning technology, and in particular to a ground-penetrating radar positioning method, apparatus, equipment and medium based on time-frequency spatial projection. Background Technology
[0002] Ground-penetrating radar (GPR) localization has gained widespread acceptance in robotics due to its ability to detect stable subsurface features. Recently, several schemes have been developed specifically for GPR imagery and have achieved good results; however, these schemes still generate a large number of mismatches in global localization tasks. The main challenges of using GPR for localization lie in the inherent sparsity of subsurface features and the variability of the subsurface dielectric constant, which complicates robust localization. Inspired by optical cameras and lidar, contrastive learning can generate compact and salient global descriptors for location identification, and matched filtering can achieve efficient offset estimation.
[0003] However, existing solutions are highly dependent on model parameters and are prone to overfitting in sparse ground-penetrating radar (GPR) data, thus their robustness and practicality for GPR data remain very limited. Summary of the Invention
[0004] Therefore, it is necessary to provide a ground-penetrating radar positioning method, device, equipment, and medium based on time-frequency spatial projection that can effectively provide matching accuracy in response to the above-mentioned technical problems.
[0005] A ground-penetrating radar (GPR) localization method based on time-frequency spatial projection, the method comprising:
[0006] Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0007] In one embodiment, in the ground-penetrating radar positioning network; The multi-scale directional texture features extracted by the encoder in the echo directional coding unit are used as the input of the offset estimation unit to perform offset estimation; After downsampling the multi-scale directional texture features, they are used as input to the location recognition unit for location recognition.
[0008] In one embodiment, the echo direction encoding unit further includes a direction selection attention mechanism block for downsampling processing and a downsampling block.
[0009] In one embodiment, the direction selection attention mechanism block performs direction enhancement of the different scale directional texture features output by the encoder along the channel dimension of the convolution kernel direction using the channel attention mechanism, and then maps the enhanced features along the channel dimension into a response distribution of the feature direction through a softmax layer.
[0010] The downsampling block comprises a convolutional layer, a max pooling layer, and a one-dimensional convolutional kernel connected in sequence, and processes the data output by the orientation selection attention mechanism block.
[0011] In one embodiment, in the location identification unit; The global descriptor is obtained by splicing and compressing the downsampled multi-scale directional texture features; Based on the global descriptors obtained from the current echo data and map data respectively, a nearest neighbor query is performed to obtain the location identification result.
[0012] In one embodiment, in the offset estimation unit: A shared convolutional layer is used as a decoder to decode the multi-scale directional texture features obtained from the current echo data and map data, respectively. A deep matching filter layer is used to obtain the matching cost curve based on the decoded data; The offset estimation result is obtained by regressing the matching cost curve using Arg Softmax with temperature.
[0013] In one embodiment, training the ground-penetrating radar positioning network includes two stages. In the first stage, contrastive learning is used to train the echo direction encoding unit and the location identification unit. In the second stage, supervised learning is used to train the offset estimation unit.
[0014] This application also provides a ground-penetrating radar positioning device based on time-frequency spatial projection, the device comprising: The current data acquisition module is used to acquire the current echo data and map data of the ground penetrating radar; A ground-penetrating radar (PTNet) positioning network processing module is used to input the current echo data and map data into a trained PTNet for matching. The PTNet is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and uses the extracted features as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, each branch is equipped with multiple filter banks of different scales, and each filter is constructed based on a learnable Gabor kernel. The ground-penetrating radar positioning result output module is used to realize the current positioning of the ground-penetrating radar based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps: Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0016] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0017] The aforementioned ground-penetrating radar (GPR) positioning method, apparatus, equipment, and medium based on time-frequency spatial projection match current echo data and map data by inputting them into a trained GPR positioning network. This GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data, respectively, and uses the extracted features as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, each with multiple filter banks of different scales constructed based on learnable Gabor kernels. This method can effectively solve the problems of numerous mismatches, overfitting, robustness, and practicality limitations in global positioning tasks using GPR. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a ground-penetrating radar positioning method based on time-frequency spatial projection in one embodiment; Figure 2 This is a schematic diagram of a learnable Gabor core structure in one embodiment; Figure 3 This is a flowchart of a ground-penetrating radar sequence descriptor-based location identification framework in one embodiment; Figure 4 This is a flowchart of a ground-penetrating radar offset estimation example; Figure 5 This is a schematic diagram of the structure of a ground-penetrating radar positioning network in one embodiment; Figure 6This is a schematic diagram of the qualitative results of an experiment on the GROUNDED dataset. Figure 7 This is a schematic diagram of the qualitative results on the CMU-GPR dataset in an experiment. Figure 8 This is a structural block diagram of a ground-penetrating radar positioning device based on time-frequency spatial projection in one embodiment; Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] In the field of ground-penetrating radar (GPR) technology, global localization is crucial for robots to accurately determine their location in complex environments, and GPR plays a key role in this task. However, current technologies still generate a large number of mismatches during global localization, mainly due to two factors. First, the inherent sparsity of underground features means that there is limited feature information available for localization matching, making accurate matching difficult. Second, the variability of the underground dielectric constant also presents a challenge; differences in dielectric constant across different regions lead to variations in GPR echo signals, thus interfering with the accuracy of localization matching and making robust localization exceptionally complex. Furthermore, the high dependence of existing technologies on model parameters and the tendency to overfit under sparse feature data further exacerbate the mismatch problem, severely limiting the robustness and practicality of GPR in global localization tasks.
[0021] To address the aforementioned issues, this application, such as Figure 1 As shown, a ground-penetrating radar (GPR) positioning method based on time-frequency spatial projection is provided, which specifically includes the following steps: Step S100: Obtain the current echo data and map data from the ground penetrating radar.
[0022] Step S110: Input the current echo data and map data into the trained ground-penetrating radar positioning network for matching. The ground-penetrating radar positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and uses the extracted features as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, each branch is equipped with multiple filter banks of different scales, and each filter is constructed based on a learnable Gabor kernel.
[0023] Step S120: Based on the current location identification result and offset estimation result output by the ground penetrating radar positioning network, the current location of the ground penetrating radar is realized.
[0024] This application addresses the two major application tasks of robot localization using ground-penetrating radar (GPR): location identification and offset estimation. The proposed method aims to design a multi-head neural network to construct a unified framework for solving both tasks. The method comprises two stages, sharing a time-frequency directional feature projection module as the encoding unit. Stage one is the location identification task, and stage two is the offset estimation module. This method effectively extracts stable underground features from channels and space, providing a high-quality, compact description of the underground scene. Furthermore, it theoretically guarantees the completeness of the time-frequency representation and the differentiability of end-to-end training. Simultaneously, the proposed neural network framework contains only a few parameters, enabling efficient training on small datasets. Combined with an open-source vector retrieval library, it achieves a runtime speed of 5ms, demonstrating potential for deployment in autonomous driving. This also addresses the issue of robust matching based on GPR under different weather conditions and sparse underground feature scenarios, as found in existing implementations.
[0025] In step S100, the acquired current echo data is a sequence of echo data obtained from real-time ground-penetrating radar (GPR) detection of the ground, and input data is generated and fed into the GPR positioning network according to a preset length. The map data is real map database data. The trained GPR positioning network matches the current data with the real map data to identify the current location and simultaneously obtain the position offset error.
[0026] In step S110, the trained ground-penetrating radar positioning network first adopts the encoder in the direction coding module, i.e., the echo direction coding unit. This structure serves as a shared encoder for the subsequent offset estimation unit and location identification unit.
[0027] In this embodiment, the encoder consists of three feature extraction branches at different scales to extract multi-scale directional texture features. Each branch also has multiple filter banks at different scales, designed based on learnable Gabor kernels.
[0028] Specifically, in each feature extraction branch, each set of directional filters contains several learnable Gabor convolutional kernels with different orientations to effectively encode directional features. The learnable Gabor kernel can be represented in the following form:
[0029] In the above formula, , Indicates the direction angle of the filter. Indicates the coordinate index of the filter. and These are the sine wave wavelength and phase deviation that constitute the Gabor filter. and These are the standard deviation of the Gaussian envelope and the aspect ratio, respectively.
[0030] In existing technologies, the parameters of a Gabor function need to be pre-specified before using it as a filter. Gabor filters generated in this way may not be optimal for some inputs. Furthermore, choosing suitable parameter values for the Gabor function is not easy. Therefore, in this embodiment, all parameters of the Gabor function can potentially be learned from the training data during training. In other words, this method uses a Gabor convolutional kernel that adaptively updates its parameters, instead of using hand-crafted parameters. Compared to classic convolutional kernels, where each pixel is learned independently, the parameters {λ, θ, φ, γ, σ} in a Gabor convolutional kernel constrain the relationships between pixels to produce the Gabor filter. Furthermore, a Gabor convolutional kernel has only five parameters, while a classic convolutional kernel of size k would be determined by k² independent pixels. Since Gabor convolutional kernels can be embedded in any network, these five parameters can be trained using gradient descent as part of a complete network to represent any number of CNN (Convolutional Neural Network) channels of the required input.
[0031] like Figure 2 The diagram shown is a schematic of a learnable Gabor core structure.
[0032] In this embodiment, the multi-scale directional texture features extracted by the encoder in the echo direction coding unit are used as input to the offset estimation unit for offset estimation. After downsampling the multi-scale directional texture features, they are used as input to the location identification unit for location identification.
[0033] like Figure 3 As shown, when performing location identification, it is also necessary to downsample the multi-scale features extracted by the encoder. Specifically, the echo direction coding unit also includes a direction selection attention mechanism block for downsampling processing and a downsampling block.
[0034] In this embodiment, the orientation selection attention mechanism is used to enhance the main orientation features of the receptive field region. The weights of the network channel dimension are redistributed by the coding branches of each Gabor filter bank and the orientation selection attention block to enhance the description of the main orientation features.
[0035] Specifically, the orientation selection attention mechanism block enhances the orientation texture features of different scales output by the encoder with the orientation attention mechanism along the channel dimension of the convolution kernel direction to filter out clutter noise in different directions. Then, the enhanced features are mapped to the response distribution of the feature direction along the channel dimension through the softmax layer.
[0036] In this embodiment, the reweighted features are dimensionality reduced by a post-processing unit. The cascaded post-processing unit module of fully connected convolutional kernels is used to achieve local feature shift invariance, so as to adapt to feature shifts caused by changes in groundwater content and undulations of the moving platform.
[0037] Specifically, the downsampling block consists of a series of convolutional layers, a max-pooling layer, and a one-dimensional convolutional kernel, which process the data output from the orientation selection attention mechanism block. The convolutional and max-pooling layers provide the network with a certain degree of feature offset invariance, while the one-dimensional convolutional kernel further enhances feature representation capabilities.
[0038] For the task of location identification, the echo direction coding unit, which consists of the encoder, the orientation selection attention mechanism block and the downsampling block mentioned above, needs to design a multi-scale space to ensure the sequential representation of the original echoes in order to generate feature vectors from the sequence echoes. On the other hand, it also needs to achieve lossy compression from the sequence echoes to compact descriptors through a global feature aggregation layer.
[0039] Therefore, in this embodiment, in the location identification unit, the multi-scale directional texture features that have undergone downsampling are spliced and compressed to obtain a global descriptor. Then, based on the global descriptors obtained from the current echo data and map data respectively, a nearest neighbor query is performed to obtain the location identification result.
[0040] Specifically, fully connected layers and NetVLAD layers are selected as supplements to global feature aggregation. The NetVLAD layer concatenates the directional texture features from different branches to obtain a multi-scale aggregated vector, which is then passed through the fully connected layer to obtain a compressed global descriptor. These two types of layers have achieved good performance verification in numerous camera and LiDAR location recognition tasks. Although these two global feature heads are simple to design, they perfectly meet the two major requirements of location recognition tasks: inversion invariance for GPR sequence acquisition and lightweight design requirements for real-time localization.
[0041] In this embodiment, during the inference-based localization retrieval process, the underground scene sequence data, processed by the aforementioned network, is first converted into a feature vector representation. The data is then organized using an index structure from a map database to facilitate rapid similarity searching. Secondly, the underground sequence scan of the query location can also be represented as a vector, and the map database is used to search for the most similar location index for rapid localization.
[0042] Specifically, during location retrieval, the Faiss feature vector retrieval open-source tool is used to perform nearest neighbor queries in the feature space between the sequence global descriptor and the map database to achieve real-time location requirements. L2 distance is selected as the feature distance for evaluation, and the L2 distance is expressed as follows:
[0043] In the above formula, and These are two feature descriptor vectors corresponding to the current echo data and map data, respectively.
[0044] In this embodiment, after selecting two frames of echo sequences as candidate matches through the location recognition task, a one-dimensional true geographic location offset of the two frames needs to be further output through a differentiable offset estimation head to recover the relative relationship between the two frames. Therefore, firstly, the learnable Gabor coding module, i.e., the encoder mentioned above, is used as the shared feature extraction layer. Secondly, a novel deep matching filter layer is proposed by introducing the idea of matched filtering, and the Arg softmax operation ensures the end-to-end differentiability of the network.
[0045] like Figure 4 As shown, in the offset estimation unit, a shared convolutional layer is used as a decoder to decode the multi-scale orientation texture features obtained from the current echo data and map data respectively. A deep matching filter layer is used to obtain the matching cost curve based on the decoded data, and a coarse-to-fine matching process is constructed. Finally, ArgSoftmax with temperature is used to regress the matching cost curve to obtain the offset estimation result.
[0046] Specifically, the depth matching filter layer generates a matching cost curve by performing sliding window correlation calculation on the decoding features of the current echo and map data, and combines it with temperature-controlled Arg Softmax to achieve differentiable continuous offset value regression, thereby completing the accurate estimation of the relative position offset between ground penetrating radar echoes.
[0047] In this embodiment, the training of the ground-penetrating radar positioning network includes two stages. In the first stage, contrastive learning is used to train the echo direction coding unit and the location identification unit. In the second stage, supervised learning is used to train the offset estimation unit.
[0048] Specifically, the loss function used when training the echo direction coding unit and the location recognition unit is expressed as follows:
[0049] In the above formula, These are the global descriptors obtained by encoding the query vector using network features. The boundary constant for the triplet loss. It is Euclidean distance. These represent the number of positive and negative samples selected for training, respectively. During training, the loss function in the feature space treats samples with similar geographical locations as intra-class positive samples, thus bringing them closer together, and treats samples with distant geographical locations as inter-class negative samples, thus widening the distance between the query sample and the negative samples.
[0050] Specifically, in the migration estimation task, the sparsity of subsurface features can lead to multi-peak or even staggered responses, which constitutes a non-convex problem in the optimization process. To address this, a peak-push loss function was developed to mitigate this non-convexity issue. By transforming the normalized cost curve into a probability integral function, it can be observed that a unimodal probability integral function with concentrated peaks exhibits a rapidly increasing probability integral curve. Based on this characteristic, the following loss function was designed.
[0051] First, by calculating the probability integral curve The kurtosis of the cost curve is indirectly defined by the difference between the maximum value and the k second-largest values.
[0052] At the same time, it forces the peak value to exceed a certain limit. Increasing the slope of the probability integral curve indirectly promotes peak generation:
[0053] Preferred, The value is 1.
[0054] Furthermore, Huber Loss is used as the supervised loss function to train the offset estimation network. It combines the advantages of Mean Square Error (MSE) and Mean Absolute Error (MAE), making it more robust to outliers. The definition of Huber Loss is as follows:
[0055] like Figure 5 The diagram shown is a schematic representation of the entire ground-penetrating radar positioning network.
[0056] This paper also demonstrates the effectiveness of the proposed method through experiments, which in turn proves the effectiveness of the ground-penetrating radar (GPR) localization network. The experiments were conducted on the large-scale public GPR datasets GROUNDED and CMU-GPR, using the Adam stochastic gradient descent optimizer. The batch size was set to 4, the weight decay was 0.0001, and the learning rate was 0.0001.
[0057] In the experiments, recall (Recall@K), a metric used in visual or lidar location recognition, was employed as a performance evaluation benchmark, and its effectiveness was validated on two public datasets, GROUNDED and CMU-GPR. The GROUNDED dataset was acquired by an array ground-penetrating radar with 11 channels operating in the 100–400 MHz frequency band; while the CMU-GPR dataset was acquired using a single-channel ground-penetrating radar operating in a similar frequency band. Each tracking sequence was aligned with map database data using RTK GPS to generate ground truth values for evaluation.
[0058] To comprehensively verify the network's performance, three sets of experiments were designed: the first set of experiments compared and analyzed with existing solutions; the second set was an ablation experiment to verify the contribution of each module; and the third set tested the model's parameter scale and running speed.
[0059] To highlight the performance of the proposed PT-Net (the location identification part of the ground penetrating radar localization network), several baseline and state-of-the-art solutions were selected for performance comparison. The baseline solutions were chosen based on the semantic segmentation models 3D ResNet18, 50, and 101 used for ground penetrating radar in recent years. State-of-the-art solutions selected included SeqVLAD for visual sequence descriptor retrieval, SeqOT for location identification of lidar sequence data, and LGPRNet, a deep learning framework for ground penetrating radar localization.
[0060] The experimental results are shown in Table 1. The proposed scheme demonstrates best performance in various scenarios across both datasets. In the CMU dataset, DEC, used for GPR recursive localization, performs poorly in location recognition tasks because it only describes energy variations in underground scenes, lacks discriminative power, and only shows some effectiveness when there are prior region constraints. CMUNet and MWSNet, based on the ResNet and U-Net frameworks respectively, are limited by their excessive focus on innovation.
[0061] Table 1 compares the performance of the best methods on public datasets with state-of-the-art methods, with the best performance highlighted in bold.
[0062] The first ablation experiment aimed to validate the effectiveness of the proposed echo direction coding module. Ablation analysis was performed on the LGF (encoder) and DAA (direction selection attention mechanism) blocks within the EDEBlock (echo direction coding unit) dataset using the GROUNDED and CMU-GPR datasets, evaluating their performance on recall@1 and recall@5 metrics. As shown in Table 2, replacing the learnable Gabor filter with a standard convolutional kernel resulted in a 6% decrease in average performance, highlighting the crucial role of the LGF in capturing directional features.
[0063] Furthermore, the DAA block (Direction Selection Attention mechanism block) was designed as a plug-and-play component; its removal resulted in a 4% performance decrease, demonstrating the importance of the adaptive attention mechanism for enhancing the output of the learnable Gabor filter. As shown in Table 2, the core function of the DAA block is to selectively enhance key features, thereby ensuring the integrity of the learned features. By combining the LGF (encoder) and DAA modules, optimal performance was achieved, further confirming the method's ability to achieve robust feature representation through directional description.
[0064] Table 2 Ablation experiments of learnable Gabor kernels and directional attention mechanisms in DAAs
[0065] Qualitative analysis (Figures 6 and 7) and comparative experiments on different datasets further validate the superiority of the proposed method. In the GROUNDED dataset, changes in weather conditions cause fluctuations in the underground dielectric constant, leading to variations in characteristic amplitude and location. Despite these adverse factors, the proposed scheme effectively captures the texture variations of underground radar echoes and accurately identifies the corresponding locations.
[0066] The second ablation experiment aimed to validate the optimal configuration of EDEBlock. By using convolutional kernels of different sizes (35×35, 23×23, 11×11, and 5×5), PTNet (i.e., the multi-scale feature extraction network) demonstrated its ability to capture diverse features. Larger convolutional kernels (such as 35×35) can capture low-frequency global information, which is particularly crucial for understanding the overall subsurface structure; while smaller convolutional kernels (such as 5×5) focus on high-frequency local details, which has a significant advantage in extracting fine features.
[0067] The combination of multi-scale convolutional kernels enables PTNet to comprehensively represent GPR data, capturing both global patterns and local details. However, as the number of scales increases, performance improvement may saturate. This is because the learnable Gabor filters are not orthogonal, which may lead to a certain degree of overlap in the feature space, introducing redundant information and consequently affecting performance due to overfitting.
[0068] Furthermore, the setting of parameter k directly affects the angular resolution of the Gabor filter in EDEBlock. A larger k value can improve the filter's ability to resolve directions, enhancing the model's ability to distinguish different locations. This is highly beneficial for GPR-based location recognition. However, an excessively high k value can also increase model complexity and potentially introduce redundancy, as the generated filter may capture similar features without providing additional useful information.
[0069] Comprehensive analysis shows that optimizing the EDEBlock configuration requires finding a balance between the selection of multi-scale convolutional kernels and the setting of the angular resolution parameter k. Multi-scale convolutional kernels are necessary for comprehensively capturing GPR features, but too many scales may lead to information redundancy. Similarly, appropriately adjusting the k value can enhance feature discrimination ability while avoiding an increase in complexity. Ultimately, the combination of 35×35, 23×23, 11×11, and 5×5 convolutional kernels with a k=64 configuration was experimentally verified as the optimal solution, effectively achieving efficient and balanced location recognition performance.
[0070] The second ablation experiment validated the selection of the optimal EDEBlock combination. By employing different kernel sizes—35×35, 23×23, 11×11, and 5×5—PTNet effectively captures a wide range of fine details. Larger kernels, such as 35×35, excel at capturing low-frequency global information, which is crucial for understanding the overall subsurface structure. In contrast, smaller kernels, such as 5×5, focus on high-frequency local details, which are essential for detecting fine features.
[0071] This multi-scale combination allows PTNet to comprehensively encode GPR data, fully leveraging the advantages of both large and small convolutional kernels. However, introducing more scales may lead to performance saturation, meaning the benefit of capturing more information diminishes due to increased redundancy. The core learnable Gabor filters in EDEBlock are not orthogonal and may overlap in the feature space, resulting in redundant information and thus degrading performance through overfitting.
[0072] The parameter k directly affects the angular resolution of the Gabor filters in each EDEBlock. A larger k value can improve angular resolution, enabling the Gabor filters to simulate a wider range of directions, which is highly advantageous for GPR-based location identification as it enhances the model's ability to distinguish different locations. However, a higher k value also increases the complexity of the filter bank and may lead to redundancy, as it may generate filters that capture similar features without adding new information.
[0073] This analysis highlights the importance of carefully balancing the EDEBlock configuration in PTNet. Multi-scale kernel combinations are crucial for capturing the full spectrum of GPR features, but appropriate combinations must be chosen to avoid performance saturation. Similarly, the parameter k should be optimized to improve angular resolution while minimizing redundancy. The final selected configurations of 35×35, 23×23, 11×11, and 5×5 EDEBlocks with k=64 have been experimentally proven to be optimal, providing a balanced and effective location recognition scheme, as shown in Table 3.
[0074] Table 3. Recall@5 rates under different scales of EDE-block module combinations and hyperparameter k
[0075] As shown in Table 4, the third experiment reports the average runtime and model size of all methods across ten trials. The results show that the proposed method is the most compact in model size while maintaining competitive performance in terms of parameter count and runtime. Notably, the proposed method does not require a dedicated function to evaluate sequence descriptor similarity and can be accelerated using the Faiss library, resulting in a runtime of 188 Hz with 100 frames of GPR sequence input, significantly exceeding the 126 Hz GPR data acquisition rate. This rate surpasses the potential for real-time deployment in practical applications. Although all selected descriptor generation schemes have the same descriptor size (dim=4096), they require more processing time due to the large number of parameters in existing backbones designed for camera / LiDAR data, thus reducing computational efficiency.
[0076] Table 4 Comparison of Model Size and Running Speed
[0077] like Figure 6 As shown, the offset estimation layer is visualized, and ablation experiments were conducted on the proposed peak-push loss function. It is clear that before using the peak-push loss function, the peak width is wide, and erroneous peaks are easily generated at the boundaries, leading to incorrect offset estimation. However, after using the peak-push loss function, the peaks become significantly more concentrated, with only a few sub-peaks appearing, ensuring the effectiveness of the offset estimation.
[0078] like Figure 7 As shown, the offset estimation error was calculated and error statistics were presented. The CMU-based approach is prone to mismatches at the boundaries, resulting in a large overall matching error variance and poor stability. The differentiable scheme proposed in this paper effectively alleviates this problem, and the offset estimation results are further improved by introducing a peak propulsion loss function.
[0079] In the aforementioned ground-penetrating radar (GPR) localization method based on time-frequency spatial projection, a multi-head neural network was designed to simultaneously address the needs of both location identification and offset estimation, providing a solid foundation for GPR-based localization retrieval and offset estimation. A time-frequency directional focusing feature encoding layer was developed based on a learnable Gabor kernel. This feature extraction backbone decouples the time-frequency and robust directional features of the GPR sequence echoes, effectively capturing subsurface features while reducing the number of network parameters. The translational equivariance of the learnable Gabor kernel was also revealed, enabling the network to maintain good physical interpretability in the downstream task of offset estimation. This method also proposes a novel network, PT-Net, specifically designed for GPR location identification. Viewing the location identification task from a novel wavelet basis projection compression perspective, the representational completeness of the network structure design was theoretically proven. PT-Net mainly consists of a multi-scale EDE module. In addition to the learnable Gabor encoding layer, the EDE module includes two sub-modules: a DAA module to promote the generation of multiple types of filter kernels from the learnable Gabor kernel, and a SIU module to mitigate feature offsets caused by changes in dielectric constant. Finally, advanced performance is achieved by integrating common feature aggregation layers. Meanwhile, the lightweight design, combined with the accelerated retrieval strategy of the existing Faiss library, ensures the real-time potential of the proposed network in real-world applications. In another branch output, this method proposes a novel differentiable offset estimation layer. First, a deep matched filtering layer is designed based on matched filtering techniques in radar signal processing. Offset estimation is transformed into a differentiable classification and regression problem to achieve offset estimation of echoes from two candidate matching sequences, ensuring the feasibility of end-to-end network training and reverse updating of decoder weights. Furthermore, to effectively train the offset estimation network, a novel loss function and specific training strategy are proposed to alleviate the multi-peak problem in the matching problem, further improving the network's accuracy and generalization ability.
[0080] Meanwhile, the method was validated on two public datasets and a self-built dataset, achieving the best performance. Ablation experiments also showed that the method can effectively improve performance based on each module.
[0081] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0082] In one embodiment, such as Figure 8 As shown, a ground-penetrating radar (GPR) positioning device based on time-frequency spatial projection is provided, comprising: a current data acquisition module 200, a GPR positioning network processing module 210, and a GPR positioning result output module 220, wherein: The current data acquisition module 200 is used to acquire the current echo data and map data of the ground penetrating radar; The ground-penetrating radar (GPR) positioning network processing module 210 is used to input the current echo data and map data into a trained GPR positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and uses the extracted features as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The ground-penetrating radar positioning result output module 220 is used to realize the current positioning of the ground-penetrating radar based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0083] Specific limitations regarding the ground-penetrating radar (GPR) positioning device based on time-frequency spatial projection can be found in the limitations of the GPR positioning method based on time-frequency spatial projection mentioned above, and will not be repeated here. Each module in the aforementioned GPR positioning device based on time-frequency spatial projection can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0084] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a ground-penetrating radar positioning method based on time-frequency spatial projection. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0085] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0086] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps: Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0087] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
[0088] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0089] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0090] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A ground-penetrating radar positioning method based on time-frequency spatial projection, characterized in that, The method includes: Acquire current echo data and map data from ground-penetrating radar; The current echo data and map data are input into a trained ground-penetrating radar (GPR) positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and the extracted features are used as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The current location of the ground-penetrating radar is achieved based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
2. The ground-penetrating radar positioning method based on time-frequency spatial projection according to claim 1, characterized in that, In the ground-penetrating radar positioning network; The multi-scale directional texture features extracted by the encoder in the echo directional coding unit are used as the input of the offset estimation unit to perform offset estimation; After downsampling the multi-scale directional texture features, they are used as input to the location recognition unit for location recognition.
3. The ground-penetrating radar positioning method based on time-frequency spatial projection according to claim 2, characterized in that, The echo direction encoding unit also includes a direction selection attention mechanism block for downsampling processing and a downsampling block.
4. The ground-penetrating radar positioning method based on time-frequency spatial projection according to claim 3, characterized in that, The direction selection attention mechanism block performs direction enhancement of the different scale directional texture features output by the encoder along the channel dimension of the convolution kernel direction, and then maps the enhanced features along the channel dimension into the response distribution of the feature direction through the softmax layer. The downsampling block comprises a convolutional layer, a max pooling layer, and a one-dimensional convolutional kernel connected in sequence, and processes the data output by the orientation selection attention mechanism block.
5. The ground-penetrating radar positioning method based on time-frequency spatial projection according to claim 4, characterized in that, In the location identification unit; The global descriptor is obtained by splicing and compressing the downsampled multi-scale directional texture features; Based on the global descriptors obtained from the current echo data and map data respectively, a nearest neighbor query is performed to obtain the location identification result.
6. The ground-penetrating radar positioning method based on time-frequency spatial projection according to claim 5, characterized in that, In the offset estimation unit: A shared convolutional layer is used as a decoder to decode the multi-scale directional texture features obtained from the current echo data and map data, respectively. A deep matching filter layer is used to obtain the matching cost curve based on the decoded data; The offset estimation result is obtained by regressing the matching cost curve using Arg Softmax with temperature.
7. The ground-penetrating radar positioning method based on time-frequency spatial projection according to any one of claims 1-6, characterized in that, The training of the ground-penetrating radar positioning network includes two stages. In the first stage, contrastive learning is used to train the echo direction encoding unit and the location identification unit. In the second stage, supervised learning is used to train the offset estimation unit.
8. A ground-penetrating radar positioning device based on time-frequency spatial projection, characterized in that, The device includes: The current data acquisition module is used to acquire the current echo data and map data of the ground penetrating radar; A ground-penetrating radar (GPR) positioning network processing module is used to input the current echo data and map data into a trained GPR positioning network for matching. The GPR positioning network is a multi-head neural network, including an echo direction encoding unit that extracts features using time-frequency directional feature projection, a location identification unit, and an offset estimation unit. The echo direction encoding unit extracts features from the current echo data and map data respectively, and uses the extracted features as inputs to the location identification unit and the offset estimation unit to obtain the current location identification result and the offset estimation result. The echo direction encoding unit includes an encoder for extracting multi-scale directional texture features. The encoder consists of multiple branches, and each branch is equipped with multiple filter banks of different scales. Each filter is constructed based on a learnable Gabor kernel. The ground-penetrating radar positioning result output module is used to realize the current positioning of the ground-penetrating radar based on the current location identification result and offset estimation result output by the ground-penetrating radar positioning network.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.