Ground penetrating radar three-dimensional inversion method and system based on multi-level neural network
Patent Information
- Application Number
- CN202410138085.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-01-30
AI Technical Summary
探地雷达容易受到地表直接反射波、发射与接收天线之间的耦合波干扰和掩盖,回波中双曲波包含杂波,以至于不能很好地对双曲线性结构进行拟合,影响的后续的反演计算
[0019] In this disclosure, the data is first converted into three-dimensional dense volume data. Based on inter-channel attention interpolation, the mutual enhancement of adjacent B-scan features is achieved, thereby preserving the target structure features during the interpolation generation of three-dimensional data. Then, conditional generation is used to suppress clutter, so that only hyperbolic wave features that can reflect the target object are retained in the three-dimensional data. Finally, a reliable three-dimensional mapping of the underground target object is obtained through inversion using a Transformer model. The three-dimensional volume data features are fused with the decoder features, automatically located, and local details are refined to generate an underground three-dimensional structure containing only the required target object, thereby obtaining a reliable three-dimensional mapping of the underground target object.
Smart Images

Figure CN117970271B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of ground-penetrating radar data inversion, specifically to a three-dimensional inversion method and system for ground-penetrating radar based on a multi-level neural network. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] Ground Penetrating Radar (GPR) data inversion plays a crucial role in GPR data interpretation, transforming raw data into a direct mapping of subsurface structures and properties. This enhances the understanding of subsurface conditions and supports subsequent decision-making. However, GPR data inversion faces challenges such as multi-parameter inversion, non-uniqueness, data noise, finite observation geometry, and computational complexity. Non-uniqueness requires interpretation and verification using prior information and additional constraints; however, prior information for GPR data is difficult to obtain. Simultaneously, data quality and noise handling are critical to the accuracy of the inversion results. Finite observation geometry limits the accurate inversion of complete subsurface structures. These challenges necessitate more efficient computational methods to improve inversion quality. With the rapid development of deep learning, its application in GPR inversion is receiving increasing attention.
[0004] The inventors discovered in their research that existing deep learning-based studies are all based on two-dimensional B-scan images, which have limited coverage and incomplete training data, leading to low accuracy in the overall identification of underground structures in the region. Furthermore, the training datasets used are mainly artificially simulated data, which lacks sufficient coverage of complex application scenarios in real-world environments, thus affecting the application of the methods to some extent. The environments based on these existing studies are often semi-idealized data in the same medium, almost free of noise, and the inversion results for complex and diverse actual collected data remain uncertain. When using ground-penetrating radar (GPR) for underground pipeline target inversion and detection, the accuracy of the results largely depends on the fitting quality of the reflected hyperbolic wave. GPR is susceptible to interference and masking from direct surface reflection waves and coupling waves between transmitting and receiving antennas. The hyperbolic wave in the echo contains clutter, making it difficult to fit hyperbolic structures well, thus affecting subsequent inversion calculations.
[0005] It is evident that the current mainstream inversion method is still based on two-dimensional slices (B-scan) for inversion calculation. Due to the non-uniqueness of the inversion method itself and the influence of data noise, the results of two-dimensional inversion have significant uncertainties. Summary of the Invention
[0006] To address the aforementioned issues, this disclosure proposes a three-dimensional inversion method and system for ground-penetrating radar based on a multi-level neural network. Through three-dimensional inversion, multi-angle observation data can be obtained, enabling a better perception of the complexity and non-uniformity of the underground environment and providing a more reliable description of underground targets.
[0007] To achieve the above objectives, the present disclosure adopts the following technical solution:
[0008] One or more embodiments provide a three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network, including the following steps:
[0009] Obtain ground-penetrating radar detection data, perform inter-channel attention interpolation on the B-scan slice set, and generate B-scan three-dimensional dense volume data;
[0010] Conditional generative adversarial networks are used to suppress redundant clutter, so that the hyperbolic wave characteristics that can reflect the target object are preserved in the 3D dense volume data, resulting in 3D volume data containing hyperbolic wave sequences for inversion.
[0011] The Transformer model is used to invert the 3D volume data, and the features of the 3D volume data are fused with the features of the decoder to automatically locate and refine local details, thereby generating the underground 3D structure of the target object.
[0012] One or more embodiments provide a ground-penetrating radar three-dimensional inversion system based on a multi-level neural network, including:
[0013] 3D Dense Volume Data Generation Module: Configured to acquire ground-penetrating radar detection data, perform inter-channel attention interpolation processing on B-scan slice sets, and generate B-scan 3D dense volume data;
[0014] 3D Denoising Module: Configured to use a conditional generative adversarial network to suppress redundant clutter, so that the hyperbolic wave features that can reflect the target object are preserved in the 3D dense volume data, resulting in 3D volume data containing a hyperbolic wave sequence for inversion;
[0015] Inversion Module: Configured to invert 3D volume data using a Transformer model, fuse 3D volume data features with decoder features, automatically locate and refine local details, and generate the underground 3D structure of the target object.
[0016] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the above-described three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network.
[0017] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps in the above-described three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network.
[0018] Compared with the prior art, the beneficial effects of this disclosure are as follows:
[0019] In this disclosure, the data is first converted into three-dimensional dense volume data. Based on inter-channel attention interpolation, the mutual enhancement of adjacent B-scan features is achieved, thereby preserving the target structure features during the interpolation generation of three-dimensional data. Then, conditional generation is used to suppress clutter, so that only hyperbolic wave features that can reflect the target object are retained in the three-dimensional data. Finally, a reliable three-dimensional mapping of the underground target object is obtained through inversion using a Transformer model. The three-dimensional volume data features are fused with the decoder features, automatically located, and local details are refined to generate an underground three-dimensional structure containing only the required target object, thereby obtaining a reliable three-dimensional mapping of the underground target object.
[0020] The advantages of this disclosure, as well as its additional advantages, will be described in detail in the following specific embodiments. Attached Figure Description
[0021] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute a limitation thereof.
[0022] Figure 1 This is a three-level structure diagram of the multi-level neural network TCCG-Net constructed in Embodiment 1 of this disclosure;
[0023] Figure 2(a) is a schematic diagram of interpolation based on B-scan interchannel attention in Embodiment 1 of this disclosure;
[0024] Figure 2(b) is a schematic diagram of the Transformer model structure with inter-channel attention module added in Embodiment 1 of this disclosure;
[0025] Figure 3 This is a schematic diagram illustrating the acquisition of time-varying features and appearance features according to Embodiment 1 of this disclosure;
[0026] Figure 4(a) is a logic diagram of the Conditional Adversarial Network (CGAN) of Embodiment 1 of this disclosure;
[0027] Figure 4(b) is a schematic diagram of the simAM attention mechanism of Embodiment 1 of this disclosure;
[0028] Figure 5 This is a schematic diagram of B-scan neighbor generation based on the GAV algorithm in Embodiment 1 of this disclosure;
[0029] Figure 6 This is a schematic diagram of the three-dimensional hyperbolic wave optimization module TRM structure of Embodiment 1 of this disclosure;
[0030] Figure 7 This is a schematic diagram of the refinement process in the three-dimensional inversion of Embodiment 1 of this disclosure;
[0031] Figure 8 This is a schematic diagram illustrating the construction of coupled data training data pairs in Embodiment 1 of this disclosure;
[0032] Figure 9 This is the result of interpolation experiments on three different simulated data using different methods in Embodiment 1 of this disclosure;
[0033] Figure 10 This is the experimental result of interpolating three real-world environmental data using different methods in Embodiment 1 of this disclosure;
[0034] Figure 11 This is a comparison chart of the volume data interpolation results in the experiment of Embodiment 1 of this disclosure;
[0035] Figure 12 This is a comparison of the three-dimensional decluttering results of dense data volumes using different methods in the experiment of Embodiment 1 of this disclosure;
[0036] Figure 13 This is a comparison diagram of the three-dimensional inversion results of dense data volumes using different methods in the experiment of Embodiment 1 of this disclosure;
[0037] Figure 14 This is a comparison diagram of the effects obtained when the feature fusion ratios before and after B-scan are different in the adjacent body in the experiment of Embodiment 1 of this disclosure. Detailed Implementation
[0038] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0040] It should be noted that the terminology used herein is for descriptive purposes only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features within those embodiments can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.
[0041] Compared to two-dimensional B-scan slice data, three-dimensional B-scan volume data offers multi-angle data observation and cross-validation opportunities, containing richer spatial information and providing abundant feature sources for complex inversion calculations. To address the inversion problem of underground pipelines, this disclosure proposes a multi-level neural network, TCGA-NET. It integrates three-dimensional dense B-scan volume data generation, clutter suppression, and three-dimensional inversion into a single multi-level neural network structure, realizing an end-to-end three-dimensional inversion model for underground pipelines. This results in a reliable three-dimensional mapping of underground target objects, allowing for a broader and more intuitive understanding of the distribution of underground pipelines. Specific embodiments are described below.
[0042] Example 1
[0043] In one or more of the technical solutions disclosed in the embodiments, such as Figures 1 to 13 As shown, a three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network includes the following steps:
[0044] Step 1: Acquire ground-penetrating radar detection data, and perform inter-channel attention interpolation processing on the B-scan slice set in the acquired data to generate B-scan three-dimensional dense volume data;
[0045] Step 2: Use a conditional generative adversarial network to suppress redundant clutter, so that the hyperbolic wave characteristics that can reflect the target object are retained in the 3D dense volume data, and obtain 3D volume data containing hyperbolic wave sequences for inversion.
[0046] Step 3: Use the Transformer model to invert the 3D volume data, fuse the 3D volume data features with the decoder features, automatically locate and refine local details, and generate the underground 3D structure of the target object.
[0047] In this embodiment, the data is first converted into three-dimensional dense volume data. Based on inter-channel attention interpolation, the mutual enhancement of adjacent B-scan features is achieved, thereby preserving the target structure features during the interpolation generation of three-dimensional data. Then, conditional generation adversarial suppression of clutter is adopted so that only hyperbolic wave features that can reflect the target object are retained in the three-dimensional data. Finally, the Transformer model is used to invert to a reliable three-dimensional mapping of the underground target object. The three-dimensional volume data features are fused with the decoder features, automatically located, and local details are refined to generate an underground three-dimensional structure containing only the required target object, thereby obtaining a reliable three-dimensional mapping of the underground target object.
[0048] The Transformer model is a deep learning model that uses an attention mechanism to improve the training speed of the model.
[0049] Steps 1 to 3 above are achieved by constructing a multi-level neural network as the TCGA-Net three-dimensional inversion network for ground-penetrating radar, including a three-level network:
[0050] 1) The first level is a 3D B-scan dense volume data generation model composed of CNN network and Transformer network: it embeds interpolation unit based on inter-channel attention, and achieves the purpose of mutual enhancement of adjacent B-scan features by calculating the similarity attention of the same region between B-scan images, thereby preserving the target structural features in the process of interpolation to generate 3D data.
[0051] 2) The second level is the clutter suppression three-dimensional denoising model (which can be simply referred to as the 3D-Declutter model): a conditional generative adversarial network is used to suppress redundant clutter, so that only hyperbolic wave features that can reflect the target object are retained in the three-dimensional data, and the hyperbolic wave sequence participating in the inversion is obtained.
[0052] 3) The third level is the three-dimensional inversion model, which uses the Transformer model to invert the three-dimensional volume data and generate a direct mapping of the underground pipeline structure.
[0053] The third-level inversion model integrates the Pyramid Vision Transformer (PVT) and the Three-Dimensional Refinement Module (TRM), fusing 3D volumetric data features with decoder features to automatically locate and refine local details, generating a subsurface 3D structure containing only the desired target object. The Pyramid Vision Transformer... Figure 1 In Chinese, it is simply referred to as PV Transformer;
[0054] Furthermore, a paired data training method with coupled learning is adopted between the second and third level networks to increase data correlation and achieve the best inversion effect.
[0055] In step 1, the obtained B-scan multi-slice data is a sparse image sequence. Interpolation is needed to generate dense 3D volumetric data. The structure for generating dense volumetric data is referenced in [reference needed]. Figure 1 The first-level network includes the following steps:
[0056] Step 11: Extract the interchannel information of the B-scan slice data and interpolate between image slices to obtain primary dense data;
[0057] Specifically, a CNN network is set as a low-level feature extractor in the first-level network, and convolution, downsampling and convolution operations are performed in sequence to obtain primary dense data.
[0058] Step 12: Embed an Inter-Bscan Attention (IBA) mechanism between slices of the primary dense data to obtain time-varying features and appearance features between slices of different channels, establish constraints and dependencies between hyperbolic waves on adjacent B-scans, and generate dense three-dimensional volume data.
[0059] To implement step 12, a Transformer network is set up as a motion appearance feature extractor in the first-level network, including a Transformer Block module, a downsampling module, and a Transformer Block module connected in sequence; the Transformer network and the CNN network are connected through a cross-scale splicing embedding module.
[0060] To address the lack of fine-grained information in the input when connecting Transformers to CNNs, low-level features derived from CNNs are reused to supplement cross-scale information. Specifically, the cross-scale concatenation embedding module is configured to use multi-scale dilated convolutions to fuse the input information together.
[0061] The extraction of time-varying features is achieved through a time-varying feature processing module set up after the Transformer network. The time-varying feature processing module includes the time-varying feature estimation module Motion Estimation and the optimization network RefineNet.
[0062] As shown in Figure 2(a), the hyperbolic waves reflected by target objects between adjacent B-scans collected in the same area can enhance each other, so as to highlight valuable features and reduce clutter interference.
[0063] Attention is embedded between slices of primary dense data to capture time-varying and appearance features between channel data;
[0064] Specifically, the inter-channel attention module is incorporated into the Transformer network, as shown in Figure 2(b). This includes an improved Transformer model that preserves the spatial-temporal structure between channel data and extracts associated features through IBA. Furthermore, depthwise convolution is used instead of positional encoding in the multilayer perceptron (MLP) to accommodate B-scan inputs of different sizes and enhance regional interactions between waveforms of different channels within the same B-scan data.
[0065] Time-varying and appearance features are obtained through the inter-channel attention module, such as Figure 3 As shown, for any region in I0, it is treated as a query, and the spatial neighbors in I1 are used as key / value pairs to generate an attention map. Then, the appearance information in I1 is aggregated using the attention map to obtain the inter-channel appearance representation of the query region, while simultaneously estimating the approximate displacement of the query region in the inter-channel data.
[0066] For appearance features, the appearance features of two adjacent B-scans are denoted as A0 and A1. For any region of a B-scan, select the same region from adjacent B-scan slices, and represent it in slice I0 as follows: In slice image I1, it is represented as N represents the size of the neighborhood window used to generate the query, keys, and values, as follows:
[0067]
[0068]
[0069]
[0070] in, It is a linear projection matrix.
[0071] Then, in and Dot products are performed between each location, and an attention map is generated using the SoftMax activation function. The value at each coordinate position (i,j) represents The formula for the appearance similarity between data and its adjacent channels is as follows;
[0072]
[0073] in, The scaling factor, representing the attention mechanism, is a scaling term used to adjust the dot product attention calculation. In the attention mechanism, the result of the dot product is divided by before calculating the SoftMax function activation. The purpose of this operation is to stabilize and scale the attention distribution.
[0074] Obtained appearance similarity It can be used to transfer morphological information and extract time-varying information. For morphological information, firstly, similar appearance information in slice image I1 is aggregated, and then the aggregated appearance features of slice image I1 are compared with the appearance information of slice image I0. To enhance the morphological information in slice image I0, such as:
[0075]
[0076] The enhanced morphological features include a mixture of appearance features of similar regions (hyperbolic wave features) in two different B-scans.
[0077] The time-varying feature processing module first creates a coordinate graph:
[0078]
[0079] Here, the value at each location represents its relative position within the entire B-scan image. The coordinates of adjacent channels are weighted to estimate the region of slice image I0. The approximate corresponding positions in adjacent images I1 are then obtained by subtracting the region of slice image I0. The original location and The estimated position in I1 is used to generate the motion vector. The calculation formula for the motion estimation module is as follows:
[0080]
[0081] Including time-varying information can provide a clear prior for motion estimation. Then, Time-varying features are generated through a linear layer. Under the assumption of local linear motion, this can be achieved by... Multiply by t to approximate the calculation from I0 to I t The time-varying vector is as follows:
[0082]
[0083] It can serve as a clue to guide the next step of inter-channel morphology estimation, and can be used for B-scan prediction with arbitrary step sizes, requiring only one calculation. Morphological characteristics Since the time period is constant, channel observation only needs to calculate the prediction of data channels for multiple arbitrary time steps once.
[0084] In step 2, clutter suppression is performed; the structure of this part is referenced. Figure 1 The second-level network uses a 3D-Declutter model for clutter suppression.
[0085] GPR data inversion is a process of extracting subsurface information by processing electromagnetic wave signals. It involves processing the electromagnetic wave signals collected by the GPR system to extract subsurface information. Due to the complexity of the subsurface medium and the influence of noise, the inversion results may contain errors.
[0086] In the inversion results, the H(·) operator can be used to represent the inversion problem between the input B-scan X and the output underground scene Y. The inversion problem between the output underground scene Y is:
[0087] X = H -1 (Y)
[0088] Here, X can be represented as X = x h +x c (x h For hyperbolic waves, x c To suppress the interference of clutter on the inversion, a second-level clutter suppression network is embedded in the multi-level neural network in this embodiment.
[0089] Declutter-GAN is a GPR data recovery network based on CGAN. It can remove noise and redundant signals from GPR data, so that the recovered data contains only the effective information related to the target object.
[0090] CGAN (Conditional Generative Adversarial Network) is a type of generative adversarial network that introduces conditional information, which can be of any type, such as class labels, specific attributes, noise vectors, etc. The introduction of conditional information allows the generator to produce samples of a specific type based on given conditions.
[0091] When using ground-penetrating radar B-scan as the training dataset, the data is transformed from its original two-dimensional structure into a two-dimensional sequence. Replacing the training data with a two-dimensional sequence can be seen as extending the data representation from a two-dimensional space to a three-dimensional space. Although no actual spatial dimension is introduced in real-world scenarios, organizing two-dimensional data into a sequence introduces an additional dimension to data representation and processing, enabling the model to better capture the information between the sequences. The benefit of this extension is increased sensitivity of the clutter-removal model to data noise.
[0092] Furthermore, in this embodiment, Declutter-GAN is extended to three-dimensional space as the basic model, and 3D-Declutter is constructed as the second level of the multi-level neural network. 3D-Declutter includes an encoder, a generator based on the SimAM attention mechanism, and a discriminator based on the SimAM attention mechanism. The SimAM modules of the generator and the discriminator are respectively added after the activation functions of the residual blocks of the generator and the discriminator. During the training process, the discriminator performs adversarial training on the output of the generator, so that the output of the generator is closer to the clutter-free ground-penetrating radar B-scan sequence.
[0093] The logical structure of 3D-Declutter is shown in Figure 4(a). For the CGAN image translation task, the generator learns to synthesize data, while the discriminator is trained to distinguish between real and simulated data. It fully utilizes the characteristic of CGAN to introduce conditional information, using this information to guide data generation. In the 3D-Declutter task, the input data volume is used as conditional information to learn the mapping between input data and output, thereby obtaining the desired output.
[0094] Figure 4(a) illustrates the training logic. On the left side, the input cluttered data and the generated simulated clutter-free data are fed together into the discriminator, which classifies them as false. On the right side, the cluttered data and the clutter-free data are fed together into the discriminator, which classifies them as true. During training, the loss function is continuously minimized until the clutter-free data generated by the generator can fool the discriminator.
[0095] During training, the second-level network pairs clutter-laden ground-penetrating radar (GPR) B-scan sequence sets with ideal hyperbolic wave B-scan sequence sets. It uses clean data from the ideal hyperbolic wave B-scan sequence set for generation guidance, thereby obtaining a clutter-free GPR B-scan sequence set.
[0096] The objective function of CGAN is:
[0097] L cGAN (G,D)=E x,y[logD(x,y)]+E x,z [log(1-D(x,G(x,z)))] (9)
[0098] In this clutter removal task, L1 loss is introduced as a constraint on the generator.
[0099] L L1 (G)=E x,y,z [yG(x,z)1] (10)
[0100] Therefore, the optimization problem can be represented as the following max-min problem:
[0101]
[0102] As shown in Figure 4(b), the SimAM attention mechanism adds the SimAM module to the generator and discriminator of the 3D-Declutter logical structure. This module can infer the three-dimensional attention weights for the feature map of the input B-scan without increasing the original network parameters.
[0103] SimAM is defined by the following energy function at each neuron level:
[0104]
[0105] in, n represents the target neuron, x i This represents the i-th neuron in the input feature single channel; λ is an additional parameter to prevent the denominator from being zero.
[0106] Based on (1) the more significant the difference between neuron n and surrounding neurons, according to It can be determined that the higher the importance of a neuron, the more it is augmented according to the definition of the attention mechanism, as shown in the following formula:
[0107]
[0108] Where X represents the input feature map tensor, and sigmoid represents the activation function. This represents the feature tensor after applying neuron attention, and ⊙ represents the dot product operation.
[0109] After performing the first and second levels of data processing, a dense B-Scan slice sequence can be obtained for three-dimensional inversion.
[0110] The third-level network is used to perform three-dimensional inversion on the obtained dense B-Scan slice sequence, such as Figure 1 The third-level network is a three-dimensional inversion.
[0111] The third-level 3D inversion model integrates the Pyramid Vision Transformer (PVT) and the 3D Refined Inverse Mapping Module (TRM), fusing 3D volumetric data features with decoder features.
[0112] The 3D inversion model consists of a neighbor generation module, a multi-level pyramid vision module, and a 3D refined inverse mapping module connected in sequence. The features obtained by the multi-level pyramid generation module are transmitted to the 3D refined inverse mapping module, and the features output by the neighbor generation module and the multi-level pyramid vision module are processed by a multilayer perceptron (MLP).
[0113] The adjacency generation module is configured to generate adjacencies;
[0114] The multi-level pyramid vision module is configured to extract multi-level features from the generated neighboring objects;
[0115] The 3D Refinement Inverse Mapping (TRM) module is configured to generate global and local semantics to guide self-correction based on the multi-level features of the generated neighboring volumes, utilizing the B-scan slice context information in the B-scan neighboring volumes. It is also fused with the decoder features to automatically locate and refine the local details of the volume data.
[0116] The 3D refined inverse mapping module uses a pre-trained pyramid Transformer as the kernel of the encoder.
[0117] To enhance the relevance between slices and describe the inherent feature consistency among them, the concept of Adjacent Volumes is proposed. Figure 5 The algorithm for generating adjacent volumes (GAV) is demonstrated.
[0118] The three-dimensional inversion method in step 3 includes the following steps:
[0119] Step 31: Using the neighbor generation algorithm, the inter-channel attention mechanism of dense volume data is generated. Weighted feature fusion is performed on each B-scan slice in the dense B-scan volume data with its neighboring slices to generate B-scan neighbor bodies with strong feature correlation attributes.
[0120] Each B-scan slice in the dense B-scan volume data is fused with front and back features at a certain ratio. The front and back feature associations of the B-scan slices in the dense B-scan volume data generated by inter-channel attention are jointly considered to form a B-scan neighbor body with strong feature association.
[0121] Specifically, the generation of adjacent bodies is as follows: Figure 5As shown, the current B-scan slice and the two sets of data before and after it are defined as B. i B i-1 B i+1 Multi-scale deep features are extracted from the corresponding data by performing multi-layer convolution (Conv) and pooling processing based on the encoder. Each scale uses a B-RFN to fuse these deep features, resulting in multi-scale feature fusion. It is fed into a fixed decoder network;
[0122] B-RFN stands for B-scan Residual Fusion Network.
[0123] Use L RFN As the loss function of B-RFN:
[0124]
[0125] Among them, L detail Indicates loss of detail retention, L feature Indicates feature enhancement loss; It is a balance parameter;
[0126] The loss of detail retention is calculated using the following formula:
[0127] L detail =1-SSIM(O,B i (15)
[0128] Among them, the SSIM function is from B i It preserves detailed information and structural features; O represents the fused multi-scale features, and the feature enhancement loss L feature The calculation formula is:
[0129]
[0130] Here, γ1 is a tradeoff parameter vector that balances the magnitude of the loss. w1 and w2 control the feature fusion ratio of the current B-scan and the previous and next B-scans in the generated neighboring data. In this embodiment, w1 is set to be greater than w2.
[0131] Step 32: Optimize the generated neighboring volumes by using the B-scan slice context information in the B-scan neighboring volumes to generate global and local semantics to guide self-correction. This semantics is then fused with decoder features to automatically locate and refine the local details of the volume data. Finally, the inversion result is obtained through the evolutionary calculation of the decoder.
[0132] In the 3D inversion process, non-target information increases the uncertainty of the inversion process. To further enhance the characteristics of the hyperbolic wave in the pipeline and reduce redundant information, a coarse-to-fine optimization strategy was adopted. A three-dimensional refinement module (TRM) was added to the inversion model, and its structure is as follows: Figure 1 As shown in the diagram, the generated neighbor data is fed into the TRM. Using the B-scan slice context information in the B-scan neighbor data, global and local semantics are generated to guide self-correction. These semantics are then fused with decoder features to automatically locate and refine the local details of the volume data.
[0133] Specifically, TRM uses a pre-trained pyramid Transformer as the kernel of the encoder. After the neighbor volume passes through the encoder, the generated features are fed into the global context branch for patch classification to obtain a low-resolution global semantic data volume. This data volume is further fused in TRM to highlight the hyperbolic localization pipeline.
[0134] Pixel Shuffle is an upsampling method that can effectively enlarge scaled-up feature maps and can replace interpolation or deconvolution methods.
[0135] The Conv-BN-ReLU layer performs convolution, normalization, and activation operations sequentially.
[0136] like Figure 6 As shown, in order to match the feature size of the data volume, pixel shuffle is applied to the global B-scan slice context feature f of the data volume. g To improve the consistency of target features within the volumetric data. Different scaling factors r are used depending on the stage of decoder computation, and the decoder features f... d Through multiple Conv-BN-ReLU layers and f g The features are then fused, and the fused features are denoted by F1. After the initial processing stage, the prediction result P1 is obtained, and its calculation process can be described as follows:
[0137] P1 = Sigmoid(F1(f d PS(f d PS(f g ,r))) (17)
[0138] Here, P1 contains regions of low confidence in the data. These regions are distributed close to 0.5, while the high-confidence values are distributed close to 0 or 1. H represents multiplying P1 by 1-P1, i.e.:
[0139] H(P1)=P1*(1-P1) (18)
[0140] Then a Conv-BN-ReLU layer is added, and the computation process is represented by F2. F2 and f... d After calculation, the Low-Confidence Region (LOR) is obtained. The calculation process is as follows:
[0141] LOR = f d *F2(H(P1))+f d (19)
[0142] Here, * represents successive multiplication. LOR highlights regions with low confidence, and these features serve as local background to guide the refinement stage, focusing on and refining low-confidence regions in P1.
[0143] In multi-head attention mechanisms, two types of tensors are typically involved: [B,N,C] and [B,C,H,W]. These two types of tensors are frequently converted to each other in the implementation of multi-head attention to represent inputs, outputs, and the results of intermediate computations.
[0144] After the above calculation process is completed, a tensor [B,N,C] is obtained, which is used to represent the output after multi-head attention calculation, where the attention results generated by each head are stacked in the channel dimension;
[0145] Where B represents the batch size; N represents the number of attention heads; and C represents the number of channels output by each attention head.
[0146] Then, the Transformer model is used to obtain the tensor [B,C,H,W]·S. This form of tensor is usually used to represent input data, such as image data; where B represents the batch size; C represents the number of output channels for each attention head; H represents the height of the image; W represents the width of the image. S indicates the number of B-scans in the dense sequence output after interpolation.
[0147] The inversion kernel consists of a Transformer model and Conv-BN-ReLU, and the refined calculation process is as follows:
[0148] P2=Sigmoid(F3(TF(LOR))) (20)
[0149] Compared to P1, P2 yields more refined results. The generated neighboring volumes are fed into the TRM, utilizing the B-scan slice context information within these volumes to generate self-correcting global and local semantic information. This information is fused with decoder features, automatically locating and refining local details of the volume data. This TRM structure allows the global context to guide the decoder's evolutionary computation, resulting in better prediction completeness and improved robustness of the inversion results.
[0150] A further technical solution is the coupled learning module, which uses a paired data training method between the second and third level networks to increase data correlation.
[0151] Obtaining completely simulated data of the real underground environment is extremely difficult during neural network training. This embodiment employs a sample dataset generation method that fuses simulated data with actually collected data to enhance the model's generalization ability.
[0152] Hyperbolic wave data of different morphologies in homogeneous media are extracted from the existing gprMax dataset. These are then fused with B-scan sequences measured in real-world underground areas without clearly visible targets to form training data pairs, such as... Figure 8 As shown in the figure, (a) is the relative permittivity model, (b) is the simulated pure hyperbolic wave structure, (c) is the real clutter data volume, and (d) is the mixed data volume.
[0153] The training data preparation method is as follows:
[0154] 1) The dielectric constant model is obtained through three-dimensional inversion; the dielectric constant model of the simulated data is generated by the simulation software gprMax, while the dielectric constant of the real data is manually calibrated.
[0155] 2) The dielectric constant model diagram is nonlinearly scaled to increase the range of grayscale values, and the relative dielectric constant values are labeled for known pipe locations in the training dataset;
[0156] Through three-dimensional inversion, a dielectric constant model can be obtained, which is represented as a sequence of grayscale dielectric constant images. Figure 7 The dielectric constant map is plotted as follows: each pixel value corresponds to the dielectric constant of the underground location. To accommodate different real-world underground models, the dielectric constant map is non-linearly scaled to increase the range of grayscale values, and relative dielectric constant values are labeled for known pipe locations in the training dataset.
[0157] In this embodiment, underground pipelines are predicted. Therefore, it is only necessary to distinguish between the cross-sectional shape of the pipeline and the dielectric constant of the surrounding environment to achieve binary classification of relative dielectric constant.
[0158] In the training phases of the second-level 3D-Declutter 3D denoising model and the third-level 3D inversion model of the multi-level neural network, the training data from the aforementioned dataset was used. This data was based on... Figure 8 The two sets of paired data generated in the dataset constitute data coupling. Data (a) and data (b) are paired for clutter removal training, while data (b) and data (d) are paired for inversion training. To improve performance and provide a more reliable solution for underground radar data processing, the data correlation between the second-level clutter suppression and the third-level three-dimensional inversion is enhanced through the coupling relationship between the datasets.
[0159] Furthermore, the effectiveness of the coupled learning training mode was evaluated through ablation experiments, and the model was adjusted based on the performance evaluation to further improve the performance of the 3D-Declutter de-cluttering and 3D inversion stages.
[0160] To illustrate the effectiveness of the method in this embodiment, an experiment was conducted, as detailed below.
[0161] The sparse data interpolation experiment is as follows:
[0162] A comparison chart of results obtained using three different simulated data interpolation methods on a simulated dataset, as shown below. Figure 9 As shown, Figure 10 A comparison chart of results obtained from three different interpolation methods for real-world data;
[0163] Figure 9 and Figure 10 In the image, (a) represents the validation data, (b) represents the bicubic interpolation result, (c) represents the wavelet transform interpolation result, (d) represents the CNN interpolation result, and (e) represents the transformer interpolation result; (f) represents the interpolation result performed by the first-level network in step 1 of this embodiment. The interpolation method proposed in this embodiment has a higher approximation to the validation data.
[0164] Figure 9 The comparison results of different interpolation methods on simulated data are presented. The intermediate B-scans generated by bicubic interpolation and wavelet transform interpolation methods differ significantly from the hyperbolic structure of the validation data. However, the ground-penetrating radar B-scan interpolated by the method in this embodiment can reconstruct the optimal hyperbolic structure. Compared to CNN and Transformer interpolation methods, the interpolation results of the proposed method show certain advantages in suppressing horizontal clutter. Furthermore, we also evaluated the method on real data and presented the results. Figure 10 In. Figure 10 In column f, row 1, it can be seen that the method of this embodiment provides the best suppression effect for dense noise at the bottom of the image. And... Figure 10In the image in column f, row 2, our method successfully removed most of the clutter in the middle of the B-scan while preserving the hyperbolic waveform structure closest to the validation data. This provides a good foundation for the next step of volume data clutter removal. To further demonstrate the continuity of the interpolation results and the clutter suppression capability...
[0165] Figure 11 The first row of images shows the original sparse B-scan data, and the second row shows the interpolated dense B-scan data. Columns (a), (b), and (c) are the interpolation results based on simulated data; columns (d), (e), and (f) are the interpolation results based on real data.
[0166] like Figure 11 As shown, a 3D visualization was presented. The interpolated dense B-scan data exhibited better continuity in both simulated and real data. Figure 11 As shown in columns (a), (b), and (c), in both single-channel and double-channel cases, the hyperbolic wave structure in the interpolation results becomes significantly more coherent and smooth, supplementing the missing information between B-scans. On real data, the method in this embodiment demonstrates a certain clutter suppression capability, such as... Figure 11 As shown in columns (d), (e), and (f), the obtained volumetric data is cleaner, and the hyperbolic wave structure is more prominent. It can reconstruct the optimal hyperbolic wave structure and has better clutter suppression capabilities. This provides more reliable and accurate data for ground-penetrating radar data processing.
[0167] Figure 12 The results of three-dimensional decluttering of dense data volumes using various methods, Figure 12 In the image, rows (a), (b), and (c) show clutter removal results based on simulated data; rows (d), (e), and (f) show clutter removal results based on real data.
[0168] Figure 12 The English explanation in the image is as follows:
[0169] Raw Data: Original image data;
[0170] RNMF stands for Robust Non-negative Matrix Factorization.
[0171] RPCA: Robust Principal Component Analysis.
[0172] CR-NET: Removes network clutter for Clutter Remove-NET;
[0173] Proposed Method: The method proposed in this embodiment is the clutter removal method based on the second-level network proposed in this embodiment.
[0174] from Figure 12 As can be seen, after interpolation, the dense B-scan data volume, processed by the Declutter module, successfully removed the direct waves generated by ground reflections, while highlighting the continuous hyperbolic wave structure originally surrounded by clutter. This processing effectively suppressed clutter in the data, making the underground target structure more clearly visible. The comparison of simulated / real data in the figure clearly shows the differences in performance of different methods in volume data processing. The RNMF method performed the worst, failing to effectively remove various clutter in the data volume, resulting in unclear information about the underground target structure and affecting the final imaging effect. The PRCA method performed well on simulated data, but failed to accurately reconstruct the continuous hyperbolic wave structure when dealing with real-world data, possibly due to the presence of more complex noise and interference in real data. The deep learning-based CR-NET method showed good performance on various data types, successfully removing most clutter and making the underground target structure more clearly visible. However, when encountering irregular clutter similar to the hyperbolic wave structure, the CR-NET method failed to effectively identify and remove it, resulting in a weakening of the hyperbolic wave structure. In addition, the CR-NET method may exhibit excessive clutter suppression in some cases, which can also affect the accuracy of underground target structure reconstruction.
[0175] Figure 13 The results of 3D inversion of dense data volumes using various methods. Figure 13 In the diagram, rows (a), (b), and (c) are the results of the 3D inversion based on simulated data; rows (d), (e), and (f) are the results of the 3D inversion based on real data.
[0176] Figure 13 The explanation is as follows:
[0177] Ground Truth model: In scientific research or data analysis, this term usually refers to real, accurate, and reliable information used to verify the accuracy of a model or algorithm.
[0178] CGAN: Conditional Generative Adversarial Network;
[0179] GPRINV-NET: GPR Inversion Network for ground-penetrating radar inversion;
[0180] PI-NET: Permittivity Inversion Network (PINet) is a network for inverting the dielectric constant.
[0181] Proposed Method: The method proposed in this embodiment.
[0182] from Figure 13 It can be seen that the subtle clutter remaining after Declutter processing affects various methods. Conditional GAN networks cannot effectively avoid the interference of these subtle clutters in the inversion results, resulting in a messy 3D structure. GPRInvNet, on the other hand, does not respond strongly to continuous hyperbolic structures, leading to numerous discontinuities in the inversion results and failing to maintain pipeline continuity. Although the improved PI-NET shows better purity and pipeline continuity in the 3D inversion results, some missing information still exists. The method in this embodiment achieves the best results in these aspects, effectively removing subtle clutters and exhibiting a clearer and more continuous 3D structure in the inversion results. Furthermore, the method in this embodiment performs excellently in maintaining data volume purity and pipeline continuity. This demonstrates the powerful performance of the proposed method in dealing with challenges such as subtle clutters and continuous hyperbolic structures.
[0183] In summary, through comparison Figure 13 The experimental results demonstrate the superiority of the proposed method in underground target imaging. The method exhibits significant advantages in resolving fine clutter and maintaining pipe continuity, demonstrating effectiveness and practicality, and providing valuable reference for research and application in the field of underground target imaging.
[0184] Figure 14 The images show the results when the feature fusion ratios before and after B-scan are different in adjacent data. In column (a), the fusion ratio is 0.5:9:0:0.5; in column (b), it is 1:8:1; in column (c), it is 2:6:2; and in column (d), it is 2.5:5:2.5. Figure 14 As shown in column (a), when the fusion of features before and after B-scan is not considered (no adjacent bodies are constructed), the pipeline model in the 3D visualization result exhibits discontinuity. After increasing the mixing ratio of features before and after B-scan, the continuity of its hyperbolic wave structural features is improved, thus making the visualization result more coherent. Figure 14 It can be seen that the addition of the GAV module has a certain corrective effect on the pipeline direction, proving the effectiveness of the proposed module in the network.
[0185] The experiments demonstrate that the method of this embodiment is better suited for structural analysis of ground-penetrating radar pipeline data. In particular, the inversion quality is significantly improved in complex field data experiments, proving the effectiveness and practicality of the proposed method.
[0186] Example 2
[0187] Based on Example 1, this example provides a ground-penetrating radar three-dimensional inversion system based on a multi-level neural network, including:
[0188] The 3D dense volume data generation module is configured to acquire ground-penetrating radar detection data, perform inter-channel attention interpolation on the acquired B-scan slice set, and generate B-scan 3D dense volume data.
[0189] 3D Denoising Module: Configured to use a conditional generative adversarial network to suppress redundant clutter, so that the hyperbolic wave features that can reflect the target object are preserved in the 3D dense volume data, resulting in 3D volume data containing a hyperbolic wave sequence for inversion;
[0190] Inversion Module: Configured to invert 3D volume data using a Transformer model, fuse 3D volume data features with decoder features, automatically locate and refine local details, and generate the underground 3D structure of the target object.
[0191] It should be noted that each module in this embodiment corresponds one-to-one with each step in embodiment 1, and their specific implementation process is the same, so it will not be repeated here.
[0192] Example 3
[0193] Based on Embodiment 1, this embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the computer instructions are executed by the processor, they complete the steps in the three-dimensional inversion method of ground penetrating radar based on a multi-level neural network described in Embodiment 1.
[0194] Example 4
[0195] Based on Embodiment 1, this embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they complete the steps in the three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network described in Embodiment 1.
[0196] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
[0197] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A three-dimensional inversion method for ground-penetrating radar based on multi-level neural networks, characterized in that: Construct the TCGA-Net three-dimensional inversion network for ground penetrating radar, which includes a three-level network: The first level is a 3D B-scan dense volume data generation model composed of CNN network and Transformer network; The second level is the 3D-Declutter model for clutter suppression, which uses a conditional generative adversarial network to suppress redundant clutter, so that only hyperbolic wave features that can reflect the target object are retained in the three-dimensional data, and the hyperbolic wave sequence participating in the inversion is obtained. The third level is the inversion model, which uses the Transformer model to invert the three-dimensional volume data and generate a mapping of the target underground pipeline structure; The aforementioned three-dimensional inversion method for ground-penetrating radar based on multi-level neural networks includes the following steps: Obtain ground-penetrating radar detection data, perform inter-channel attention interpolation on the B-scan slice set, and generate B-scan three-dimensional dense volume data; Conditional generative adversarial networks are used to suppress redundant clutter, so that the hyperbolic wave characteristics that can reflect the target object are preserved in the 3D dense volume data, resulting in 3D volume data containing hyperbolic wave sequences for inversion. The Transformer model is used to invert the 3D volume data, and the features of the 3D volume data are fused with the features of the decoder to automatically locate and refine local details, thereby generating the underground 3D structure of the target object.
2. The three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in claim 1, characterized in that, The method for generating B-scan 3D dense volume data includes the following steps: Interchannel information is extracted from B-scan slice data, and interpolation is performed between image slices to obtain primary dense data; An attention mechanism is embedded between slices of primary dense data to obtain time-varying and appearance features between slices of different channels. Constraints and dependencies are established between hyperbolic waves on adjacent B-scans to generate dense three-dimensional volume data.
3. The three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in claim 1, characterized in that, The inversion of 3D volume data using the Transformer model specifically includes the following steps: The GAV algorithm is used to generate an inter-channel attention mechanism for dense volume data. Weighted feature fusion is performed on each B-scan slice in the dense B-scan volume data with its neighboring slices to form a B-scan neighbor body with strong feature correlation attributes. The generated neighboring volumes are optimized by using the B-scan slice context information in the B-scan neighboring volumes to generate global and local semantics to guide self-correction. These semantics are then fused with decoder features to automatically locate and refine the local details of the volume data. Finally, through the evolutionary calculation of the decoder, the inversion result is obtained.
4. The three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in claim 1, characterized in that: In the first-level network, a CNN network is set up as a low-level feature extractor. Convolution, multiple downsampling and convolution operations are performed sequentially to obtain primary dense data. An attention mechanism is embedded between slices of primary dense data to obtain time-varying and appearance features between slices of different channels. Constraints and dependencies are established between hyperbolic waves on adjacent B-scans to generate dense three-dimensional volume data. In the first-level network, a Transformer network is set up as a motion appearance feature extractor, which includes a Transformer Block module, a downsampling module, and another Transformer Block module connected in sequence. The Transformer Block module is connected to the CNN network through a cross-scale concatenation embedding module. The cross-scale concatenation embedding module is configured to use multi-scale dilated convolution to fuse the input information together. The extraction of time-varying features is achieved through a time-varying feature processing module set up after the Transformer network. The time-varying feature processing module includes a motion estimation module and an optimization network.
5. The three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in claim 1, characterized in that: The 3D-Declutter model includes an encoder, a generator based on the SimAM attention mechanism, and a discriminator based on the SimAM attention mechanism. The SimAM modules for the generator and discriminator are added after the activation functions of the residual blocks of the generator and discriminator, respectively.
6. The three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in claim 1, characterized in that: The inversion model integrates pyramid vision and 3D refined inverse mapping modules, fusing 3D volumetric data features with decoder features; The inversion model consists of a neighbor generation module, a multi-level pyramid vision module, and a 3D refined inverse mapping module connected in sequence. The features obtained by the multi-level pyramid generation module are transmitted to the 3D refined inverse mapping module, and the features output by the neighbor generation module and the multi-level pyramid vision module are processed by a multilayer perceptron. The adjacency generation module is configured to generate adjacencies; The multi-level pyramid vision module is configured to extract multi-level features from the generated neighboring objects; The 3D fine-grained inverse mapping module is configured to generate global and local semantics to guide self-correction based on the multi-level features of the generated neighboring volumes, using the B-scan slice context information in the B-scan neighboring volumes, and fused with the decoder features to automatically locate and refine the local details of the volume data. The 3D refined inverse mapping module uses a pre-trained pyramid Transformer as the kernel of the encoder.
7. A three-dimensional inversion system for ground-penetrating radar based on a multi-level neural network, wherein the three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in any one of claims 1-6 is characterized in that, include: 3D Dense Volume Data Generation Module: Configured to acquire ground-penetrating radar detection data, perform inter-channel attention interpolation processing on B-scan slice sets, and generate B-scan 3D dense volume data; 3D Denoising Module: Configured to use a conditional generative adversarial network to suppress redundant clutter, so that the hyperbolic wave features that can reflect the target object are preserved in the 3D dense volume data, resulting in 3D volume data containing hyperbolic wave sequences for inversion; Inversion Module: Configured to invert 3D volume data using a Transformer model, fuse 3D volume data features with decoder features, automatically locate and refine local details, and generate the underground 3D structure of the target object.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, complete the steps in the ground-penetrating radar three-dimensional inversion method based on a multi-level neural network as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, complete the steps in the three-dimensional inversion method for ground-penetrating radar based on a multi-level neural network as described in any one of claims 1-6.
Citation Information
Patent Citations
Ground penetrating radar image enhancement method based on window self-attention neural network
CN115345790A
Dense time-varying array construction method and system based on controllable variational auto-encoder
CN116449305A