A method for cross-slice ROI enhancement for ground penetrating radar

By aligning and cropping multi-slice ROIs using self-supervised learning and deformable attention mechanisms, combined with shared weight multi-scale encoding and multi-scale aggregation, the problem of insufficient integrity and accuracy of hyperbolic target fusion in ground penetrating radar (GPR) detection of building exterior walls is solved, thereby enhancing GPR across slice ROIs.

CN122487408APending Publication Date: 2026-07-31BEIJING ZHONGJIAN CONSTR RES INST CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGJIAN CONSTR RES INST CO LTD
Filing Date
2026-05-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing ground-penetrating radar technology has difficulty in simultaneously preserving the shallow texture details and deep geometric semantics of hyperbolic targets in the detection of the bonding status of building exterior wall insulation layers, and it cannot effectively utilize the complementary information between adjacent slices, resulting in insufficient integrity and accuracy of the fused hyperbolic targets.

Method used

We employ a fusion method guided by deformable uncertainty, which aligns and cuts multi-slice ROIs through self-supervised learning and deformable attention mechanism, and combines shared weight multi-scale encoding and multi-scale aggregation to enhance cross-slice ROIs.

Benefits of technology

It significantly improves the integrity and accuracy of the fused hyperbolic target, effectively suppresses misalignment, blurring and noise interference, and preserves the shallow details and deep morphology of the hyperbolic target.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122487408A_ABST
    Figure CN122487408A_ABST
Patent Text Reader

Abstract

This invention discloses a method for cross-slice ROI enhancement of ground-penetrating radar (GPR), belonging to the field of radar signal processing technology. The method includes: Step S11, obtaining a first multi-slice group of GPR slices; wherein the first multi-slice group includes multiple spatially adjacent slices; Step S12, aligning and cropping the hyperbolic target regions of each slice within the first multi-slice group to obtain a first multi-slice ROI group; Step S13, inputting the first multi-slice ROI group into a pre-trained target enhancement model for processing to obtain target-enhanced slice ROIs; wherein the target enhancement model is used to perform shared-weight multi-scale encoding on the first multi-slice ROI group, perform cross-slice attention fusion and multi-scale aggregation processing at each scale, and output the target-enhanced slice ROI. This method can improve the integrity and accuracy of the fused hyperbolic target, achieving enhancement of cross-slice ROIs of GPR.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar signal processing technology, specifically relating to a method for cross-slice ROI enhancement of ground penetrating radar. Background Technology

[0002] Ground Penetrating Radar (GPR) is a geophysical exploration technology widely used in non-destructive testing of civil engineering and building quality assessment. Its working principle involves emitting high-frequency electromagnetic waves and receiving reflected signals from the target area to obtain information on the structure and media distribution of the target region. Due to its non-destructive testing advantage, it has found a core application in detecting the bonding status of building exterior wall insulation layers. The bonding status of building exterior wall insulation layers directly affects the safety, durability, and energy efficiency of buildings, including three main bonding conditions: top separation (insulation board-air-adhesive-wall structure), base separation (insulation board-adhesive-air-wall structure), and normal bonding (insulation board-adhesive-wall structure). Accurately identifying these bonding conditions can promptly detect separation defects, prevent safety hazards such as the detachment of exterior wall insulation systems, and provide crucial information for building maintenance, reinforcement, and quality control. However, in actual engineering testing, obtaining high-quality, complete hyperbolic echo images faces multiple challenges. Existing methods mostly focus on feature fusion at a single scale, making it difficult to simultaneously preserve the features of hyperbolic targets at different scales, such as shallow texture details (e.g., edge sharpness) and deep geometric semantics (e.g., overall shape). Furthermore, existing methods struggle to fully utilize complementary information between adjacent slices, resulting in poor suppression of hyperbolic target misalignment, blurring, and noise interference in ground-penetrating radar (GPR) data. This leads to insufficient integrity and accuracy of the fused hyperbolic targets. Therefore, there is an urgent need for a method for cross-slice ROI enhancement in GPR to improve the integrity and accuracy of the fused hyperbolic targets and achieve cross-slice ROI enhancement.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0005] This disclosure provides a method for enhancing cross-slice ROIs of ground penetrating radar, which can improve the integrity and accuracy of the fused hyperbolic target and achieve enhancement of cross-slice ROIs of ground penetrating radar.

[0006] In some embodiments, a method for cross-slice ROI enhancement for ground-penetrating radar includes: Step S11: Obtain the first multi-slice group of the ground penetrating radar; wherein, the first multi-slice group includes multiple spatially adjacent slices; Step S12: Align and crop the hyperbolic target regions of each slice in the first multi-slice group to obtain the first multi-slice ROI group; wherein, the first multi-slice ROI group includes multiple first slice ROIs, and the first slice ROI represents the ROI of the slice in the first multi-slice group. Step S13: Input the first multi-slice ROI group into the pre-trained target augmentation model for processing to obtain the target augmented slice ROI; wherein, the target augmentation model is used to encode the first multi-slice ROI group with shared weights at multiple scales, perform cross-slice attention fusion and multi-scale aggregation processing at each scale, and output the target augmented slice ROI.

[0007] The beneficial effects of this invention are as follows: By acquiring the first multi-slice group of ground-penetrating radar (GPR) data and aligning and cropping the hyperbolic target regions of each slice within this group, a first multi-slice ROI group is obtained to reduce spatial offset between the first slice ROIs. This first multi-slice ROI group is then input into a pre-trained target enhancement model. This model performs shared-weight multi-scale encoding on the first multi-slice ROI group and performs cross-slice attention fusion at each scale to fully utilize complementary cross-slice information and effectively suppress misalignment, blurring, and noise interference. Further multi-scale aggregation processing preserves both shallow details and deep morphology of the hyperbolic target, avoiding multi-scale feature loss, resulting in enhanced target-enhanced slice ROIs. Thus, by processing the first multi-slice ROI group using the pre-trained target enhancement model, complementary cross-slice information is fully utilized, effectively suppressing misalignment, blurring, and noise interference, while simultaneously preserving both shallow details and deep morphology of the hyperbolic target. This improves the integrity and accuracy of the fused hyperbolic target (i.e., the enhanced target-enhanced slice ROI), achieving enhancement of the GPR's cross-slice ROIs.

[0008] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0009] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a flowchart illustrating a method for cross-slice ROI enhancement of ground-penetrating radar provided by the present invention; Figure 2 This is a schematic diagram of a B-scan data preprocessing and ROI alignment process provided by the present invention; Figure 3 This is a flowchart of a training method for a target enhancement model provided by the present invention; Figure 4 This is a schematic diagram of the structure of a shared-weight multi-scale convolutional encoder provided by the present invention; Figure 5 This is a schematic diagram of the structure of a deformable cross-slice attention module provided by the present invention; Figure 6 This is a schematic diagram of the structure of a multi-scale aggregation module provided by the present invention; Figure 7 This is a schematic diagram of the structure of a decoder provided by the present invention; Figure 8 This is a flowchart illustrating the training method for a target enhancement model provided by the present invention.

[0010] Figure 9 This is a schematic diagram of a loss value calculation process provided by the present invention; Figure 10 This is a visualization result diagram of the data input model training process for thermal insulation layer detection provided by the present invention; Figure 11 This is a result diagram provided by the present invention when using a target augmentation model to infer new data; Figure 12 This is a schematic diagram of the Tenengrad self-supervised ablation result provided by the present invention; Figure 13 This is a schematic diagram of a deformable alignment ablation result provided by the present invention; Figure 14 This is a schematic diagram of the uncertainty-guided ablation result provided by the present invention; Figure 15 This is a schematic diagram of a multi-scale FPN fusion ablation result provided by the present invention; Figure 16 This is a schematic diagram of the ablation result of adversarial training provided by the present invention; Figure 17 This is a schematic diagram of an embodiment provided by the present invention; Figure 18This is a schematic diagram of another embodiment provided by the present invention. Detailed Implementation

[0011] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0012] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0013] Unless otherwise stated, the term "multiple" means two or more.

[0014] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0015] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0016] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0017] In the detection of the bonding status of building exterior wall insulation layer, the hyperbolic echo patterns of different bonding states are highly similar, the information of a single slice is insufficient, and the feature is weakened due to the offset of the measuring line.

[0018] Existing related technologies are mainly studied from three levels. At the traditional signal processing level, researchers use methods such as weighted averaging, wavelet / Shearlet transform, PCA, and multi-scale transforms (e.g., NSCT, Curvelet) to fuse multi-slice images. They separate high- and low-frequency information and enhance detail features using maximum selection or energy criteria, while leveraging signal-to-noise ratio (SNR) driven methods like SVD and VMD to improve weak signal resolution. Sparse representation and dictionary learning techniques (e.g., CSC, ICM-ODL) achieve weak feature recovery under low SNR conditions through adaptive dictionary selection. At the deep learning level, CNNs, Transformers, and their hybrid models achieve adaptive signal-to-noise differentiation through multi-scale feature extraction and attention mechanisms. Channel / spatial attention mechanisms can dynamically focus on weak signal regions and suppress environmental noise, while GANs and perceptual loss functions are used to improve the texture realism of the fusion results. For the problem of micro-displacement between slices, high-precision registration techniques such as OPDR, Zernike moments, and mutual information methods are used to correct sub-pixel-level spatial misalignment. However, existing technologies face significant limitations: Traditional methods rely heavily on precise alignment between slices, which can easily lead to feature blurring and signal-to-noise ratio degradation in scenarios with micro-displacement caused by survey line offset. While deep learning methods have strong feature extraction capabilities, their "black box" nature lacks physical interpretability, and they are sensitive to spatial misalignment and rely on large amounts of labeled data. Registration algorithms have limited performance under extreme deformation, making it difficult to balance global structure preservation and local detail enhancement. Furthermore, existing fusion strategies fail to effectively achieve intelligent information aggregation that "complements the weak with the strong," cannot dynamically select the optimal information source based on local clarity, and lack a unified framework for self-supervised learning in the absence of clear annotations, thus restricting the enhancement of target signals in complex backgrounds.

[0019] Furthermore, current methods for processing and enhancing information complementarity across multiple slices are still imperfect. Traditional fusion techniques (such as mean fusion and simple weighted fusion) do not consider local quality differences and geometric misalignments between slices, only superimposing global information, resulting in weak features not being effectively enhanced, and noise and interference being amplified simultaneously. Some deep learning-based fusion methods rely on a large number of clearly labeled ground truth values, which are difficult to obtain in engineering applications, and the model decision-making process is black-boxed, lacking interpretability and failing to meet the requirement of traceability of detection results. At the same time, existing methods mostly focus on feature fusion at a single scale, and cannot simultaneously preserve the shallow texture details (such as edge sharpness) and deep geometric semantics (such as overall shape) of hyperbolas, resulting in insufficient integrity and accuracy of the fused features.

[0020] Therefore, there is an urgent need for a cross-slice ROI enhancement method that combines physical interpretability with intelligent adaptive capabilities. This method should be tailored to adjacent B-scan slice sequences from ground-penetrating radar, capable of adaptive alignment, intelligent selection and fusion of multi-slice information, and possessing both high fidelity and strong interpretability. The method should effectively overcome problems such as incomplete information in a single slice, the submergence of weak features, and micro-displacements between slices, thereby steadily improving the signal-to-noise ratio and structural fidelity of hyperbolic targets. This will provide a high-quality, interpretable data foundation for subsequent high-precision classification and condition diagnosis, ultimately enhancing the accuracy and engineering reliability of building exterior wall insulation layer bonding condition detection.

[0021] This invention provides a ground-penetrating radar (GPR) cross-slice ROI enhancement method based on deformable uncertainty-guided fusion (i.e., a method for GPR cross-slice ROI enhancement). Addressing the difficulties in identifying weak feature hyperbolic curves caused by incomplete information in a single slice, weak target signals, and spatial misalignment between adjacent slices in GPR detection of building exterior walls, this invention proposes an intelligent fusion algorithm combining self-supervised learning and deformable attention mechanisms.

[0022] First, a ROI extraction module based on physical coordinate mapping is constructed. This module accurately crops multiple ROI slices based on the spatial location information of the hyperbolic target. Innovatively, Tenengrad gradient energy is introduced as a no-reference sharpness metric, automatically selecting the sharpest ROI slice as the pseudo-target ROI. Random degradation perturbations are applied to the remaining ROI slices, forming a self-supervised "degradation-reconstruction" training task to address the engineering challenge of missing truly sharp annotations. Next, a deformable cross-slice attention fusion module is designed. This module computes slice-level global weights and pixel-level uncertainty confidence maps in parallel, and predicts deformable offset fields to achieve sub-pixel-level spatial correction. Through weighted aggregation, fusion features are generated, achieving an intelligent multi-focusing effect of "automatic selection of sharp regions, suppression of blurred regions, and correction of misaligned regions," making weak-feature hyperbolic features more prominent. Finally, a multi-scale fusion strategy in the style of Feature Pyramid Network (FPN) is adopted to ensure the collaborative optimization of shallow texture details and deep semantic features, balancing hyperbolic realism and detail preservation. Finally, progressive adversarial training optimization was implemented. After the L1 and SSIM losses were stabilized and optimized, the PatchGAN discriminator and VGG perceptual loss were gradually introduced to drive the generated results to approach the characteristics of high-quality radar images in terms of texture realism and edge sharpness, and to enhance important details of the echo.

[0023] This invention introduces a deformable sampling and uncertainty-guided mechanism in GPR fusion, achieving end-to-end fusion of "slice-level selection - pixel-level alignment - multi-scale aggregation". Without requiring clear ground truth annotation, it achieves precise alignment and detail enhancement of weak feature hyperbolas, significantly improving the detail and edge clarity of the fusion results, providing high-quality input for subsequent feature extraction and state classification. The embodiments disclosed in this invention are mainly based on ground-penetrating radar multi-slice signal fusion and weak feature enhancement technology. Addressing the core problems in the detection of the bonding status of building exterior wall insulation layers, such as incomplete information from a single slice, the susceptibility of weak feature hyperbolas to noise masking, fusion distortion caused by cross-slice micro-displacement, and insufficient model interpretability, this invention proposes a ground-penetrating radar cross-slice ROI enhancement method based on deformable uncertainty-guided fusion.

[0024] Having introduced the overall concept of the present invention, to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0025] Combination Figure 1 As shown, this disclosure provides a method for cross-slice ROI enhancement for ground-penetrating radar, including: Step S11: Obtain the first multi-slice group of the ground penetrating radar; wherein, the first multi-slice group includes multiple spatially adjacent slices.

[0026] Step S12: Align and crop the hyperbolic target regions of each slice within the first multi-slice group to obtain the first multi-slice ROI group; wherein, the first multi-slice ROI group includes multiple first slice ROIs, and the first slice ROI represents the ROI of the slice within the first multi-slice group. The Chinese name for ROI is Region of Interest, and the English name is Region of Interest.

[0027] For example, combined Figure 2 As shown, Figure 2This paper presents a schematic diagram of B-scan data preprocessing and ROI alignment. Based on the physical spatial location information of hyperbolic targets recorded during ground-penetrating radar (GPR) detection (including the depth direction t_cut time window parameter and the horizontal detection x_range lateral range parameter), an explicit mapping relationship between physical coordinates and image pixel coordinates is established to ensure a one-to-one correspondence between each physical location and a pixel location. Following the slice acquisition order, K1 spatially adjacent B-scan slices (in actual GPR detection, the number of survey lines K1 is generally 4-25) are grouped together. The hyperbolic target regions in each group of slices are then aligned and cropped to obtain ROI regions of uniform size. Subsequently, the cropped ROI regions are normalized (mapping pixel values ​​to the [0,1] interval) to form an input set containing K1 ROI regions (i.e., obtaining the first slice ROI group), which is mapped to a unified image pixel coordinate system to ensure the spatial consistency of the echo signal of the same hyperbolic target in different slices.

[0028] Step S13: Input the first multi-slice ROI group into the pre-trained target augmentation model for processing to obtain the target augmented slice ROI; wherein, the target augmentation model is used to encode the first multi-slice ROI group with shared weights at multiple scales, perform cross-slice attention fusion and multi-scale aggregation processing at each scale, and output the target augmented slice ROI.

[0029] This disclosure provides a method for cross-slice ROI enhancement using ground-penetrating radar (GPR). The method involves acquiring a first multi-slice group from the GPR and aligning and cropping the hyperbolic target regions of each slice within this group to obtain a first multi-slice ROI group, thereby reducing spatial offset between the first slice ROIs. The first multi-slice ROI group is then input into a pre-trained target enhancement model. This model performs shared-weight multi-scale encoding on the first multi-slice ROI group and performs cross-slice attention fusion at each scale to fully utilize complementary cross-slice information and effectively suppress misalignment, blurring, and noise interference. Furthermore, multi-scale aggregation processing is used to simultaneously preserve the shallow details and deep morphology of the hyperbolic target, avoiding multi-scale feature loss, thus obtaining the enhanced target-enhanced slice ROI. In this way, by processing the first multi-slice ROI group through the pre-trained target enhancement model, the complementary information across slices can be fully utilized, effectively suppressing misalignment, ambiguity and noise interference, and the target enhancement slice ROI can simultaneously retain the shallow details and deep morphology of the hyperbolic target, thereby improving the integrity and accuracy of the fused hyperbolic target (i.e. the target enhancement slice ROI) and realizing the enhancement of the ground penetrating radar cross-slice ROI.

[0030] Preferably, combined with Figure 3 As shown, Figure 3A flowchart illustrating a training method for an object augmentation model is provided. The object augmentation model is obtained through the following steps: Step S21: Obtain the training sample set; wherein each sample in the training sample set includes a second multi-slice ROI group and a pseudo-target slice ROI, and the second multi-slice ROI group includes multiple second slice ROIs.

[0031] Understandably, the second slice ROI represents the ROI of the slices within the second multi-slice group. The second multi-slice group comprises multiple spatially adjacent slices. The first and second multi-slice groups are ground-penetrating radar slice groups with different application scenarios. The first multi-slice group represents the slice group to be augmented, meaning its corresponding first multi-slice ROI group serves as the input to the target augmentation model. The second multi-slice group, on the other hand, represents the slice group used for model training, meaning its corresponding second multi-slice ROI group serves as the input to the original augmentation model.

[0032] Step S22: Train the original augmentation model using the training sample set. The original augmentation model includes a shared-weight multi-scale convolutional encoder, a deformable cross-slice attention module, a multi-scale aggregation module, and a decoder; the training process is as follows: Step S221: For each second multi-slice ROI group, a shared-weight multi-scale convolutional encoder is used to process each second slice ROI to obtain multi-level feature maps corresponding to each second slice ROI, forming a multi-level feature map set; wherein, the multi-level feature map includes feature maps of multiple levels. One second multi-slice ROI group corresponds to one multi-level feature map set, which represents the set of multi-level feature maps corresponding to the second slice ROIs within the second multi-slice ROI group.

[0033] It is understandable that each level corresponds to a scale.

[0034] For example, combined Figure 4 As shown, Figure 4A schematic diagram of a shared-weight multi-scale convolutional encoder is provided. For each second multi-slice ROI group, K2 ROI regions of the second multi-slice ROI group (i.e., K2 second-slice ROIs) are input into the shared-weight multi-scale convolutional encoder. This shared-weight multi-scale convolutional encoder consists of three cascaded convolutional blocks. Each convolutional block contains double 3×3 convolution, batch normalization (BN), and ReLU activation function, interspersed with 2×2 max pooling operations. Finally, it outputs a three-level feature map (dimensions of K2×C×H×W, where C is the number of feature channels, and H and W are the height and width of the feature map, i.e., a multi-level feature map) for each ROI region (i.e., each second-slice ROI). The output multi-level feature map is used as the input of a deformable cross-slice attention module. The first-level convolutional block outputs a shallow feature map (denoted as Enc1), which focuses on capturing the detailed texture information of the hyperbolic target within the ROI region (i.e., the second slice ROI). The number of feature channels is set to 32, which can completely preserve the subtle texture changes of the echo signal. The second-level convolutional block outputs a mid-level feature map (denoted as Enc2), which corresponds to the mid-level semantic information and can represent the local morphological features of the hyperbolic target. The number of feature channels is set to 64. The third-level convolutional block outputs a deep feature map (denoted as Enc3), which focuses on extracting the overall geometric semantic information of the hyperbolic target and can reflect the global structural features of the target. The number of feature channels is set to 128. The above three layers of feature maps together constitute a multi-scale feature pyramid (i.e., multi-level feature maps) covering shallow texture, mid-level local semantics, and deep global semantics.

[0035] Step S222: For each multi-level feature map set, the deformable cross-slice attention module is used to fuse all feature maps of the same level to obtain the fused feature map corresponding to each level, so as to form a multi-level fused feature map; wherein, the multi-level fused feature map includes fused feature maps of multiple levels. Step S223: For each multi-level fusion feature map, the fusion feature maps corresponding to each level are aggregated using the multi-scale aggregation module to obtain a multi-scale aggregated feature map, and the multi-scale aggregated feature map is decoded using the decoder to obtain the enhanced slice ROI. In step S224, based on the loss between each augmented slice ROI and the corresponding pseudo-target slice ROI, adjust the network parameters of the original augmentation model, and repeat steps S221 to S224 until training is completed and the target augmentation model is obtained.

[0036] Preferably, training is considered complete when the loss value converges or the number of iterations reaches the third preset number of iterations.

[0037] The original augmentation model is trained using a training sample set. For each sample (i.e., each second multi-slice ROI group), the shared weight multi-scale encoder of the original augmentation model is used to extract multi-level feature maps of each second-slice ROI within the sample, thus obtaining the multi-level feature map set corresponding to the second multi-slice ROI group. This multi-level feature map set is then used as input to a deformable cross-slice attention module. This module performs cross-slice attention fusion on the feature maps of all second-slice ROIs within the sample at the same level to obtain a multi-level fused feature map. This fully utilizes the complementary information across slices and effectively suppresses misalignment, blurring, and noise interference. The multi-level fused feature map is then used as input to a multi-scale aggregation module. This module performs multi-scale aggregation on the multi-level fused feature map to simultaneously preserve the shallow details and deep morphology of the hyperbolic target, thereby improving the integrity and accuracy of the fused hyperbolic target and achieving augmentation of the ground-penetrating radar cross-slice ROI. Based on the loss between each multi-scale aggregated feature map and the corresponding pseudo-target slice ROI, the network parameters of the original augmentation model are adjusted (i.e., the network parameters of the shared-weight multi-scale convolutional encoder, the deformable cross-slice attention module, and the multi-scale aggregation module are adjusted) until the target augmentation model is obtained. In this way, by using this training sample set to train the original augmentation model, a target augmentation model can be obtained to improve the integrity and accuracy of the fused hyperbolic target, thereby achieving the augmentation of ground-penetrating radar cross-slice ROIs.

[0038] Preferably, obtaining the training sample set includes: Obtain multiple third-level multi-slice ROI groups; where each third-level multi-slice ROI group includes multiple third-level ROIs. For each third multi-slice ROI group, the third slice ROI with the highest Tenengrad value is taken as the pseudo-target slice ROI, and random degradation is applied to the other third slice ROIs except for the pseudo-target slice ROI to obtain the second multi-slice ROI group. A second multi-slice ROI group and its corresponding pseudo-target ROI are used as a sample to obtain the training sample set.

[0039] In this way, for each third multi-slice ROI group, the third slice ROI with the highest Tenengrad value in the third multi-slice ROI group is used as the pseudo target slice ROI corresponding to that third multi-slice ROI group. Random degradation is applied to the remaining third slice ROIs in each third multi-slice ROI group (i.e., the fourth slice ROI group) to simulate the quality degradation of the ROI region caused by signal attenuation, environmental interference, equipment noise, and gain fluctuation during ground penetrating radar detection, generating low-quality degradation samples (i.e., the second multi-slice ROI group). This allows for the provision of supervision signals for model training without the need for manually annotating clear ground truth values.

[0040] For example, obtaining multiple third multi-slice ROI groups includes: obtaining multiple second multi-slice groups of ground-penetrating radar, aligning and cropping the hyperbolic target regions of each slice within the second multi-slice group to obtain the third multi-slice ROI group. In this embodiment, the ROI of each slice within each second multi-slice group is extracted based on physical coordinate-pixel coordinate mapping to obtain the third multi-slice ROI group corresponding to each second multi-slice group.

[0041] Preferably, the third multi-slice ROI group includes multiple third slice ROIs, where the third slice ROI represents the ROI of the slices within the third multi-slice group, and the slices within the third multi-slice group are spatially adjacent slices.

[0042] Preferably, the third slice ROI with the highest Tenengrad value is used as the pseudo-target slice ROI, which includes: calculating the Tenengrad value of each third slice ROI in the third multi-slice ROI group using the following formula, and using the third slice ROI with the highest Tenengrad value as the pseudo-target slice ROI: in, These are the selected pseudo-target slice ROIs (i.e., the ROI of the slice with the largest gradient energy, which serves as a reference for the self-supervised "degeneration-reconstruction" task, i.e., as a label). This is the ROI image matrix of the k-th slice within the third multi-slice ROI group (i.e., the ROI image matrix of the k-th slice, also known as the image matrix of the k-th third slice ROI). pixel coordinates ( , , (Height and width of the ROI). This is the gradient energy (i.e., the no-reference sharpness index) of the ROI of the k-th slice within the third multi-slice ROI group. A larger value indicates sharper image edges. For pixels The horizontal gradient value (reflecting the change in grayscale in the vertical direction). For pixels Vertical gradient value (reflecting horizontal grayscale changes): in, This represents a two-dimensional convolution operation (the boundaries are symmetrically padded to avoid loss of edge information). The horizontal Sobel convolution kernel is used to detect vertical edges (the longitudinal contour of the GPR hyperbola). A vertical Sobel convolution kernel is used to detect horizontal edges (lateral details of the GPR hyperbola): In this way, by introducing Tenengrad gradient energy as a no-reference sharpness metric, we can quickly filter the ROI regions with the richest details and clearest textures in each input set (i.e., each third multi-slice ROI group) as pseudo-targets (i.e. pseudo-sharp targets). This enables the automatic selection of the unit with the highest energy as the "pseudo-sharp target", thereby reducing the workload of manual annotation, improving the efficiency of sample acquisition, and solving the core problem of the lack of "sharp" ground truth values ​​in training data.

[0043] Preferably, random degradation includes at least one of random Gaussian blur, random additive Gaussian noise, and random contrast perturbation. The number of slices in the third multi-slice ROI group is K3, the number of slices in the second multi-slice ROI group is K2, and K3 = K2 + 1.

[0044] For example, random degradation is applied to the remaining third slice ROIs (excluding the pseudo-target slice ROI) to obtain a second multi-slice ROI group. This includes applying random Gaussian blur, random additive Gaussian noise, and random contrast perturbation to the remaining third slice ROIs (excluding the pseudo-target slice ROI) to obtain the second multi-slice ROI group, as shown in the following formula: in, The kth third slice ROI (i.e. the kth multi-slice ROI) is represented within the third slice ROI group, excluding the pseudo-target slice ROI. The k-th third slice ROI (i.e. the k-th multi-slice ROI) is represented after applying random Gaussian blur, excluding the pseudo-target slice ROI. The standard deviation is Gaussian kernel, This refers to the kth third slice ROI (i.e., the kth multi-slice ROI) excluding the pseudo-target slice ROI, after applying random additive Gaussian noise. The kth third slice ROI (i.e. the kth multi-slice ROI) is characterized after applying random contrast perturbation, excluding the pseudo-target slice ROI. nuclear size , , In this data experiment, each transformation is expressed with probability. Standalone application: .

[0045] In this way, for each third multi-slice ROI group, random Gaussian blur, additive Gaussian noise, and brightness / contrast perturbation are applied to the remaining third slice ROIs except for the pseudo-target slice ROI. This enables the construction of a self-supervised reconstruction task of "multiple degenerate low-quality inputs → single high-quality pseudo-targets". This allows the model to learn to aggregate effective information from neighboring slices with uneven quality without the guidance of clear fusion ground truth without any manual annotation.

[0046] Preferably, combined with Figure 5 As shown, Figure 5 A schematic diagram of a deformable cross-slice attention module is provided. This module fuses all feature maps of the same level to obtain fused feature maps for each level, including: Perform the following steps using the deformable cross-slice attention module: For each level, the slice-level global attention weights, pixel-level confidence maps, and deformable offset fields of each feature map are calculated, as well as... Based on the global attention weights at each slice level and the confidence maps at each pixel level, the fusion weight matrix for each feature map is determined, and Based on each deformable migration field, spatial position correction is performed on each feature map to obtain each corrected feature map, and Based on each fusion weight matrix, all corrected feature maps are weighted and summed to obtain the fusion feature map corresponding to the current level.

[0047] In this way, by utilizing the deformable cross-slice attention module to compute slice-level global weights and pixel-level confidence maps in parallel and predict deformable offset fields, misalignment correction and adaptive selection of high-confidence features are achieved. This mechanism realizes a three-in-one process of "global importance assessment → local confidence weighting → active correction of geometric misalignment", resulting in intelligent feature selection and fusion effects of "prioritizing the retention of clear regions, automatically suppressing blurred regions, and accurately correcting misaligned regions".

[0048] Preferably, calculating the slice-level global attention weights of each feature map includes: obtaining the reference feature vector corresponding to the current multi-level feature map set; and performing global interaction between the reference feature vector and each feature map to obtain the slice-level global attention weights corresponding to each feature map.

[0049] For example, obtaining the reference feature vector corresponding to the current multi-level feature map set includes: obtaining the reference feature corresponding to the current multi-level feature map set, and determining the corresponding reference feature vector based on the reference feature, as shown in the following formula: in, The reference features that characterize the current multi-level feature map set (i.e., the reference features corresponding to the current sample). The corresponding reference feature vector is represented (randomly initialized in the early stage of training and then optimized iteratively with the model).

[0050] In this way, during the training of the target augmentation model, each group (i.e., each sample) uses an independent reference feature vector to calculate the slice-level global attention weights of each second slice ROI within the group, based on each adjacent slice ROI within the batch (i.e., each sample within the batch). Different groups do not share reference feature vectors to ensure the independence and accuracy of cross-slice fusion for each group.

[0051] Understandably, each second multi-slice ROI group (i.e., each sample, and each sample corresponding to a multi-level feature map set) independently calculates its own reference feature vector; the reference feature vectors of different samples within a batch are independent of each other and are not shared; only within a single second multi-slice ROI group do the slices share the group's own reference feature vector. That is, within the same batch, each group of K2 adjacent slices (one second multi-slice ROI group) shares a reference feature vector; different groups and different samples use their own independent reference feature vectors.

[0052] A sample consists of a set of K2 second multi-slice ROIs (i.e., a second multi-slice ROI group), and a batch B represents the simultaneous processing of B samples (B groups). Each group shares a reference vector, but the reference vectors are not shared between groups.

[0053] For example, the reference feature vector is globally interacted with each feature map to obtain the slice-level global attention weights corresponding to each feature map, including: Global average pooling is performed on each feature map to obtain the slice feature embedding vector, as shown in the following formula: in, The slice feature embedding vector representing the feature map of the k-th ROI of the second slice within the second multi-slice group. The feature map representing the k-th ROI of the second slice within the second multi-slice group; B is the Batch Size (the number of samples processed at one time, 16 during training and 1 during inference). The similarity between each feature map and the reference feature vector is calculated by vector dot product operation, and then normalized by the Softmax function to obtain the slice-level global attention weights corresponding to each feature map, as shown in the following formula: in, The slice-level global attention weights represent the feature maps of the kth second slice ROI within the second multi-slice group (i.e., the slice-level global attention weights of the feature maps of the kth second slice ROI within the current sample). The feature embedding vector representing the feature map of the k-th ROI of the second slice within the second slice group. With reference feature vector Dot product similarity.

[0054] Preferably, calculating the pixel-level confidence map of each feature map includes: compressing and mapping each feature map channel to 1 through a lightweight convolutional layer, and activating it with Sigmoid to obtain the pixel-level confidence map corresponding to each feature map.

[0055] In this way, a lightweight convolutional layer compresses and maps the feature map channels, adjusting the number of feature channels to 1. Then, a sigmoid activation function is used to output a pixel-level confidence map (i.e., pixel-level uncertainty confidence, also known as independent prediction confidence). For each spatial location (H, W) of the feature map of each second slice ROI, a confidence value between 0 and 1 is predicted. This value quantifies the reliability of the feature information at that location. Low confidence usually corresponds to noisy regions or areas with blurred signals, while a higher confidence value indicates that the feature information at that pixel location is more reliable.

[0056] For example, for each second slice ROI, its pixel-level confidence map is calculated using the following formula: in, A pixel-level confidence map representing the feature map corresponding to the k-th ROI of the second slice within the second slice ROI group. This is the Sigmoid function, with an output range of (0,1).

[0057] Preferably, calculating the deformable offset field of each feature map includes: based on the reference features corresponding to the current multi-level feature map set, using offset prediction convolutional layers to predict the deformable offset field corresponding to each feature map, as shown in the following formula: in, The deformable offset field characterizes the feature map corresponding to the k-th second-slice ROI within the second multi-slice ROI group. The offset field of each slice is independent, which can specifically correct the micro-displacement of its own hyperbola, achieving slice-level sub-pixel level spatial misalignment correction. The reference features that represent the current multi-level feature map set (i.e., the reference features corresponding to the current sample, or the reference features corresponding to the current second multi-slice group). The feature map representing the ROI of the k-th second slice within the second multi-slice group. The channel-dimensional concatenation operation is used to generate joint features between the reference features and the feature map of the k-th second slice ROI. The offset prediction convolutional layer is characterized, preferably a 3x3 convolutional layer.

[0058] In this way, the feature map is processed by the offset prediction convolutional layer, and a deformable offset field is output (with dimensions of B×2×H×W, where "2" represents the sub-pixel offset in the horizontal and vertical directions, used to correct spatial misalignment between adjacent slices).

[0059] Preferably, based on the global attention weights at each slice level and the confidence maps at each pixel level, the fusion weight matrix for each feature map is determined, including: For each feature map, its slice-level global weights are multiplied element-wise by the pixel-level confidence map, and then normalized using the Softmax function to obtain the fused weight matrix: in, The fusion weight matrix represents the feature map corresponding to the k-th second slice ROI within the second slice ROI group.

[0060] Preferably, spatial position correction is performed on each feature map based on each deformable offset field to obtain each corrected feature map, including: Based on each deformable offset field, the sampling grid corresponding to each feature map is determined, as shown in the following formula: in, This represents the deformable offset field corresponding to the feature map of the k-th second-slice ROI within the second multi-slice ROI group. The offset field of each slice is independent, allowing for targeted correction of micro-displacements of its own hyperbola, achieving slice-level sub-pixel level spatial misalignment correction. The sampling network representing the feature map of the k-th second-slice ROI within the second multi-slice ROI group. Represents a regular grid (range [−1,1]). This is the offset range parameter, with a default value of 2.0.

[0061] Based on the sampling grid corresponding to each feature map, deformable sampling is performed on each feature map through bilinear sampling operation to obtain the corrected feature map, as shown in the following formula: in, Characterization The corresponding correction feature map (i.e., the correction feature map corresponding to the feature map of the kth second slice ROI in the second multi-slice ROI group).

[0062] This achieves pixel-level adaptive spatial alignment to correct for minute misalignments between slices caused by acquisition geometry.

[0063] Preferably, based on each fusion weight matrix, all corrected feature maps are weighted and summed to obtain the fusion feature map corresponding to the current level, including the following calculation formula: in, The fusion feature map represents the current level, and ⊙ represents element-wise multiplication.

[0064] In this way, the features (i.e. the corrected feature maps) after deformable sampling alignment are weighted and summed using the various fusion weight matrices, and the fused single feature map is output, thus obtaining the fused feature map of the current level.

[0065] Preferably, combined with Figure 6 As shown, Figure 6 A schematic diagram of a multi-scale aggregation module is provided. The multi-scale aggregation module aggregates the fused feature maps corresponding to each level to obtain a multi-scale aggregated feature map, including: Perform the following steps using the multi-scale aggregation module: Channel mapping and downsampling are performed on each low-level fusion feature map to obtain each processed fusion feature map. Among them, the low-level fusion feature map represents the fusion feature map in the multi-level fusion feature map except for the high-level fusion feature map, and the high-level fusion feature map represents the fusion feature map with the most channels in the multi-level fusion feature map. The number of channels and spatial size of the processed fusion feature map are the same as those of the high-level fusion feature map. The aggregated feature map is obtained by adding the individual processed fusion feature maps and the high-level fusion feature map element by element. Add positional encoding to the aggregated feature map to obtain a multi-scale aggregated feature map.

[0066] In this way, the multi-scale aggregation module aggregates the multi-level fused feature maps at multiple scales, fusing high-resolution shallow features containing rich echo textures and details with low-resolution deep features containing robust target structure and semantic information to obtain the target fusion features corresponding to each sample (i.e., the second multi-slice ROI group). This target fusion feature preserves the original texture authenticity of the hyperbolic target while ensuring that subtle texture details are not lost, achieving synergistic optimization of shallow detail texture features and deep semantic features.

[0067] Preferably, a multi-scale fusion strategy in the FPN style is adopted to achieve feature collaborative optimization. The multi-level feature map in this embodiment consists of a feature pyramid composed of feature maps at three different levels; therefore, the multi-level fused feature map also includes fused feature maps at three different levels. See again... Figure 6 The multi-level fusion feature map consists of low-level fusion feature maps (32 channels) and mid-level fusion feature maps (64 channels), and high-level fusion feature maps (deep-level fusion feature maps). Lateral convolution is used to map the shallow fusion feature maps, mapping the 32-channel fusion feature map (Shallow fusion feature map) for Enc1 and the 64-channel fusion feature map (Mid-level fusion feature map) for Enc2 to 128 channels, ensuring consistency with the 128-channel fusion feature map (Deep fusion feature map) for Enc3, thus guaranteeing dimensionality matching across different levels. Then, downsampling is performed on the mapped fusion feature maps for Enc1 and Enc2 proportionally to align their spatial dimensions perfectly with those of the fusion feature map for Enc3. Finally, the three fusion feature maps are added element-wise to complete the initial aggregation of multi-scale features. To enhance the representation of spatial location information of features, a location code is added to the aggregated feature map. This location code is generated based on sine and cosine functions. By using trigonometric functions of different frequencies to represent feature information of different spatial locations, it can effectively compensate for the weakening of location information by convolution operations.

[0068] This multi-scale path, combining top-down and lateral connections, does not transmit information between different scales of a single image. Instead, it performs cross-scale information aggregation after cross-slice fusion has been completed at each scale. This ensures that the fusion process utilizes complementary information from multiple slices and integrates visual cues from multiple resolutions. Systematically fusing high-resolution shallow features containing rich echo textures and details with low-resolution deep features containing robust target structure and semantic information preserves the original texture authenticity of the hyperbolic target while ensuring that subtle texture details are not lost. This achieves synergistic optimization of shallow detail texture features and deep semantic features, providing core support for the detail richness and texture authenticity of subsequent enhancement results.

[0069] Preferably, the decoder is used to decode the multi-scale aggregated feature map to obtain the enhanced slice ROI, including: the decoder performs the following steps: the multi-scale aggregated feature map is upsampled at two layers to restore the spatial resolution of the multi-scale aggregated feature map to the same size as the original input ROI (i.e. the second slice ROI) to obtain multi-channel features; and then the multi-channel features are mapped to a single-channel enhanced radar echo image to obtain the enhanced slice ROI.

[0070] In this way, by using a decoder to decode the multi-scale aggregated feature map, it is possible to obtain the enhanced slice ROI.

[0071] Combination Figure 7 As shown, Figure 7 A schematic diagram of a decoder structure is provided. The decoder includes two upsampling units, a 1×1 convolutional layer, and a Tanh activation function. Each upsampling unit includes a bilinear upsampling operation (magnifying the feature map size by a factor of 2), a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function. The multi-scale aggregated feature map is upsampled through two layers to restore its spatial resolution to the same size as the original input ROI (i.e., the second slice ROI), obtaining multi-channel features. These multi-channel features are then mapped to a single-channel enhanced radar echo image to obtain the enhanced slice ROI. This process includes: performing two upsampling layers on the multi-scale aggregated feature map through the decoder's two upsampling units to restore its spatial resolution to the same size as the original input ROI (i.e., the second slice ROI), obtaining multi-channel features; and then processing these features with a 1×1 convolutional layer and a Tanh activation function to map the multi-channel features to a single-channel enhanced radar echo image to obtain the enhanced slice ROI. The decoder does not use skip connections throughout the process to avoid introducing inconsistent raw information between slices, ensuring that the final output is generated entirely based on cross-slice fused features.

[0072] Combination Figure 8 As shown, Figure 8 This paper presents a flowchart illustrating the training method for a target augmentation model. The overall process is divided into four stages: input preprocessing (i.e., acquiring the training sample set), feature encoding (using a shared-weight multi-scale convolutional encoder for feature encoding), cross-slice fusion (using a deformable cross-slice attention module and a multi-scale aggregation module for cross-slice fusion), and decoding output. Simultaneously, pseudo-target supervision and composite loss constraints are introduced.

[0073] The process includes input, ROI preprocessing and alignment, and pseudo-target construction stages: The input receives K3 adjacent B-scan slice ROIs as raw data, converts their physical coordinates to pixel coordinates, and performs normalization to ensure spatial consistency and uniform numerical range between different slices. Through Tenengrad pseudo-target selection, the highest quality reference slice ROI is selected as the pseudo-target slice ROI, used as a supervision signal for subsequent loss calculation. This involves obtaining a training sample set and using it to train the original augmentation model.

[0074] The feature encoding stage includes a multi-scale convolutional encoder, also known as a shared-weight multi-scale encoding module. It constructs a feature pyramid Enc1→Enc2→Enc3, extracts multi-scale features through convolutions at different levels, captures ROI structural information from fine-grained to coarse-grained, and obtains multi-level feature maps.

[0075] The cross-slice feature fusion stage includes Cross-Slice Variable Attention (CSDA, or Deformable Cross-Slice Attention Module) and FPN Multi-Scale Fusion (i.e., Multi-Scale Aggregation Module). Cross-Slice Variable Attention (CSDA) dynamically models the correlation between adjacent slices through slice-level weight allocation, pixel-level confidence modeling, and deformable sampling, adaptively aggregating effective information and suppressing noise interference. FPN Multi-Scale Fusion utilizes lateral convolution, feature aggregation, and positional encoding to efficiently fuse feature maps of different scales, preserving spatial location information and enhancing feature representation capabilities.

[0076] The decoding and output stage includes a decoder and a composite loss constraint. The decoder uses upsampling and convolution operations to progressively restore the fused multi-scale features to a single-channel output, obtaining the final enhanced ROI (i.e., the enhanced slice ROI). The composite loss constraint, during training, combines L1 loss, SSIM structural similarity loss, perceptual loss, and GAN adversarial loss to constrain the model, ensuring the enhancement result performs well in terms of pixel accuracy, structural consistency, and visual realism.

[0077] Final output: The model finally outputs the enhanced ROI (i.e., the enhanced slice ROI), which improves the quality and enhances the information of the original B-scan slice ROI.

[0078] Preferably, based on the loss between each augmented slice ROI and its corresponding pseudo-target slice ROI, the network parameters of the original augmentation model are adjusted, including: If the current iteration round is less than the first preset round, then the network parameters of the original augmentation model are adjusted based on the basic reconstruction loss between all augmented slice ROIs and the corresponding pseudo-target slice ROIs in the current batch, on a batch basis. If the first preset round ≤ the current iteration round < the second preset round, then fix the network parameters of the original augmentation model, and adjust the network parameters of the discriminator based on the discriminator loss between all augmented slice ROIs and the corresponding pseudo-target slice ROIs in the current batch, in batches. If the current iteration round is greater than the second preset round, then on a batch basis, based on all augmented slice ROIs and their corresponding pseudo-target slice ROIs within the current batch, the basic reconstruction loss, perceptual loss, adversarial loss, and discriminator loss are determined. The network parameters of the original augmentation model are adjusted based on the basic reconstruction loss, perceptual loss, and adversarial loss, and the network parameters of the discriminator are adjusted based on the discriminator loss. Wherein, the first preset round < the second preset round < the third preset round.

[0079] The entire training process employs a progressive strategy. In the initial warm-up phase, only the L1+SSIM loss is used to train the generator. After the warm-up, the generator is fixed, and the discriminator is trained separately for several epochs. Subsequently, the model is trained end-to-end using all four types of losses (L1, SSIM, VGG perceptual loss, and adversarial loss) in an alternating optimization manner. That is, in the initial training phase, the parameters of the original augmentation model are adjusted using the basic reconstruction loss function. After the original augmentation model has initially stabilized (i.e., the current iteration reaches the second preset epoch), more complex loss components are gradually introduced, using the basic reconstruction loss, perceptual loss, adversarial loss, and discriminator loss as the total loss to adjust the parameters of the original augmentation model. In this way, the phased, multi-constraint optimization mechanism, while ensuring the stability of the training process, drives the model output to be not only accurate in pixel values ​​but also richer in texture details and visually closer to the fusion results of high-quality radar echoes, significantly enhancing the clarity and recognizability of hyperbolic features. By adopting a progressive training strategy of "first ensuring fidelity, then improving quality", the visual realism and edge sharpness of hyperbolic textures are improved while ensuring that the fusion result (i.e., the enhanced slice ROI) and the pseudo target slice ROI are consistent at the pixel level. At the same time, through multi-loss collaborative optimization, pixel fidelity, structural consistency and perceptual quality are balanced to meet the dual requirements of interpretability and visual clarity of GPR engineering detection.

[0080] The first preset round is any one of rounds from 5 to 20.

[0081] Preferably, to ensure training efficiency and quality, the training sample set for each training round is divided into multiple batches. For each batch, the loss of all samples in that batch is calculated, and the network parameters of the original augmentation model are updated based on the loss. After the parameter update of the current batch is completed, the next batch of samples is processed until all batches of the current round are processed, and then the next round of training begins.

[0082] Preferably, the network parameters of the original augmentation model are adjusted based on the basic reconstruction loss between all augmented slice ROIs and their corresponding pseudo-target slice ROIs in the current batch, including: Based on all enhanced slice ROIs and corresponding pseudo-target slice ROIs in the current batch, determine the L1 loss and SSIM structural similarity loss; Based on L1 loss and SSIM structural similarity loss, the basic reconstruction loss is determined; Based on the basic reconstruction loss, the network parameters of the original enhancement model are adjusted.

[0083] L1 loss measures the absolute difference in pixel values ​​between the enhanced slice ROI feature map and the pseudo-target ROI feature map (i.e., the pseudo-target slice ROI). SSIM structural similarity loss is constructed based on the structural similarity index, measuring the structural consistency between the two feature maps (enhanced slice ROI and pseudo-target ROI) by comparing their local mean, local standard deviation, and local covariance. Thus, using L1 loss and SSIM structural similarity loss as the basic reconstruction losses ensures pixel-level similarity and structural consistency between the enhanced result and the pseudo-target.

[0084] Preferably, based on all enhanced slice ROIs and corresponding pseudo-target slice ROIs within the current batch, the L1 loss and SSIM structural similarity loss are determined, including: For the enhanced slice ROI and the corresponding pseudo-target slice ROI within the current batch, the L1 loss and SSIM structural similarity loss are calculated using the following formulas: in, Characterizing L1 loss, To enhance the ROI of the slice (i.e., the ROI of the model-predicted fusion slice); It is a pseudo-target slice ROI (pseudo-target reference slice ROI) ); B stands for Batch Size, which is the batch size of the current batch. Characterizing structural similarity loss, It is the local mean of the enhanced slice ROI (i.e., the local mean of the fused image). It is a pseudo-target slice ROI (i.e., a pseudo-target slice) (local mean) To enhance the local variance of the sliced ​​ROI; The local variance of the ROI (region of interest) in the pseudo-target slice; To enhance the local covariance between the sliced ​​ROI and the pseudo-target slice ROI. It is the smoothing constant ( This refers to the dynamic range of grayscale values. hour, ; hour, ).

[0085] L1 loss measures the absolute difference in pixel values ​​between the enhanced ROI feature map (i.e., enhanced slice ROI) and the pseudo-target ROI feature map (i.e., pseudo-target slice ROI). It is calculated by averaging the absolute values ​​of the differences across all pixel locations. The Structural Similarity Index (SSIM) is built upon this index, measuring the structural similarity between the two feature maps (i.e., enhanced slice ROI and pseudo-target slice ROI) by comparing their local mean, local standard deviation, and local covariance.

[0086] Preferably, the basic reconstruction loss is determined based on L1 loss and SSIM structural similarity loss, including: Characterizing the basic reconstruction loss, This is the weight value.

[0087] In this way, by using L1 loss and SSIM structural similarity loss as the basic losses, we can ensure the pixel-level similarity and structural consistency between the augmented result and the pseudo-target.

[0088] Preferably, based on the discriminator loss between all enhanced slice ROIs and their corresponding pseudo-target slice ROIs in the current batch, the network parameters of the discriminator are adjusted, including: Based on all enhanced slice ROIs and their corresponding pseudo-target slice ROIs within the current batch, the discriminator loss is determined using the following formula: in, Characterize discriminator loss, This is the output of the PatchGAN discriminator.

[0089] Based on the discriminator loss, adjust the discriminator's network parameters.

[0090] In this way, the discriminator no longer judges the authenticity of the entire image, but classifies each local sub-region in the image, thereby capturing the authenticity of local textures more precisely.

[0091] Preferably, the perceptual loss and adversarial loss are determined in the following manner: The perceptual loss employs a VGG16 network pre-trained on ImageNet as a feature extractor. The output feature maps of layers 3, 8, and 15 of this network are selected, and the mean Euclidean distance between the enhanced result (i.e., the enhanced slice ROI) and the pseudo-target (i.e., the pseudo-target slice ROI) at the high-level semantic level is calculated. The perceptual loss does not directly constrain pixel values ​​but guides the generated image to maintain perceptual consistency with the target image in terms of texture, edge, and structural patterns. It also ensures similarity to high-quality images in high-level semantic features, effectively improving the texture realism and semantic consistency of the enhanced result. The formula is as follows: in, Representational perception loss, For pre-training the VGG-16 network Feature map of the layer .

[0092] The adversarial loss is implemented using a spectral normalized PatchGAN discriminator. This discriminator no longer judges the authenticity of the entire image (i.e., the sliced ​​ROI), but instead classifies each local sub-region in the image, thus capturing the authenticity of local textures more precisely. The adversarial loss adopts the Hinge Loss form, as shown in the following formula: in, Characterizing adversarial loss, This is the output of the PatchGAN discriminator.

[0093] In this way, adversarial training drives the enhancement results to approximate the texture realism and edge sharpness characteristics of high-quality radar images.

[0094] Preferably, the network parameters of the original augmentation model are adjusted based on the basic reconstruction loss, perception loss, and adversarial loss, including: determining the total loss value based on the basic reconstruction loss, perception loss, and adversarial loss, using the following formula: in, Characterizing the total loss value, Characterizing the basic reconstruction loss, Representational perception loss, Characterize the adversarial loss. The perceived loss weight is 0.01 when enabled. To counteract the loss weight (0.01 when enabled).

[0095] Adjust the network parameters of the original augmentation model based on the total loss value.

[0096] Combination Figure 9 As shown, Figure 9 A schematic diagram of the loss value calculation process is provided.

[0097] Preferably, a method for cross-slice ROI enhancement for ground-penetrating radar further includes: Obtain the test sample set; each sample in the test sample set includes a test slice ROI group and a test pseudo-target slice ROI; Input each test slice ROI group into the target augmentation model to obtain the corresponding test augmentation slice ROI; Based on all test enhanced slice ROIs and corresponding test pseudo-target slice ROIs, quantitative metrics are calculated; among them, quantitative metrics include at least one of peak signal-to-noise ratio, structural similarity index, spatial frequency, and gradient energy. The enhancement effect of the target enhancement model is determined based on quantitative indicators.

[0098] In this way, by acquiring a test sample set and calculating the quantitative indicators of the test sample set, it is possible to quantify the effect of the target enhancement model on the cross-slice ROI enhancement of ground penetrating radar.

[0099] PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), SF (Spatial Frequency), and Tenengrad (Tenengrad Gradient Energy) are used as quantitative indicators to measure the effect of the model on the cross-slice enhancement of the B-scan adhesive echo in the insulation layer.

[0100] For ease of understanding, this disclosure illustrates the calculation process for the quantitative indicators between a single test slice ROI group and its corresponding test pseudo-target slice ROI. For the entire test sample set, the quantitative indicators between all individual test slice ROI groups and their corresponding test pseudo-target slice ROIs are summed or averaged to obtain the quantitative indicators corresponding to the test sample set.

[0101] Peak Signal-to-Noise Ratio (PSNR) measures the fidelity and distortion at the pixel level between the enhanced ROI image (i.e., the enhanced slice ROI) and the reference image (i.e., the pseudo-target slice ROI), measured in decibels (dB). It quantifies the noise level by calculating the mean square error (MSE) between the two images. A higher PSNR value indicates a smaller difference between the enhanced result and the high-quality reference image, lower image distortion, a higher signal-to-noise ratio, and better restoration of weak feature hyperbolic curves.

[0102] in, This represents the maximum dynamic range of image pixel values ​​(in this method, due to normalization, the value is set to 1). For the enhanced image The mean squared error (MSE) between the enhanced ROI and the pseudo-target ROI is calculated using the following formula: The Structural Similarity Index (SSIM) measures the overall similarity between the enhanced image (i.e., the enhanced region of interest, ROI) and the reference image (i.e., the pseudo-target region of interest, ROI) in terms of brightness, contrast, and structure. It focuses on the fidelity of structural features such as hyperbolic texture and edge contours. Unlike PSNR, SSIM better aligns with the human visual system's perception of image quality. Its value range is [0, 1]. The closer the value is to 1, the smaller the structural difference between the enhanced image and the pseudo-target, the stronger the ability of the fused image to preserve the original signal structure, and the better the realism of the hyperbolic texture.

[0103] in, and Representing the fused images (i.e., enhanced slice ROI) and pseudo-target slice ROI The local mean (brightness); and These represent their local standard deviations (contrast); This represents the local covariance (structural correlation) between the two. Small positive numbers are used to maintain computational stability.

[0104] Spatial Frequency (SF) reflects the overall activity and detail richness of a fused image (i.e., an enhanced region of interest). It is evaluated by calculating the gradient changes in the row (RF) and column (CF) directions of the image. A higher SF value indicates that the image contains more high-frequency detail information, and that subtle hyperbolic textures (such as echo intensity variations and local morphological differences) are preserved more completely. Its calculation formula is: The formulas for calculating RF (Row Frequency) and CF (Column Frequency) are as follows: Tenengrad gradient energy is a no-reference metric used to evaluate image sharpness and edge strength. It measures the edge sharpness and detail richness of an enhanced ROI image (i.e., an enhanced slice ROI), directly related to the sharpness of hyperbolic edges and the prominence of subtle textures. It calculates the sum of squared gradient magnitudes for each pixel in the image based on the Sobel operator. A higher Tenengrad value indicates sharper edges, more prominent hyperbolic details, stronger discernibility of weak features, and better focus quality. This metric is also the core measure for selecting pseudo-targets during self-supervised training in this embodiment.

[0105] in, and The pixel coordinates of the enhanced image (i.e., the enhanced slice ROI) are respectively. The horizontal and vertical gradient values ​​are used to capture the intensity of texture variations in hyperbolic edges and details.

[0106] in, The horizontal Sobel convolution kernel is used to detect vertical edges (the longitudinal contour of the GPR hyperbola). A vertical Sobel convolution kernel is used to detect horizontal edges (lateral details of the GPR hyperbola): In one embodiment, to verify the effectiveness of the proposed method for cross-slice ROI enhancement using ground-penetrating radar, this method is applied to the processing of data collected in a laboratory setting of an external insulation layer scenario and data detected in a real-world building exterior wall insulation layer scenario. In the detection of the insulation layer scenario, a stepped-frequency ground-penetrating radar is used to detect the wall, and the collected data is processed using the method of this invention.

[0107] Combination Figure 10 As shown, Figure 10 This is a visualization of the training process of a thermal insulation layer detection data input model. Here, 'ref' represents the selection of the clearest / most informative pseudo-target slice (ROI) using Tenengrad energy as a supervision reference; 'slice' represents random degradation (blurring, noise, contrast perturbation) of the remaining slices excluding pseudo-target slices; and 'fused' represents the slice result image after the model has learned and aggregated the information.

[0108] Combination Figure 11 As shown, Figure 11 This is a graph showing the results of using a target augmentation model to infer new data.

[0109] Combination Figure 12 As shown, Figure 12 This is a schematic diagram of a Tenengrad self-supervised ablation result. It is a result of a method for enhancing ROI across slices of ground penetrating radar (GPR), in which Tenengrad self-supervised ablation (-Tenengrad) replaces the Tenengrad energy criterion with a simple image variance, and then trains and enhances the GPR B-scan data.

[0110] Combination Figure 13 As shown, Figure 13 This is a schematic diagram of a deformable alignment ablation (-Deformable). It is a result of a method for enhancing ground penetrating radar (GPR) cross-slice ROIs by removing deformable offsets, keeping the offsets constant at zero, and then training and enhancing GPR B-scan data.

[0111] Combination Figure 14 As shown, Figure 14 This is a schematic diagram of uncertainty-guided ablation (-Uncertainty). It illustrates the result of removing the confidence prediction module and then retraining and enhancing the ground-penetrating radar B-scan data in a method for cross-slice ROI enhancement using ground-penetrating radar. Combination Figure 15 As shown, Figure 15This is a schematic diagram of a multi-scale FPN fusion ablation (-FPN) result. It is a method for enhancing the ROI across slices of ground penetrating radar, in which only the deepest fused3 features are directly fed into the decoder, lateral connections such as lateral1 and lateral2 are removed, and feature pyramid addition operations are performed before training and enhancing the ground penetrating radar B-scan data.

[0112] Combination Figure 16 As shown, Figure 16 This is a schematic diagram of an adversarial training ablation (-GAN) result. It is a method for enhancing ground penetrating radar (GPR) across slice ROIs by completely removing the discriminator and adversarial loss, training only with L1+SSIM, and then training and enhancing the GPR B-scan data.

[0113] Combination Figure 17 As shown, Figure 17 This is a schematic diagram of one embodiment. Specifically, it shows the result of visualizing B-scan data of the actual scene of enhancing the thermal insulation layer of a building exterior wall in the field of ground penetrating radar during the training process of a method for enhancing cross-slice ROI of ground penetrating radar. Combination Figure 18 As shown, Figure 18 This is a schematic diagram of another embodiment. Specifically, it is a result diagram of B-scan data of a real-world scenario of a ground-penetrating radar (GPR) external wall insulation layer during the training process of a method for GPR cross-slice ROI enhancement.

[0114] As shown in Table 1, Table 1 presents the ablation experiment results of different algorithm modules combined for building exterior wall insulation layer data, as follows: Through experimental comparison, the proposed cross-slice ROI enhancement method for ground-penetrating radar exhibits significant weak feature enhancement and texture fidelity in both laboratory and real-world outdoor scenarios for detecting the bonding status of building exterior wall insulation layers. In the B-scan image visualization results, during laboratory scene data processing, as shown... Figure 10 and Figure 11 As shown, the ROI image enhanced by the method of this invention exhibits realistic and coherent hyperbolic texture, significantly enhanced weak feature echoes, and effective suppression of background interference; in the field measurement data, as... Figure 17 and Figure 18 As shown, even in the face of signal attenuation and interference in complex environments, this method can still accurately align the micro-displacements of adjacent slices, aggregate the high-confidence features of each slice, make the originally blurry hyperbolic edges clear, and completely preserve the subtle texture details (such as echo intensity fluctuations).

[0115] The ablation experiment results further validated the indispensability of each mechanism, such as Figures 12 to 16 As shown, each mechanism plays a unique and crucial role, and the absence of any module will lead to a significant decrease in enhancement effect: The Tenengrad self-supervised mechanism (i.e., acquiring the training sample set) is the core of the model to obtain reliable supervision signals. It accurately captures hyperbola edge and texture information through Sobel gradient energy, and can filter out truly clear pseudo-target slice ROIs. If replaced by traditional variance slice selection, it will be unable to distinguish between texture and noise, resulting in a significant decrease in the stability of pseudo-target selection. The enhancement result will have jagged edge artifacts, and the transition between layered echoes will also become unnatural. Its Tenengrad value is only 8.0138, far lower than the 9.1069 of the complete model, which fully demonstrates the superiority of this module in terms of sharpness measurement; The deformable alignment mechanism is the key to solving cross-slice micro-displacement. It achieves accurate alignment by predicting sub-pixel level offset fields. If this module is removed, the small misalignment between slices cannot be corrected, and obvious misalignment blur will appear after hyperbola enhancement, which not only affects the visual effect. The lack of coherence also led to a drop in PSNR to 19.6455 and SF to only 0.3838, highlighting signal distortion and loss of detail. The uncertainty guidance mechanism, which intelligently filters local information, automatically reduces the weight of noisy and blurred regions. If removed, the model cannot identify unreliable information in low-quality slices, and clutter will remain and cover part of the hyperbolic contour, resulting in an SSIM of only 0.7108 and a significant decrease in structural consistency. The FPN multi-scale fusion mechanism is the core that balances detail and semantics, integrating shallow texture details and deep global semantics. If removed, the two features cannot be optimized synergistically, resulting in a smoother enhancement and loss of many subtle texture details. The adversarial training mechanism is key to improving visual realism. It guides the model to generate textures and edges that are closer to high-quality radar images through the PatchGAN discriminator. If removed, the enhanced texture lacks realism, and the edge sharpness is insufficient, making it difficult to achieve the visual effect required for engineering applications. Only by fully combining the four modules of Tenengrad self-supervision, deformable cross-slice attention, FPN multi-scale fusion, and progressive adversarial training can the optimal enhancement effect be achieved.

[0116] Quantitative results show that the method of this invention outperforms the ablation group in all evaluation metrics: the PSNR of the complete model reaches 20.4461dB, higher than all ablation groups (-Tenengrad: 20.1250dB, -Deformable: 19.6455dB, etc.), indicating higher pixel-level fidelity and less signal distortion; the SSIM is 0.7284, closer to 1, indicating stronger consistency of the hyperbolic structure (it should be noted that the SSIM value is not close to 1 because the enhancement process aims to amplify and enhance weak signals and details on the pseudo-target slices, so the output image and the original pseudo-target will inevitably differ in pixel distribution. However, the higher SSIM value clearly indicates that the enhancement is carried out on the basis of preserving the original hyperbolic texture structure to the greatest extent, rather than introducing distortion or changing its inherent pattern, which is in line with the engineering goal of "enhancement" rather than "distortion"); the SF reaches 0.4347, higher than other groups, reflecting better detail richness; Tenengrad The gradient energy was 9.1069, significantly higher than that of the ablation group, demonstrating stronger edge sharpness and greater discernibility of weak features. These quantitative results fully illustrate that the method of this invention can effectively balance noise suppression and texture fidelity while enhancing weak features and preserving details, significantly improving image quality.

[0117] The superior performance of this invention is primarily attributed to the synergistic effect of four mechanisms: the Tenengrad self-supervised mechanism accurately filters out false targets, providing reliable supervision signals for model training and addressing the engineering challenge of lacking clear annotations; the deformable cross-slice attention mechanism simultaneously achieves slice-level global weight evaluation, pixel-level uncertainty weighting, and sub-pixel-level alignment, intelligently aggregating high-confidence features and resolving the blindness issues in misalignment correction and local information selection; the FPN multi-scale fusion strategy integrates shallow textures and deep semantics, balancing detail integrity and texture realism; and progressive adversarial training, through an optimization logic of "first ensuring fidelity, then improving quality," enhances texture realism and edge clarity while maintaining pixel-level consistency. These four modules form a complete chain of "data-driven - intelligent alignment - multi-scale aggregation - quality optimization," making the method both engineering interpretable and robust, providing high-quality data support for subsequent feature extraction and adhesive state classification.

[0118] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A method for ground penetrating radar cross-slice ROI enhancement, characterized in that, include: Step S11: Obtain the first multi-slice group of the ground penetrating radar; wherein, the first multi-slice group includes multiple spatially adjacent slices; Step S12: Align and crop the hyperbolic target regions of each slice in the first multi-slice group to obtain the first multi-slice ROI group; wherein, the first multi-slice ROI group includes multiple first slice ROIs, and the first slice ROI represents the ROI of the slice in the first multi-slice group. Step S13: Input the first multi-slice ROI group into the pre-trained target augmentation model for processing to obtain the target augmented slice ROI; wherein, the target augmentation model is used to encode the first multi-slice ROI group with shared weights at multiple scales, perform cross-slice attention fusion and multi-scale aggregation processing at each scale, and output the target augmented slice ROI.

2. The method of claim 1, wherein, The target augmentation model is obtained through the following methods: Step S21: Obtain the training sample set; wherein each sample in the training sample set includes a second multi-slice ROI group and a pseudo-target slice ROI, and the second multi-slice ROI group includes multiple second slice ROIs; Step S22: Train the original augmentation model using the training sample set. The original augmentation model includes a shared-weight multi-scale convolutional encoder, a deformable cross-slice attention module, a multi-scale aggregation module, and a decoder; the training process is as follows: Step S221: For each second multi-slice ROI group, the shared weight multi-scale convolutional encoder is used to process each second slice ROI to obtain the multi-level feature map corresponding to each second slice ROI, so as to form a multi-level feature map set; wherein, the multi-level feature map includes feature maps of multiple levels. Step S222: For each multi-level feature map set, the deformable cross-slice attention module is used to fuse all feature maps of the same level to obtain the fused feature map corresponding to each level, so as to form a multi-level fused feature map; wherein, the multi-level fused feature map includes fused feature maps of multiple levels. Step S223: For each multi-level fusion feature map, the fusion feature maps corresponding to each level are aggregated using the multi-scale aggregation module to obtain a multi-scale aggregated feature map, and the multi-scale aggregated feature map is decoded using the decoder to obtain the enhanced slice ROI. In step S224, based on the loss between each augmented slice ROI and the corresponding pseudo-target slice ROI, adjust the network parameters of the original augmentation model, and repeat steps S221 to S224 until training is completed and the target augmentation model is obtained.

3. The method of claim 2, wherein, The acquisition of the training sample set includes: Obtain multiple third-level multi-slice ROI groups; where each third-level multi-slice ROI group includes multiple third-level ROIs. For each third multi-slice ROI group, the third slice ROI with the highest Tenengrad value is taken as the pseudo-target slice ROI, and random degradation is applied to the other third slice ROIs except for the pseudo-target slice ROI to obtain the second multi-slice ROI group. A second multi-slice ROI group and its corresponding pseudo-target ROI are used as a sample to obtain the training sample set.

4. The method of claim 2, wherein, The method of fusing all feature maps of the same level using a deformable cross-slice attention module to obtain fused feature maps corresponding to each level includes: Perform the following steps using the deformable cross-slice attention module: For each level, the slice-level global attention weights, pixel-level confidence maps, and deformable offset fields of each feature map are calculated, as well as... Based on the global attention weights at each slice level and the confidence maps at each pixel level, the fusion weight matrix for each feature map is determined, and Based on each deformable migration field, spatial position correction is performed on each feature map to obtain each corrected feature map, and Based on each fusion weight matrix, all corrected feature maps are weighted and summed to obtain the fusion feature map corresponding to the current level.

5. The method of claim 2, wherein, The process of aggregating the fused feature maps corresponding to each level using a multi-scale aggregation module to obtain a multi-scale aggregated feature map includes: Perform the following steps using the multi-scale aggregation module: Channel mapping and downsampling are performed on each low-level fusion feature map to obtain each processed fusion feature map. Among them, the low-level fusion feature map represents the fusion feature map in the multi-level fusion feature map except for the high-level fusion feature map, and the high-level fusion feature map represents the fusion feature map with the most channels in the multi-level fusion feature map. The number of channels and spatial size of the processed fusion feature map are the same as those of the high-level fusion feature map. The aggregated feature map is obtained by adding the individual processed fusion feature maps and the high-level fusion feature map element by element. Add positional encoding to the aggregated feature map to obtain a multi-scale aggregated feature map.

6. The method according to any one of claims 2 to 5, characterized in that, The process of adjusting the network parameters of the original augmentation model based on the loss between each augmented slice ROI and its corresponding pseudo-target slice ROI includes: If the current iteration round is less than the first preset round, then the network parameters of the original augmentation model are adjusted based on the basic reconstruction loss between all augmented slice ROIs and the corresponding pseudo-target slice ROIs in the current batch, on a batch basis. If the first preset round ≤ the current iteration round < the second preset round, then fix the network parameters of the original augmentation model, and adjust the network parameters of the discriminator based on the discriminator loss between all augmented slice ROIs and the corresponding pseudo-target slice ROIs in the current batch, in batches. If the current iteration round is greater than the second preset round, then on a batch basis, based on all augmented slice ROIs and corresponding pseudo-target slice ROIs in the current batch, determine the basic reconstruction loss, perception loss, adversarial loss and discriminator loss, and adjust the network parameters of the original augmentation model based on the basic reconstruction loss, perception loss and adversarial loss, and adjust the network parameters of the discriminator based on the discriminator loss.

7. The method according to claim 6, characterized in that, The adjustment of the network parameters of the original augmentation model based on the fundamental reconstruction loss between all augmented slice ROIs and their corresponding pseudo-target slice ROIs in the current batch includes: Based on all enhanced slice ROIs and corresponding pseudo-target slice ROIs in the current batch, determine the L1 loss and SSIM structural similarity loss; Based on L1 loss and SSIM structural similarity loss, the basic reconstruction loss is determined; Based on the basic reconstruction loss, the network parameters of the original enhancement model are adjusted.

8. The method according to claim 2, characterized in that, Also includes: Obtain the test sample set; each sample in the test sample set includes a test slice ROI group and a test pseudo-target slice ROI; Input each test slice ROI group into the target augmentation model to obtain the corresponding test augmentation slice ROI; Based on all test enhanced slice ROIs and corresponding test pseudo-target slice ROIs, quantitative metrics are calculated; among them, quantitative metrics include at least one of peak signal-to-noise ratio, structural similarity index, spatial frequency, and gradient energy. The enhancement effect of the target enhancement model is determined based on quantitative indicators.