Complex environment adaptive training method for three-dimensional ground penetrating radar underground disease recognition model
Through the pseudo-label supervision learning method optimized by the three-dimensional codec network and self-supervising module, the disease detection problem of three-dimensional ground penetrating radar in complex environments is solved, and efficient and accurate underground disease recognition is achieved.
Patent Information
- Application Number
- CN202510759586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Due to the limitation of two-dimensional data processing or lack of three-dimensional spatial modeling capabilities, the prior art is difficult to adapt to the challenges of complex environments such as parameter differences in multi-device and geological noise interference in underground disease detection.
A three-dimensional codec network is used to combine self-supervised modules and pseudo-tagged supervision learning, and a hybrid domain data set is constructed for training through InfoNCE loss and MC Dropout optimization model to improve the generalization ability of the model in complex environments.
It significantly improves the accuracy and robustness of underground disease recognition of three-dimensional ground penetrating radar, reduces the computational complexity, and adapts to the cross-domain detection needs of different geological conditions and equipment parameters.
Smart Images

Figure CN120279543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of specific computer models, and particularly to a complex environment adaptive training method for a three-dimensional ground penetrating radar underground disease identification model. Background Art
[0002] Hidden structural defects and medium anomalies caused by natural geological actions or human engineering activities in urban underground spaces (such as structural layers and surrounding soils of roads, tunnels, pipelines, foundations, etc.) are characterized by the concealment of their positions and the danger of evolution. With the continuous increase in the development intensity of urban underground spaces, the original stable strata under roads are gradually structurally damaged or the properties of the medium deteriorate under the combined influence of multiple factors such as the repeated action of vehicle loads, the vibration disturbance of underground projects, and the dynamic changes in groundwater flow (such as pipeline leakage and rain erosion), forming typical diseases including cavities (substance-free areas formed by soil loss), voids (separation between structural layers), loose bodies (loose soils with insufficient compaction, divided into severely loose and generally loose), water-rich bodies (seepage areas saturated with water), and fracture networks (fracture zones in rock and soil bodies or structures). These diseases often have no obvious surface signs in the initial stage, but will continue to expand through stress redistribution (such as the longitudinal extension of cracks and the formation of honeycomb erosion channels in cavities), resulting in the gradual loss of the bearing capacity of rock and soil bodies; when the diseases develop upward and break through the critical state, they may trigger sudden disasters such as road collapses, tunnel deformations, pipeline bursts, and building foundation instabilities, posing a serious threat to urban public safety.
[0003] The three-dimensional ground penetrating radar technology is a non-destructive detection method that obtains three-dimensional information of underground structures through electromagnetic wave reflection signals and is widely used in underground disease detection. Traditional detection methods rely on manual experience to analyze radar images, resulting in low efficiency and high false detection rates. In recent years, the combination of deep learning and three-dimensional ground penetrating radar has provided new ideas for intelligent disease identification. Deep learning can automatically extract multi-scale features, learn complex feature expressions and potential laws through data training, and can accurately locate the morphology and distribution of diseases from complex radar data, significantly improving the detection efficiency and accuracy. In addition, the three-dimensional ground penetrating radar data has the characteristics of spatial continuity and high resolution, can accurately obtain detailed information such as the morphology, size, and depth of underground diseases, and combined with deep learning models (such as the encoder-decoder architecture), can realize three-dimensional analysis of diseases, meet the cross-domain generalization requirements of different geological conditions and equipment parameters, and provide reliable support for disease detection in complex environments.
[0004] However, in practical applications, different geological conditions, differences in radar equipment parameters, and environmental interference can cause data distribution deviations, seriously affecting the model's ability to generalize across scenarios. Domain adaptation is a transfer learning technology that aims to solve the problem of inconsistent data distribution between the source domain (training data) and the target domain (test data). Domain adaptation adjusts the model or data representation so that the knowledge learned in the source domain can be effectively transferred to the target domain, thereby improving the model's generalization ability.
[0005] As mentioned in the paper "Apple Leaf Disease Recognition Based on Self-Supervised Domain Adaptation Network", for the model of apple leaf disease recognition, domain adaptation methods and self-supervision modules are introduced, and joint training of source and target domain data sets is used to reduce domain bias, enhance the generalization ability of the pre-trained model, and enhance the model's ability to distinguish similar diseases by shortening the feature distance of positive and negative samples through contrast loss, so that the model can learn more detailed representation information of the diseased area, improve classification performance, and enhance the generalization of the model in real agricultural scenarios. However, this method is mainly aimed at two-dimensional image data and lacks the ability to model continuous features in three-dimensional space. Underground diseases (such as the longitudinal extension of cracks and the honeycomb structure of hollows) need to be analyzed stereoscopically through three-dimensional radar data, while the self-supervision module of this paper only focuses on two-dimensional transformations such as image rotation and brightness enhancement, and cannot effectively capture long-range dependencies in three-dimensional space. Secondly, in the source domain training stage, only cross entropy loss is used for optimization, and no dynamic adjustment is made to the problem of class imbalance, resulting in the model's limited ability to recognize sparse diseases (such as hollows and cracks). In addition, its data mixing strategy does not combine the uncertainty of the target domain data with the local noise assessment. For example, it does not introduce a three-dimensional neighborhood confidence screening mechanism, which leads to insufficient reliability of the pseudo-labels generated under complex geological interference (such as metal pipeline reflection artifacts) and is prone to false detection due to noise propagation.
[0006] As the paper "Hierarchical Attention-based Domain Adaptive Remote Sensing Image Segmentation" focuses on the semantic segmentation task of remote sensing images, it adopts a hierarchical multi-head self-attention mechanism and multi-scale feature fusion, combines a self-training domain adaptive framework, and optimizes cross-domain knowledge transfer through local pseudo-label quality screening (LPQ) and dynamic class weights (DCW), significantly improving the accuracy and robustness of remote sensing image segmentation. Although hierarchical attention and multi-scale fusion are introduced, its domain adaptive framework is designed based on two-dimensional images and is difficult to directly transfer to three-dimensional radar data. For example, the local pseudo-label quality (LPQ) algorithm uses two-dimensional neighborhood convolution to evaluate confidence, ignoring the spatial continuity between three-dimensional voxels, resulting in large errors in pseudo-label screening in boundary fuzzy regions (such as the transition zone between cracks and normal media). At the same time, its self-training strategy relies on two-dimensional sliding window slicing, with low efficiency in extracting longitudinal stratification features of three-dimensional radar data and being unable to effectively handle high-dimensional computational complexity problems, restricting real-time detection capabilities. In addition, its data mixing strategy also does not fully consider the target domain data, resulting in pseudo-labels being vulnerable to false signals under complex geological interference, further reducing the stability of cross-domain knowledge transfer.
[0007] Therefore, due to the limitation of two-dimensional data processing or the lack of three-dimensional space modeling ability in the existing technology, it is difficult to adapt to the challenges of complex environments such as multi-device parameter differences and geological noise interference in underground disease detection.
[0008] Glossary: InfoNCE (Information Noise Contrastive Estimation) is a loss function for contrastive learning, used to maximize the mutual information between positive samples (matching samples) while pulling negative samples (non-matching samples) apart.
[0009] The three-dimensional neighborhood confidence calculation formula of the LPQ (Local Phase Quantization) function is as follows: , This formula means that in three-dimensional space, within the neighborhood centered at position (d, h, w), the phase quantization value Q at each position is averaged to obtain the LPQ value at that position, and K is the neighborhood radius.
[0010] Monte Carlo Dropout (MC Dropout for short) is a deep learning inference method that combines the Monte Carlo method and Dropout regularization technology, mainly used to quantify the prediction uncertainty of neural network models. The key innovation of Monte Carlo Dropout is that it still keeps the Dropout activation during the test phase and performs multiple forward propagations on the same input. Each time during forward propagation, different neurons are randomly dropped, resulting in different outputs. Through the results of multiple forward propagations, the prediction distribution of the model can be estimated, thereby quantifying the prediction uncertainty. Its workflow is as follows: (1) Enable Dropout for inference: Keep the Dropout layer activated during the test phase, and randomly drop some neurons during each forward propagation; (2) Perform multiple forward propagations on each input sample to obtain multiple prediction results; (3) Statistical analysis: By analyzing the outputs of multiple forward propagations, calculate the mean and variance of the prediction results. The mean represents the prediction of the model, and the variance represents the uncertainty of the model about this prediction. Summary of the Invention
[0011] The object of the present invention is to provide a complex environment adaptive training method for a three-dimensional ground penetrating radar underground disease recognition model, which solves the problems that the prior art is difficult to adapt to complex environments such as differences in multiple device parameters and geological noise interference in underground disease detection due to being limited to two-dimensional data processing or lacking three-dimensional space modeling capabilities.
[0012] To achieve the above object, the technical solution adopted by the present invention is as follows: A complex environment adaptive training method for a three-dimensional ground penetrating radar underground disease recognition model, including the following steps; S1, Obtain several samples with category annotation maps to form a source domain dataset D s , and several unannotated samples to form a target domain dataset D t , where the samples are three-dimensional ground penetrating radar data and the category is the disease category; D s For a sample X s in it, the category annotation map is P s , D t For a sample in it is X t , X s ∈R D×H×W×Cin , X t ∈R D×H×W×Cin , P s ∈R D×H×W×C , D, H, and W are the depth, length, and width of the sample resolution respectively, Cin is the number of channels of the radar reflection intensity, Cin = 1, C is the total number of disease categories, and the value at the position (d, h, w) in P s is X sThe true disease category at the position (d, h, w), where d, h, and w are the depth, length, and width of the position (d, h, w) respectively; S2. Construct a 3D encoding and decoding network, including a 3D encoder, a 3D decoder, and a self-supervised module; The 3D encoder is used to encode the input samples to obtain encoded features, and the 3D decoder is used to decode the encoded features and then perform boundary enhancement, and output the predicted class probabilities at each position to form a probability prediction map The self-supervised module is used to construct positive and negative samples based on the encoded features of a batch of samples, and calculate the InfoNCE loss; S3. Use D S To minimize the total source domain training loss L s Train the 3D encoding and decoding network to obtain a basic model, L s = λ cn1 L cn1 + λ cls L cls L, cn1 The InfoNCE loss, L cls Is the classification loss, λ cn1 λ, cls Are the weights of L cn1 And L cls Respectively; S4. Duplicate two basic models and use them as the teacher model and the student model respectively; S5. Use D t Train the student model, and optimize the student model based on self-supervised contrast learning and double-screening pseudo-label supervised learning in each iteration, where one optimization includes steps S51~S54; S51. Use the teacher model to output the probability prediction map of each sample in D t As the pseudo-label, the pseudo-label of X t Is Set the quality threshold λ and the uncertainty threshold ϵ; S52. Select a batch of samples from D t Among them, the b-th sample X b Use the student model to output the probability prediction map of each sample, and use the self-supervised module to calculate the InfoNCE loss; S53. Calculate the pseudo-label supervised loss L b Of X p,b ; For the pseudo-label b Of X Based on local phase quantization, calculate the LPQ value LPQ(d, h, w) of the position (d, h, w), for X b, calculate the uncertainty U(d, h, w) of the position (d, h, w) based on MC Dropout, and form X with the positions where LPQ(d, h, w) > λ and U(d, h, w) < ϵ b Region of interest , calculate L p,b ; , In the formula, is the total number of positions within, is the normalized value of U(d, h, w), is X b 's probability prediction map value at the position (d, h, w); S54, calculate the total loss L of the target domain t , and minimize L t to adjust the parameters of the student model; , In the formula, B is the batch size, and L cn2 is the InfoNCE loss of the self-supervised module in the student model, and λ cn2 is the weight of L cn2 ; S6, use D s and D t to construct a mixed-domain dataset with multiple batches; S7, train the student model with the mixed-domain dataset to minimize the joint loss L total to update the student model, and update the teacher model using the EMA method, where L total = + L s + λ t L t + λ c L consist , and L consist is the consistency loss of a batch of samples.
[0013] Preferably, the disease categories include background, cavity, void, loose body, water-rich body, and fracture network.
[0014] Preferably, the method for the self-supervised module to calculate the InfoNCE loss is; Obtain the encoded features of a batch of samples, sequentially labeled as F1~F B , calculate the InfoNCE loss of the b-th encoded feature and the InfoNCE loss L of a batch of samples cn1 ; , , In the formula, is to randomly add data augmentation to F b to obtain positive samples, is the v-th negative sample of F b , 1 ≤ v ≤ 2B - 2, sim(⋅) is to calculate the cosine similarity, τ is the temperature coefficient, and exp(⋅) is the exp function.
[0015] Preferably, in S3, the classification loss is calculated according to the following formula; , In the formula, α c is the weight of the c-th disease category, 1 ≤ c ≤ C, γ is the modulation factor for reducing the weight of easily classified samples. For the position (d, h, w) of the b-th sample, , are respectively the normalized value of the predicted class probability at this position and the true disease class annotation.
[0016] Preferably, LPQ(d, h, w) in S53 is calculated according to the following formula: , In the formula, K is the neighborhood size, i, j, and k are the offsets of the position (d, h, w) in the depth, length, and width directions respectively, and Q(d + i, h + j, w + k) is the pseudo label at the position (d + i, h + j, w + k).
[0017] Preferably, the method for calculating U(d, h, w) in S53 is: Use the MC Dropout method to perform M forward propagations on X t , and generate a probability prediction map for each forward propagation. Calculate the uncertainty U(d, h, w) at the position (d, h, w) according to the following formula; , In the formula, is the value of the probability prediction map obtained from the m-th forward propagation at the position (d, h, w), is the mean value of the M probability prediction maps at the position (d, h, w).
[0018] Preferably, in S6, the method for constructing a batch of mixed-domain datasets is: Mix D s and D t , and then randomly select a batch of samples, or select a batch of samples from D s and D t according to the mixing weights. The mixing weights are preset or obtained according to steps S61~S63; S61, from Ds , D t Any batch of samples in each respectively constitute a set , ; S62, calculate the LPQ value and uncertainty at each position of each sample in, and the average of all LPQ values is used to obtain the batch-average local quality score , and the average of all uncertainties is used to obtain the batch-average uncertainty ; S63, calculate the mixing weight of ; , α and β are respectively and the adjustment coefficients of, and Sigmoid(⋅) is the Sigmoid function; According to , generate a batch of mixed-domain datasets .
[0019] Preferably, in S7, L consist is calculated according to the following formula: , , In the formula, is the square of the L2 norm. For the b-th sample in the mixed domain, is the probability prediction map output by the student model, is the pseudo label generated by the teacher model, is the consistency loss of the b-th sample.
[0020] Preferably, in the 3D encoding and decoding network, the 3D encoder includes a fixed chunk embedding module, a first multi-scale sampling module, and a second multi-scale sampling module connected in sequence; The fixed chunk embedding module includes a 3D convolutional layer and a 3D position encoding layer; The 3D convolutional layer is used to perform 3D convolutional dimensionality reduction on the input samples to obtain the dimensionality-reduced feature F cn , where the convolutional kernel of the 3D convolution is 3×3×3 and the stride is 2; The 3D position encoding layer is used to generate 3D sine position encoding for each position of F cn to obtain the position encoding feature F p , and then add F cn and F p element-wise to obtain the embedding feature F0. The scale of the sample is D×H×W×C in , F cn , Fp , the scale of F0 is ; The first multi-scale sampling module includes a global attention layer, a local convolutional layer, a layer normalization layer, and a three-dimensional downsampling layer; The multi-scale sampling module is used to perform global attention on the input features to obtain the global feature F global ; The local convolutional layer is used to perform depthwise separable three-dimensional convolution on F global to obtain the local feature F local , the convolution kernel is 3×3×3, and the stride is 1; The layer normalization layer is used to obtain the mixed feature F mix according to the formula F global = LayerNorm(F local + F mix ); The three-dimensional downsampling layer is used to perform three-dimensional convolutional downsampling on F mix to obtain the output feature F d1 ; The second multi-scale sampling module has the same structure as the first multi-scale sampling module, the input feature is F d1 , the output scale is of the output feature F d2 , and F d2 is the encoded feature of the three-dimensional encoder.
[0021] Preferably, in the three-dimensional encoding and decoding network, the three-dimensional decoder includes a first upsampling module, a second upsampling module, a third upsampling module, and a boundary enhancement layer connected in sequence; The first upsampling module includes a transposed convolutional upsampling layer, a cross-layer splicing layer, a dilated residual module, and a residual connection activation layer; The transposed convolutional upsampling layer is used to perform transposed three-dimensional convolution on the encoded feature to obtain the first feature F u1 , the convolution kernel of the transposed three-dimensional convolution is 2×2×2, and the stride is 2; the cross-layer splicing layer is used to splice F u1 and the cross-layer feature along the channel dimension to obtain the second feature F u2 , where the cross-layer feature is F d1 ; The dilated residual module performs dilated convolution with dilation rates of 2 and 4 on F u2 respectively, and then adds them to obtain the third feature F u3 ; The residual connection activation layer is used to combine F u2 and F u3Perform residual connection and then activate it through the GELU function to obtain the output F of the first upsampling module uout1 ; Duplicate a first upsampling module, and add a channel compression layer at the input end to compress the channels of the input features by half and then send them into the transposed convolutional upsampling layer, modify the cross-layer features to F0, and modify the dilation rate of the dilated residual module to 1 and 2 to obtain a second upsampling module; Duplicate a second upsampling module and modify the cross-layer features to feature F 01 ; the dilated residual module only performs dilated convolution with a dilation rate of 1 on the input features to obtain a third upsampling module, and feature F 01 is obtained by transposed convolutional upsampling of F0, with a convolutional kernel of 2×2×2 and a stride of 2; The classifier is used to classify F uout3 position by position to obtain the initial predicted class probability P of each position base ; The boundary enhancement layer performs edge detection on F uout3 using a three-dimensional Sobel operator, outputs a boundary response map P Sobel , and then generates a probability prediction map according to the formula , where λ is the weight of P Sobel . Sobel
[0022] In the present invention, the types of underground disease bodies mainly include cavities (substance-free areas formed by soil loss), voids (detachment between structural layers), loose bodies (loose soils with insufficient compaction degree, divided into severely loose and generally loose), water-rich bodies (seepage areas saturated with water), and fracture networks (fracture zones in rock and soil bodies or structures). Of course, it is not limited to these several types, and can be defined and marked according to actual needs.
[0023] The present invention proposes a complex environment adaptive training method for a three-dimensional ground penetrating radar underground disease identification model aiming at the problem of the decline in the model generalization performance caused by the cross-domain data distribution differences (such as different geological conditions and equipment parameters) of the ground penetrating radar. This method is mainly trained in three stages to achieve cross-domain knowledge transfer, namely source domain training, target domain training, and joint optimization.
[0024] The first stage: source domain training, which combines supervised learning and self-supervised contrast learning to enhance the model's generalization ability to sparse diseases, and uses Focal Loss to alleviate the class imbalance problem in the data.
[0025] The supervised learning in the first stage refers to the process of training a three-dimensional encoding and decoding network with the source domain dataset, inputting the samples into the model, and outputting a probability prediction map to calculate the classification loss L cls , in L cls In the formula, to address the class imbalance problem, Focal Loss is adopted to replace the cross-entropy loss. By dynamically adjusting the weights of easy and hard samples, the classification accuracy of sparse classes is improved, significantly alleviating the model bias problem caused by class imbalance. Generally, set such that is inversely proportional to the class frequency of class c, which is used to reduce the weight of easy-to-classify samples and focus on hard samples (such as sparse diseases).
[0026] The first-stage self-supervised contrast learning refers to the process of using a self-supervised module to construct positive and negative samples based on the InfoNCE method according to a batch of source domain samples input, and calculating the InfoNCE loss through contrast learning. This is because ground-penetrating radar data will have domain shift problems due to differences in acquisition environments. For example, changes in soil moisture will cause domain shift. Contrast learning through the InfoNCE loss is used to enhance feature discriminability, forcing the model to pull the enhanced positive samples closer in the feature space while pushing the negative samples farther away, thereby enhancing the robustness of features to geometric deformations and noises and alleviating the distribution differences between the source domain and the target domain.
[0027] The second stage: target domain training, optimizing the student model based on self-supervised contrast learning and double-screening pseudo-label supervised learning.
[0028] The self-supervised contrast learning in the second stage is similar to that in the first stage, except that a batch of samples input comes from the target domain dataset.
[0029] The double-screening pseudo-label supervised learning in the second stage refers to the process of double-screening the region of interest through LPQ and MC Dropout to calculate the pseudo-label supervised loss L p,b of the sample, and calculating the batch pseudo-label supervised loss L p,b based on L ped In this stage, an LPQ value is first calculated for each position of the pseudo-label of the sample, and then the uncertainty of each position of the sample is calculated using the MC Dropout method. Then, through the condition LPQ(d, h, w) > λ and U(d, h, w) < ϵ, the positions are double-screened. The advantage of LPQ screening is to filter low-quality pseudo-label regions through three-dimensional neighborhood confidence evaluation and suppress the error propagation of noisy labels. The advantage of MCDropout screening is to quantify the model prediction uncertainty and exclude high-variance regions to ensure the reliability of knowledge transfer. In this stage, through the double-filtering mechanism, the reliability of the pseudo-labels and the robustness of cross-domain knowledge transfer are improved, and the feature consistency is enhanced through contrast learning, autonomously enhancing the discriminability for similar diseases.
[0030] The third stage: Mixed-domain training refers to the process of constructing a student model for mixed-domain training and updating the parameters of the teacher model through Exponential Moving Average (EMA). In this stage, a mixed weight is introduced, which is calculated based on the and in a batch of target-domain samples, and has the advantage of dynamically adapting to data quality: emphasizing the target domain (γ→1) to accelerate adaptation for high-quality target data, and emphasizing the source domain (γ→0) to avoid negative transfer for low-quality data. This can enhance the adaptive ability of the model in complex environments. And a consistency constraint is introduced in this stage to reduce noise interference.
[0031] Compared with the prior art, the advantages of the present invention are as follows: (1) The constructed 3D encoding and decoding network has a significant improvement in 3D semantic segmentation accuracy and boundary recognition ability: When traditional methods deal with 3D complex distributions of underground diseases such as the reticular expansion of cracks and the longitudinal extension of cavities, they often result in blurred boundaries and misjudgments due to insufficient local feature extraction or inadequate global dependency modeling. The 3D encoding and decoding network proposed in the present invention models the long-range global dependency through a 3D multi-head self-attention mechanism, combines depthwise separable convolution to extract local detailed features, and effectively fuses multi-scale context information; at the same time, a progressive dilated residual upsampling and a 3D Sobel operator boundary enhancement module are introduced to enhance the distinguishability of disease boundaries through gradient response. Theoretical derivation shows that this method can accurately capture the three-dimensional distribution characteristics of underground diseases, significantly improve the positioning ability of small targets and the robustness of boundary segmentation, and is especially suitable for the refined identification of cracks and cavities in complex geological structures.
[0032] (2) A breakthrough enhancement in cross-domain environmental adaptability and generalization performance: Existing models often suffer from performance degradation due to data distribution shift in cross-domain scenarios such as different geological conditions and differences in radar equipment parameters. The training method proposed in the present invention is divided into three stages, which can autonomously adapt to noise interference and domain shift in complex environments, reduce the dependence on labeled data, and still maintain high-precision detection ability in unlabeled target-domain data, significantly superior to models trained in a traditional single domain.
[0033] (3) Optimization of computational efficiency and resource consumption: For the problem of high computational complexity in 3D data processing, the model adopts depthwise separable 3D convolution and progressive dilated residual upsampling techniques, which not only reduce the computational complexity but also maintain the efficient fusion of multi-scale features. In addition, the teacher-student framework combines the exponential moving average (EMA) parameter update mechanism, and reduces the training fluctuations through the smooth iteration of pseudo-labels, further optimizing the consumption of training resources. Theoretical derivation shows that compared with the frequent parameter adjustment and high computing power requirements in traditional domain adaptation methods, the present invention achieves more efficient utilization of computing resources while ensuring accuracy, and is applicable to the real-time processing requirements of large-scale 3D radar data in practical engineering.
[0034] In summary, the present invention can maintain high-precision detection ability in complex scenarios with different geological conditions and equipment parameters, significantly reduce the dependence on labeled data in the target domain, and provide reliable support for the intelligent and universal detection of underground diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of the present invention; Figure 2 is an architecture diagram during the training of the present invention; Figure 3 is a schematic diagram of the 3D encoder structure; Figure 4 is a schematic diagram of the 3D decoder structure; Figure 5 is a schematic diagram of the structure of the first upsampling layer; Figure 6 is a schematic diagram of the structure of the second upsampling layer; Figure 7 is a schematic diagram of cross-layer feature access. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present invention will be further described below in conjunction with embodiments and the drawings.
[0037] Embodiment 1: Refer to Figure 1 and Figure 2 , a method for adaptively training a 3D ground-penetrating radar underground disease recognition model in a complex environment, comprising the following steps; S1, obtaining several samples with class labels to form a source domain dataset D s , several unlabeled samples to form a target domain dataset D t , the samples being 3D ground-penetrating radar data and the class being disease classes; D s a class label map of a sample X s in is P s , D t a sample in is X t , X s ∈RD×H×W×Cin , X t ∈R D×H×W×Cin , P s ∈R D×H×W×C , where D, H, and W are the depth, length, and width of the sample resolution respectively, Cin is the number of channels of the radar reflection intensity, Cin = 1, C is the total number of disease categories, and the value at position (d, h, w) in P s is X s The true disease category at position (d, h, w) in, where d, h, and w are the depth, length, and width of position (d, h, w) respectively; S2, construct a 3D encoder-decoder network, including a 3D encoder, a 3D decoder, and a self-supervised module; The 3D encoder is used to encode the input samples to obtain encoded features, and the 3D decoder is used to decode the encoded features and then perform boundary enhancement, and output the predicted class probabilities at each position to form a probability prediction map , and the self-supervised module is used to construct positive and negative samples based on the encoded features of a batch of samples and calculate the InfoNCE loss; S3, use D S to minimize the total source domain training loss L s Train the 3D encoder-decoder network to obtain a basic model, L s = λ cn1 L cn1 + λ cls L cls , L cn1 is the InfoNCE loss, L cls is the classification loss, and λ cn1 , λ cls are the weights of L cn1 and L cls respectively; S4, copy two basic models and use them as the teacher model and the student model respectively; S5, use D t to train the student model, and optimize the student model based on self-supervised contrast learning and double-screening pseudo-label supervised learning in each iteration, where one optimization includes steps S51~S54; S51, use the teacher model to output the probability prediction map of each sample in D t as the pseudo-label, and the pseudo-label of X t is , and preset the quality threshold λ and the uncertainty threshold ϵ; S52, randomly select a batch of samples from D t , where the b-th sample is X b , use the student model to output the probability prediction map of each sample, and use the self-supervised module to calculate the InfoNCE loss; S53, Calculate X based on double screening and calculate the pseudo-label supervision loss L of b ; b For the pseudo-label of X, calculate the LPQ value LPQ(d, h, w) of the position (d, h, w) based on local phase quantization. For X, calculate the uncertainty U(d, h, w) of the position (d, h, w) based on MC Dropout, and form the region of interest of X with positions where LPQ(d, h, w) > λ and U(d, h, w) < ϵ, and calculate L; p,b ; For X b pseudo-label , calculate the LPQ value LPQ(d, h, w) of the position (d, h, w) based on local phase quantization, for X b , calculate the uncertainty U(d, h, w) of the position (d, h, w) based on MC Dropout, and form the region of interest of X with positions where LPQ(d, h, w) > λ and U(d, h, w) < ϵ b region of interest , calculate L p,b ; , wherein, is the total number of positions within, is the normalized value of U(d, h, w), is the probability prediction map of X b at the position (d, h, w); value at the position (d, h, w); S54, Calculate the total loss L of the target domain, and adjust the parameters of the student model by minimizing L; t , and adjust the parameters of the student model by minimizing L t ; , wherein, B is the batch size, L cn2 is the InfoNCE loss of the self-supervised module in the student model, and λ cn2 is the weight of L cn2 ; S6, Construct a mixed-domain dataset of multiple batches using D s and D t ; S7, Train the student model with the mixed-domain dataset to update the student model by minimizing the joint loss L, and update the teacher model using the EMA method, L total = + L total + λ s L t + λ t L c + λ consist L consist , and L is the consistency loss of a batch of samples.
[0038] The disease categories include background, cavity, void, loose body, water-rich body, and fracture network.
[0039] The method for the self-supervised module to calculate the InfoNCE loss is; Obtain the encoded features of a batch of samples, and label them as F1~F in sequenceB , calculate the InfoNCE loss of the b-th encoded feature and the InfoNCE loss L of a batch of samples cn1 ; , ,
[0040] wherein, is to randomly add data augmentation to F b to obtain positive samples, is the v-th negative sample of F b , 1 ≤ v ≤ 2B - 2, sim(⋅) is to calculate the cosine similarity, τ is the temperature coefficient, and exp(⋅) is the exp function.
[0041] In S3, the classification loss is calculated according to the following formula; , wherein, α c is the weight of the c-th disease category, 1 ≤ c ≤ C, γ is the modulation factor for reducing the weight of easily classified samples, for the position (d, h, w) of the b-th sample, , are the normalized value of the predicted class probability and the true disease class annotation at this position, respectively.
[0042] In S53, LPQ(d, h, w) is calculated according to the following formula: , wherein, K is the neighborhood size, i, j, and k are the offsets of the position (d, h, w) in the depth, length, and width directions respectively, and Q(d + i, h + j, w + k) is the pseudo label value at the position (d + i, h + j, w + k).
[0043] The method for calculating U(d, h, w) in S53 is: Use the MC Dropout method to perform M forward propagations on X t , and generate a probability prediction map for each forward propagation. Calculate the uncertainty U(d, h, w) at the position (d, h, w) according to the following formula; , wherein, is the value of the probability prediction map obtained from the m-th forward propagation at the position (d, h, w), is the mean value of the M probability prediction maps at the position (d, h, w).
[0044] In S6, the method for constructing a batch of mixed-domain datasets is: Combine D sand D t After mixing, select any batch of samples, or select a batch of samples from D according to the mixing weight s and D t The mixing weight is preset or obtained according to steps S61 to S63; S61, from D s , D t Select any batch of samples respectively to form sets , ; S62, calculate The LPQ value and uncertainty at each position of each sample in, and the average value of all LPQ values is used to obtain the batch average local quality score , and the average value of all uncertainties is used to obtain the batch average uncertainty ; S63, calculate The mixing weight of ; , α and β are respectively and The adjustment coefficients of, and Sigmoid(⋅) is the Sigmoid function; According to , generate a batch of mixed-domain data sets .
[0045] In S7, L consist Calculate according to the following formula: , , In the formula, Is the square of the L2 norm. For the b-th sample in the mixed domain, Is the probability prediction map output by the student model, Is the pseudo-label generated by the teacher model, Is the consistency loss of the b-th sample.
[0046] Example 2: See Figures 1 to 7 , on the basis of Example 1, a more specific three-dimensional encoder structure is given as Figure 3 Shown, a more specific three-dimensional decoder structure is as Figure 4 Shown. Of course, the three-dimensional encoder structure and the three-dimensional decoder are not limited to the structures in this embodiment, and any structure that can encode and decode three-dimensional samples is acceptable.
[0047] In this embodiment, the three-dimensional encoder includes a fixed chunk embedding module, a first multi-scale sampling module, and a second multi-scale sampling module connected in sequence; The fixed chunk embedding module includes a three-dimensional convolutional layer and a three-dimensional position encoding layer; The three-dimensional convolutional layer is used to perform three-dimensional convolutional dimensionality reduction on the input samples to obtain the dimensionality-reduced feature F cn , where the convolutional kernel of the three-dimensional convolution is 3×3×3 and the stride is 2; The three-dimensional position encoding layer is used to generate three-dimensional sine position encoding for each position of F cn to obtain the position encoding feature F p , and then add F cn and F p element-wise to obtain the embedding feature F0. The scale of the sample is D×H×W×C in , F cn , F p , and the scales of F0 are ; The first multi-scale sampling module includes a global attention layer, a local convolutional layer, a layer normalization layer, and a three-dimensional downsampling layer; The multi-scale sampling module is used to perform global attention on the input feature to obtain the global feature F global ; The local convolutional layer is used to perform depthwise separable three-dimensional convolution on F global to obtain the local feature F local , with a convolutional kernel of 3×3×3 and a stride of 1; The layer normalization layer is used to obtain the mixed feature F mix =LayerNorm(F global +F local ) mix ; The three-dimensional downsampling layer is used to perform three-dimensional convolutional downsampling on F mix to obtain the output feature F d1 ; The structure of the second multi-scale sampling module is the same as that of the first multi-scale sampling module. The input feature is F d1 , and the output scale is the output feature F d2 . And F d2 is the encoded feature of the three-dimensional encoder.
[0048] The working principle of the three-dimensional encoder is as follows: In the downsampling process of the three-dimensional encoder, global attention and local convolution are alternately used for feature extraction and then downsampling. Among them, global attention obtains the global feature F global , and then performs depthwise separable three-dimensional convolution on F global to obtain the local feature F local , and then fuses F global and F local, the fused features take into account both long-range dependencies and local details while reducing computational complexity. After downsampling, the resolution is reduced and the number of channels is increased. In the present invention, the input sample scale is D×H×W×C in , and the feature F is output after passing through the first multi-scale sampling module d1 , with a scale of , and the feature F is output after passing through the second multi-scale sampling module d2 , with a scale of . The output of the second multi-scale sampling module is also used as the encoded feature, which is the fundamental depth representation
[0049] In this embodiment, the three-dimensional decoder includes a first upsampling module, a second upsampling module, a third upsampling module, and a boundary enhancement layer connected in sequence The first upsampling module includes a transposed convolutional upsampling layer, a cross-layer splicing layer, a dilated residual module, and a residual connection activation layer The transposed convolutional upsampling layer is used to perform transposed three-dimensional convolution on the encoded feature to obtain the first feature F u1 , and the convolutional kernel of the transposed three-dimensional convolution is 2×2×2, with a stride of 2 The cross-layer splicing layer is used to splice F u1 and the cross-layer feature along the channel dimension to obtain the second feature F u2 , where the cross-layer feature is F d1 ; The dilated residual module performs dilated convolutions with dilation rates of 2 and 4 on F u2 respectively, and then adds them together to obtain the third feature F u3 ; The residual connection activation layer is used to perform residual connection on F u2 and F u3 , and then activate through the GELU function to obtain the output F of the first upsampling module uout1 ; Duplicate a first upsampling module, and add a channel compression layer at the input end, which is used to compress the channels of the input feature by half and then send it to the transposed convolutional upsampling layer. Modify the cross-layer feature to F0, and modify the dilation rate of the dilated residual module to 1 and 2 to obtain the second upsampling module Duplicate a second upsampling module, and modify the cross-layer feature to the feature F 01 , and the dilated residual module only performs dilated convolution with a dilation rate of 1 on the input feature to obtain the third upsampling module. The feature F 01 is obtained by transposed convolutional upsampling of F0, with a convolutional kernel of 2×2×2 and a stride of 2 The classifier is used to classify F uout3 position by position to obtain the initial predicted class probability P base ; The boundary enhancement layer processes F uout3 to perform edge detection using a 3D Sobel operator and output a boundary response map P Sobel , and then according to the formula , generate a probability prediction map , where λ Sobel is the weight of P Sobel .
[0050] The working principle of the 3D decoder is as follows: The first upsampling module to the third upsampling module are used to perform progressive upsampling operations and then boundary enhancement. Each upsampling module involves cross-layer concatenation and dilated residuals. Cross-layer concatenation can double the number of channels while retaining fine-grained structural features; dilated residuals can effectively expand the receptive field of features while retaining local details, thereby enhancing the features.
[0051] In the first upsampling module, the input encoded feature is , first perform transposed 3D convolutional upsampling and adjust the number of channels to 128 to obtain F u1 with a scale of , and then concatenate it with F with a scale of d1 along the channel dimension to obtain F u2 with a scale of , and then pass it through a dilated residual module to obtain F u3 , use a residual connection activation layer to process F u2 and F u3 to obtain F uout1 with a scale of .
[0052] In the second upsampling module, first compress the channels of F uout1 to obtain a feature with a scale of , perform transposed 3D convolutional upsampling to obtain , and then concatenate it with the cross-layer feature F0 with a scale of along the channel dimension to obtain a feature with a scale of , after passing through dilated residuals and residual connection activation, obtain the output F uout2 of the second upsampling module with a scale of .
[0053] In the third upsampling module, first compress the channels of F uout2 to obtain a feature with a scale of , perform transposed 3D convolutional upsampling to obtain , and then concatenate it with the cross-layer feature F with a scale of 01 along the channel dimension to obtain a feature with a scale of , after passing through dilated residuals and residual connection activation, obtain the output Fuout3 , with a scale of .
[0054] Through the three - time upsampling module, it has the characteristics of progressive resolution restoration, symmetric skip connections and multi - scale feature fusion. Then, through the classifier, P is obtained base , perform boundary enhancement on F uout3 to obtain P Sobel , and then according to the formula calculate the probability prediction map , λ Sobel is generally set to 0.5. The 3D decoder restores the resolution through multi - level upsampling, fuses the context through dilated convolutions, and optimizes the details through boundary enhancement, finally obtaining a high - precision probability prediction map.
[0055] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A complex environment adaptive training method for a three-dimensional ground penetrating radar underground disease identification model, characterized in that Including the following steps; S1. Obtain a source domain dataset D consisting of several samples with category-annotated images s , and a target domain dataset D consisting of several unannotated samples t , where the samples are 3D ground penetrating radar data and the category is the disease category; D s The class annotation map of a sample X in s is P s , D t One sample in is X t , X s ∈R D×H×W×Cin , X t ∈R D×H×W×Cin , P s ∈R D ×H×W×C , where D, H, and W are the depth, length, and width of the sample resolution respectively, Cin is the number of channels of the radar reflection intensity, Cin = 1, C is the total number of disease classes, and the value at position (d, h, w) in P s is X s is the true disease class at position (d, h, w), and d, h, and w are the depth, length, and width of position (d, h, w) respectively; S2. Construct a 3D encoding and decoding network, including a 3D encoder, a 3D decoder, and a self-supervised module; The 3D encoder is used to encode the input samples to obtain encoded features, and the 3D decoder is used to decode the encoded features and then perform boundary enhancement to output the predicted class probabilities at each position to form a probability prediction map. , and the self-supervised module is used to construct positive samples and negative samples based on the encoded features of a batch of samples and calculate the InfoNCE loss. S3, use D S to minimize the total source domain training loss L s Train the 3D encoding and decoding network to obtain the basic model, L s = λ cn1 L cn1 + λ cls L cls where L cn1 is the InfoNCE loss, and L cls is the classification loss, and λ cn1 and λ cls are the weights of L cn1 and L cls respectively; S4. Duplicate two basic models, serving as the teacher model and the student model respectively; S5, using D t Train the student model, and optimize the student model based on self-supervised contrastive learning and double-screening pseudo-label supervised learning in each iteration, where one optimization includes steps S51 to S54; S51, use the teacher model to output the probability prediction graph D t of each sample in as the pseudo-label, X t The pseudo-label of is , preset the quality threshold λ and the uncertainty threshold ϵ; S52, from D t Select any batch of samples, where the b-th sample is X b , use the student model to output the probability prediction map of each sample, and use the self-supervised module to calculate the InfoNCE loss; S53, Calculate X based on double screening b Pseudo-label supervised loss L p,b ; For X b pseudo-label of , calculate the LPQ value LPQ(d, h, w) of the position (d, h, w) based on local phase quantization for X b , calculate the uncertainty U(d, h, w) of the position (d, h, w) based on MC Dropout, and form X with the positions where LPQ(d, h, w) > λ and U(d, h, w) < ϵ b region of interest , calculate L p,b ; , In the formula, is the total number of internal positions, is the normalized value of U(d, h, w), is b the probability prediction graph of X at the value of the position (d, h, w); S54, calculate the total loss L of the target domain t , and minimize L t to adjust the parameters of the student model; , where B is the batch size, and L cn2 is the InfoNCE loss of the self-supervised module in the student model, and λ cn2 is the weight of L cn2 ; S6, using D s and D t Construct multiple batches of hybrid-domain datasets; S7. Train the student model with the hybrid-domain dataset to minimize the joint loss L total Update the student model and update the teacher model using the EMA method, L total = + L s + λ t L t + λ c L consist where L consist is the consistency loss for a batch of samples.
2. The complex environment adaptive training method for the 3D ground penetrating radar underground disease identification model according to claim 1, wherein The disease categories include background, cavity, void, loose body, water-rich body, and crack network.
3. The complex environment adaptive training method for the 3D ground penetrating radar underground disease identification model according to claim 1, characterized in that, The method for the self-supervised module to calculate the InfoNCE loss is; Obtain the encoded features of a batch of samples, sequentially labeled as F1~F B , calculate the InfoNCE loss of the b-th encoded feature and the InfoNCE loss L of a batch of samples cn1 ; , , In the formula, is the positive sample obtained by randomly adding data augmentation to F b , is the v-th negative sample of F b , 1 ≤ v ≤ 2B - 2, sim(⋅) is to calculate the cosine similarity, τ is the temperature coefficient, and exp(⋅) is the exp function.
4. The adaptive training method for the three-dimensional ground penetrating radar underground disease identification model in complex environments according to claim 1, characterized in that In S3, the classification loss is calculated according to the following formula; , where α c is the weight of the c-th disease category, 1 ≤ c ≤ C, γ is the modulation factor for reducing the weight of easily classifiable samples, and for the position (d, h, w) of the b-th sample, and are the normalized value of the predicted category probability and the true disease category annotation at this position, respectively.
5. The adaptive training method for a three-dimensional ground penetrating radar underground disease identification model in a complex environment according to claim 1, wherein In S53, LPQ(d, h, w) is calculated according to the following formula: , Where K is the neighborhood size, i, j, and k are the offsets in the depth, length, and width of the position (d, h, w) respectively, and Q(d + i, h + j, w + k) is the pseudo-label at the position (d + i, h + j, w + k).
6. The complex environment adaptive training method for the 3D ground penetrating radar underground disease identification model according to claim 1, wherein The method for calculating U(d, h, w) in S53 is: Perform M forward propagations on X using the MC Dropout method t Generate a probability prediction map for each forward propagation. Calculate the uncertainty U(d, h, w) at position (d, h, w) according to the following formula; , In the formula, is the probability prediction map obtained from the m-th forward propagation The value at position (d, h, w), is the mean of the M probability prediction maps at position (d, h, w).
7. The adaptive training method for the complex environment of the 3D ground penetrating radar underground disease identification model according to claim 1, wherein In S6, the method for constructing a batch of mixed-domain datasets is as follows: Mix D s and D t , and then randomly select a batch of samples, or select a batch of samples from D s and D t according to the mixing weights. The mixing weights are preset or obtained according to steps S61 to S63; S61, from D s , D t Any batch of samples from each of them respectively constitutes a set , ; S62, Calculate the LPQ value and uncertainty at each position of each sample in , and calculate the batch-average local quality score by averaging all LPQ values and the batch-average uncertainty by averaging all uncertainties ; S63, calculate the mixed weight of ; , α and β are respectively and adjustment coefficients, and Sigmoid(⋅) is the Sigmoid function; According to , generate a batch of mixed-domain datasets .
8. The complex environment adaptive training method for the three-dimensional ground penetrating radar underground disease identification model according to claim 1, characterized in that In S7, L consist It is calculated according to the following formula: , , In the formula, is the square of the L2 norm. For the b-th sample in the mixed domain, is the probability prediction map output by the student model, is the pseudo-label generated by the teacher model, is the consistency loss of the b-th sample.
9. The adaptive training method for the three-dimensional ground penetrating radar underground disease identification model in a complex environment according to claim 1, wherein In the 3D encoding and decoding network, the 3D encoder includes a fixed chunk embedding module, a first multi-scale sampling module, and a second multi-scale sampling module connected in sequence; The fixed chunk embedding module includes a 3D convolutional layer and a 3D position encoding layer; The three-dimensional convolutional layer is used to perform three-dimensional convolutional dimensionality reduction on the input samples to obtain dimensionality-reduced features F cn , where the convolutional kernel of the three-dimensional convolution is 3×3×3 and the stride is 2; The three-dimensional position encoding layer is used to generate three-dimensional sine position encoding for each position in F cn to obtain position encoding feature F p . Then, F cn and F p are added element-wise to obtain embedding feature F0. The scale of the sample is D×H×W×C in . The scales of F cn , F p , and F0 are ; The first multi-scale sampling module includes a global attention layer, a local convolutional layer, a layer normalization layer, and a 3D downsampling layer; The multi-scale sampling module is used to perform global attention on the input features to obtain the global feature F global ; The local convolutional layer is used for F global to perform depthwise separable 3D convolution to obtain local feature F local , with a convolution kernel of 3×3×3 and a stride of 1; The layer normalization layer is used to obtain the mixed feature F according to the formula F mix =LayerNorm(F global +F local ) mix ; The three-dimensional downsampling layer is used to perform three-dimensional convolutional downsampling on F mix to obtain the output feature F d1 ; The structure of the second multi-scale sampling module is the same as that of the first multi-scale sampling module, and the input feature is F d1 , and the output scale is of the output feature F d2 , and F d2 is the encoded feature of the three-dimensional encoder.
10. The complex environment adaptive training method for the three-dimensional ground penetrating radar underground disease identification model according to claim 9, wherein In the 3D encoding and decoding network, the 3D decoder includes a first upsampling module, a second upsampling module, a third upsampling module, and a boundary enhancement layer connected in sequence; The first upsampling module includes a transposed convolutional upsampling layer, a cross-layer splicing layer, a dilated residual module, and a residual connection activation layer; The transposed convolutional upsampling layer is used to perform transposed 3D convolution on the encoded features to obtain the first feature F u1 , and the convolutional kernel of the transposed 3D convolution is 2×2×2, and the stride is 2; The cross-layer splicing layer is used to splice F u1 and the cross-layer feature along the channel dimension to obtain the second feature F u2 , where the cross-layer feature is F d1 ; The hollow residual module performs dilated convolutions on F u2 with dilation rates of 2 and 4 respectively, and then adds them together to obtain the third feature F u3 ; The residual connection activation layer is used to perform a residual connection on F u2 and F u3 and then activate it through the GELU function to obtain the output F uout1 ; Duplicate a first upsampling module, add a channel compression layer at the input end to compress the channels of the input features by half and then send them into the transposed convolutional upsampling layer, modify the cross-layer features to F0, and modify the dilation rate of the dilated residual module to 1 and 2 to obtain a second upsampling module; Duplicate a second upsampling module to modify the cross-layer features into feature F 01 The dilated residual module only performs dilated convolution with a dilation rate of 1 on the features input thereto to obtain a third upsampling module, feature F 01 It is obtained by transposed convolution upsampling of F0, with a convolution kernel of 2×2×2 and a stride of 2; The classifier is used to classify F uout3 position by position to obtain the initial predicted class probability P for each position base ; The boundary enhancement layer pairs with F uout3 perform edge detection using a 3D Sobel operator and output a boundary response map P Sobel , and then according to the formula , generate a probability prediction map , where λ Sobel is the weight of P Sobel .
Citation Information
Patent Citations
Cross-domain multi-modal remote sensing image classification method based on comparative learning
CN116912595A
Ground penetrating radar target identification method and system based on semi-supervised learning
CN118209954A
Medical image segmentation method based on semi-supervised domain self-adaption and cross-domain collaborative learning
CN119785033A
Cited By
Water supply pipeline leakage detection method and system based on radar image and deep learning
CN121599966A
Water supply pipeline leakage detection method and system based on radar image and deep learning
CN121599966B