A Smart Identification Method for the Upper Limit of Permafrost Based on Multi-Source Data and Deep Clustering Fusion

By integrating multi-source data and using deep learning algorithms, a dual-branch convolutional autoencoder network was constructed, which solved the problem of low efficiency in manual experience interpretation during permafrost exploration. This enabled automated and high-precision identification of the upper limit of permafrost, improving exploration efficiency and accuracy.

CN121071440BActive Publication Date: 2026-03-13CHINA RAILWAY DESIGN GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing permafrost exploration technologies rely on manual experience for interpretation, which is inefficient and makes it difficult to achieve automated and high-precision identification of the upper limit depth of permafrost.

Method used

By fusing drilling results and geophysical data from multiple sources, and employing deep learning and clustering algorithms, a dual-branch convolutional autoencoder network is constructed. Combining unsupervised and semi-supervised learning, intelligent identification of the upper limit of permafrost is achieved.

Benefits of technology

It improves the efficiency and accuracy of permafrost exploration, provides a reliable intelligent solution, and enhances the objectivity and accuracy of exploration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071440B_ABST
    Figure CN121071440B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion, belonging to the field of permafrost exploration technology. The method first collects and preprocesses ground-penetrating radar and high-density electrical resistivity tomography (EDS) data, and registers them with borehole data; it then constructs a sample dataset using a sliding window; a bi-branch convolutional autoencoder is constructed for unsupervised pre-training to fuse features from multiple data sources; based on this, a clustering layer is introduced, and a joint loss function combining reconstruction loss, clustering loss, and supervised loss is used for semi-supervised deep clustering optimization, dividing the subsurface medium into three categories: overburden, transition zone, and permafrost; finally, the trained model is used to automatically identify the upper limit depth of permafrost at each point on the survey line and output the confidence score. This invention overcomes the problems of traditional methods relying on manual experience and being inefficient, achieving intelligent, high-precision, and high-efficiency identification of the upper limit of permafrost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of frozen soil detection technology in railway engineering, and in particular to an intelligent identification method for the upper limit of frozen soil based on the fusion of multi-source data and deep clustering. Background Technology

[0002] With the continuous improvement of the national high-speed railway backbone network and the basic formation of the inland trunk railway system, my country's railway construction is advancing towards the western plateau and high-altitude regions. In this process, the permafrost problem has become a core engineering geological challenge affecting railway line site selection, structural stability, and long-term safe operation. When constructing new railway projects traverse plateau permafrost areas, it is essential to first accurately determine the development characteristics and spatial distribution of permafrost along the route. Among these, the burial depth of the upper limit of permafrost is one of the most critical parameters determining the subgrade design, bridge and culvert foundation type, and engineering treatment measures. The upper limit of permafrost refers to the upper boundary of the permafrost layer, i.e., the boundary between the active layer and the permafrost layer; it is mainly used to identify the initial depth of the permafrost layer. The detection of the upper limit of permafrost is actually to determine the specific depth or thickness of the active layer. Therefore, identifying the upper limit of permafrost based on exploration data is a crucial foundation for railway construction.

[0003] Currently, permafrost exploration mainly relies on drilling and geophysical methods. Drilling, as the most direct and reliable traditional method, provides decisive evidence for accurately determining the upper limit depth of permafrost. Geophysical methods, such as high-density electrical resistivity tomography, ground-penetrating radar, seismic exploration, and transient electromagnetic methods, utilize the significant differences in physical properties such as resistivity, dielectric constant, and wave velocity between permafrost and thawed soil to achieve rapid, non-destructive, and large-scale detection, effectively compensating for the lack of spatial continuity in point drilling.

[0004] However, the existing exploration technology system suffers from a fundamental bottleneck: whether it's the direct method centered on drilling or the indirect method represented by geophysical exploration, the interpretation of the final data and the identification of the permafrost layer—especially the precise location of the critical interface of the upper limit of permafrost—still heavily rely on the personal experience and professional knowledge of geological engineers. Faced with the massive amounts of diverse information generated by exploration, such as borehole data, resistivity profiles, and radar images, the traditional manual interpretation process is not only highly subjective and difficult to guarantee consistency, but also inefficient and time-consuming, becoming a prominent shortcoming restricting the rapid exploration and design of plateau railways.

[0005] For example, existing technologies, such as the patent application number CN202310556604 entitled "A method, device, equipment and medium for permafrost exploration in high-latitude forest areas", although it proposes a comprehensive exploration approach from surface to point, and comprehensively uses remote sensing, geophyte mapping and geophysical methods to macroscopically delineate the permafrost range, its core interpretation link is still based on traditional manual experience, and the data processing model mostly assumes a linear relationship, failing to achieve automated and intelligent high-precision identification and output of the upper limit depth of permafrost.

[0006] Therefore, to address the aforementioned issues, this invention aims to propose an intelligent identification method for the upper limit of permafrost depth based on the fusion of multi-source data and deep clustering. This method integrates precise sample control points provided by deep drilling with large-scale, continuous physical field data obtained through geophysical methods. Through deep learning and clustering algorithms, it ultimately achieves automated, high-precision intelligent identification and output of the upper limit depth of permafrost, thereby fundamentally improving the exploration efficiency and quality of results in plateau permafrost regions. Summary of the Invention

[0007] Therefore, the purpose of this invention is to provide an intelligent identification method for the upper limit of permafrost based on the fusion of multi-source data and deep clustering. This method addresses the problem of traditional interpretation methods being inefficient due to heavy reliance on manual interpretation. By cross-integrating drilling results, geophysical data, and machine learning algorithms, it achieves intelligent and high-precision exploration of the upper limit of permafrost.

[0008] To achieve the above objectives, this invention discloses an intelligent identification method for the upper limit of permafrost based on the fusion of multi-source data and deep clustering, comprising the following steps:

[0009] S1. For each measuring point on the survey line, multi-source geophysical data are acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed. At the same time, borehole data are collected and the upper limit depth of frozen soil is marked, and spatial registration is performed with the geophysical survey line.

[0010] S2. The preprocessed ground-penetrating radar data and high-density electrical resistivity tomography data are divided into data blocks using a sliding window, and a label is generated for each data block.

[0011] S3. Construct a dual-branch convolutional autoencoder network to extract features from the preprocessed ground-penetrating radar data and high-density electrical resistivity tomography (EDT) data, and perform unsupervised pre-training. Introduce a clustering layer on the pre-trained model, where the joint loss function is formed by a weighted combination of the reconstruction loss function, clustering loss function and supervision loss function as shown in the following formula.

[0012]

[0013] in, For joint losses, This represents the total reconstruction loss. Clustering loss; To monitor losses, , and These are the weight parameters for each loss function;

[0014] S4. Using the trained model, predict the data acquired in real time along the entire survey line, output the upper limit depth of frozen soil and confidence level of each measuring point, and finally obtain the upper limit of frozen soil along the entire measurement profile.

[0015] In a further preferred embodiment, in S1, when acquiring ground penetrating radar data and high-density electrical resistivity tomography (EDT) data, data is collected along the same survey line with the same spacing between measurement points; RTK is used to accurately locate the measurement points to ensure the synchronization of the two sets of data in terms of location.

[0016] Further preferably, in S1, multi-source geophysical data is acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed, including:

[0017] After zero-point correction, gain adjustment and filtering, the ground penetrating radar data is converted into a ground penetrating radar time profile result map.

[0018] The ground-penetrating radar time profile results are converted into depth profile results based on velocity calibration.

[0019] A resistivity profile is obtained by inverting high-density electrical resistivity data.

[0020] The locations of the measuring points in the depth profile results map and the resistivity profile results map are matched; the ground-penetrating radar amplitude value and high-density electrical resistivity value of each measuring point are normalized to the interval [0, 1].

[0021] More preferably, in S2, the step of using a sliding window to cut the preprocessed ground-penetrating radar data and high-density electrical resistivity tomography (EDT) data into data blocks and generating a label for each data block includes:

[0022] Each sample set contains a pair of registered data blocks: ground-penetrating radar data. and high-density electrical resistivity data ;

[0023] When generating labels for each data block, the following is included:

[0024] For geophysical measurement points without borehole markers, the labels corresponding to the sample data are empty, which is used for unsupervised learning;

[0025] For geophysical measuring points with borehole markings, the sample data are labeled as labeled samples; the labeled samples are represented by the binarized labels shown in the following formula. ;

[0026]

[0027] in, d The value indicates the actual depth, with 0 indicating non-frozen soil and 1 indicating frozen soil.

[0028] More preferably, in S3, the dual-branch convolutional autoencoder network includes:

[0029] Encoder: consists of an input layer and a fully connected layer;

[0030] The input layer has two ports, one for extracting ground-penetrating radar data. Data characteristics A tool for extracting high-density electrical resistivity data. Data characteristics ;

[0031] Fully connected layers are used to integrate data features and The process involves fusion and dimensionality reduction to output the final latent vector. ;

[0032] The decoder is configured as a deconvolution symmetric to the encoder, to obtain the latent vector. As input, the reconstructed ground-penetrating radar data are obtained respectively. and reconstructed high-density electrical resistivity data .

[0033] More preferably, in S3, the loss function of the dual-branch convolutional autoencoder network adopts mean squared error loss;

[0034] Calculate the reconstructed ground-penetrating radar data Compared with the original ground-penetrating radar data Mean square error and reconstructed high-density electrical resistivity data Compared with the original high-density electrical resistivity data The mean squared error is weighted and summed as the total reconstruction loss function;

[0035]

[0036] in, Original ground-penetrating radar data Or the original high-density electrical resistivity data ; For the reconstructed ground-penetrating radar data Or reconstructed high-density electrical resistivity data , where i is the detection point.

[0037] More preferably, the clustering layer is positioned after the encoder outputs the feature vector, and the clustering layer is used to calculate the cluster probability of a data point belonging to a certain cluster category using the following formula:

[0038]

[0039] in, This represents the probability that sample i belongs to cluster j; It is the feature vector extracted by the encoder for the i-th input data block; It is the center of the j-th cluster and belongs to a trainable parameter in the network; It is an eigenvector With cluster center The square of the Euclidean distance between them Let j be the reconstructed clusters.

[0040] Further preferred embodiments include calculating the target distribution probability according to the following formula, and emphasizing the correlation strength between data points and cluster centers by introducing the target assignment P to increase the probability of high confidence assignments in Q;

[0041]

[0042] in, It is the probability that sample i belongs to cluster j in the target distribution; It is the sum of all data points assigned to cluster j. To balance cluster sizes, high-confidence classifications are amplified through squaring, and by dividing by... This is to prevent a cluster from becoming too large and swallowing up other clusters.

[0043] More preferably, the clustering loss function is expressed by the following formula:

[0044]

[0045] This is the Kullback-Leibler divergence, which has a value of 0 when the two distributions are identical; the greater the difference, the larger the value; by minimizing... This makes the network's output distribution Q closer to the target distribution P.

[0046] Further preferably, the total reconstruction loss It is expressed by the following formula:

[0047]

[0048] in, and This is a weighting hyperparameter that balances the contributions of the two data sources to the total loss.

[0049] This application discloses an intelligent identification method for the upper limit of permafrost based on the fusion of multi-source data and deep clustering. This method significantly overcomes the core bottlenecks of traditional exploration methods, which rely on manual experience interpretation, are inefficient and highly subjective, and are limited by the expensive borehole label data of existing supervised learning methods. It innovatively fuses ground-penetrating radar (GPR) and high-density electrical resistivity tomography (ERT) data at the feature level and constructs a semi-supervised deep clustering framework for permafrost identification. This framework enables the model to autonomously extract essential features from multi-source data through unsupervised pre-training, and then introduces a small amount of borehole label data as physical constraints. By jointly optimizing reconstruction loss, clustering loss, and supervised loss, the model is collaboratively driven to achieve accurate separation of permafrost, thawed soil, and transition zones in the feature space. The final trained model possesses excellent generalization ability and physical interpretability, and can efficiently, accurately, and automatically identify the spatial distribution of the upper limit of permafrost under different geological conditions, and quantitatively output the uncertainty of the prediction results. This provides a reliable and efficient intelligent solution for large-scale cold-region engineering exploration and geological hazard assessment, significantly improving the objectivity, accuracy, and efficiency of exploration results. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating the intelligent identification method for the upper limit of permafrost based on the fusion of multi-source data and deep clustering according to the present invention.

[0051] Figure 2 This is the probability curve of the frozen soil layer obtained in the present invention. Detailed Implementation

[0052] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] like Figure 1 As shown, one embodiment of the present invention provides an intelligent identification method for the upper limit of permafrost based on the fusion of multi-source data and deep clustering, comprising the following steps:

[0054] S1. For each measuring point on the survey line, multi-source geophysical data are acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed. At the same time, borehole data are collected and the upper limit depth of frozen soil is marked, and spatial registration is performed with the geophysical survey line.

[0055] In a further preferred embodiment, this application uses the Canadian Pulse EKKO series ground penetrating radar and the Chongqing Benteng high-density electrical resistivity tomography (RTK) instrument along the same survey line. In S1, when acquiring ground penetrating radar data and high-density RTK data, data is collected along the same survey line with the same spacing between measurement points. RTK is used to accurately locate the measurement points to ensure the synchronization of the two sets of data in terms of position.

[0056] Further preferably, in S1, multi-source geophysical data is acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed, including:

[0057] After zero-point correction, gain adjustment and filtering, the ground penetrating radar data is converted into a ground penetrating radar time profile result map.

[0058] The ground-penetrating radar time profile results are converted into depth profile results based on velocity calibration.

[0059] A resistivity profile is obtained by inverting high-density electrical resistivity data.

[0060] The locations of the measuring points in the depth profile map and the resistivity profile map are matched; the ground-penetrating radar amplitude value and high-density electrical resistivity value of each measuring point are normalized to the interval [0, 1] to accelerate model convergence.

[0061] Simultaneously, borehole data was collected and the upper limit depth of frozen soil was marked, and the upper limit depth value of frozen soil for each borehole was accurately measured and recorded. Record the precise coordinates of the borehole for spatial registration with geophysical survey lines; process the ground-penetrating radar (GPR) data and high-density electrical resistivity (HERS) data to obtain corresponding profile maps; locate the GPR and HERS measurement points corresponding to the borehole locations on the profile maps; the upper limit depth of permafrost revealed by the boreholes can be effectively used as a label in subsequent steps.

[0062] S2. The preprocessed ground-penetrating radar data and high-density electrical resistivity tomography data are divided into data blocks using a sliding window, and a label is generated for each data block.

[0063] Furthermore, the entire survey line data was divided into numerous small data blocks using a sliding window method. The window size was set to 128×32, with 128 pixels in the depth direction and 32 pixels in the horizontal direction, and a horizontal sliding step size of 1, to obtain the sample dataset. Each sample set contains a pair of registered data blocks: ground-penetrating radar data. and high-density electrical resistivity data .

[0064] Generate labels for each data block, including:

[0065] Each sample set contains a pair of registered data blocks: ground-penetrating radar data. and high-density electrical resistivity data ;

[0066] When generating labels for each data block, the following is included:

[0067] For geophysical measurement points without borehole markers, the labels corresponding to the sample data are empty, which is used for unsupervised learning;

[0068] For geophysical measuring points with borehole markings, the sample data are labeled as labeled samples; the labeled samples are represented by the binarized labels shown in the following formula. .

[0069]

[0070] in, d The value indicates the actual depth, with 0 indicating non-frozen soil and 1 indicating frozen soil.

[0071] S3. Construct a dual-branch convolutional autoencoder network to extract features from preprocessed ground-penetrating radar data and high-density electrical resistivity tomography (EDT) data, and perform unsupervised pre-training. Introduce a clustering layer on the pre-trained model and optimize the model based on the established joint loss function. The joint loss function is formed by a weighted combination of the reconstruction loss function, clustering loss function, and supervision loss function as shown in the following formula.

[0072]

[0073] in, For joint losses, This represents the total reconstruction loss. Clustering loss; To monitor losses, , and These are the weight parameters for each loss function;

[0074] Furthermore, in S3, the dual-branch convolutional autoencoder network includes:

[0075] Encoder: consists of an input layer and a fully connected layer;

[0076] The input layer has two ports, one for extracting ground-penetrating radar data. Data characteristics A tool for extracting high-density electrical resistivity data. Data characteristics ;

[0077] The input layer has two ports, which receive signals respectively. as well as The encoder design for feature extraction is as follows:

[0078] It consists of three 2D convolutional layers (Conv2D) and a max-pooling layer (Maxpooling2D), and finally outputs a one-dimensional feature vector through global average pooling. ;

[0079] It consists of three one-dimensional convolutional layers (Conv1D) and a max pooling layer (Maxpooling1D), and finally outputs a one-dimensional feature vector through global average pooling. 。

[0080] Fully connected layers are used to integrate data features and The process involves fusion and dimensionality reduction to output the final latent vector. ;

[0081] The decoder is configured as a deconvolution symmetric to the encoder, to obtain the latent vector. As input, the reconstructed ground-penetrating radar data are obtained respectively. and reconstructed high-density electrical resistivity data .

[0082] More preferably, in S3, the loss function of the dual-branch convolutional autoencoder network adopts mean squared error loss;

[0083] Calculate the reconstructed ground-penetrating radar data Compared with the original ground-penetrating radar data Mean square error and reconstructed high-density electrical resistivity data Compared with the original high-density electrical resistivity data The mean squared error is weighted and summed as the total reconstruction loss function;

[0084]

[0085] in, Original ground-penetrating radar data Or the original high-density electrical resistivity data ; For the reconstructed ground-penetrating radar data Or reconstructed high-density electrical resistivity data , where i is the detection point.

[0086] Further optimization involves introducing a clustering objective on top of the pre-trained model to optimize the feature space, giving it clear class separability. The decoder portion of the pre-trained network is removed, and the original data is processed by the encoder to output feature vectors. Next, a clustering layer is added with 3 neurons and the softmax function is selected as the activation function. (Setting the number of neurons to 3 means that there are 3 classification categories, corresponding to the three parts: cover soil, transition zone, and permafrost.)

[0087] The clustering layer is positioned after the encoder outputs the feature vector. This clustering layer is used to calculate the probability of a data point belonging to a specific cluster category using the following formula:

[0088]

[0089] in, This represents the probability that sample i belongs to cluster j; It is the feature vector extracted by the encoder for the i-th input data block; It is the center of the j-th cluster and belongs to a trainable parameter in the network; It is an eigenvector With cluster center The square of the Euclidean distance between them Let j be the reconstructed clusters.

[0090] Further optimization involves finding a better target distribution to guide network learning in order to improve clustering performance. This is achieved by introducing a target assignment P to emphasize the correlation between data points and cluster centers by increasing the probability of high-confidence assignments in Q. Specifically, the target distribution probability is calculated using the following formula: by introducing a target assignment P, the correlation between data points and cluster centers is emphasized by increasing the probability of high-confidence assignments in Q.

[0091]

[0092] in, It is the probability that sample i belongs to cluster j in the target distribution; It is the sum of all data points assigned to cluster j. To balance cluster sizes, high-confidence classifications are amplified through squaring, and by dividing by... This is to prevent a cluster from becoming too large and swallowing up other clusters.

[0093] Further preferably, the total reconstruction loss It is expressed by the following formula:

[0094]

[0095] in, and This is a weighting hyperparameter that balances the contributions of the two data sources to the total loss.

[0096] This is the clustering loss function, aiming to make the clustering assignment Q output by the network as close as possible to the calculated "better" target distribution P. This invention uses KL divergence to measure the difference between two probability distributions. The formula is shown below:

[0097]

[0098] This is the Kullback-Leibler divergence. Its value is 0 when the two distributions are identical. The greater the difference, the larger the value. This is achieved by minimizing... This makes the network's output distribution Q closer to the target distribution P.

[0099] This is a supervised loss function, used only for labeled borehole data, directly guiding the training of the clustering layer in the neural network. For each labeled data block, it takes the probability distribution Q of the clustering layer output and its true label vector. Calculate the cross-entropy loss. The formula is as follows:

[0100]

[0101] After designing the error loss function, we are ready to begin training. We extract the features Z of all samples using a pre-trained encoder, and then initialize K cluster centers using the K-Means algorithm. For unlabeled data, we only calculate... and For labeled data, calculate all three losses. , , Through total loss Backpropagation is performed, updating both encoder weights and cluster centers simultaneously. Model training stops when the change in cluster centers between two iterations is less than a threshold, or when the maximum number of iterations is reached.

[0102] S4. Using the trained model, predict the data acquired in real time along the entire survey line, output the upper limit depth of frozen soil and confidence level of each measuring point, and finally obtain the upper limit of frozen soil along the entire measurement profile.

[0103] Using a trained model, predictions are made for the entire survey line, automatically identifying the upper limit depth of permafrost. Ground-penetrating radar and high-density electrical resistivity resistivity data of the survey line are divided into blocks of the same size using sliding windows and input into the trained depth clustering model. The model outputs a probability distribution vector for each depth sampling point at the center point of each data block. , This indicates the probability that the depth represents a cover layer. This represents the probability that the depth is a transition zone. This represents the probability that the soil at that depth is frozen.

[0104] Therefore, for each geophysical measurement point, there are three probability curves that vary with depth. —Probability curve of the covering layer, —Probability curve of transition zone and —Permafrost probability curve. The physical essence of the upper limit of permafrost is the core region of the freeze-thaw transition; therefore, the probability curve of the transition zone... The peak point best represents the location of this core interface. (Regarding the curve...) Perform Savitzky-Golay filtering for smoothing and find the depth corresponding to its global maximum value. This depth is the best estimated depth for the upper limit of the permafrost.

[0105] like Figure 2 The diagram shows the probability curves for shallow permafrost. Taking a specific geophysical measuring point as an example, the diagram illustrates the changes in the probabilities of overburden, transition zone, and permafrost depths within the 0-3m range. The blue curve represents the overburden probability, which decreases rapidly from the surface, ranging from 0-1.39m. In this range, the overburden probability dominates, essentially confirming the area as overburden. As depth increases further, the overburden probability decreases sharply, while the transition zone probability increases dramatically, indicating that this area may be a transition zone from surface overburden to permafrost. When the transition zone probability reaches its peak, this is considered the location of the most drastic freeze-thaw changes, thus representing the upper limit of permafrost depth. After passing the peak, the transition zone probability gradually decreases, while the permafrost probability gradually increases, indicating the gradual transition into permafrost. The borehole at this geophysical measuring point revealed an upper limit of permafrost depth of 1.65m. Based on the probability curves in the diagram, the upper limit of permafrost is inferred to be 1.52m, with an error of 0.13m, which allows for a basic estimation of the approximate location of the upper limit of permafrost.

[0106] To further quantify the uncertainty, a confidence level (CL) is output at the finally determined upper limit of permafrost depth. The formula for the confidence level is as follows:

[0107]

[0108] Finally, the upper limit of frozen soil corresponding to each of the predicted measuring points is connected to obtain the upper limit of frozen soil for the entire measuring profile.

[0109] In summary, this application focuses on the intelligent identification of the upper limit of permafrost. Through the fusion of multi-source data (ground penetrating radar and high-density electrical resistivity tomography) and deep learning clustering algorithms, combined with borehole data annotation, it achieves automated and high-precision prediction of the upper limit of permafrost depth and outputs confidence scores to improve the objectivity and efficiency of exploration. The method proposed in this application is based on intelligent algorithms, employing a bi-branch convolutional autoencoder and deep clustering, combined with drilling results to achieve unsupervised and semi-supervised learning. Compared with existing traditional methods based on traditional comprehensive exploration, which rely heavily on human experience interpretation, have relatively linear data processing, and lack complex models, this application achieves high-precision identification through intelligent algorithms, highlighting the advantages of data-driven and automated approaches. This application represents the future direction of permafrost exploration technology. Its advantage lies in the deep integration of artificial intelligence and geophysics, which not only improves accuracy and efficiency but also provides quantifiable reliability indicators, making it particularly suitable for large-scale, high-requirement cold-region engineering projects.

[0110] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for intelligent identification of the upper limit of permafrost based on the fusion of multi-source data and deep clustering, characterized in that, Includes the following steps: S1. For each measuring point on the survey line, multi-source geophysical data are acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed. At the same time, borehole data are collected and the upper limit depth of frozen soil is marked, and spatial registration is performed with the geophysical survey line. S2. The preprocessed ground-penetrating radar data and high-density electrical resistivity tomography data are divided into data blocks using a sliding window, and a label is generated for each data block. S3. Construct a dual-branch convolutional autoencoder network to extract features from preprocessed ground-penetrating radar data and high-density electrical resistivity data, and perform unsupervised pre-training. Introduce a clustering layer on the pre-trained model; The dual-branch convolutional autoencoder network includes: Encoder: consists of an input layer and a fully connected layer; The input layer has two ports, one for extracting ground-penetrating radar data. Data characteristics A tool for extracting high-density electrical resistivity data. Data characteristics ; Fully connected layers are used to integrate data features and The process involves fusion and dimensionality reduction to output the final latent vector. ; The clustering layer is set in the latent vector output by the encoder. after; The decoder is configured as a deconvolution symmetric to the encoder, to obtain the latent vector. As input, the reconstructed ground-penetrating radar data are obtained respectively. and reconstructed high-density electrical resistivity data ; The model is optimized based on the established joint loss function; the joint loss function is formed by a weighted combination of the reconstruction loss function, clustering loss function and supervision loss function as shown in the following formula; in, For joint losses, This represents the total reconstruction loss. Clustering loss; To monitor losses, , and These are the weight parameters for each loss function; S4. Using the trained model, predict the data acquired in real time along the entire survey line, and output the upper limit depth of frozen soil and confidence level for each measuring point, thereby obtaining the upper limit of frozen soil along the entire measurement profile.

2. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 1, characterized in that, In S1, When acquiring ground-penetrating radar data and high-density electrical resistivity tomography (ERT) data, data are collected along the same survey line at the same distance between measurement points. RTK is used to accurately locate the measurement points to ensure the synchronization of the two sets of data in terms of location.

3. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 1, characterized in that, In S1, multi-source geophysical data were acquired using ground-penetrating radar and high-density electrical resistivity tomography, and preprocessed, including: After zero-point correction, gain adjustment and filtering, the ground penetrating radar data is converted into a ground penetrating radar time profile result map. The ground-penetrating radar time profile results are converted into depth profile results based on velocity calibration. A resistivity profile is obtained by inverting high-density electrical resistivity data. The locations of the measuring points in the depth profile results map and the resistivity profile results map are matched; the ground-penetrating radar amplitude value and high-density electrical resistivity value of each measuring point are normalized to the interval [0, 1].

4. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 1, characterized in that, In S2, the preprocessed ground-penetrating radar data and high-density electrical resistivity tomography (EDT) data are divided into data blocks using a sliding window, and a label is generated for each data block, including: Each sample set contains a pair of registered data blocks: ground-penetrating radar data. and high-density electrical resistivity data ; When generating labels for each data block, the following is included: For geophysical measurement points without borehole markers, the labels corresponding to the sample data are empty, which is used for unsupervised learning; For geophysical measuring points with borehole markings, the sample data are labeled as labeled samples; the labeled samples are represented by the binarized labels shown in the following formula. ; in, d The value indicates the actual depth, with 0 indicating non-frozen soil and 1 indicating frozen soil.

5. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 1, characterized in that, In S3, the reconstruction loss function of the dual-branch convolutional autoencoder network adopts the mean squared error loss. Calculate the reconstructed ground-penetrating radar data Compared with the original ground-penetrating radar data Mean square error and reconstructed high-density electrical resistivity data Compared with the original high-density electrical resistivity data The mean squared error is weighted and summed as the total reconstruction loss function; in, Original ground-penetrating radar data Or the original high-density electrical resistivity data ; For the reconstructed ground-penetrating radar data Or reconstructed high-density electrical resistivity data , where i is the detection point.

6. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 1, characterized in that, The clustering layer is used to calculate the cluster probability of a data point belonging to a certain cluster category using the following formula. in, This represents the probability that sample i belongs to cluster j; It is the feature vector extracted by the encoder for the i-th input data block; It is the center of the j-th cluster and belongs to a trainable parameter in the network; It is an eigenvector With cluster center The square of the Euclidean distance between them Let j be the reconstructed clusters.

7. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 6, characterized in that, It also includes calculating the target distribution probability according to the following formula, and emphasizing the correlation between data points and cluster centers by introducing the target assignment P to increase the probability of high confidence assignments in Q; in, It is the probability that sample i belongs to cluster j in the target distribution; It is the sum of all data points assigned to cluster j. To balance cluster sizes, high-confidence classifications are amplified through squaring, and by dividing by... This is to prevent a cluster from becoming too large and swallowing up other clusters.

8. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 7, characterized in that, The clustering loss function is expressed by the following formula: This is the Kullback-Leibler divergence, which has a value of 0 when the two distributions are identical; the greater the difference, the larger the value; by minimizing... This makes the network's output distribution Q closer to the target distribution P.

9. The intelligent identification method for the upper limit of permafrost based on multi-source data and deep clustering fusion as described in claim 8, characterized in that, The total reconstruction loss It is expressed by the following formula: in, and This is a weighting hyperparameter that balances the contributions of the two data sources to the total loss.

Citation Information

Patent Citations

  • Method, device and equipment for surveying frozen soil in high-latitude forest region and medium

    CN116597321A

  • Unsupervised radar signal sorting method based on deep clustering

    CN113971440A

  • Deep convolution embedding clustering method based on Resnet50 improved auto-encoder, storage medium and terminal equipment

    CN118115769A