Multi-source data stratigraphic stratigraphic prediction method based on random field and deep learning

By adopting a multi-source data stratification prediction method based on random field and deep learning in geotechnical engineering, the problem of uncertainty in the strata division process is solved, more efficient and robust stratification prediction is achieved, and the identification ability of high-risk areas is improved.

CN120030884AActive Publication Date: 2025-05-23SHANGHAI GEOTECHN INVESTIGATIONS & DESIGN INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510085932.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing technology has great uncertainty in the process of strata division in geotechnical engineering, especially when the on-site survey data is limited, which may lead to unforeseeable problems during the construction process, affecting the progress of the project and the performance of the geotechnical structure.

Method used

A multi-source data stratification prediction method based on random field and deep learning is adopted. By extracting drilling data, static exploration data and hierarchical information from geotechnical engineering survey database, continuous random field of drilling physical and mechanical parameters and continuous random field of static exploration PS value are constructed, and a Transformer-LSTM-CRF hybrid deep learning model is used to accurately predict stratigraphic stratification and boundary accuracy. At the same time, combined with Monte Carlo random sampling, the uncertainty of continuous random fields is introduced to provide confidence quantification of stratified results.

Benefits of technology

It significantly improves the applicability of stratigraphic stratification prediction, capture ability of nonlinear relationships and quantify uncertainties, provides a more robust and efficient solution for soil layer division, and improves the ability to identify high-risk areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030884A_ABST
    Figure CN120030884A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data stratigraphic stratigraphic prediction method based on random fields and deep learning, and the method comprises the steps: extracting drilling data, static sounding data and stratigraphic information, capturing a depth trend through an HGB, fitting a complex nonlinear residual error between a predicted value and a true value of an HGB model in combination with a Gaussian process, and finally generating two continuous random fields; based on the generated continuous random field, a multi-task deep learning model is constructed, deep semantic features are extracted by using Transform, LSTM modeling depth direction long-distance dependence, a multi-head attention mechanism are combined to reinforce local and global feature learning, and a CRF layer accurately optimizes layered boundary prediction, so that unification of stratum layering and accurate boundary prediction is realized. Random field uncertainty is introduced through Monte Carlo random sampling, and confidence calculation and risk quantification of stratigraphic stratification are achieved. The method has the advantages that stratum characteristics and boundaries are accurately captured, and a stable and reliable stratum layering solution is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geotechnical engineering strata stratification, and in particular to a multi-source data strata stratification prediction method based on random fields and deep learning. Background Art

[0002] Stratigraphic division plays an important fundamental role in geotechnical engineering construction. It is not only the key to designing foundations and supporting structures, but also provides an important basis for soil behavior and deformation analysis. However, the stratigraphic division process is usually accompanied by large uncertainties, especially when field survey data is limited. If the uncertainty of stratigraphic division and classification is not handled properly, it may lead to unforeseen problems during construction, thus affecting the progress of the project and the performance of geotechnical structures. A survey of 28 construction projects in the UK showed that 22% of geotechnical engineering problems were related to the uncertainty of stratigraphic division, highlighting the need to reasonably evaluate and quantify stratification uncertainty.

[0003] CN113945975A discloses a method for joint inversion of stratigraphic structure based on Love wave and Rayleigh wave, which belongs to the field of seismic exploration. The technical scheme is to convert the obtained surface wave seismic record to obtain Love wave and Rayleigh wave dispersion curves respectively, and bring the dispersion curves into the target equation for joint inversion. Compared with the prior art, the present invention has the following obvious advantages: the comprehensive use of Love wave and Rayleigh wave signals greatly improves the inversion accuracy and effectively curbs the multi-solution of the inversion problem. The inversion result of the method of the present invention matches well with the real model of the stratigraphic layer. It is completely different from the data and method used in this application, and only one kind of data is used for stratigraphic layering, and more comprehensive information of the stratigraphic layer is not integrated.

[0004] CN 118131344 A discloses a method for stratifying a survey area, comprising the steps of: S10, determining the resistivity of the strata in the survey area according to the electromagnetic observation point data collected in the survey area, determining the preliminary stratification of the strata in the survey area and the upper and lower interfaces of each layer according to the resistivity logging data or geological drilling lithology cataloging data of the survey area; S20, dividing the stratum into a plurality of unit blocks, and filling the unit blocks in the stratum with different colors according to the resistivity values ​​according to the resistivity values ​​of the strata determined in S10, so as to form a resistivity mosaic profile; S30, determining the resistivity values ​​of the upper and lower interfaces of the resistivity mosaic profile according to the preliminary stratification and the resistivity mosaic profile; S40, determining the stratification of the resistivity mosaic profile according to the resistivity values ​​of the upper and lower interfaces, i.e., the stratification of the survey area. The method of the embodiment of the present invention can improve the accuracy and reliability of stratum stratification. This method is completely different from the method used in this application, and the only data source used is resistivity, without using multi-source fusion data to comprehensively characterize the soil layer.

[0005] CN114880950A discloses a soil layer quantitative stratification method, device, equipment and medium based on XGBoost pore pressure static penetration data, the method comprising the following steps: S1, collection and arrangement of pore pressure static penetration data and soil layer classification information; S2, conversion of original static penetration data; S3, establishment of a pore pressure static penetration soil layer quantitative stratification XGBoost prediction model using a set of debugged optimal hyperparameters; S4, training the soil layer quantitative stratification XGBoost prediction model; S5, using the trained soil layer quantitative stratification XGBoost prediction model to predict the soil layer type of another site, using the softprob objective function to output the probability of the soil layer corresponding to each soil type; S6, determining the division accuracy, merging the stratification results, and finally obtaining the soil layer quantitative stratification results. This method only makes a very rough stratification, and does not divide into sub-layers. It relies on only one data source for prediction, and does not use multi-source data for stratification. The uncertainty of the formation random field is not calculated, the confidence is not calculated, and the reliability of the prediction is questionable. This method only discloses a small part of the technical features in this project plan, and the main technical principles and details are different.

[0006] Previous soil layer divisions mostly used single data and did not use deep learning to explore the relationship between features and soil layers, so previous stratigraphic stratification methods lacked regional adaptability. Summary of the invention

[0007] The purpose of the present invention is to provide a multi-source data stratigraphic stratification prediction method based on random fields and deep learning according to the deficiencies of the above-mentioned prior art. First, by extracting drilling data, static exploration data and stratification information from a geotechnical engineering survey database, two continuous random fields are constructed: one captures the overall trend of drilling data along the depth through HGB, and combines Gaussian process to model nonlinear residuals to generate a continuous random field of drilling physical and mechanical parameters; the other uses static exploration data to generate a continuous random field of static exploration PS values ​​through optimized HGB and Gaussian process to make up for the problem of missing static exploration data at the drilling position. Then, based on the data input of these two continuous random fields, a Transformer-LSTM-CRF hybrid deep learning model is constructed. Transformer is used to capture the contextual semantic features of drilling and static exploration data, LSTM further models long-distance dependencies in the depth direction, and the multi-head attention mechanism enhances the model's learning ability for local and global features. The CRF layer optimizes the stratification boundary prediction, thereby achieving the unification of stratigraphic stratification and accurate boundary prediction. At the same time, this method combines Monte Carlo random sampling to introduce the uncertainty of continuous random fields into the model prediction process, providing confidence quantification for stratification results and improving the ability to identify high-risk areas. Compared with the traditional Roberson method and hidden Markov method (HMM) based on static exploration data, this method significantly improves the applicability of stratigraphic stratification prediction, the ability to capture nonlinear relationships, and the ability to quantify uncertainty through multi-source data fusion and deep modeling, providing a more robust and efficient solution for soil stratification, which is very important in geotechnical engineering design and construction.

[0008] The purpose of the present invention is achieved by the following technical solutions:

[0009] A method for predicting stratigraphic stratification using multi-source data based on random fields and deep learning, the method comprising the following steps:

[0010] S1: Acquire drilling data, static exploration data and stratification data from the geotechnical engineering investigation database, and generate drilling data table, static exploration data table and stratification data table;

[0011] S2: Based on the drilling data, a continuous random field of drilling physical and mechanical parameters is generated through the HGB model and Gaussian process;

[0012] S3: Based on the static sounding data, a continuous random field of static sounding PS values ​​is generated through the HGB model and Gaussian process;

[0013] S4: Based on the continuous random field of the drilling physical and mechanical parameters and the continuous random field of the static exploration PS value, the Transformer model is used to capture the contextual semantic features of the drilling data and the static exploration data, the bidirectional LSTM layer is used to model the long-distance dependency in the depth direction, the multi-head attention mechanism is used to enhance the model's learning ability for local and global features, and the CRF layer is used to optimize the stratification boundary prediction, thereby achieving the unification of stratum stratification and accurate boundary prediction;

[0014] S5: using Monte Carlo random sampling to introduce the uncertainty of the continuous random field into the model prediction process, providing confidence quantification for the stratified results;

[0015] S6: Use evaluation indicators for soil layer division and visual plotting to evaluate model training and prediction results.

[0016] In step S1, the StandardScaler function is used to pre-process the drilling data, the static exploration data and the layered data to ensure that the feature scales are consistent.

[0017] In step S2, the specific steps are as follows:

[0018] S2.1: Using the hyperparameter-optimized HGB model to capture the trend of the drilling physical and mechanical parameters along the depth;

[0019] S2.2: Use Gaussian process to capture the residual between the trend of the drilling physical and mechanical parameters along the depth and the true value.

[0020] In step S3, the specific steps are as follows:

[0021] S3.1: Use the HGB model with optimized hyperparameters to capture the trend of the static PS value along the depth;

[0022] S3.2: For each depth point of the static exploration data, the nearest drilling point is selected to construct a Kriging model, and the residual of each depth is compensated by a Gaussian process to capture the geological characteristics of different depth levels, and the overall prediction results are obtained by combining them.

[0023] In step S4, the specific steps are as follows:

[0024] S4.1: aligning the continuous random field of the drilling physical and mechanical parameters generated in step S2 and the continuous random field of the static sounding PS value generated in step S3, and matching the drilling physical and mechanical parameters, the static sounding PS value and the layered data according to the drilling position and depth information to ensure the consistency and integrity of the data;

[0025] S4.2: standardizing the numerical features of the drilling data and the static exploration data and converting them into character strings; performing two-step processing on the layered data, firstly mapping the stratum labels, and then further standardizing the individual geological layer labels to reduce the diversity and redundancy of the labels;

[0026] S4.3: Encode the drilling data and the static exploration data using a pre-trained Transformer model to obtain a contextual embedding representation of the input sequence to extract deep semantic information and capture the contextual relationship of geological parameters;

[0027] S4.4: Add a bidirectional LSTM layer to further extract the contextual information of the sequence and capture the long-distance dependencies in depth;

[0028] S4.5: A multi-head attention mechanism is introduced to enable the model to better capture the diversity and complexity of geological data by focusing on different features in different attention heads;

[0029] S4.6: Add layer normalization and batch normalization to ensure the stability of data distribution during feature extraction at different levels and reduce the gradient vanishing problem;

[0030] S4.7: Add a boundary prediction module, use a linear layer to map the output of the bidirectional LSTM layer to the emission score of the boundary label, use a CRF layer to perform sequence labeling on the boundary label, and capture the dependency relationship between the boundary and the drilling data and the static exploration data;

[0031] S4.8: Introduce L2 regularization and Dropout mechanism to prevent model overfitting and ensure the generalization ability of the model on different geological data sets;

[0032] S4.9: Combine classification loss, CRF loss and L2 regularization to obtain the loss function;

[0033] S4.10: The drilling number division strategy is used when dividing the data set to improve the generalization ability of the model and avoid overfitting;

[0034] S4.11: Use Gradient Clipping to control the gradient norm and avoid gradient explosion; introduce ReduceLROnPlateau learning rate scheduler to automatically reduce the learning rate when the validation loss does not decrease significantly, so as to achieve more refined model optimization and improve the convergence ability of the model in the later training stage;

[0035] S4.12: During the model training process, record the model status and validation loss of the last few training rounds, compare them, and select the optimal model status to save, thus ensuring the final performance of the model.

[0036] In step S5, the specific steps are as follows:

[0037] Each test example is sampled multiple times to simulate different possibilities of the input features, estimate the model prediction distribution, and calculate the prediction probability of each category, thereby providing uncertainty measures at different levels.

[0038] In step S6, the specific steps are as follows:

[0039] The f1 value, accuracy, and recall rate are used to evaluate the model prediction effect, and the loss function curve with the number of training rounds, ROC curve, confusion matrix, and prediction measurement with depth distribution are used to evaluate the model prediction results.

[0040] The advantages of the present invention are:

[0041] 1) Advantages of multi-source data fusion: Roberson method and HMM usually perform stratigraphic stratification based only on static exploration data, and it is difficult to comprehensively consider information from other data sources, which may cause the model to ignore the physical and mechanical parameters contained in the drilling data (such as density, porosity, shear strength, etc.), thereby limiting the comprehensive utilization of information related to soil layer division; this method combines drilling data with static exploration data, and realizes the alignment and complementarity of multi-source data by constructing two continuous random fields (drilling physical and mechanical parameter random field and static exploration PS value random field), thereby improving the ability to characterize soil layer properties and fully utilizing effective information for soil layer division;

[0042] 2) The separation of trends and residuals enables the model to learn both macro trends and micro local changes, which has better fitting ability and interpretability;

[0043] 3) Parameter search is performed through Bayesian optimization, which improves the efficiency of hyperparameter selection and avoids the blindness of manual parameter adjustment, especially for models with large parameter space;

[0044] 4) The uncertainty range of each prediction point is given through the standard deviation of the Gaussian process, which can clearly characterize the confidence level of the model at different depths and help understand the reliability of the prediction; this uncertainty estimation based on the prediction standard deviation provides an important reference for geological engineering decision-making, allowing engineers to better weigh risks;

[0045] 5) The GEO-MULTI underground space multi-source data multi-index prediction model obtained through supervised training was used to review the original data. By comparing the prediction results of the model with the measured values, the statistical box plot method was used to identify potential outliers based on outliers;

[0046] 6) The results show that the overall accuracy of classification prediction is 1.00, which achieves very good prediction results; the regression prediction R 2 Between 0.91 and 0.92, a relatively ideal prediction accuracy is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flow chart of the method for predicting stratigraphic stratification based on multi-source data based on random field and deep learning of the present invention;

[0048] Figure 2 It is a schematic diagram of the overall structure of the deep learning model of the present invention;

[0049] Figure 3 It is a distribution diagram of the drilling holes and static exploration holes of the present invention;

[0050] Figure 4 Schematic diagram of the compression modulus continuous random field of the present invention;

[0051] Figure 5 It is a schematic diagram of a heavy continuous random field of the present invention;

[0052] Figure 6 Schematic diagram of the continuous random field of cohesion force of the present invention;

[0053] Figure 7 Schematic diagram of the continuous random field of the internal friction angle of the present invention;

[0054] Figure 8 It is a schematic diagram of the continuous random field of water content of the present invention;

[0055] Fig. 9 It is a schematic diagram of the continuous random field of statically detecting PS value of the present invention;

[0056] Fig.10 A schematic diagram of the confusion matrix of the predicted classification and the measured classification of the present invention;

[0057] Fig.11 It is a schematic diagram of the ROC curve of the classification prediction results of each soil layer of the present invention;

[0058] Fig.12 A schematic diagram of the comparison between the predicted soil layer and the artificial soil layer and the prediction confidence of the drilling method of the present invention;

[0059] Fig.13It is a table diagram of the calculation results of the prediction evaluation indexes for each stratum category of the present invention. DETAILED DESCRIPTION

[0060] The features of the present invention and other related features are further described in detail below through embodiments in conjunction with the accompanying drawings to facilitate understanding by those skilled in the art:

[0061] Example: Figure 1 As shown, this embodiment relates to a multi-source data stratigraphic layer prediction method based on random field and deep learning, and the method mainly includes the following steps:

[0062] S1: Figure 3 As shown in the figure, the original drilling data (soil layer parameters), static exploration data (soil layer parameters) and stratification data of each survey project are obtained from the geotechnical engineering survey database, and the drilling data table, static exploration data table and stratification data table are generated. The StandardScaler function is used to preprocess the drilling data, static exploration data and stratification data to ensure the consistency of feature scales.

[0063] Among them, the drilling data include compression modulus, density, cohesion, internal friction angle, water content and coordinates of the drilling hole, and the static exploration data include static exploration PS value and coordinates of the static exploration hole.

[0064] S2: Figure 2 as well as Figures 4 to 8 As shown, based on the drilling data, the continuous random field of drilling physical and mechanical parameters is generated through the HGB (HistGradientBoosting) model and the Gaussian Process (GP), that is, the continuous random field of drilling physical and mechanical parameters with continuous distribution in the depth direction is generated by utilizing the non-continuous drilling data in the depth direction.

[0065] In step S2, the specific steps are as follows:

[0066] S2.1: Use the HGB model with optimized hyperparameters to capture the trends of drilling physical and mechanical parameters along the depth.

[0067] Among them, the HGB model learns complex functions through incremental construction of trees, which is suitable for capturing the main trends of data. At the same time, for drilling data, the Optuna hyperparameter optimization framework is used to optimize the hyperparameters of the HGB model. This optimization method is used to determine the optimal hyperparameters to ensure that the trend model achieves optimal performance on drilling data.

[0068] S2.2: Use Gaussian process to capture the residual between the trend of drilling physical and mechanical parameters along the depth and the true value.

[0069] Among them, the parameter search is performed through Bayesian optimization, which improves the efficiency of hyperparameter selection and avoids the blindness of manual parameter adjustment, especially for models with large parameter space. In view of the nonlinear characteristics of drilling data, different kernel functions are used to capture more complex nonlinear local features to ensure that the Gaussian process can capture complex local residuals that the trend model cannot explain.

[0070] Specifically, the HGB model learns complex functions through incremental construction of trees:

[0071]

[0072] Where M is the total number of weak learners (regression trees), η is the learning rate, which controls the contribution of each tree and is optimized through the Optuna hyperparameter optimization framework. m (x) is the output of the mth tree, i.e., the main trend of soil layer parameters.

[0073] Hyperparameter optimization of the HGB model: learning_rate, the value range is [0.3, 0.5]; max_depth, controls the maximum depth of the tree, used to limit the model complexity, the value range is [15, 20]; min_samples_leaf, the minimum number of samples for each leaf node, the value range is [20, 25]; max_iter, the number of iterations (i.e. the maximum number of trees), the value range is [500, 5000]; l2_regularization, L2 regularization coefficient, used to prevent overfitting, the value range is [0.0, 1.0].

[0074] In view of the nonlinear characteristics of drilling data, different kernel functions are used to capture more complex nonlinear local features to ensure that the Gaussian process can capture the complex local residuals that the trend model cannot explain:

[0075]

[0076] Where K is the kernel matrix, which is composed of kernel functions.

[0077] This embodiment uses the following kernel function combination:

[0078] k(x,x')=C·RBF(x,x')+WhiteKernel(x,x');

[0079] ConstantKernel, used to define the global amplitude of the predicted value, and RBF radial basis function kernel, used to capture the smooth correlation of input features:

[0080]

[0081] Where ||x-x'|| is the Euclidean distance of the input feature, and length_scale is the length scale, which determines the decay rate of the correlation.

[0082] WhiteKernel is used to represent the observation noise. The parameters of the Gaussian process are optimized using maximum likelihood estimation (MLE).

[0083] S3: Figure 2 and Fig. 9 As shown, based on the static exploration data, a continuous random field of static exploration PS values ​​is generated through the HGB model and Gaussian process, that is, the non-continuous static exploration data in the horizontal direction is used to generate a continuous random field of static exploration PS values ​​continuously distributed in the horizontal direction at each depth, which makes up for the problem of missing static exploration data at the drilling position, so as to better realize the fusion and alignment of multi-source data.

[0084] In step S3, the specific steps are as follows:

[0085] S3.1: Use the HGB model with optimized hyperparameters to capture the trend of static PS values ​​along depth.

[0086] Among them, the Optuna hyperparameter optimization framework is used to optimize the hyperparameters of HGB based on static detection data. In the process of optimizing the HGB model, KFold cross-validation is used in combination with parallel computing (through joblib's Parallel) to ensure the robustness and generalization ability of hyperparameters while greatly accelerating the model training speed according to the characteristics of static detection data.

[0087] S3.2: For each depth point of the static exploration data, the nearest drilling point is selected to construct the Kriging model. The residual of each depth is compensated by Gaussian process to capture the geological characteristics of different depth levels and combine them to obtain the overall prediction results.

[0088] The Gaussian process model is trained by using the Adam optimizer and using ExactMarginalLogLikelihood (used to calculate the marginal log-likelihood of the Gaussian process) as the loss function. This method enables the model to better fit the nonlinear residuals of static exploration data and show strong adaptability in strata with different depths.

[0089] S4: Figure 2As shown in the figure, based on the continuous random field of drilling physical and mechanical parameters and the continuous random field of static exploration PS values, the Transformer model is used to capture the contextual semantic features of drilling data and static exploration data, the bidirectional LSTM (Long Short-Term Memory) layer is used to model the long-distance dependencies in the depth direction, the multi-head attention mechanism is used to enhance the model's learning ability for local and global features, and the CRF (Conditional Random Field Layer) layer is used to optimize the stratification boundary prediction, thereby realizing the unification of stratigraphic stratification and precise boundary prediction.

[0090] In step S4, the specific steps are as follows:

[0091] S4.1: Align the continuous random field of drilling physical and mechanical parameters generated in step S2 and the continuous random field of static sounding PS values ​​generated in step S3. According to the drilling position and depth information, match the drilling physical and mechanical parameters, static sounding PS values ​​and layered data to ensure the consistency and integrity of the data.

[0092] S4.2: The numerical features of drilling data and static exploration data are standardized and converted into strings. For layered data, a two-step process is performed, firstly the stratigraphic labels are preliminarily mapped, and then the individual geological layer labels are further standardized to reduce the diversity and redundancy of the labels.

[0093] S4.3: Use the pre-trained Transformer model to encode the drilling data and static exploration data to obtain the contextual embedding representation of the input sequence to extract deep semantic information and capture the contextual relationship of geological parameters. The output of the Transformer model is used as the input of the subsequent layers.

[0094] S4.4: Add a bidirectional LSTM layer to further extract the contextual information of the sequence and capture the long-distance dependencies in depth. The bidirectional LSTM layer can effectively process the relationship between adjacent time points in the sequence and enhance the understanding of the strata.

[0095] S4.5: A multi-head attention mechanism is introduced to enable the model to better capture the diversity and complexity of geological data by focusing on different features at different attention heads. Combining the multi-head attention mechanism with the bidirectional LSTM layer enhances the balance between capturing global information and local features.

[0096] S4.6: Add layer normalization and batch normalization to ensure the stability of data distribution during feature extraction at different levels and reduce the gradient vanishing problem.

[0097] S4.7: Add a boundary prediction module, use a linear layer to map the output of the bidirectional LSTM layer to the emission score of the boundary label (used to reflect the model's confidence in each label), and use a CRF layer to sequence the boundary labels to capture the dependency between the boundary and the drilling data and static exploration data.

[0098] S4.8: Introduce L2 regularization and Dropout mechanism to prevent model overfitting and ensure the generalization ability of the model on different geological data sets.

[0099] S4.9: Combine the classification loss, CRF loss and L2 regularization to obtain the loss function.

[0100] S4.10: The drill hole number partitioning strategy is adopted when dividing the data set. This ensures that the model does not see the same drill hole data in the training and test sets, thereby improving the generalization ability of the model and avoiding overfitting.

[0101] S4.11: Use Gradient Clipping to control the gradient norm and avoid gradient explosion, especially when training deep models, to ensure the stability of the model training process; introduce the ReduceLROnPlateau learning rate scheduler, which automatically reduces the learning rate when the validation loss does not decrease significantly, to achieve more refined model optimization and improve the convergence ability of the model in the later training stage.

[0102] S4.12: During model training, record the model status and validation loss of the last few training epochs, compare them, and select the optimal model status to save, thus ensuring the final performance of the model.

[0103] S5: Monte Carlo random sampling is used to introduce the uncertainty of continuous random fields into the model prediction process, providing confidence quantification for the stratified results.

[0104] In step S5, the specific steps are as follows:

[0105] Each test sample is sampled multiple times to simulate different possibilities of input features, estimate the model prediction distribution, and calculate the prediction probability of each category, thereby providing uncertainty measures for different layer numbers. This probability distribution can not only identify the most likely layer number, but also reflect the degree of uncertainty of the model in the formation classification, helping engineers to conduct risk assessment and decision optimization.

[0106] S6: Use evaluation indicators for soil layer division and visual plotting to evaluate model training and prediction results.

[0107] In step S6, the specific steps are as follows:

[0108] The f1 value, accuracy, and recall rate are used to evaluate the model prediction effect, and the loss function curve with the number of training rounds, ROC curve, confusion matrix, and predicted measured depth distribution are used to evaluate the model prediction results. The confusion matrix of the predicted soil name is as follows: Fig.10 The ROC curves of the prediction results of each soil layer classification are shown in Fig.11 The classification prediction accuracy evaluation score bar chart of each soil layer is shown in Fig.12 shown.

[0109] like Fig.13 As shown in the figure, the weighted average calculation results of the prediction evaluation indicators of each layer category are: accuracy rate is 0.91, precision rate is 0.82, recall rate is 0.85, and f1 value is 0.83.

[0110] Specifically, as a traditional spatial statistical method, Kriging interpolation is often used for spatial interpolation of data from limited survey points (such as boreholes and static penetration points), but it has significant limitations. The Kriging method mainly assumes the stationarity and linear correlation of variables, so it performs poorly when dealing with areas with complex geological conditions or significant nonlinear changes. In addition, Kriging interpolation focuses more on generating "optimal estimates" and is relatively weak in quantifying and predicting interpolation uncertainties, making it difficult to meet the needs of probabilistic analysis. More importantly, Kriging interpolation only relies on spatial location correlation and fails to effectively combine regional historical survey experience and prior knowledge, thus limiting its prediction accuracy.

[0111] To make up for these shortcomings, the kriging method combining machine learning technology and Gaussian process has become a more promising alternative. This method can generate more refined random fields and achieve more accurate quantification and prediction of uncertainty in stratigraphic division by leveraging the ability of machine learning to capture nonlinear relationships and the ability of Gaussian process to model local features.

[0112] In stratified prediction, relying solely on static sounding data (such as PS values) will have certain limitations. Although static sounding data are densely distributed, they mainly reflect the resistance characteristics of the soil layer and cannot directly characterize the physical and mechanical properties. Drilling data contains key parameters such as density, porosity and shear strength, which can provide a more comprehensive characterization of soil layer characteristics. In addition, the empirical classification charts of static sounding data (such as SBT charts) are developed based on global data and may not perform well in specific sites, while drilling data can provide regional information to optimize the stratification results. Therefore, combining drilling data and static sounding data, constructing a continuous random field, and aligning the two can significantly improve the accuracy of the model's characterization of stratum characteristics.

[0113] Combining machine learning and Kriging Gaussian process, two continuous random fields can be constructed respectively: one is the physical and mechanical parameter random field based on drilling data, which captures the overall trend in the depth direction through HistGB and combines Gaussian process to compensate for nonlinear residuals, thereby aligning with the layered data; the other is the PS value random field based on static exploration data, which predicts the static exploration value of the drilling position through static exploration data, thereby making up for the problem of missing static exploration data at the drilling point. Ultimately, the combination of the two random fields can generate a complete drilling-static exploration data set, which, combined with the stratigraphic layering data, forms a comprehensive characterization of the stratigraphic formation.

[0114] In order to achieve more accurate stratigraphic stratification prediction, the Transformer-LSTM-CRF model based on deep learning is introduced. Transformer can capture the implicit complex relationships in drilling data and static exploration data through context embedding and extract deep semantic information; LSTM further models long-distance dependencies in the depth direction, which is particularly suitable for reflecting the continuous changes of strata along the depth; CRF can effectively solve the problem of fuzzy stratigraphic boundaries by globally optimizing sequence annotations, and significantly improve the prediction accuracy of stratification boundaries. The accuracy of stratification boundaries is crucial for stratigraphic division, because only accurate boundaries can ensure correct stratification results.

[0115] In addition, in the prediction process based on Monte Carlo random sampling, the Transformer-LSTM-CRF hybrid model showed superior capabilities. The model can capture the residual of the input continuous random field, reflect the uncertainty of the continuous random field in the layered prediction, provide prediction probability for each category, and calculate the confidence of the final layered result as the depth changes. This not only helps to identify the most likely layer number, but also quantifies the uncertainty of the model in the formation classification, providing engineers with a basis for decision optimization.

[0116] Capturing uncertainty is of great significance for soil layer division. On the one hand, it can reduce construction risks and avoid unexpected problems caused by incorrect soil layer division; on the other hand, it improves the robustness of the model under different site conditions, making the prediction results more reliable. In addition, the quantification of uncertainty provides key input for probabilistic analysis, which helps to further conduct reliability analysis in slope stability assessment or foundation design.

[0117] In summary, combining drilling and static exploration data, generating continuous random fields through machine learning and Kriging Gaussian process, and using Transformer-LSTM-CRF model for stratified prediction not only significantly improves the accuracy and reliability of soil layer division, but also provides a powerful tool for quantifying uncertainty in stratum division. This method can comprehensively optimize the decision-making process in geotechnical engineering design and construction, laying a solid foundation for improving engineering quality and reducing risks.

[0118] The beneficial technical effects of this embodiment are:

[0119] 1) Advantages of multi-source data fusion: Roberson method and HMM usually perform stratigraphic stratification based only on static exploration data, and it is difficult to comprehensively consider information from other data sources, which may cause the model to ignore the physical and mechanical parameters contained in the drilling data (such as density, porosity, shear strength, etc.), thereby limiting the comprehensive utilization of information related to soil layer division; this method combines drilling data with static exploration data, and realizes the alignment and complementarity of multi-source data by constructing two continuous random fields (drilling physical and mechanical parameter random field and static exploration PS value random field), thereby improving the ability to characterize soil layer properties and fully utilizing effective information for soil layer division;

[0120] 2) The separation of trends and residuals enables the model to learn both macro trends and micro local changes, which has better fitting ability and interpretability;

[0121] 3) Parameter search is performed through Bayesian optimization, which improves the efficiency of hyperparameter selection and avoids the blindness of manual parameter adjustment, especially for models with large parameter space;

[0122] 4) The uncertainty range of each prediction point is given through the standard deviation of the Gaussian process, which can clearly characterize the confidence level of the model at different depths and help understand the reliability of the prediction; this uncertainty estimation based on the prediction standard deviation provides an important reference for geological engineering decision-making, allowing engineers to better weigh risks;

[0123] 5) The GEO-MULTI underground space multi-source data multi-index prediction model obtained through supervised training was used to review the original data. The statistical box plot method was used to compare the prediction results of the model with the measured values, and potential outliers were identified based on outliers;

[0124] 6) The results show that the overall accuracy of classification prediction is 1.00, which achieves very good prediction results; the regression prediction R 2 Between 0.91 and 0.92, a relatively ideal prediction accuracy is obtained.

[0125] Although the above embodiments have described in detail the concepts and embodiments of the present invention with reference to the accompanying drawings, ordinary technicians in this field can recognize that various improvements and changes can still be made to the present invention without departing from the scope of the claims, so they are not described one by one here.

Claims

1. A multi-source data stratigraphic layer prediction method based on random field and deep learning, characterized in that The method comprises the following steps: S1: Acquire drilling data, static exploration data and stratification data from the geotechnical engineering investigation database, and generate drilling data table, static exploration data table and stratification data table; S2: Based on the drilling data, a continuous random field of drilling physical and mechanical parameters is generated through the HGB model and Gaussian process; S3: Based on the static sounding data, a continuous random field of static sounding PS values ​​is generated through the HGB model and Gaussian process; S4: Based on the continuous random field of the drilling physical and mechanical parameters and the continuous random field of the static exploration PS value, the Transformer model is used to capture the contextual semantic features of the drilling data and the static exploration data, the bidirectional LSTM layer is used to model the long-distance dependency in the depth direction, the multi-head attention mechanism is used to enhance the model's learning ability for local and global features, and the CRF layer is used to optimize the stratification boundary prediction, thereby achieving the unification of stratum stratification and accurate boundary prediction; S5: using Monte Carlo random sampling to introduce the uncertainty of the continuous random field into the model prediction process, providing confidence quantification for the stratified results; S6: Use evaluation indicators for soil layer division and visual plotting to evaluate model training and prediction results.

2. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 1, characterized in that In step S1, the StandardScaler function is used to pre-process the drilling data, the static exploration data and the layered data to ensure that the feature scales are consistent.

3. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 1, characterized in that In step S2, the specific steps are as follows: S2.1: Using the hyperparameter-optimized HGB model to capture the trend of the drilling physical and mechanical parameters along the depth; S2.2: Use Gaussian process to capture the residual between the trend of the drilling physical and mechanical parameters along the depth and the true value.

4. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 3, characterized in that In step S3, the specific steps are as follows: S3.1: Use the HGB model with optimized hyperparameters to capture the trend of the static PS value along the depth; S3.2: For each depth point of the static exploration data, the nearest drilling point is selected to construct a Kriging model, and the residual of each depth is compensated by a Gaussian process to capture the geological characteristics of different depth levels, and the overall prediction results are obtained by combining them.

5. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 4, characterized in that In step S4, the specific steps are as follows: S4.1: aligning the continuous random field of the drilling physical and mechanical parameters generated in step S2 and the continuous random field of the static sounding PS value generated in step S3, and matching the drilling physical and mechanical parameters, the static sounding PS value and the layered data according to the drilling position and depth information to ensure the consistency and integrity of the data; S4.2: standardizing the numerical features of the drilling data and the static exploration data and converting them into character strings; performing two-step processing on the layered data, firstly mapping the stratum labels, and then further standardizing the individual geological layer labels to reduce the diversity and redundancy of the labels; S4.3: Encode the drilling data and the static exploration data using a pre-trained Transformer model to obtain a contextual embedding representation of the input sequence to extract deep semantic information and capture the contextual relationship of geological parameters; S4.4: Add a bidirectional LSTM layer to further extract the contextual information of the sequence and capture the long-distance dependencies in depth; S4.5: A multi-head attention mechanism is introduced to enable the model to better capture the diversity and complexity of geological data by focusing on different features in different attention heads; S4.6: Add layer normalization and batch normalization to ensure the stability of data distribution during feature extraction at different levels and reduce the gradient vanishing problem; S4.7: Add a boundary prediction module, use a linear layer to map the output of the bidirectional LSTM layer to the emission score of the boundary label, use a CRF layer to perform sequence labeling on the boundary label, and capture the dependency relationship between the boundary and the drilling data and the static exploration data; S4.8: Introduce L2 regularization and Dropout mechanism to prevent model overfitting and ensure the generalization ability of the model on different geological data sets; S4.9: Combine classification loss, CRF loss and L2 regularization to obtain the loss function; S4.10: The drilling number division strategy is used when dividing the data set to improve the generalization ability of the model and avoid overfitting; S4.11: Use Gradient Clipping to control the gradient norm and avoid gradient explosion; introduce ReduceLROnPlateau learning rate scheduler to automatically reduce the learning rate when the validation loss does not decrease significantly, so as to achieve more refined model optimization and improve the convergence ability of the model in the later training stage; S4.12: During the model training process, record the model status and validation loss of the last few training rounds, compare them, and select the optimal model status to save, thus ensuring the final performance of the model.

6. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 5, characterized in that In step S5, the specific steps are as follows: Each test example is sampled multiple times to simulate different possibilities of the input features, estimate the model prediction distribution, and calculate the prediction probability of each category, thereby providing uncertainty measures at different levels.

7. A method for predicting strata based on multi-source data based on random fields and deep learning as claimed in claim 6, characterized in that In step S6, the specific steps are as follows: The f1 value, accuracy, and recall rate are used to evaluate the model prediction effect, and the loss function curve with the number of training rounds, ROC curve, confusion matrix, and prediction measurement with depth distribution are used to evaluate the model prediction results.

Citation Information

Patent Citations

  • Method for joint inversion of stratigraphic layered structure based on Loff wave and Rayleigh wave

    CN113945975A

  • Method for stratigraphic stratification of survey area

    CN118131344A

  • Tibetan word segmentation method based on Transform-CRF

    CN114330328A

  • Coal mine gas drilling track prediction method and system driven by measurement-while-drilling data

    CN117266836A

  • Stratum lithology prediction method based on deep learning model and while-drilling information

    CN118820715A