Unsampled point geotechnical parameter evaluation method based on random forest machine learning algorithm

By integrating multi-source data and optimizing the model using the random forest machine learning algorithm, the problems of data integration and model optimization in the evaluation of soil and rock parameters at unsampled points were solved, achieving efficient and accurate evaluation of soil and rock parameters and providing reliable data support.

CN121959022APending Publication Date: 2026-05-01CHONGQING THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING THREE GORGES UNIV
Filing Date
2025-12-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies lack effective multi-source geological data integration mechanisms and efficient model parameter optimization methods in the assessment of geotechnical parameters at unsampled points. This results in low data correlation, excessive redundant information, and difficulty in accurately reflecting the characteristics of geotechnical parameters, thus failing to meet the accuracy requirements of engineering assessment.

Method used

A random forest-based machine learning algorithm was adopted. A multi-source geological data fusion computing platform was constructed to perform spatial matching and attribute integration. Key features were screened by combining a high-dimensional data dimensionality reduction mapping model. A spatiotemporal feature random forest evaluation model was used for training, and a stochastic gradient descent optimizer was introduced to iteratively optimize the model parameters, outputting the evaluation results of soil and rock parameters for unsampled points.

Benefits of technology

It achieves efficient integration and accurate feature extraction of multi-source data, improves the accuracy and reliability of geotechnical parameter assessment, and meets the safety and economic requirements of engineering survey and design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121959022A_ABST
    Figure CN121959022A_ABST
Patent Text Reader

Abstract

The invention discloses an unsampled point rock and soil parameter evaluation method based on a random forest machine learning algorithm, and the method comprises the steps: obtaining multi-source geological data of a sampled point, constructing a multi-source geological data fusion calculation platform, achieving the data space matching and attribute integration, and forming a multi-dimensional geological data set; retaining feature vectors with high correlation degree with rock and soil parameters of unsampled points through a high-dimensional data dimension reduction mapping model to obtain a low-dimensional feature data set; dividing the low-dimensional feature data set into sample sets, training a spatial-temporal feature random forest evaluation model and adjusting parameters; introducing a stochastic gradient descent optimizer to iteratively optimize model parameters until a prediction error is converged; and inputting space coordinates of unsampled points and surrounding geological data, and outputting an evaluation result through the optimization model. The method solves the problems that multi-source data integration is difficult, the model does not consider spatial-temporal characteristics and optimization is insufficient, data utilization efficiency and evaluation precision are improved, and reliable data support is provided for geotechnical engineering investigation and design.
Need to check novelty before this filing date? Find Prior Art

Description

A Method for Evaluating Geotechnical Parameters at Unsampled Points Based on Random Forest Machine Learning Algorithm Technical Field

[0001] This invention relates to the field of geotechnical parameter assessment technology, and in particular to a method for assessing geotechnical parameters at unsampled points based on a random forest machine learning algorithm. Background Technology

[0002] In geotechnical engineering investigation and design, the assessment of geotechnical parameters at unsampled sites is a crucial step in ensuring project safety and economic efficiency. In actual engineering projects, multi-source geological data, including geotechnical physical and mechanical parameters, geological structure data, topographic data, stratigraphic lithology data, and geophysical exploration data from sampled sites, are scattered and exhibit complex data types and high dimensionality. Traditional assessment methods rely on manual experience or inferences from single data sources, making it difficult to fully integrate the effective information from multiple sources and accurately reflect the spatial and temporal variations of geotechnical parameters. With the expanding application of machine learning technology in engineering, using algorithms to deeply process multi-source geological data and achieve scientific assessment of geotechnical parameters at unsampled sites has become an important direction for overcoming the limitations of traditional methods and improving assessment accuracy. Providing more reliable data support for engineering investigation and design has become an industry requirement.

[0003] Existing technologies suffer from two significant drawbacks in assessing geotechnical parameters at unsampled sites. Firstly, they lack effective integration mechanisms for processing multi-source geological data, failing to achieve efficient spatial matching and attribute integration of different types of geological data. This results in low data correlation, excessive redundancy, and difficulty in forming a dataset that accurately reflects the characteristics of geotechnical parameters. Consequently, this affects the subsequent extraction of effective information by the assessment model, reducing the reliability of the assessment results. Secondly, in the model construction and optimization phase, the influence of spatiotemporal characteristics on geotechnical parameters is not fully considered during model training, and efficient parameter optimization methods are lacking. The inability to dynamically adjust model parameters based on data characteristics makes it difficult for model prediction errors to converge to the ideal range, resulting in inaccurate output of geotechnical parameter assessment results for unsampled sites, and failing to meet the accuracy requirements of engineering projects. Summary of the Invention

[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a method for evaluating unsampled soil and rock parameters based on a random forest machine learning algorithm. The method for evaluating unsampled soil and rock parameters based on a random forest machine learning algorithm is characterized by the following steps: S1, acquiring multi-source geological data, including soil and rock physical and mechanical parameters, geological structural data, topographic data, stratigraphic lithology data, and geophysical exploration data from sampled points; S2, constructing a multi-source geological data fusion computing platform, importing the multi-source geological data acquired in S1 into the platform, and performing spatial matching and attribute integration of different types of geological data through data association analysis to form a multi-dimensional geological data set; S3, deploying a high-dimensional data dimensionality reduction mapping model on the multi-source geological data fusion computing platform, performing dimensionality reduction processing on the multi-dimensional geological data set formed in S2, retaining feature vectors in the dataset with a correlation degree higher than a preset threshold with the soil and rock parameters of the unsampled points, obtaining a low-dimensional feature dataset; S4, ... The spatiotemporal feature random forest evaluation model is invoked. The low-dimensional feature dataset obtained in S3 is divided into a training sample set and a validation sample set. The spatiotemporal feature random forest evaluation model is trained using the training sample set, and the number of decision trees, node splitting threshold, and spatiotemporal weight coefficients of the model are adjusted using the validation sample set. In S5, a stochastic gradient descent optimizer is introduced to iteratively optimize the parameters of the spatiotemporal feature random forest evaluation model trained in S4. The gradient value of the model prediction error is calculated in each iteration, and the feature weights and decision tree branch parameters of the model are adjusted according to the gradient value until the model prediction error converges to a preset range. In S6, the spatial coordinates of the unsampled points and the surrounding geological environment data are determined and input into the spatiotemporal feature random forest evaluation model optimized in S5. The geotechnical parameter evaluation results of the unsampled points are output through the model's multi-decision tree voting mechanism.

[0005] Furthermore, the expression for the spatiotemporal feature random forest evaluation model is:

[0006]

[0007] Where F(x,y,t) is the geotechnical parameter evaluation value of the last sample point (x,y) at time t, K is the total number of decision trees in the spatiotemporal feature random forest, and ω k Let α be the weight coefficient of the k-th decision tree, M be the number of feature split nodes in each decision tree, and α be the weight coefficient of the k-th decision tree. k,m Let φ be the feature sensitivity coefficient of the m-th split node in the k-th decision tree. k,m (x,y,t) represents the spatiotemporal feature value of the unsampled point (x,y) at time t corresponding to the m-th split node of the k-th decision tree, θ k,m Let η be the feature threshold of the m-th split node in the k-th decision tree. k(x,y,t) represents the predicted geotechnical parameters of the unsampled point (x,y) at time t by the k-th decision tree.

[0008] Furthermore, the expression for the high-dimensional data dimensionality reduction mapping model is:

[0009]

[0010] Among them, Z p Let β be the p-th feature vector in the reduced-dimensional feature dataset, Q be the number of feature types in the original high-dimensional dataset, and β be the feature vector. p,q Let γ be the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature, where I is the number of data samples included in the q-th original high-dimensional feature, and γ is the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature. q,i X is the weight factor of the i-th sample in the original high-dimensional features of class q. q,i Let be the feature value of the i-th sample in the original high-dimensional features of class q, R be the number of spatiotemporal correction factors, and δ be the feature value of the i-th sample. p,r Let T be the correlation coefficient between the p-th low-dimensional feature and the r-th spatiotemporal correction factor. r Let r be the r-th spatiotemporal correction factor.

[0011] Furthermore, the expression for the stochastic gradient descent optimizer is:

[0012]

[0013] Where, θ t+1 To optimize the model parameters for round t+1, θ t To optimize the model parameters in the first t rounds, η t Let B be the learning rate in round t, B be the number of samples in each batch during each iteration, and b be the sample index within the batch. F(x) is the gradient of the model loss function with respect to the parameter θ. b ,y b ,t b ;θ) represents the model's pair of samples (x) b ,y b ,t b The predicted value of y b,true For sample (x) b ,y b ,t b The actual geotechnical parameter values, λ t Let be the regularization coefficient for round t.

[0014] Furthermore, the data fusion expression of the multi-source geological data fusion computing platform is as follows:

[0015]

[0016] Among them, D fusion (x,y,t) represents the geological data value at (x,y,t) after fusion, S represents the number of data source types for the multi-source geological data, and μ s Let D be the weight of the s-th type of data source. s (x,y,t) represents the original data value at (x,y,t) of the s-th data source, U represents the number of metrics for data consistency verification, and ξ s,u Let be the deviation coefficient of the u-th verification index corresponding to the s-th type of data source. Let be the average data value of the u-th verification index in the region surrounding (x,y,t).

[0017] Furthermore, the corrected expression for the geotechnical parameter evaluation results of the unsampled points is as follows:

[0018]

[0019] Among them, P final (x0,y0,t0) represents the final geotechnical parameter assessment value of the unsampled point (x0,y0,t0), F(x0,y0,t0) represents the initial predicted value of the spatiotemporal characteristic random forest assessment model, V represents the number of sampled points surrounding the unsampled point, and ε represents the final geotechnical parameter assessment value. v Let P be the influence coefficient of the v-th surrounding sampled point on the unsampled point. v (x v ,y v ,t v ) represents the v-th surrounding sampled point (x) v ,y v ,t v The actual geotechnical parameters.

[0020] Further, S2 includes the following sub-steps: S21, extracting spatial reference system, data precision, and attribute field information of each data source from the metadata of multi-source geological data, establishing a data source information comparison table, and clarifying the spatial coordinate transformation relationship and attribute mapping rules between different data sources; S22, based on the spatial coordinate transformation relationship, uniformly transforming the geological data from different data sources to the same spatial coordinate system, correcting the spatial position deviation of the transformed data, and ensuring that the spatial alignment accuracy of the data meets the preset requirements; S23, according to the attribute mapping rules, standardizing the attribute fields of different data sources, merging semantically similar attribute fields, and supplementing missing attribute values ​​using an interpolation method based on neighborhood similarity to form preliminary integrated geological data; S24, performing logical consistency verification on the preliminary integrated geological data, eliminating data records with obvious logical contradictions, marking data outliers using a statistical distribution-based identification method, and correcting the marked outliers through comparative analysis with surrounding data, ultimately forming a multi-dimensional geological data set.

[0021] Further, S3 includes the following sub-steps: S31, calculate the Pearson correlation coefficient and mutual information value between each feature vector in the multi-dimensional geological data set and the soil and rock parameters of the unsampled points, and determine the correlation degree between each feature vector and the soil and rock parameters of the unsampled points based on the weighted sum of the correlation coefficient and mutual information value; S32, set a correlation degree threshold, filter out feature vectors with a correlation degree higher than the threshold, and form an initial low-dimensional feature dataset from the filtered feature vectors; S33, perform feature redundancy analysis on the initial low-dimensional feature dataset, calculate the cosine similarity between each feature vector, remove feature vectors with a similarity higher than the redundancy threshold, and retain feature vectors with independent information to obtain the final low-dimensional feature dataset; S34, record the feature selection rules and mapping parameters of the high-dimensional data dimensionality reduction mapping model in the dimensionality reduction process, and form a dimensionality reduction processing report for subsequent model reproduction and parameter adjustment.

[0022] Further, step S4 includes the following sub-steps: S41, according to a preset sample division ratio, the low-dimensional feature dataset is randomly divided into a training sample set and a validation sample set, wherein the training sample set is used for model training and the validation sample set is used for model performance validation; S42, the parameters of the spatiotemporal feature random forest evaluation model are initialized, including the number of decision trees, the depth of the decision trees, the number of features for node splitting, and the initial values ​​of the spatiotemporal weight coefficients, and the maximum number of iterations and the error convergence threshold are set for model training; S43, the training sample set is input into the spatiotemporal feature random forest evaluation model, and an independent training sample subset is generated for each decision tree using the bootstrap sampling method. Decision trees are constructed based on the sample subsets. During the decision tree construction process, the information gain of each split node is calculated in combination with spatiotemporal features, and the feature with the largest information gain is selected as the split feature; S44, the validation sample set is input into the trained spatiotemporal feature random forest evaluation model, and the prediction error of the model is calculated. If the prediction error is greater than the error convergence threshold, the number of decision trees, the node splitting threshold, and the spatiotemporal weight coefficients are adjusted, and the process from S43 to S44 is repeated until the model prediction error is less than or equal to the error convergence threshold.

[0023] Further, S5 includes the following sub-steps: S51, determining the initial learning rate, batch sample size, and initial values ​​of the regularization coefficient for the stochastic gradient descent optimizer, and setting the termination condition for the optimization iteration, wherein the termination condition includes the number of iterations reaching the maximum number of iterations or the change in model prediction error being less than a preset error change threshold; S52, randomly selecting batch samples from the training sample set, inputting them into the spatiotemporal feature random forest evaluation model, calculating the loss function value of the model on the batch samples, and solving for the gradient value of the model parameters based on the loss function value; S53, updating the feature weights and decision tree branch parameters of the model according to the gradient value and the initial learning rate, and introducing a regularization term to constrain the model parameters to prevent overfitting; S54, calculating the prediction error of the optimized model on the validation sample set. If the termination condition is not met, adjusting the learning rate and regularization coefficient, returning to S52 to continue iterative optimization; if the termination condition is met, stopping the optimization and saving the optimized model parameters.

[0024] Beneficial Effects: This invention proposes a method for evaluating geotechnical parameters at unsampled points based on a random forest machine learning algorithm. By constructing a multi-source geological data fusion computing platform, it achieves spatial matching and attribute integration of multi-source data such as geotechnical physical and mechanical parameters and geological structural data from sampled points. This solves the problems of low data correlation and redundant information caused by the dispersion of multi-source data and the lack of an effective integration mechanism in traditional technologies, forming a multi-dimensional dataset that accurately reflects the characteristics of geotechnical parameters, providing a high-quality data foundation for subsequent evaluation. A high-dimensional data dimensionality reduction mapping model is used to reduce the dimensionality of multi-dimensional data, retaining feature vectors with high correlation to geotechnical parameters at unsampled points, reducing redundant information interference, and improving data utilization efficiency. A spatiotemporal feature random forest evaluation model is combined with spatiotemporal feature training and validation, while a stochastic gradient descent optimizer is introduced to iteratively optimize model parameters, dynamically adjusting feature weights and decision tree branch parameters, so that the model prediction error converges to a preset range. This compensates for the shortcomings of traditional models that do not fully consider spatiotemporal features and lack efficient parameter optimization methods, ultimately outputting accurate evaluation results for geotechnical parameters at unsampled points, providing more reliable data support for geotechnical engineering investigation and design, and balancing engineering safety and economic requirements. Attached Figure Description

[0025] Figure 1 is a flowchart of the overall steps of the method of the present invention;

[0026] Figure 2 is a flowchart of method step S2 of the present invention;

[0027] Figure 3 is a flowchart of method step S3 of the present invention;

[0028] Figure 4 is a flowchart of method step S4 of the present invention;

[0029] Figure 5 is a flowchart of method step S5 of the present invention; Detailed Implementation

[0030] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] As shown in Figure 1, the method for evaluating soil and rock parameters at unsampled points based on the random forest machine learning algorithm includes the following steps:

[0032] S1, acquire multi-source geological data, which includes the physical and mechanical parameters of soil and rock at the sampled points, geological structure data, topographic data, stratigraphic lithology data, and geophysical exploration data;

[0033] Specifically, step S1 involves acquiring multi-source geological data. During implementation, comprehensive collection of various geological data from the sampled points is required. This includes geophysical and mechanical parameters such as cohesion, internal friction angle, and compression modulus. At least three core parameters should be collected from each sampled point, and the measurement error must be controlled within ±5%. Geological structural data, including fault strike, fold morphology, and joint density, must be obtained through geological mapping and drilling records, with a mapping accuracy of no less than 1:5000. Topographic data is extracted using UAV aerial photography or high-precision topographic maps, with elevation accuracy controlled within ±0.5 meters and horizontal position accuracy no less than ±1 meter. Stratigraphic and lithological data must clearly define the name, thickness, and distribution range of each rock layer, obtained through borehole core logging, with a core logging interval of no more than 0.5 meters between each borehole. Geophysical exploration data includes seismic wave velocity and resistivity. The spacing between exploration points is set according to the size of the exploration area, generally not exceeding 50 meters. Measurement time and ambient temperature must be recorded simultaneously during data acquisition. This step requires ensuring that the collected data covers all sampled points within the survey area, and that the sample size for each data type is no less than 100 sets. This provides sufficient and accurate basic data for subsequent data fusion and model training, and avoids affecting the reliability of subsequent evaluation results due to missing data or excessive errors.

[0034] S2, construct a multi-source geological data fusion computing platform, import the multi-source geological data obtained in S1 into the platform, and perform spatial matching and attribute integration of different types of geological data through data correlation analysis to form a multi-dimensional geological data set;

[0035] Specifically, step S2 involves constructing a multi-source geological data fusion computing platform and integrating the data. This is achieved by first building a computing platform based on a distributed architecture. The platform's processor uses a server-grade CPU with at least 8 cores, a memory capacity of at least 64GB, and a storage capacity of at least 2TB to ensure it can handle the processing needs of large-scale multi-source data. Then, the multi-source geological data acquired in S1 is imported into the platform according to data type. During the import process, XML format is used for data encapsulation to ensure the uniformity of the data structure. Next, the platform's built-in spatial matching algorithm is used to uniformly convert the spatial coordinates of different data sources to the 2000 National Geodetic Coordinate System, with the conversion error controlled within ±0.3 meters, achieving precise spatial alignment of the data. Next, based on the attribute mapping rules, semantically similar attribute fields are merged. For example, "cohesion (kPa)" and "cohesion (kPa)" are merged into a unified field. Missing attribute values ​​are supplemented using the K-nearest neighbor interpolation method, selecting the five closest surrounding sample points as references during interpolation to ensure the rationality of the supplemented data. Finally, the integrity of the integrated data is verified to ensure that the core data fields of each sampled point are not missing, forming a multi-dimensional geological data set including at least 10 dimensions. This step, through a standardized data integration process, eliminates format differences and spatial deviations of multi-source data, providing a data foundation with a unified structure and meeting accuracy standards for subsequent high-dimensional data dimensionality reduction and model training.

[0036] S3. Deploy a high-dimensional data dimensionality reduction mapping model on the multi-source geological data fusion computing platform to perform dimensionality reduction processing on the multi-dimensional geological data set formed in S2, retain the feature vectors in the dataset whose correlation with the soil and rock parameters of unsampled points is higher than a preset threshold, and obtain a low-dimensional feature dataset.

[0037] Specifically, step S3 involves deploying a high-dimensional data dimensionality reduction mapping model and processing the data. During implementation, the high-dimensional data dimensionality reduction mapping model module is first installed on a multi-source geological data fusion computing platform. This module supports custom correlation calculation methods and threshold settings. Then, the multi-dimensional geological data set formed in S2 is input into the model. The model first calculates the correlation between each feature vector and the soil and rock parameters of unsampled points. The correlation calculation uses a weighted summation of the Pearson correlation coefficient and mutual information values, with the Pearson correlation coefficient weight set at 0.6 and the mutual information value weight set at 0.4 to ensure the comprehensiveness of the correlation assessment. Next, a correlation threshold is set. Based on the degree of variation of soil and rock parameters in the exploration area, the threshold is generally set between 0.6 and 0.8, with the upper limit used for areas with high variation. For regions with low correlation, the lower limit is used. The model selects feature vectors with a correlation higher than this threshold. For example, if the correlation of the cohesion feature vector is 0.75 and the correlation of the terrain slope feature vector is 0.58, the cohesion feature vector is retained and the terrain slope feature vector is removed. At the same time, redundancy detection is performed on the selected feature vectors, and the similarity between feature vectors is calculated. The similarity threshold is set to 0.85, and redundant feature vectors with similarity higher than this threshold are removed. Finally, a low-dimensional feature dataset with a dimension controlled at 3-5 is obtained. This step reduces the interference of high-dimensional data on model training and reduces the computational complexity of the model through precise feature selection and redundancy removal. At the same time, it retains the feature information that plays a key role in the evaluation of soil and rock parameters of unsampled points, thereby improving the efficiency and accuracy of subsequent model training.

[0038] S4, call the spatiotemporal feature random forest evaluation model, divide the low-dimensional feature dataset obtained in S3 into training sample set and validation sample set, use the training sample set to train the spatiotemporal feature random forest evaluation model, and adjust the number of decision trees, node splitting threshold and spatiotemporal weight coefficient of the model through the validation sample set.

[0039] Specifically, step S4 involves calling and training a spatiotemporal feature random forest evaluation model. During implementation, the spatiotemporal feature random forest evaluation model is first called from the model library of the multi-source geological data fusion computing platform. The initial model setting is 100-200 decision trees, with a maximum depth of 15 layers. The number of features selected during node splitting is the square root of the dimension of the low-dimensional feature dataset. Then, the low-dimensional feature dataset obtained in S3 is randomly divided into a training sample set and a validation sample set in a 7:3 ratio. The partitioning process uses stratified sampling to ensure consistent distribution characteristics of feature vectors in both sample sets. Next, the training sample set is input into the model. The model uses bootstrap sampling to generate independent training sample subsets for each decision tree, with each sample having a 63.2% probability of being selected. During decision tree construction, the model combines spatiotemporal features to calculate... The information gain and spatiotemporal feature weights of each split node are set according to the sample collection time span and spatial distance. The time weight of samples with a time span of less than one year is 0.9, and the spatial weight of samples with a spatial distance of less than 100 meters is 0.8. After the initial training of the model using the training sample set, the validation sample set is input into the model, and the prediction error of the model is calculated. If the error is higher than the preset 5% error threshold, the number of decision trees (increase or decrease by 20 each time), the node split threshold (adjusted by ±0.05 each time), and the spatiotemporal weight coefficient (adjusted by ±0.1 each time) are adjusted. The training and validation process is repeated until the model prediction error is lower than the error threshold. This step, through scientific sample division and model parameter adjustment, enables the spatiotemporal feature random forest evaluation model to initially have the ability to predict soil and rock parameters, while incorporating spatiotemporal features to improve the model's adaptability to the spatiotemporal variation law of soil and rock parameters.

[0040] S5 introduces a stochastic gradient descent optimizer to iteratively optimize the parameters of the spatiotemporal feature random forest evaluation model trained in S4. It calculates the gradient value of the model prediction error in each iteration and adjusts the feature weights and decision tree branch parameters of the model according to the gradient value until the model prediction error converges to the preset range.

[0041] Specifically, step S5 involves introducing a stochastic gradient descent optimizer and iteratively optimizing it. First, the stochastic gradient descent optimizer is connected to the spatiotemporal feature random forest evaluation model trained in S4. The optimizer's initial learning rate is set to 0.01-0.05, the batch sample size for each iteration is 32-64 groups, the regularization coefficient is set to 0.001, and the iteration termination condition is set to 500 iterations or the change in model prediction error being less than 0.001 for 10 consecutive iterations. Then, the optimizer randomly selects batch samples from the training sample set and inputs them into the spatiotemporal feature random forest evaluation model. The model's prediction error on that batch of samples is calculated, and the gradient values ​​of the model parameters are calculated based on the error. The gradient calculation uses a combination of batch gradient descent and stochastic gradient descent. This approach balances computational accuracy and efficiency. It adjusts the model's feature weights (each adjustment not exceeding 10% of the initial weights) and decision tree branch parameters (such as node splitting thresholds) based on the gradient value and learning rate. Simultaneously, it constrains the parameter variation range through regularization to prevent overfitting. After each iteration, the prediction error of the optimized model on the validation set is calculated. If the termination condition is not met, the learning rate is reduced to 0.8 times its original value and the regularization coefficient is increased by 0.0005 every 100 iterations, continuing the iteration. If the termination condition is met, optimization stops, and the optimized model parameters are saved. This step, through continuous parameter iteration optimization, further reduces the model's prediction error, improves its stability and generalization ability, and ensures that the model outputs accurate results in subsequent evaluations.

[0042] S6 determines the spatial coordinates of the unsampled points and the surrounding geological environment data, inputs them into the spatiotemporal characteristic random forest evaluation model optimized by S5, and outputs the geotechnical parameter evaluation results of the unsampled points through the model's multi-decision tree voting mechanism.

[0043] Specifically, step S6 involves determining the data of unsampled points and outputting the evaluation results. During implementation, the spatial coordinates of the unsampled points are first determined using GPS positioning, with a positioning accuracy controlled within ±0.2 meters. Simultaneously, geological environmental data within a 100-meter radius of the unsampled points are collected, including soil and rock parameters, lithological distribution, and topographic relief of surrounding sampled points. The data collection time should not exceed two years from the current evaluation time. Subsequently, the spatial coordinates of the unsampled points and the surrounding geological environmental data are standardized according to the data format in S2 to ensure consistency with the format of the low-dimensional feature dataset. The standardized data is then input into the spatiotemporal feature random forest evaluation model optimized in S5. The model calls all trained decision trees to process the data in parallel, with each decision tree outputting a predicted value for a soil and rock parameter. The model determines the final evaluation result through a multi-decision tree voting mechanism. During voting, the weight of each decision tree is set based on its prediction accuracy on the validation sample set: a decision tree with an accuracy higher than 95% has a weight of 1.2, an accuracy between 90% and 95% has a weight of 1.0, and an accuracy lower than 90% has a weight of 0.8. The final model outputs the geotechnical parameter evaluation values ​​for the unsampled points, along with the confidence level of the evaluation results. The confidence level is calculated based on the standard deviation of all decision tree predictions: a standard deviation less than 5% indicates a confidence level of 95%, and a standard deviation between 5% and 10% indicates a confidence level of 90%. This step, through precise data collection from unsampled points and a scientific model prediction mechanism, ensures the accuracy and reliability of the output geotechnical parameter evaluation results, providing direct data support for subsequent geotechnical engineering investigation and design.

[0044] Preferably, the expression for the spatiotemporal feature random forest evaluation model is:

[0045]

[0046] Where F(x,y,t) is the geotechnical parameter evaluation value of the last sample point (x,y) at time t, K is the total number of decision trees in the spatiotemporal feature random forest, and ω k Let a be the weight coefficient of the k-th decision tree, M be the number of feature split nodes in each decision tree, and a be the weight coefficient of the k-th decision tree. k,m Let φ be the feature sensitivity coefficient of the m-th split node in the k-th decision tree. k,m (x,y,t) represents the spatiotemporal feature value of the unsampled point (x,y) at time t corresponding to the m-th split node of the k-th decision tree, θ k,m h is the feature threshold of the m-th split node in the k-th decision tree. k (x,y,t) represents the predicted geotechnical parameters of the unsampled point (x,y) at time t by the k-th decision tree.

[0047] Specifically, when implementing the spatiotemporal feature random forest evaluation model, the value range and setting basis of the key model parameters should be clearly defined first: the total number of decision trees should be set according to the sample size of the low-dimensional feature dataset. For datasets with fewer than 500 samples, 100-150 trees should be used; for datasets with more than 500 samples, 150-200 trees should be used to ensure the model has both computational efficiency and prediction accuracy. The weight coefficient of each decision tree should be allocated according to its prediction accuracy on the validation sample set. Decision trees with an accuracy higher than 95% should have a weight of 1.2, those with an accuracy between 90% and 95% should have a weight of 1.0, and those with an accuracy lower than 90% should have a weight of 0.8 to avoid low-precision decision trees from interfering too much with the evaluation results. The number of feature splitting nodes in each decision tree should match the dimension of the low-dimensional feature dataset. For a dimension of 3, 5-8 nodes are set; for a dimension of 4-5, 8-12 nodes are set to ensure sufficient feature information is extracted. The feature sensitivity coefficient is adjusted according to the correlation between the feature and the soil and rock parameters. The coefficient for features with a correlation higher than 0.8 is set to 1.5-2.0, and for features with a correlation between 0.6 and 0.8, it is set to 1.0-1.5 to enhance the influence of key features on model prediction. Spatiotemporal feature values ​​need to be calculated by combining the spatial coordinates of unsampled points and the data acquisition time. Spatially, a distance-weighted method is used, and temporally, a time decay method with higher weight for recent data is used. The feature threshold is determined by statistically analyzing the distribution range of feature values ​​of sampled points, taking the range of ±10% of the median of the distribution to ensure that node splitting can effectively distinguish different levels of soil and rock parameters. By accurately setting the model parameters, the model can fully integrate spatiotemporal features, improve the accuracy of soil and rock parameter assessment for unsampled points, and solve the prediction bias problem caused by traditional models ignoring spatiotemporal differences.

[0048] Preferably, the expression for the high-dimensional data dimensionality reduction mapping model is:

[0049]

[0050] Among them, Z p Let β be the p-th feature vector in the reduced-dimensional feature dataset, Q be the number of feature types in the original high-dimensional dataset, and β be the feature vector. p,q Let γ be the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature, where I is the number of data samples included in the q-th original high-dimensional feature, and γ is the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature. q,i X is the weight factor of the i-th sample in the original high-dimensional features of class q. q,i Let be the feature value of the i-th sample in the original high-dimensional features of class q, R be the number of spatiotemporal correction factors, and δ be the feature value of the i-th sample. p,r Let T be the correlation coefficient between the p-th low-dimensional feature and the r-th spatiotemporal correction factor. r Let r be the r-th spatiotemporal correction factor.

[0051] Specifically, the high-dimensional data dimensionality reduction mapping model strictly controls the parameter values ​​and calculation logic during the mapping process: the number of low-dimensional feature vectors is determined based on the number of feature types in the original high-dimensional data; 3 low-dimensional feature vectors are used when there are 10-15 original feature types, and 4-5 are used when there are 15-20 original feature types, ensuring that core information is preserved while reducing dimensionality; the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature is calculated using the least squares method, so that the low-dimensional feature can restore the effective information of the original feature to the greatest extent, and the mapping error must be less than 5% during the calculation; the weight factor of the i-th sample in the q-th original high-dimensional feature is set according to the sample acquisition accuracy, with an accuracy of ±2. The weight of samples within a certain percentage is set to 1.2, between ±2% and ±5% to 1.0, and above ±5% to 0.8, filtering out interference from low-precision samples. The number of spatiotemporal correction factors is set according to the spatiotemporal complexity of the surveyed area. For complex terrain or data collection spanning more than 3 years, 5-8 factors are used; for simple terrain or data collection spanning less than 1 year, 3-5 factors are used to correct for the impact of spatiotemporal differences on the data. The correlation coefficient between the p-th low-dimensional feature and the r-th spatiotemporal correction factor is determined through correlation analysis. Correlation coefficients with absolute values ​​higher than 0.7 are set to 0.8-1.0, and those lower than 0.7 are set to 0.3-0.8, highlighting the role of strongly correlated correction factors. By scientifically setting dimensionality reduction parameters and mapping rules, efficient dimensionality reduction of high-dimensional data is achieved, reducing data redundancy and computational load while avoiding the loss of core feature information, providing high-quality low-dimensional data for subsequent model training.

[0052] Preferably, the expression for the stochastic gradient descent optimizer is:

[0053]

[0054] Where, θ t+1 To optimize the model parameters for round t+1, θ t To optimize the model parameters in the first t rounds, η t Let B be the learning rate in round t, B be the number of samples in each batch during each iteration, and b be the sample index within the batch. F(x) is the gradient of the model loss function with respect to the parameter θ. b ,y b ,t b ;θ) represents the model's pair of samples (x) b ,y b ,t b The predicted value of y b,true For sample (x) b ,y b ,t b The actual geotechnical parameter values, λ t Let be the regularization coefficient for round t.

[0055] Specifically, the stochastic gradient descent optimizer, when implemented, clearly defines the parameter configuration and iteration logic during the optimization process: the number of iteration rounds for model parameters is set according to the initial prediction error; 300-400 rounds are set when the initial error is higher than 10%, 200-300 rounds are set when the initial error is between 5% and 10%, and 100-200 rounds are set when the initial error is lower than 5%, ensuring sufficient optimization without wasting computational resources; the learning rate in each round adopts a dynamic decay strategy, with the initial learning rate set at 0.03-0.05 when the sample size is less than 300 sets, and at 0.01-0.03 when the sample size exceeds 300 sets, decaying to 0.8 times the original rate every 100 rounds, balancing rapid convergence in the early stage with accurate optimization in the later stage; the batch sample size in each iteration is set according to the server memory capacity. For 64GB of memory, 48-64 groups are set; for 32-64GB of memory, 32-48 groups are set to avoid memory overflow and ensure the representativeness of batch data. The gradient calculation of the loss function needs to consider the deviation between the model's predicted values ​​and the true values, using the mean squared error as the loss function. During gradient calculation, outlier values ​​(exceeding the mean ± 3 standard deviations) are weighted with a weight of 1.5 to reduce the misleading influence of outliers on the gradient direction. The regularization coefficient is adjusted according to the degree of model overfitting: 0.0015-0.002 when the difference between the training and validation set errors exceeds 5%, 0.001-0.0015 when the difference is between 3% and 5%, and 0.0005-0.001 when the difference is below 3%, effectively suppressing overfitting. By reasonably configuring optimization parameters and iterative strategies, efficient optimization of model parameters is achieved, enabling the model's prediction error to quickly converge to the preset range, improving model stability and generalization ability.

[0056] Preferably, the data fusion expression of the multi-source geological data fusion computing platform is:

[0057]

[0058] Among them, D fusion (x,y,t) represents the geological data value at (x,y,t) after fusion, S represents the number of data source types for the multi-source geological data, and μ s Let D be the weight of the s-th type of data source. s (x,y,t) represents the original data value at (x,y,t) of the s-th data source, U represents the number of metrics for data consistency verification, and ξ s,u Let be the deviation coefficient of the u-th verification index corresponding to the s-th type of data source. Let be the average data value of the u-th verification index in the region surrounding (x,y,t).

[0059] Specifically, the data fusion process of the multi-source geological data fusion computing platform involves precisely setting fusion parameters and verification rules during implementation: the number of data source types is determined according to the needs of the exploration project; basic exploration projects include 5-8 types of data sources (such as geotechnical parameters, geological structures, topography, etc.), while detailed exploration projects include 8-12 types, ensuring comprehensive data coverage; the weight of the s-th type of data source is determined through expert scoring and data reliability analysis, with field-measured data (such as drilling core data) weighted at 0.8-1.0, and indirectly inferred data (such as geophysical exploration data) weighted at 0.5-0.8, highlighting the role of highly reliable data; the number of data consistency verification indicators is set according to the data type, with numerical data (such as geotechnical parameters, geological structures, topography, etc.) weighted at 0.5-0.8. For categorical data (such as stratigraphy and lithology), 4-6 indicators (e.g., mean, standard deviation) are set, and for classified data (such as stratigraphy and lithology), 2-3 indicators (e.g., classification consistency, distribution rationality) are set to comprehensively verify data consistency. The deviation coefficient of the u-th verification indicator corresponding to the s-th data source is adjusted according to the importance of the indicator. The deviation coefficient of core indicators (such as cohesion and internal friction angle) is set at 0.8-1.0, and that of secondary indicators is set at 0.3-0.8 to strengthen the verification of core indicators. The average data value of the u-th verification indicator in the surrounding area is calculated using the sliding window method. The window size is set according to the data density. A 50-100 meter window is set in dense data areas, and a 100-200 meter window is set in sparse data areas to ensure that the average value can reflect the true level of the area. By scientifically setting fusion parameters and verification rules, efficient integration and quality control of multi-source geological data are achieved, solving the problem of poor data quality caused by unreasonable weight allocation and lack of consistency verification in traditional data fusion.

[0060] Preferably, the corrected expression for the geotechnical parameter evaluation results of the unsampled points is:

[0061]

[0062] Among them, P final (x0,y0,t0) represents the final geotechnical parameter assessment value of the unsampled point (x0,y0,t0), F(x0,y0,t0) represents the initial predicted value of the spatiotemporal characteristic random forest assessment model, V represents the number of sampled points surrounding the unsampled point, and ε represents the final geotechnical parameter assessment value. v Let P be the influence coefficient of the v-th surrounding sampled point on the unsampled point. v (x v ,y v ,t v ) represents the v-th surrounding sampled point (x) v ,y v ,t v The actual geotechnical parameters.

[0063] Specifically, the correction of the geotechnical parameter assessment results for unsampled points should be implemented with a clear definition of the basis for parameter values ​​and calculation logic: the number of sampled points around the unsampled point should be set according to the density of sampled points in the survey area. When the density is higher than 2 points / 100 square meters, 5-8 surrounding points should be selected; when the density is between 1-2 points / 100 square meters, 3-5 points should be selected; and when the density is lower than 1 point / 100 square meters, 2-3 points should be selected to ensure sufficient reference samples and avoid excessive redundancy. The influence coefficient of the vth surrounding sampled point on the unsampled point is calculated using the distance attenuation formula. The influence coefficient within 50 meters of the unsampled point is set to 0.8-1.0, 50-100 meters to 0.5-0.8, and 100-150 meters to 0. The influence coefficient is set at 3-0.5, with the influence decreasing as the distance increases, consistent with the continuous spatial distribution of geotechnical parameters. The deviation between the initial predicted value and the actual values ​​of surrounding sampled points must be calculated using absolute deviation to avoid inaccurate corrections caused by the cancellation of positive and negative deviations. The deviation percentage is calculated based on the actual values ​​of surrounding sampled points; if the actual value is 0 (a special case), the average value of surrounding points is used instead, ensuring rigorous calculation logic. The correction coefficient is adjusted according to the deviation percentage: 0.05-0.1 for deviations below 5%, 0.1-0.2 for 5%-10%, and 0.2-0.3 for 10%-15%. The larger the deviation, the greater the correction, making the final evaluation result closer to the actual situation. By incorporating reference information from surrounding sampled points to correct the initial predicted value, the model prediction error is further reduced, improving the reliability of geotechnical parameter evaluation results for unsampled points and meeting the engineering requirements for high-precision data.

[0064] Preferably, as shown in Figure 2, step S2 includes the following sub-steps: S21, extracting spatial reference system, data precision, and attribute field information of each data source from the metadata of multi-source geological data, establishing a data source information comparison table, and clarifying the spatial coordinate transformation relationship and attribute mapping rules between different data sources; S22, based on the spatial coordinate transformation relationship, uniformly transforming the geological data from different data sources to the same spatial coordinate system, correcting the spatial position deviation of the transformed data, and ensuring that the spatial alignment accuracy of the data meets the preset requirements; S23, according to the attribute mapping rules, standardizing the attribute fields of different data sources, merging semantically similar attribute fields, and supplementing missing attribute values ​​using an interpolation method based on neighborhood similarity to form preliminary integrated geological data; S24, performing logical consistency verification on the preliminary integrated geological data, removing data records with obvious logical contradictions, marking data outliers using a statistical distribution-based identification method, and correcting the marked outliers through comparative analysis with surrounding data, ultimately forming a multi-dimensional geological data set.

[0065] Specifically, step S2 includes sub-steps S21-S24: S21 extracts information from the metadata of multi-source geological data. The data source information lookup table must include spatial reference systems such as the 2000 National Geodetic Coordinate System and the 1985 National Height Datum, with data accuracy of ±1 meter for plane and ±0.5 meters for elevation, as well as attribute fields such as the name and unit of soil and rock parameters. Spatial coordinate transformation uses the seven-parameter transformation method with an error ≤ ±0.3 meters, and clarifies the merging criteria for semantically similar fields such as "compression modulus" and "deformation modulus." S22 uniformly transforms data from different data sources to the same coordinate system, and verifies the transformation using the root mean square error to ensure accuracy. Over 95% of the data points have a spatial deviation of less than ±0.3 meters; S23 standardizes the attribute fields, and missing values ​​are supplemented using the K-nearest neighbor interpolation method, selecting the 5 most similar sample points in the surrounding area, with the interpolation error ≤ 10% of the sample standard deviation; S24 performs logical consistency verification, using the standard deviation multiple method of ±3 times the standard deviation to identify outliers, and corrects them by comparing data within 100 meters of the surrounding area, ensuring that core fields such as cohesion and internal friction angle are not missing, ultimately forming a multi-dimensional geological data set with a dimension ≥ 10. Through step-by-step standardization, the problems of chaotic formats and large spatial deviations of multi-source data are solved, providing a high-quality data foundation for subsequent processing.

[0066] Preferably, as shown in Figure 3, step S3 includes the following sub-steps: S31, calculating the Pearson correlation coefficient and mutual information value between each feature vector in the multi-dimensional geological data set and the soil and rock parameters of the unsampled points, and determining the correlation degree between each feature vector and the soil and rock parameters of the unsampled points based on the weighted sum of the correlation coefficient and mutual information value; S32, setting a correlation degree threshold, filtering out feature vectors with a correlation degree higher than the threshold, and forming an initial low-dimensional feature dataset from the filtered feature vectors; S33, performing feature redundancy analysis on the initial low-dimensional feature dataset, calculating the cosine similarity between each feature vector, removing feature vectors with a similarity higher than the redundancy threshold, and retaining feature vectors with independent information to obtain the final low-dimensional feature dataset; S34, recording the feature selection rules and mapping parameters of the high-dimensional data dimensionality reduction mapping model in the dimensionality reduction process, forming a dimensionality reduction processing report for subsequent model reproduction and parameter adjustment.

[0067] Specifically, step S3 includes sub-steps S31-S34: S31 calculates the correlation between the feature vector and the soil and rock parameters of the unsampled points. The Pearson correlation coefficient is calculated based on ≥100 samples, and the mutual information value is estimated using the K-nearest neighbor method. The correlation is the weighted sum of the two (Pearson coefficient weight 0.6, mutual information value weight 0.4), and the result is rounded to 3 decimal places; S32 sets the correlation threshold. The threshold is set to 0.8 when the coefficient of variation of the soil and rock parameters is >0.3, and 0.6 when it is <0.2. After filtering, the dimensionality of the initial low-dimensional feature dataset is reduced by 60% compared to the original high-dimensional data. Step 1: S33 performs feature redundancy analysis, calculates cosine similarity based on standardized feature vectors, sets the redundancy threshold to 0.85, and removes feature vectors with similarity exceeding the threshold. The final low-dimensional feature dataset has a dimension of 3-5. Step 2: S34 generates a dimensionality reduction report, including feature selection rules such as correlation threshold and redundancy threshold, mapping parameters such as feature weights and dimension transformation matrices, and a comparison of dimensions before and after data processing. Through step-by-step selection and redundancy removal, the data dimension is reduced while retaining core information, reducing the computational complexity of subsequent model training and improving processing efficiency.

[0068] Preferably, as shown in Figure 4, step S4 includes the following sub-steps: S41, according to a preset sample division ratio, the low-dimensional feature dataset is randomly divided into a training sample set and a validation sample set, wherein the training sample set is used for model training and the validation sample set is used for model performance validation; S42, the parameters of the spatiotemporal feature random forest evaluation model are initialized, including the number of decision trees, the depth of the decision trees, the number of features for node splitting, and the initial values ​​of the spatiotemporal weight coefficients, and the maximum number of iterations and the error convergence threshold for model training are set; S43, the training sample set is input into the spatiotemporal feature random forest evaluation model, and an independent training sample subset is generated for each decision tree using the bootstrap sampling method. Decision trees are constructed based on the sample subsets. During the decision tree construction process, the information gain of each split node is calculated in combination with spatiotemporal features, and the feature with the largest information gain is selected as the split feature; S44, the validation sample set is input into the trained spatiotemporal feature random forest evaluation model, and the prediction error of the model is calculated. If the prediction error is greater than the error convergence threshold, the number of decision trees, the node splitting threshold, and the spatiotemporal weight coefficients are adjusted, and the process from S43 to S44 is repeated until the model prediction error is less than or equal to the error convergence threshold.

[0069] Specifically, step S4 includes sub-steps S41-S44: S41 uses stratified sampling to divide the low-dimensional feature dataset into training and validation sample sets at a 7:3 ratio, ensuring that the mean and standard deviation differences of each geotechnical parameter between the two sample sets are <5%. The sample sets are labeled with the collection time accurate to the day and the spatial coordinates accurate to the meter. S42 initializes the model parameters, initially setting the number of decision trees to 150, the maximum depth to 15 layers, the number of split features per node to the square root (integer) of the low-dimensional feature dimension, the maximum number of iterations to 200, and the error convergence threshold to 5%. S43 inputs the training sample set into the model and uses bootstrap sampling (sample selection probability 63.2%). A training subset for the decision tree is generated. When constructing the decision tree, the information gain calculation incorporates spatiotemporal features. The time weight of samples within 1 year is 0.9, 1-3 years is 0.7, and more than 3 years is 0.5. The spatial weight of samples within 100 meters is 0.8, 100-200 meters is 0.6, and more than 200 meters is 0.4. S44 uses the validation sample set to calculate the model prediction error (root mean square error). If the error exceeds the threshold, 20 decision trees are added or removed each time, the node splitting threshold is adjusted by ±0.05, and the spatiotemporal weight coefficient is adjusted by ±0.1 until the error is lower than the threshold. Through step-by-step sample processing and parameter optimization, the model initially has the ability to make accurate predictions and improves its adaptability to the spatiotemporal variation of soil and rock parameters.

[0070] Preferably, as shown in Figure 5, step S5 includes the following sub-steps: S51, determining the initial learning rate, batch sample size, and initial value of the regularization coefficient of the stochastic gradient descent optimizer, and setting the termination condition for the optimization iteration, wherein the termination condition includes the number of iterations reaching the maximum number of iterations or the change in the model prediction error being less than a preset error change threshold; S52, randomly selecting batch samples from the training sample set, inputting them into the spatiotemporal feature random forest evaluation model, calculating the loss function value of the model on the batch samples, and solving the gradient value of the model parameters based on the loss function value; S53, updating the feature weights and decision tree branch parameters of the model according to the gradient value and the initial learning rate, and introducing a regularization term to constrain the model parameters to prevent the model from overfitting; S54, calculating the prediction error of the optimized model on the validation sample set. If the termination condition is not met, adjusting the learning rate and regularization coefficient, returning to S52 to continue iterative optimization; if the termination condition is met, stopping the optimization and saving the optimized model parameters.

[0071] Specifically, step S5 includes sub-steps S51-S54: S51 Configure optimizer parameters: the initial learning rate is set to 0.05 for samples < 300 groups, 0.03 for 300-500 groups, and 0.01 for > 500 groups; the batch sample size is set to 64 groups for 64GB memory and 32 groups for 32GB memory; the regularization coefficient is initially set to 0.001; the termination condition is that the error change is < 0.001 after 500 iterations or 10 consecutive iterations; S52 Draw batch samples without replacement from the training sample set to input into the model, use the mean squared error as the loss function, calculate the gradient using the backpropagation method, and add a weight of 1 to outlier values ​​exceeding ±3 times the standard deviation. 5. Avoid interfering with the gradient direction; S53. Adjust model parameters based on the gradient and learning rate. The adjustment range of feature weights is ≤10% of the initial value each time, and the branch parameters of decision tree such as node splitting threshold are adjusted by ±0.02 each time. L2 regularization is used to prevent overfitting; S54. Calculate the validation set error after each iteration. If the termination condition is not reached, the learning rate is reduced to 0.8 times the original value and the regularization coefficient is increased by 0.0005 every 100 iterations. After the termination condition is reached, the model parameters such as feature weights and decision tree structure are saved. Through step-by-step iterative optimization, the model prediction error is further reduced, the stability and generalization ability are improved, and the subsequent evaluation results are ensured to be accurate and reliable.

[0072] The spatiotemporal feature random forest evaluation model in this invention is a machine learning model that integrates spatiotemporal dimensional information and uses voting by multiple decision trees to predict soil and rock parameters at unsampled points. The implementation process requires first dividing the dimensionality-reduced low-dimensional feature dataset into a training sample set and a validation sample set in a 7:3 ratio to ensure that the feature distribution of the two types of samples is consistent; initializing model parameters, setting the number of decision trees (initially 150), the upper limit of depth (15 layers), and the number of node split features (square root of the low-dimensional feature dimension), while incorporating spatiotemporal weights, i.e., in terms of time, the weight of samples within 1 year is 0.9, 1-3 years is 0.7, and more than 3 years is 0.5; in terms of space, the weight of samples within 100 meters is 0.8, 100-200 meters is 0.6, and more than 200 meters is 0.4; generating an independent training subset for each decision tree through bootstrap sampling, calculating the node information gain in combination with spatiotemporal features to construct the decision tree, and then using the validation sample set to verify the model. If the prediction error exceeds the 5% threshold, the number of decision trees (increasing or decreasing by 20 each time), the node split threshold (±0.05), and the spatiotemporal weights (±0.1) are adjusted until the error reaches the target. This model has the initial ability to predict soil and rock parameters. It can adapt to the changes in soil and rock parameters over time and in spatial distribution, avoid the prediction bias caused by the traditional model ignoring spatiotemporal differences, provide a basic model framework for subsequent iterative optimization of the optimizer, improve the initial accuracy of soil and rock parameter assessment at unsampled points, and provide a more realistic spatiotemporal prediction basis for engineering surveys.

[0073] The high-dimensional data dimensionality reduction mapping model is a processing model used to reduce the dimensionality of high-dimensional, multi-dimensional geological data sets formed by the fusion of multi-source geological data while preserving core features. The implementation first calculates the correlation between each feature vector and the soil and rock parameters of unsampled points. Based on ≥100 sets of samples, the Pearson correlation coefficient (weight 0.6) and mutual information value (weight 0.4) are calculated, and their weighted sum is taken as the correlation degree, rounded to three decimal places. Then, a correlation degree threshold is set according to the coefficient of variation of soil and rock parameters (0.8 for coefficient of variation > 0.3, 0.6 for < 0.2). Feature vectors with correlation degrees exceeding the threshold are selected to form an initial low-dimensional dataset, reducing its dimensionality by more than 60% compared to the original high-dimensional data. Subsequently, the cosine similarity between the feature vectors of the initial low-dimensional dataset is calculated, and redundant features are removed according to a redundancy threshold of 0.85, ultimately obtaining a 3-5 dimensional low-dimensional feature dataset. Finally, a dimensionality reduction processing report is generated, recording the selection rules, mapping parameters, and dimensionality comparison. This model removes redundant information and invalid features from high-dimensional data, reduces data dimensionality, and decreases the computational load and interference factors in subsequent model training. It provides high-quality, low-dimensional training data for the spatiotemporal feature random forest evaluation model, avoids the "curse of dimensionality" caused by high-dimensional data, improves model training efficiency and initial prediction accuracy, and ensures the smooth progress of the subsequent evaluation process.

[0074] The stochastic gradient descent optimizer is an optimization tool used to iteratively adjust the parameters of a trained spatiotemporal feature random forest evaluation model and reduce prediction error. Implementation involves configuring parameters first: setting the initial learning rate based on the sample size (0.05 for <300 groups, 0.03 for 300-500 groups, and 0.01 for >500 groups); setting the batch sample size based on server memory (64 groups for 64GB, 32 groups for 32GB); setting the initial regularization coefficient to 0.001; and terminating when the error change is <0.001 after 500 iterations or 10 consecutive iterations. Batch samples are then drawn from the training set without replacement and input into the model, using mean squared error as the loss function. The gradient is calculated using the backpropagation method, and outlier values ​​exceeding ±3 standard deviations are weighted by 1.5 to avoid interference. Model parameters are adjusted based on the gradient and learning rate, with feature weights adjusted by ≤10% of their initial values ​​each time, and decision tree branch parameters (such as node splitting thresholds) adjusted by ±0.02 each time. L2 regularization is used to prevent overfitting. After each iteration, the validation set error is checked. If the termination condition is not met, the learning rate is decayed to 0.8 times and the regularization coefficient is increased by 0.0005 every 100 iterations. Once the condition is met, the optimized parameters are saved. This optimizer continuously optimizes the model parameters, bringing the model prediction error to a preset range, improving model stability and generalization ability, solving the problems of insufficient initial model prediction accuracy and easy overfitting, further reducing the evaluation error, and ensuring that the final output of the geotechnical parameter evaluation results for unsampled points is accurate and reliable, meeting the engineering requirements for high-precision data.

[0075] The multi-source geological data fusion computing platform is a distributed computing architecture platform used to integrate various types of geological data and provide a unified data foundation for subsequent model processing. The implementation first involves building a hardware architecture with at least 8-core server-grade CPUs, at least 64GB of memory, and at least 2TB of storage to ensure the capacity for large-scale data processing. Spatial reference systems (2000 National Geodetic Coordinate System, 1985 National Height Datum), data accuracy (±1 meter for horizontal plane, ±0.5 meters for vertical plane), and attribute field information are extracted from multi-source geological data metadata. A comparison table is established, and a seven-parameter transformation method (error ≤ ±0.3 meters) is used to unify the data coordinate system. The merging standard for semantically similar fields (such as "compression modulus" and "deformation modulus") is clarified. The data is packaged according to the specified format and imported into the platform. After being uniformly converted to the same coordinate system, root mean square error is used for verification to ensure that over 95% of the data has a spatial deviation of < ±0.3 meters. Attribute fields are standardized, and missing values ​​are supplemented using the K-nearest neighbor interpolation method (selecting 5 similar samples from the surrounding area, with an error ≤ 10% of the sample standard deviation). Outliers are identified using the ±3 standard deviation method, and corrected using data within 100 meters of the surrounding area. The core fields (cohesion, internal friction angle) are verified to be free of missing values, forming a multi-dimensional geological data set of at least 10 dimensions. This platform eliminates differences in multi-source data formats, spatial deviations, and quality issues, achieving unified data integration and quality control. It addresses the shortcomings of traditional technologies, such as scattered multi-source data, low correlation, and excessive redundancy. It provides structurally unified and accurate basic data for dimensionality reduction of high-dimensional data, model training, and optimization, serving as the core data support carrier for the smooth implementation of the entire evaluation method.

[0076] The method for evaluating geotechnical parameters at unsampled points based on the random forest machine learning algorithm constructs a multi-source geological data fusion computing platform. This platform enables spatial matching and attribute integration of geotechnical physical and mechanical parameters, geological structural data, topographic and geomorphological data, stratigraphic and lithological data, and geophysical exploration data from sampled points. This solves the problems of low data correlation and excessive redundant information caused by the dispersion of multi-source data and the lack of effective integration mechanisms in traditional techniques. Simultaneously, the platform can eliminate contradictory data and correct outliers through logical consistency checks, forming a high-quality, multi-dimensional geological data set, laying a solid foundation for subsequent evaluation. Furthermore, by using a high-dimensional data dimensionality reduction mapping model to reduce the dimensions of multi-dimensional data, it can accurately retain feature vectors with high correlation to geotechnical parameters at unsampled points, eliminate invalid and redundant information, and avoid the problems of excessive information interference and low computational efficiency in traditional high-dimensional data processing, further improving data utilization efficiency and its supporting role in evaluation results.

[0077] This method utilizes a spatiotemporal random forest evaluation model, combining spatiotemporal features for training and validation. This fully considers the spatial and temporal variations of geotechnical parameters, addressing the problem of large discrepancies between evaluation results and actual conditions caused by traditional models' insufficient consideration of spatiotemporal characteristics. Simultaneously, the introduction of a stochastic gradient descent optimizer iteratively adjusts the model's feature weights and decision tree branch parameters, enabling the model's prediction error to converge to a preset range. This overcomes the shortcomings of traditional models, such as lack of efficient parameter optimization methods, low prediction accuracy, and difficulty in error control. Ultimately, the optimized model can accurately output geotechnical parameter evaluation results for unsampled points, providing more reliable data support for geotechnical engineering investigation and design, while balancing engineering safety and economic requirements.

[0078] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0079] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating soil and rock parameters at unsampled points based on a random forest machine learning algorithm, characterized in that, Includes the following steps: S1. Acquire multi-source geological data, including geophysical and mechanical parameters of sampled points, geological structural data, topographic and geomorphological data, stratigraphic and lithological data, and geophysical exploration data. S2. Construct a multi-source geological data fusion computing platform, import the multi-source geological data acquired in S1 into this platform, and perform spatial matching and attribute integration of different types of geological data through data association analysis to form a multi-dimensional geological data set. S3. Deploy a high-dimensional data dimensionality reduction mapping model on the multi-source geological data fusion computing platform to perform dimensionality reduction processing on the multi-dimensional geological data set formed in S2, retaining feature vectors in the dataset whose correlation with geophysical parameters of unsampled points is higher than a preset threshold, thus obtaining a low-dimensional feature dataset. S4. Call the spatiotemporal feature random forest evaluation model to evaluate the data obtained in S3. The low-dimensional feature dataset is divided into a training sample set and a validation sample set. The spatiotemporal feature random forest evaluation model is trained using the training sample set, and the number of decision trees, node splitting threshold, and spatiotemporal weight coefficients of the model are adjusted using the validation sample set. In S5, a stochastic gradient descent optimizer is introduced to iteratively optimize the parameters of the spatiotemporal feature random forest evaluation model trained in S4. The gradient value of the model prediction error is calculated in each iteration, and the feature weights and decision tree branch parameters of the model are adjusted according to the gradient value until the model prediction error converges to the preset range. In S6, the spatial coordinates of the unsampled points and the surrounding geological environment data are determined and input into the spatiotemporal feature random forest evaluation model optimized in S5. The geotechnical parameter evaluation results of the unsampled points are output through the model's multi-decision tree voting mechanism.

2. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, The expression for the spatiotemporal feature random forest evaluation model is as follows: Where F(x,y,t) is the geotechnical parameter evaluation value of the last sample point (x,y) at time t, K is the total number of decision trees in the spatiotemporal feature random forest, and ω k Let α be the weight coefficient of the k-th decision tree, M be the number of feature split nodes in each decision tree, and α be the weight coefficient of the k-th decision tree. k,m Let φ be the feature sensitivity coefficient of the m-th split node in the k-th decision tree. k,m (x,y,t) represents the spatiotemporal feature value of the unsampled point (x,y) at time t corresponding to the m-th split node of the k-th decision tree, θ k,m Let η be the feature threshold of the m-th split node in the k-th decision tree. k (x,y,t) represents the predicted geotechnical parameters of the unsampled point (x,y) at time t by the k-th decision tree.

3. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, The expression for the high-dimensional data dimensionality reduction mapping model is: Among them, Z p Let β be the p-th feature vector in the reduced-dimensional feature dataset, Q be the number of feature types in the original high-dimensional dataset, and β be the feature vector. p,q Let γ be the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature, where I is the number of data samples included in the q-th original high-dimensional feature, and γ is the mapping coefficient between the p-th low-dimensional feature and the q-th original high-dimensional feature. q,i X is the weight factor of the i-th sample in the original high-dimensional features of class q. q,i Let be the feature value of the i-th sample in the original high-dimensional features of class q, R be the number of spatiotemporal correction factors, and δ be the feature value of the i-th sample. p,r Let T be the correlation coefficient between the p-th low-dimensional feature and the r-th spatiotemporal correction factor. r Let r be the r-th spatiotemporal correction factor.

4. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, The expression for the stochastic gradient descent optimizer is: Where, θ t+1 To optimize the model parameters for round t+1, θ t To optimize the model parameters in the first t rounds, η t Let B be the learning rate in round t, B be the number of samples in each batch during each iteration, and b be the sample index within the batch. F(x) is the gradient of the model loss function with respect to the parameter θ. b ,y b ,t b ;θ) represents the model's pair of samples (x) b ,y b ,t b The predicted value of y b,true For sample (x) b ,y b ,t b The actual geotechnical parameter values, λ t Let be the regularization coefficient for round t.

5. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, The data fusion expression of the multi-source geological data fusion computing platform is as follows: Among them, D fusion (x,y,t) represents the geological data value at (x,y,t) after fusion, S represents the number of data source types for the multi-source geological data, and μ s Let D be the weight of the s-th type of data source. s (x,y,t) represents the original data value at (x,y,t) of the s-th data source, U represents the number of metrics for data consistency verification, and ξ s,u Let be the deviation coefficient of the u-th verification index corresponding to the s-th type of data source. Let be the average data value of the u-th verification index in the region surrounding (x,y,t).

6. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, The corrected expression for the geotechnical parameter evaluation results of the unsampled points is: Among them, P final (x0,y0,t0) represents the final geotechnical parameter assessment value of the unsampled point (x0,y0,t0), F(x0,y0,t0) represents the initial predicted value of the spatiotemporal characteristic random forest assessment model, V represents the number of sampled points surrounding the unsampled point, and ε represents the final geotechnical parameter assessment value. v Let P be the influence coefficient of the v-th surrounding sampled point on the unsampled point. v (x v ,y v ,t v ) represents the v-th surrounding sampled point (x) v ,y v ,t v The actual geotechnical parameters.

7. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, S2 includes the following sub-steps: S21, extracting spatial reference system, data precision, and attribute field information of each data source from the metadata of multi-source geological data, establishing a data source information comparison table, and clarifying the spatial coordinate transformation relationship and attribute mapping rules between different data sources; S22, based on the spatial coordinate transformation relationship, uniformly transforming the geological data from different data sources to the same spatial coordinate system, correcting the spatial position deviation of the transformed data, and ensuring that the spatial alignment accuracy of the data meets the preset requirements; S23, according to the attribute mapping rules, standardizing the attribute fields of different data sources, merging semantically similar attribute fields, and supplementing missing attribute values ​​using an interpolation method based on neighborhood similarity to form preliminary integrated geological data; S24, performing logical consistency verification on the preliminary integrated geological data, removing data records with obvious logical contradictions, marking data outliers using a statistical distribution-based identification method, and correcting the marked outliers through comparative analysis with surrounding data, ultimately forming a multi-dimensional geological data set.

8. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, S3 includes the following sub-steps: S31, calculate the Pearson correlation coefficient and mutual information value between each feature vector in the multi-dimensional geological data set and the soil and rock parameters of the unsampled points, and determine the correlation degree between each feature vector and the soil and rock parameters of the unsampled points based on the weighted sum of the correlation coefficient and mutual information value; S32, set a correlation degree threshold, filter out feature vectors with a correlation degree higher than the threshold, and form an initial low-dimensional feature dataset from the filtered feature vectors; S33, perform feature redundancy analysis on the initial low-dimensional feature dataset, calculate the cosine similarity between each feature vector, remove feature vectors with a similarity higher than the redundancy threshold, and retain feature vectors with independent information to obtain the final low-dimensional feature dataset; S34, record the feature selection rules and mapping parameters of the high-dimensional data dimensionality reduction mapping model in the dimensionality reduction process, and form a dimensionality reduction processing report for subsequent model reproduction and parameter adjustment.

9. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, S4 includes the following steps: S41, according to a preset sample division ratio, the low-dimensional feature dataset is randomly divided into a training sample set and a validation sample set, wherein the training sample set is used for model training and the validation sample set is used for model performance validation; S42, the parameters of the spatiotemporal feature random forest evaluation model are initialized, including the number of decision trees, the depth of the decision trees, the number of features for node splitting, and the initial values ​​of the spatiotemporal weight coefficients, and the maximum number of iterations and the error convergence threshold for model training are set; S43, the training sample set is input into the spatiotemporal feature random forest evaluation model, and an independent training sample subset is generated for each decision tree using the bootstrap sampling method. Decision trees are constructed based on the sample subsets. During the decision tree construction process, the information gain of each splitting node is calculated in combination with spatiotemporal features, and the feature with the largest information gain is selected as the splitting feature; In step S44, the validation sample set is input into the trained spatiotemporal feature random forest to evaluate the model. The prediction error of the model is calculated. If the prediction error is greater than the error convergence threshold, the number of decision trees, the node splitting threshold, and the spatiotemporal weight coefficient are adjusted. The process from S43 to S44 is repeated until the model prediction error is less than or equal to the error convergence threshold.

10. The method for evaluating unsampled soil and rock parameters based on the random forest machine learning algorithm according to claim 1, characterized in that, S5 includes the following sub-steps: S51, determining the initial learning rate, batch size, and initial values ​​of the regularization coefficient for the stochastic gradient descent optimizer, and setting the termination condition for the optimization iteration, including the number of iterations reaching the maximum number of iterations or the change in model prediction error being less than a preset error change threshold; S52, randomly selecting batch samples from the training sample set, inputting them into the spatiotemporal feature random forest evaluation model, calculating the loss function value of the model on the batch samples, and solving for the gradient value of the model parameters based on the loss function value; S53, updating the feature weights and decision tree branch parameters of the model according to the gradient value and the initial learning rate, while introducing a regularization term to constrain the model parameters to prevent overfitting; S54, calculating the prediction error of the optimized model on the validation sample set. If the termination condition is not met, adjusting the learning rate and regularization coefficient, returning to S52 to continue iterative optimization; if the termination condition is met, stopping the optimization and saving the optimized model parameters.