A tunnel large deformation prediction method based on geological feature adaptation and PSO-RF algorithm

The tunnel large deformation prediction method based on geological feature adaptation and PSO-RF algorithm solves the problems of accuracy and stability in tunnel large deformation prediction under complex geological conditions, realizes high-precision and low-cost tunnel large deformation prediction, and provides data-driven basis for key risk factors.

CN121562429BActive Publication Date: 2026-05-08中国水利水电第七工程局有限公司 +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for predicting large tunnel deformations lack accuracy and stability under complex geological conditions, and have poor adaptability to geological features, resulting in limited model generalization ability.

Method used

A method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm is adopted. Through data preprocessing, geological scene clustering, improved PSO joint optimization and model training, a weighted feature matrix is ​​constructed, the hyperparameters and feature weights of the RF model are optimized, and the large tunnel deformation is predicted in combination with geological features.

Benefits of technology

It improves the accuracy and generalization ability of tunnel large deformation prediction, reduces reliance on human experience, lowers implementation costs, provides data-driven basis for key risk factors, and is applicable to tunnel engineering under complex geological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562429B_ABST
    Figure CN121562429B_ABST
Patent Text Reader

Abstract

The application discloses a tunnel large deformation prediction method based on geological feature adaptation and PSO-RF algorithm and belongs to the technical field of tunnel engineering safety monitoring. The method aims at the problems of poor geological condition adaptability and insufficient prediction accuracy of the prior art, and a whole-process technical framework of "data preprocessing-geological scene clustering-PSO parameter optimization-RF prediction-result verification" is constructed. Firstly, nine core indexes are collected, and after missing value processing and standardization, the geological scene is divided into three categories of soft rock water-rich section, fault broken section and hard rock stable section by K-means clustering. Then, the improved PSO algorithm is used to simultaneously optimize the RF hyperparameters and feature weights, and the parameter self-adaptation is realized through the scene inertia weight adjustment and the fitness function constraint. Finally, the weighted feature matrix is constructed to train the PSO-RF model, and the I-V grade large deformation grade prediction results and feature importance ranking are output. Experiments show that the application has high accuracy and low misjudgment rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tunnel engineering safety monitoring technology, and in particular to a method for predicting large deformations in tunnels based on geological feature adaptation and the PSO-RF algorithm. Background Technology

[0002] Tunnel engineering is a core component of transportation infrastructure. However, large deformation is a common problem in tunnel construction and operation under complex geological conditions such as soft rock and high ground stress. It can easily lead to damage to the support structure, construction stagnation, and even safety accidents. Therefore, accurate prediction of the large deformation level is of great significance for risk prevention and control. Early tunnel deformation prediction mainly relied on traditional empirical mathematical models, such as regression analysis and time series analysis. In "Prediction of Surrounding Rock Deformation Combination Based on Regression Analysis and Grey Theory" (Journal of Underground Space and Engineering, 2017, Vol. 13, No. S1, pp. 48-51, DOI:10.20174), Wang Tao, Sun Wenlong, and Li Lei used a method combining regression analysis and grey theory to carry out short- and medium-term combination prediction of tunnel surrounding rock deformation, aiming to improve prediction accuracy. However, this type of method relies on linear assumptions and has limited generalization ability in nonlinear and high-dimensional geological data.

[0003] With the development of intelligent detection technology, machine learning methods such as Random Forest (RF), Support Vector Machine (SVM), and Convolutional Neural Network (CNN) have been introduced into the field of tunnel deformation prediction. Huang Agang et al., in "Analysis of Tunnel Operation Safety Status Based on Chaos-RF-SVM Deformation Prediction Model" (Surveying and Mapping Engineering, 2022, Vol. 31, No. 4, pp. 52-56, DOI: 10.19349), constructed a prediction model based on operational tunnel monitoring data, combining chaos theory, the RF algorithm, and SVM, and verified its reliability through the MK test, demonstrating the advantages of intelligent algorithms in improving accuracy. Lin Guangdong, in "Prediction of Cumulative Settlement in the Early Stage of Tunnel Construction Based on Random Forest" (Computational Technology and Automation, 2022, Vol. 41, No. 1, pp. 160-163, DOI: 10.16339), further confirmed that the RF model outperforms deep neural networks in predicting cumulative settlement in the early stage of tunnel construction, highlighting its nonlinear processing capabilities and anti-overfitting characteristics. However, these models are still limited by the subjectivity of hyperparameter selection, and inappropriate parameters can restrict prediction accuracy and generalization ability.

[0004] In recent years, the integration of machine learning and intelligent optimization algorithms has become a research hotspot. Du Junsheng et al. (patent CN202510683628.4, "Method for Identifying Rock Deformation in Tunnels in Water-Rich Composite Strata") used a Random Forest (RF) model to classify rock deformation. By leveraging a trained rock deformation classification model, they analyzed the current rock data to be identified, obtaining the first deformation data. The powerful classification capability of the Random Forest model enabled a preliminary judgment of the rock deformation tendency, improving the accuracy and efficiency of identification. However, when dealing with water-rich composite strata, the weights and thresholds rely heavily on historical experience, making parameter determination difficult. Wu Xianguo et al. (patent CN116050603A, "Method and Equipment for Predicting and Optimizing Deformation of Cut-and-Cut Tunnels Based on Hybrid Intelligent Methods") used Bayesian optimization (BO) to adjust RF hyperparameters. However, for small-spacing cut-and-cut tunnels, the parameter constraints relied on manual setting, exhibiting strong subjectivity. Yao Zhixiong et al. (patent CN202311096268.5) proposed a prediction method based on PSO-LSTM, verifying the effectiveness of the Particle Swarm Optimization (PSO) algorithm in optimizing time series models; this confirms the effectiveness of the PSO algorithm in optimizing time series prediction models. Bo Yin et al. used an improved Black-winged Kite algorithm to optimize hyperparameters, reflecting the technological trend of metaheuristic algorithms. Furthermore, Wang Shudong pointed out in "Research on Large Deformation Control Technology of Tunnel in Weak Surrounding Rock in Complex Geostress Zones" ([D]. Beijing Jiaotong University, 2010.) that due to various uncertainties in geological conditions, tunnel construction in squeezing weak surrounding rock is prone to deformation and collapse. It is evident that geological conditions can lead to deformation.

[0005] While significant progress has been made in combining intelligent optimization algorithms with machine learning models, existing technologies still have shortcomings in adapting to complex geological conditions. On the one hand, whether the model's input features can fully and effectively reflect key geological attributes directly determines the engineering applicability of the prediction model. Gao Chengbo et al. (CN202510188244.5), in their method for predicting large deformations in tunnels, emphasized the importance of feature selection through mutual information and mRMR algorithms to screen out the most representative indicators and improve model performance. Mu Linlong et al. (CN202510475938.7) constructed combined features with clear physical meaning through physical-guided feature engineering to enhance the model's predictive ability and physical interpretability under limited sample conditions. On the other hand, the prediction method developed by Tan Zhongsheng et al. (CN202311126073.0) emphasizes the necessity of dynamic and rapid prediction by combining geostress characteristics with surrounding rock properties. These studies all indicate that an effective prediction model must match the geological characteristics of a specific project.

[0006] Existing methods either focus on complex deep learning models, which have high training costs and require large amounts of data; or fail to fully consider the intrinsic relationship between geological features and model parameters, resulting in insufficient prediction stability and accuracy under complex and variable geological conditions.

[0007] Therefore, there is an urgent need in this field for a prediction method that can integrate geological feature adaptation, automatic parameter optimization, and strong interpretability to overcome the above-mentioned shortcomings. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for predicting large deformation of tunnels based on geological feature adaptation and PSO-RF algorithm.

[0009] The objective of this invention is achieved through the following technical solution: a method for predicting large deformations of tunnels based on geological feature adaptation and the PSO-RF algorithm, comprising the following steps:

[0010] Data preprocessing stage: Collect core indicator data of tunnel engineering and perform standardization processing, including missing value imputation, outlier correction and data transformation, and then output a standardized feature matrix;

[0011] Geological scene clustering stage: Based on the encoded lithology and geostress as the core clustering features, the K-means algorithm is used for cluster analysis to eliminate dimensional differences, determine the optimal number of clusters, and divide the geological scenes into three categories: soft rock water-rich section, fault fracture section, and hard rock stable section. Scene labels are assigned to each sample.

[0012] Improve the PSO joint optimization stage: Construct a multi-dimensional particle vector, corresponding to the hyperparameters and feature weights of the core index data of the RF model, and output the optimal hyperparameters and optimal feature weight vectors through scenario-based inertia weight adjustment and fitness function optimization.

[0013] Model training and large deformation prediction stage: The preprocessed standardized feature matrix is ​​multiplied by the optimal feature weight vector to construct a weighted feature matrix. The RF model is trained using the optimal hyperparameters and cross-validation to obtain the PSO-RF model. The trained PSO-RF model is used to predict the level of large deformation of the tunnel and output the feature importance ranking.

[0014] Preferably, the core indicator data include lithology, burial depth, integrity, rock strength, weathering degree, rock mass strength, geostress, strength-stress ratio, and deformation rate.

[0015] Preferably, the missing value imputation in the data preprocessing stage includes the following steps: filling missing values ​​of categorical variables with "unknown" and encoding them as 0; filling missing values ​​of numerical variables with the arithmetic mean of the same lithological segment.

[0016] Outlier correction uses the 3σ criterion to identify outliers, and for the 3σ criterion... j Numerical feature indicators X j Calculate its average value within the same lithological section. with standard deviation When the sample value satisfies Values ​​that are identified as outliers are replaced by the 95th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as upper-side outliers, and by the 5th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as lower-side outliers.

[0017] Data transformation includes the following steps: converting categorical variables into numerical vectors using one-hot encoding, and mapping numerical variables to the [0,1] interval using min-max normalization.

[0018] Preferably, the classification variables include lithology, integrity, and weathering degree; the numerical variables include burial depth, rock strength, rock mass strength, geostress, strength-stress ratio, and deformation rate.

[0019] Preferably, in the geological scene clustering stage, the lithology coding value and the normalized value of geostress are standardized using StandardScaler to eliminate dimensional differences.

[0020] The elbow method was used to analyze the trend of the sum of squared clustering errors (SSE) with the number of clusters K. When K=3, the SSE curve showed an obvious inflection point, and the optimal number of clusters was determined to be 3.

[0021] The K-means algorithm was used for scene clustering. Geological scenes were divided according to the lithology code and geostress value of the cluster center and named as soft rock water-rich segment, fault fracture segment and hard rock stable segment.

[0022] Preferably, in the improved PSO joint optimization stage, the particle vector is 13-dimensional, with the first 4 dimensions corresponding to the hyperparameters of the random forest model and the last 9 dimensions corresponding to the feature weights of the 9 core indicator data.

[0023] The fitness function is: ,in MSE To predict the mean squared error; The number of decision trees; Weights are the core features of the scene.

[0024] Preferably, the hyperparameters of the random forest model include the number of decision trees, the maximum depth of the decision trees, the minimum number of samples required for node splits, and the maximum number of features per tree.

[0025] Preferably, in the model training and large deformation prediction stage, the large deformation level is divided into multiple levels according to the amount of deformation.

[0026] Preferably, the model validation stage is also included, in which a confusion matrix is ​​used to visually evaluate the prediction results, with the true level label as the vertical axis and the predicted level as the horizontal axis, to analyze the number of correctly predicted samples on the main diagonal and the misjudgment situation on the off-diagonal.

[0027] The beneficial effects of this invention are:

[0028] 1) By combining the excellent parameter optimization capability of the PSO algorithm with the robust prediction performance of the RF model, and by introducing targeted geological feature analysis and adaptation mechanisms, an intelligent prediction model with high prediction accuracy, strong generalization ability and clear engineering relevance was constructed.

[0029] 2) This invention utilizes the K-means clustering algorithm to classify geological conditions into three scenarios based on core features of lithology and geostress: soft rock water-rich sections, fault fracture sections, and hard rock stable sections. Combined with scenario-based parameter constraints (such as dynamic adjustment of inertia weights), the model can adapt to different geological environments. This refined classification avoids the limitations of a "one-size-fits-all" approach and significantly enhances the model's generalization ability and engineering applicability under complex geological conditions. For example, in soft rock water-rich sections, the model automatically strengthens the weight allocation for lithology encoding and geostress, thereby more accurately capturing deformation risks.

[0030] 3) The improved PSO algorithm uses 13-dimensional particle vectors to simultaneously optimize RF hyperparameters and feature weights, and effectively balances global search and local convergence through scenario-based inertia weight adjustment (e.g., linearly reducing from the initial 0.8 to 0.45) and fitness function constraints. Figure 3 As shown, the fitness curve drops rapidly in the early stages of iteration and stabilizes after 30 iterations, demonstrating the efficiency of the optimization process and avoiding the subjectivity and time-consuming problems of traditional manual parameter tuning.

[0031] 4) Through training with a weighted feature matrix and 5-fold cross-validation, the PSO-RF model can accurately classify the level IV large deformation. By jointly optimizing the adaptive learning of feature weights using PSO-RF, the risk of overfitting is reduced. Validation on the test set shows that the model has a high percentage of correct predictions on the main diagonal of the confusion matrix across 593 samples (e.g., 50 cases for level I and 59 cases for level III), and most misclassifications are for adjacent levels (e.g., only 1 case of level II being misclassified as level I), with no serious cross-level errors. The overall accuracy is better than traditional regression methods or single machine learning models.

[0032] 5) The model output not only includes deformation level predictions but also provides a ranking of feature importance based on the reduction in the Gini coefficient (such as the contribution of core indicators like strength-stress ratio and in-situ stress). Figure 4As shown, the feature importance bar chart intuitively displays the degree of influence of various geological indicators on the prediction results, helping engineers identify key risk factors and providing data-driven basis for targeted prevention and control measures (such as support design).

[0033] 6) This invention reduces reliance on human experience through standardized preprocessing (such as missing value imputation and outlier correction) and automated processes (such as elbow method for determining cluster numbers and PSO adaptive stopping), making it suitable for new tunnel projects lacking historical data. Furthermore, the method is based on nine easily accessible core indicators, eliminating the need for complex sensor deployments and reducing implementation costs. Attached Figure Description

[0034] Figure 1 This is a flowchart of the method of the present invention;

[0035] Figure 2 A scatter plot of the clustering results for geological scenes;

[0036] Figure 3 A fitness change curve for the PSO optimization process;

[0037] Figure 4 A bar chart showing the importance of features in a random forest.

[0038] Figure 5 Heatmap of the confusion matrix for model prediction results. Detailed Implementation

[0039] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] See Figures 1-5 This invention provides a technical solution: a method for predicting large deformation of tunnels based on geological feature adaptation and the PSO-RF algorithm, comprising the following steps:

[0041] Data preprocessing stage: Collect core indicator data of tunnel engineering and perform standardization processing, including missing value imputation, outlier correction and data transformation, and then output a standardized feature matrix;

[0042] Geological scene clustering stage: Based on the encoded lithology and geostress as the core clustering features, the K-means algorithm is used for cluster analysis to eliminate dimensional differences, determine the optimal number of clusters, and divide the geological scenes into three categories: soft rock water-rich section, fault fracture section, and hard rock stable section. Scene labels are assigned to each sample.

[0043] Improve the PSO joint optimization stage: Construct a multi-dimensional particle vector, corresponding to the hyperparameters and feature weights of the core index data of the RF model, and output the optimal hyperparameters and optimal feature weight vectors through scenario-based inertia weight adjustment and fitness function optimization.

[0044] Model training and large deformation prediction stage: The preprocessed standardized feature matrix is ​​multiplied by the optimal feature weight vector to construct a weighted feature matrix. The RF model is trained using the optimal hyperparameters and cross-validation to obtain the PSO-RF model. The trained PSO-RF model is used to predict the level of large deformation of the tunnel and output the feature importance ranking.

[0045] In some embodiments, the core indicator data include lithology, burial depth, integrity, rock strength, weathering degree, rock mass strength, geostress, strength-stress ratio, and deformation rate.

[0046] In this embodiment, core geological-engineering parameters are collected through multiple channels and standardized to construct a high-quality dataset, providing a foundation for subsequent model training. The specific execution process is as follows: Statistical analysis of measured data is conducted to select nine key prediction indicators, covering four dimensions: lithological characteristics, stress conditions, rock mass properties, and deformation characteristics. Static indicators such as lithology, integrity, and weathering degree are obtained from geological survey reports; parameters such as in-situ stress and rock strength are obtained through in-situ stress testing; and indicators such as lithology and integrity are updated through tunnel face logging. Standardized preprocessing, missing value imputation: missing values ​​of categorical variables (lithology, integrity, weathering degree) are filled with "unknown" and coded as 0; missing values ​​of numerical variables (burial depth, rock strength, etc.) are filled with the mean of the same lithological segment (e.g., missing burial depth values ​​of soft rock segments are taken as the arithmetic mean of the measured burial depths of that lithological segment); outlier correction: outliers are identified using the 3σ criterion, and the 95th or 5th quantile of the index is used to replace and correct outliers; data transformation: categorical variables are converted into numerical vectors using one-hot encoding; numerical variables are mapped to the [0,1] interval using min-max normalization.

[0047] The raw data for tunnel engineering covers nine core indicators, including lithology, burial depth, and integrity. Data preprocessing involves encoding the classification indicators and handling outliers and missing values ​​to output a standardized feature matrix. Geological scene segmentation uses encoded lithology and geostress as core clustering features to segment geological scenes, achieving refined classification of complex geological conditions. Model training first optimizes the four hyperparameters and nine-axis feature weights of the random forest using an improved PSO algorithm, then trains the model with the optimized parameters, incorporating 5-fold cross-validation to enhance generalization ability.

[0048] In some embodiments, missing value imputation in the data preprocessing stage includes the following steps: filling missing values ​​of categorical variables with "unknown" and encoding them as 0; filling missing values ​​of numerical variables with the arithmetic mean of the same lithological segment.

[0049] Outlier correction uses the 3σ criterion to identify outliers. For each numerical variable (denoted as the 3σ criterion), outliers are identified. j Numerical feature indicators X j ), calculate its mean value within the same lithological section. with standard deviation When the sample value satisfies Values ​​that are identified as outliers are replaced by the 95th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as upper-side outliers, and by the 5th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as lower-side outliers.

[0050] Data transformation includes the following steps: converting categorical variables into numerical vectors using one-hot encoding, and mapping numerical variables to the [0,1] interval using min-max normalization.

[0051] In some embodiments, the classification variables include lithology, integrity, and weathering degree; the numerical variables include burial depth, rock strength, rock mass strength, geostress, strength-stress ratio, and deformation rate.

[0052] In some embodiments, during the geological scene clustering stage, the lithology coding value and the normalized value of geostress are standardized using StandardScaler to eliminate dimensional differences.

[0053] The elbow method was used to analyze the trend of the sum of squared clustering errors (SSE) with the number of clusters K. When K=3, the SSE curve showed an obvious inflection point, and the optimal number of clusters was determined to be 3.

[0054] The K-means algorithm was used for scene clustering. Geological scenes were divided according to the lithology code and geostress value of the cluster center and named as soft rock water-rich segment, fault fracture segment and hard rock stable segment.

[0055] In this embodiment, scene clustering is performed based on core geological features to achieve refined classification of complex geological conditions. Encoded lithology and geostress are selected as core clustering features to provide a basis for subsequent scene-based parameter optimization. Lithology directly determines the deformation resistance of rock masses, while geostress is a key dynamic factor inducing large deformations. The coupling effect of the two dominates the core differences in geological scenes, and the data is easy to obtain and highly stable. Feature standardization: The coded lithology values ​​and normalized geostress values ​​are standardized using StandardScaler to eliminate dimensional differences. Determination of the optimal number of clusters: The Elbow Method is used to analyze the trend of the sum of squared clustering errors (SSE) with the number of clusters K. When K=3, the SSE curve shows a clear inflection point, and the optimal number of clusters is determined to be 3.

[0056] Scene clustering and naming: The K-means clustering algorithm is executed (iterations = 100, initial cluster centers are randomly generated). Based on the lithology code and geostress value of the cluster centers, the geological scenes are named into three categories:

[0057] Scenario 1: Water-rich section of soft rock (lithology code 1-2, normalized geostress value ≥ 0.6);

[0058] Scenario 2: Fault fracture segment (lithological code 1-3, normalized geostress value 0.3-0.6);

[0059] Scenario 3: Stable section of hard rock (lithology code 3, normalized geostress value ≤ 0.3);

[0060] Scene label assignment: Assign a corresponding scene label to each sample to form structured data of "feature matrix + scene label".

[0061] like Figure 2 As shown, to improve the accuracy of model training and reduce training time, multiple samples were first divided into scenarios based on indicators. K-means clustering was used combined with feature criteria to classify geological scenarios, selecting "lithology" and "geological stress" as core clustering features (reflecting the core differences in geological scenarios), and standardization was performed using StandardScaler. The elbow method was used to determine the optimal number of clusters, and K=3 was finally determined. Clustering was then performed with K=3 to obtain the scenario label for each sample. Finally, based on the lithology code and geostress value of the cluster center, the three scenarios were named soft rock water-rich segment, fault fracture segment, and hard rock stable segment, respectively.

[0062] In some embodiments, in the improved PSO joint optimization stage, the particle vector is 13-dimensional, with the first 4 dimensions corresponding to the hyperparameters of the random forest model and the last 9 dimensions corresponding to the feature weights of the 9 core index data.

[0063] The fitness function is: ,in MSE To predict the mean squared error; The number of decision trees; Weights are the core features of the scene.

[0064] In this embodiment, a scenario-based improved PSO algorithm is adopted to simultaneously optimize the hyperparameters and 9 feature weights of the RF model, construct an optimization mechanism adapted to different geological scenarios, and optimize the output training set feature matrix and scene label vector to obtain the optimal particle vector: 13-dimensional parameters, the first 4 being hyperparameters and the last 9 being feature weights.

[0065] The particle encoding consists of a 13-dimensional particle vector, with the following structure:

[0066] The optimization process records include: fitness change curves (horizontal axis = number of iterations, vertical axis = optimal fitness value, with the number of convergence iterations marked), and tables of inertia weight changes for each scenario (e.g., ω decreases from 0.9 to 0.4 in water-rich soft rock sections). This simultaneously addresses the problems of "difficulty in selecting hyperparameters and imbalance in feature weights" in RF, and improves the model's prediction accuracy by optimizing parameters to better match geological conditions through scenario-based optimization.

[0067] like Figure 3 As shown, when optimizing RF parameters using the improved PSO algorithm, the "Best Fitness" value changes with the number of iterations, with a maximum of 50 iterations. Before optimization, parameter encoding and fitness function definition are performed. The particle encoding rule is that the particle vector has a 13-dimensional dimension; the first 4 dimensions correspond to the 4 hyperparameters of the RF, and the last 9 dimensions correspond to the weights of the 9 core geological features. The fitness function comprehensively considers three objectives: prediction accuracy, model efficiency, and feature importance. During iterations 1-10, which are the initialization and rapid decline phases, the initial fitness value is generally high (approximately 0.075). A scenario-based dynamic adjustment strategy is adopted, with the initial inertia weight calculated based on the sample proportions of three geological scenarios (soft rock water-rich sections, fault fracture sections, and hard rock stable sections). As the number of iterations increases, the inertia weight linearly decreases from 0.8 to 0.45. In the final stage, after 30 iterations, the inertia weight decreases to 0.45, the particle velocity approaches 0, and the position remains essentially unchanged, avoiding parameter oscillations. In the later stages, the feature variance decreases (particle weight allocation converges), the mutation probability drops to 0.1, and only minor parameter adjustments are made to ensure the stability of the optimal solution.

[0068] In some embodiments, the hyperparameters of the random forest model include the number of decision trees, the maximum depth of the decision trees, the minimum number of samples required for node splits, and the maximum number of features per tree.

[0069] In some embodiments, the large deformation level in the model training and large deformation prediction stage is divided into multiple levels according to the amount of deformation.

[0070] In this embodiment, the preprocessed feature matrix is ​​multiplied by the optimal feature weights to construct a weighted feature matrix (highlighting the influence of core features); the RF model is trained using optimal hyperparameters and 5-fold cross-validation; large deformation levels are divided into I-V according to the amount of deformation (Level I ≤ 50mm, Level II 50-100mm, Level III 100-200mm, Level IV 200-300mm, Level V > 300mm); the large deformation level prediction results of the test set are output, and the feature importance (based on the average reduction of the Gini coefficient of the decision tree node) is calculated and the contribution ranking is output.

[0071] like Figure 4 As shown, after the model training is completed, the feature importance is calculated using the reduction in node impurity (reduction in Gini coefficient). For each step of the decision tree in the RF process, all nodes used for feature splitting are traversed. The reduction in node impurity due to the use of a particular feature during splitting is calculated. If a feature is used for splitting multiple times in the decision tree, and each split significantly reduces node impurity (e.g., the parent node's Gini coefficient decreases from 0.6 to the child node's 0.2 after a strength-stress ratio split), then the feature has a higher importance score in that single tree. The average importance score of the same feature across all trees is then used to obtain the "global importance score" for that feature.

[0072] In some embodiments, a model validation stage is also included, in which a confusion matrix is ​​used to visually evaluate the prediction results, with the true level label as the vertical axis and the predicted level as the horizontal axis, to analyze the number of correctly predicted samples on the main diagonal and the misjudgment situation on the off-diagonal.

[0073] In this embodiment, as Figure 5 As shown, the confusion matrix of the optimized PSO-RF model is plotted with the true level label as the ordinate and the predicted level as the abscissa. Labels 0-4 correspond to large deformation levels I-V. The matrix size represents the sample size. 593 sets of samples from tunnels including the right mileage of the No. 1 cross passage of the Moxi Tunnel, the left mileage of the No. 2 inclined shaft of the Moxi Tunnel, the Chaluo Tunnel, and the Wagang Tunnel were used as the dataset for model training and validation. The results show that the values ​​on the main diagonal, i.e., the number of correct predictions, are significantly higher than those on the off-diagonal. For example, 50 cases were correctly predicted for level I, 59 for level III, and 27 for level IV, indicating high overall classification accuracy. The off-diagonal model only exhibits a small number of cross-level misclassifications (e.g., 1 case of level II being misclassified as level I), with no serious cross-level errors. Comparison with the actual validation levels reveals that the model can accurately predict the levels of multiple large deformations.

[0074] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for predicting large deformation of tunnels based on geological feature adaptation and PSO-RF algorithm, characterized in that: Includes the following steps: Data preprocessing stage: Collect core indicator data of tunnel engineering and perform standardization processing, including missing value imputation, outlier correction and data transformation, and then output a standardized feature matrix; Geological scene clustering stage: Based on the encoded lithology and geostress as the core clustering features, the K-means algorithm is used for cluster analysis to eliminate dimensional differences, determine the optimal number of clusters, and divide the geological scenes into three categories: soft rock water-rich section, fault fracture section, and hard rock stable section. Scene labels are assigned to each sample. Improve the PSO joint optimization stage: Construct a multi-dimensional particle vector, corresponding to the hyperparameters and feature weights of the core index data of the RF model, and output the optimal hyperparameters and optimal feature weight vectors through scenario-based inertia weight adjustment and fitness function optimization. Model training and large deformation prediction stage: The preprocessed standardized feature matrix is ​​multiplied by the optimal feature weight vector to construct a weighted feature matrix. The RF model is trained using the optimal hyperparameters and cross-validation to obtain the PSO-RF model. The trained PSO-RF model is used to predict the level of large deformation of the tunnel and output the feature importance ranking. Missing value imputation in the data preprocessing stage includes the following steps: filling missing values ​​of categorical variables with "unknown" and encoding them as 0; filling missing values ​​of numerical variables with the arithmetic mean of the same lithological segment. Outlier correction uses the 3σ criterion to identify outliers, and for the 3σ criterion... j Numerical feature indicators X j Calculate its average value within the same lithological section. with standard deviation When the sample value satisfies Values ​​that are identified as outliers are replaced by the 95th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as upper-side outliers, and by the 5th percentile of the numerical characteristic index within the same lithological segment for values ​​identified as lower-side outliers. Data transformation includes the following steps: converting categorical variables into numerical vectors using one-hot encoding, and mapping numerical variables to the [0,1] interval using min-max normalization; In the geological scene clustering stage, the lithology coding value and the normalized value of geostress are standardized using StandardScaler to eliminate dimensional differences. The elbow method was used to analyze the trend of the sum of squared clustering errors (SSE) with the number of clusters K. When K=3, the SSE curve showed an obvious inflection point, and the optimal number of clusters was determined to be 3. The scene clustering uses the K-means algorithm to divide the geological scenes according to the lithology code and geostress value of the cluster center, and name them as soft rock water-rich section, fault fracture section and hard rock stable section; In the improved PSO joint optimization stage, the particle vector is 13-dimensional, with the first 4 dimensions corresponding to the hyperparameters of the random forest model and the last 9 dimensions corresponding to the feature weights of the 9 core index data. The fitness function is: ,in MSE To predict the mean squared error; The number of decision trees; Weights are the core features of the scene.

2. The method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm according to claim 1, characterized in that: The core indicator data include lithology, burial depth, integrity, rock strength, weathering degree, rock mass strength, geostress, strength-stress ratio, and deformation rate.

3. The method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm according to claim 1, characterized in that: The classification variables include lithology, integrity, and weathering degree; the numerical variables include burial depth, rock strength, rock mass strength, geostress, strength-stress ratio, and deformation rate.

4. The method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm according to claim 1, characterized in that: The hyperparameters of the random forest model include the number of decision trees, the maximum depth of the decision trees, the minimum number of samples required for node splits, and the maximum number of features per tree.

5. The method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm according to claim 1, characterized in that: In the model training and large deformation prediction stage, the large deformation level is divided into multiple levels according to the amount of deformation.

6. The method for predicting large tunnel deformation based on geological feature adaptation and PSO-RF algorithm according to any one of claims 1-5, characterized in that: It also includes a model validation phase, which uses a confusion matrix to visually evaluate the prediction results. The true level label is used as the vertical axis and the predicted level is used as the horizontal axis to analyze the number of correctly predicted samples on the main diagonal and the misjudgment situation on the off-diagonal.

Citation Information

Patent Citations

  • Dynamic rapid prediction method for large deformation grade of high-crustal-stress soft rock tunnel

    CN116861704A

  • Tunnel surrounding rock deformation prediction method based on PSO-LSTM model

    CN117113842A

  • A hierarchical prediction method for large tunnel deformation and related products

    CN119669882B

  • Few-sample tunnel deformation prediction method based on MAML and feature engineering guidance

    CN120316445A

  • Water-rich composite stratum tunnel surrounding rock deformation identification method

    CN120579079A