Hyperparameter optimization method based on optuna-xgboost tunnel stability prediction

By using the Optuna-XGBoost hyperparameter optimization method, a tunnel stability sample database was constructed and an XGBoost classifier was trained. This solved the problems of multi-source data fusion and hyperparameter dependence in tunnel stability prediction, and achieved efficient, robust and reliable intelligent decision support for tunnel stability prediction.

CN122154040APending Publication Date: 2026-06-05SOUTHWEST JIAOTONG UNIV +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2026-04-14
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies for tunnel stability prediction suffer from problems such as insufficient fusion of multi-source heterogeneous data, hyperparameter dependence on empirical settings, insufficient model adaptability, instability of the decision-making layer, and lack of online update mechanisms, making it difficult to meet the needs for rapid and robust on-site early warning and decision-making.

Method used

A hyperparameter optimization method based on Optuna-XGBoost was adopted to construct a sample database of stability of jointed and fractured rock mass tunnels. An XGBoost classifier was trained using the Optuna Bayesian optimization framework of TPE. Combining class weights and data source weights, an early stopping mechanism and dynamic threshold decision were adopted to achieve intelligent prediction for multiple tasks.

Benefits of technology

It improves the applicability and robustness of tunnel stability prediction, enhances generalization stability and repeatability, reduces the risk of false alarms and false negatives, and supports the fusion and online updating of multi-source heterogeneous data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122154040A_ABST
    Figure CN122154040A_ABST
Patent Text Reader

Abstract

The application discloses an Optuna-XGBoost tunnel stability prediction-based hyperparameter optimization method and relates to the technical field of tunnel construction, and solves the problem that the existing technology is difficult to capture the complex deformation mode by using the posterior probability criterion as the deformation anomaly point strategy, is prone to misjudgment, and influences the accuracy and generalization ability of prediction; the application comprises the following steps: according to the proportion of real data and simulation data in the stability comprehensive index category and the joint fissure rock mass tunnel stability sample database, a training set and a verification set are divided to obtain a sample space; an XGBoost classifier is trained by using an Optuna Bayesian optimization framework based on TPE and the training set; category weights and data source weights are constructed and are distributed to the training samples; after multiple rounds of training, the performance effect is verified by using the verification set to determine the optimal hyperparameters of the XGBoost classifier; and the application improves the applicability and robustness of the prediction model under the condition of complex joint fissure surrounding rock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tunnel construction technology, and specifically to a hyperparameter optimization method for tunnel stability prediction based on Optuna-XGBoost. Background Technology

[0002] The construction environment for tunnel projects in Southwest China is extremely complex. Many projects traverse the boundary zone between the Eurasian and Indian Ocean plates, and are subject to intense crustal tectonics, resulting in a dense network of active fault zones and joints. The rock mass is characterized by highly developed faults, weak interlayers, and fractured structural surfaces, giving the surrounding rock significant heterogeneity, anisotropy, and nonlinear mechanical characteristics. Under these jointed and fractured rock mass conditions, the surrounding rock has weak self-stabilizing capacity, making it prone to collapses, crown settlement, sidewall compression, and support failure, seriously threatening construction safety and schedule. Due to the randomness and multi-scale coupling of structural surfaces, the deformation mechanism and stability assessment of the surrounding rock exhibit high uncertainty. Traditional experience-based grading or single information sources are insufficient to accurately characterize the evolution of deformation around the tunnel and the distribution of anomalous points (abnormal locations).

[0003] Faced with the increasing demand for transportation tunnel construction and the growing geological challenges, the shortcomings of traditional drill-and-blast methods in terms of mechanization, intelligence, and information loop are becoming increasingly prominent. This is especially true in jointed and fractured rock masses, where construction disturbances coupled with complex structures lead to delays in risk identification and mitigation. While existing analytical and numerical methods have made progress in both theory and engineering, they suffer from high computational complexity, strong dependence on parameter calibration, and insufficient adaptability to real-time monitoring and dynamic operating conditions under conditions of strong heterogeneity, long-term creep, and complex hydro-mechanical-thermal coupling. These limitations make it difficult to meet the requirements for rapid and robust on-site early warning and decision-making.

[0004] In recent years, data- and image-driven intelligent methods have provided new technical pathways for tunnel stability prediction. Regarding data-driven information, preliminary results have been achieved in time-series prediction (e.g., convergence deformation trends), regression, and classification (e.g., surrounding rock grade, key deformation quantities) based on monitoring and construction process data; some work has introduced ensemble learning and hyperparameter optimization to improve prediction accuracy and generalization ability. Regarding image-driven information, the accuracy of crack, leakage, and weak interlayer identification, as well as face structure segmentation, based on target detection, semantic segmentation, and point cloud segmentation has been significantly improved, making it possible to rapidly obtain joint geometric parameters and spatial distribution.

[0005] However, existing technologies still have common shortcomings: (1) The integration of multi-source heterogeneous data (monitoring time series, construction parameters, numerical simulation samples and images / point clouds) is insufficient, making it difficult to simultaneously utilize "global statistics of data-type information" and "local structure of image-type information"; (2) The model is not adaptable to class imbalance and differences between field / simulation samples, and hyperparameters mostly rely on experience settings, resulting in low parameter tuning efficiency and limited stability; (3) The decision-making level generally adopts a single maximum a posteriori (MAP) judgment, which is prone to unstable output when facing multi-peak or flat probability distributions, making it difficult to provide "Top-N credible labels" for engineering handling guidance for different parts around the tunnel; (4) There is a lack of online update and version management mechanisms that match field applications, making it difficult to maintain long-term reliability under data distribution drift and working condition changes. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a hyperparameter optimization method for tunnel stability prediction based on Optuna-XGBoost. This method solves the problem that the existing technology, which uses only posterior probability criteria as a deformation anomaly strategy, is unable to capture such complex deformation modes, is prone to misjudgment, and affects the accuracy and generalization ability of the prediction.

[0007] The hyperparameter optimization method for tunnel stability prediction based on Optuna-XGBoost includes constructing a sample database of jointed and fractured rock mass tunnel stability. Based on the comprehensive stability index categories and the ratio of real to simulated data in the sample database, a training set and a validation set are divided to obtain the sample space. Then, the Optuna Bayesian optimization framework based on TPE is used to train the XGBoost classifier on the training set. During training, class weights and data source weights are constructed based on the proportion of samples at each stability level, and the proportion of real and simulated samples in the sample space, and assigned to the training samples. After multiple rounds of training, the performance is verified using a validation set to determine the optimal hyperparameters of the XGBoost classifier. The XGBoost classifier with the optimal hyperparameters is used as the final output model for multi-task intelligent prediction of tunnel stability, for risk warning and decision support in practical engineering.

[0008] Furthermore, the construction of the jointed fractured rock mass tunnel stability sample database includes: First, based on the complexity of the number of joint groups and their combination forms, a four-level progressively more complex jointed fractured rock mass tunnel sample database is constructed, forming type I to IV sample databases; then, a four-layer progressive data structure model is designed using the information abundance layering concept, constructing four information layers for type I to IV sample databases respectively: basic layer, extended layer, enhanced layer, and holographic layer; finally, each database type-information layer is combined into a sample database, and a total of 16 sample databases are cross-trained during training.

[0009] Furthermore, the type I to IV sample library includes:

[0010] Type I library: contains only a single-dip joint system, representing the tunnel working condition under a single set of joints;

[0011] Type II joints: contain two sets of joints with different inclinations, representing a double-set cross-joint condition;

[0012] Type III: Contains three sets of complex joint combinations with different tendencies, representing multiple sets of cross-joint systems;

[0013] Type IV library: It does not limit the number of joint groups, but directly collects actual complex joint combination conditions to simulate the real complex joint fracture rock mass tunnel environment in engineering.

[0014] Furthermore, the base layer, extension layer, enhancement layer, and holographic layer each include:

[0015] Basic layer: Contains only the geometric orientation information of each joint group, used to characterize the geometric features of the simplest joint;

[0016] Extended layer: Based on the base layer, geometric and statistical parameters of each joint group are added, including average density and trace length distribution, to characterize the spatial distribution characteristics of the joints;

[0017] Enhanced layer: Based on the extended layer, it further integrates joint structure and weathering characteristic indicators, including Jv, SR, and SCR, extracted from the joint distribution and weathering degree of the face, to reflect the integrity of the jointed rock mass and the weathering filling status;

[0018] Holographic layer: Based on the reinforcement layer, it integrates the geostress field parameters and construction support parameters to form a multi-source heterogeneous full-element information layer for predicting the deformation and distortion points around the tunnel.

[0019] Furthermore, the sample space is represented as follows:

[0020]

[0021] in, This represents the tunnel working conditions and surrounding rock characteristic vectors. For stability category labels; the stability category labels include the tunnel perimeter convergence deformation value, the location of the maximum deformation around the tunnel, and the safety factor of the support structure.

[0022] Furthermore, the construction of category weights and data source weights includes:

[0023] Based on the sample proportions of each stability level or anomaly category in the sample space The ratio of real samples to simulated samples Construct class weights With data weight And define the comprehensive weight for each sample; for the first sample... Class, denoted by its sample size. The total number of samples is The number of categories is To alleviate class imbalance, inverse frequency weights are used.

[0024]

[0025] in, For unnormalized weights, The normalized class weights ensure that the effective contributions of different classes to gradient statistics are roughly on the same order of magnitude. This explicitly improves the first and second-order gradient statistics for minority class samples. The weights in the calculation of the split gain (Gain) are used to increase the probability of the minority class split being selected.

[0026] For different data sources, let the number of real samples be denoted as . The number of simulation samples is Then, the data source is weighted:

[0027]

[0028] And normalization yields:

[0029]

[0030] This setting adjusts the total weight of the field samples. and the total weight of the simulation samples Within the same order of magnitude: Therefore, the influence of field samples will not be overwhelmed by their small number in the first and second gradient statistics;

[0031] The final overall weight for each sample is:

[0032]

[0033] in, The category to which sample i belongs Let i be the data source to which sample i belongs; during the decision tree growth process, for any candidate split point, the first-order and second-order gradient statistics are replaced with weighted forms:

[0034]

[0035] This yields the weighted split gain.

[0036]

[0037] This represents the weight of the i-th sample; This represents the second derivative of the loss function corresponding to the i-th sample with respect to the current predicted value; This represents the first derivative of the loss function corresponding to the nth sample with respect to the current predicted value; , These represent the weighted sums of the first-order gradients of all samples within the left and right child nodes, respectively. , Let represent the weighted sum of the second-order gradients of all samples within the left and right child nodes, respectively; λ represents the weight of the leaf node. Regularization coefficient; γ represents the node splitting penalty term; Gain represents the decrease in the objective function caused by the current candidate split.

[0038] Furthermore, the evaluation function during the training process employs macro-average AUC and multi-class log loss;

[0039]

[0040]

[0041] In the formula, Let N represent the average accuracy of the i-th class of samples, and N and C be the number of samples and the number of classes, respectively. and These are the actual label and predicted probability of the i-th sample in category c, respectively.

[0042] Furthermore, the TPE-based Optuna Bayesian optimization framework includes: introducing the Optuna hyperparameter tuning framework, treating each training iteration as an evaluation sample of hyperparameters and objective function values, and iteratively accumulating the results; employing a Tree-structured Parzen Estimator (TPE) as a surrogate model to model subsets of hyperparameters that perform well and poorly, respectively, and guiding the search to prioritize potentially superior regions through strategies such as likelihood ratio; using the TPE as a surrogate model, Gaussian kernel density estimation is used to describe the distribution of hyperparameter configurations: for combinations that perform well... , The threshold is used to assign a higher probability to TPE and a lower probability to poor performance, thus dividing the parameter space into two regions: good performance and poor performance, and establishing probability models for each region.

[0043] Furthermore, the training process employs an early stopping mechanism and dynamic threshold decision-making, including:

[0044] Set the maximum number of training rounds. With tolerance rounds Calculate the validation set evaluation index after the t-th iteration. And record the current optimal value. If continuous Wheel satisfies:

[0045]

[0046] This triggers early stopping, terminating the training process for that set of hyperparameters. This indicates the improvement tolerance or minimum effective improvement threshold, used to avoid mistaking extremely small numerical fluctuations for real performance improvements;

[0047] The probability vector of multi-class output Calculate entropy

[0048]

[0049] And calculate the entropy ratio.

[0050]

[0051] Simultaneously calculate the kurtosis of the probability distribution. And construct the threshold function.

[0052]

[0053] When category The predicted probability satisfies When, include it in the sample Trusted tag set

[0054]

[0055] Indicates the first The sample was predicted by the model to be the th sample. The probability of a class; Indicates the first The information entropy of the predicted probability distribution of a sample is used to measure the uncertainty of the model's classification result for that sample. Indicates the first The entropy ratio of each sample; logC represents the maximum information entropy across C categories; because when the probability distribution is perfectly uniform, i.e. Entropy reaches its maximum value. therefore, The value of is usually located in the interval [0,1]. Indicates the first kurtosis of the predicted probability distribution for each sample; Indicates the first An adaptive probability threshold for each sample; Let represent the set of trustworthy labels for the i-th sample; if the true label If the prediction is successful, the threshold function parameters are iteratively updated based on the overall performance of the validation set, thereby improving the stability and fault tolerance of the model decision under imbalanced class and multi-source data conditions.

[0056] Furthermore, the multi-round training includes:

[0057] In the In the round, the TPE agent model uses historical datasets. With objective function Select a new hyperparameter configuration:

[0058]

[0059] In the formula, Configure candidate hyperparameters The corresponding acquisition function value; This indicates selecting the acquisition function from the hyperparameter search space. The largest candidate configuration is selected as the optimal search result for the current round;

[0060] And use this hyperparameter configuration to train a weighted XGBoost classifier. The evaluation results of multiple tasks and multiple indicators are calculated on the validation set through early stopping mechanism and dynamic threshold decision. ,Will Feedback sent to Optuna for an update. With the proxy model; after completing the preset total number of iterations. Afterwards, from Select the configuration with the optimal objective function value. :

[0061]

[0062] and the corresponding trained XGBoost classifier As the final output model for multi-task intelligent prediction of tunnel stability, it is used for risk warning and decision support in practical engineering.

[0063] The beneficial effects of this invention include:

[0064] This invention improves upon the traditional approach of identifying anomalous points based on a single information source or experience-based grading. It constructs a unified sample space and feature specifications for key parts around the tunnel, supporting multi-source heterogeneous fusion of monitoring data, construction method parameters, and numerical simulation samples, thereby enhancing applicability and robustness under complex jointed and fractured surrounding rock conditions.

[0065] This invention improves upon the problems of hyperparameter configuration relying on experience and low tuning efficiency by introducing Bayesian hyperparameter optimization (Optuna-TPE) with macro-average AUC as the target and combined with an early stopping mechanism. It automatically seeks optimization within the joint value space of "tree structure - regularization - learning rate - sampling ratio", replacing manual grid or trial-and-error tuning, and significantly improving generalization stability and repeatability.

[0066] This invention improves the handling of class imbalance and source bias by proposing an inverse proportional weight based on "class ratio + source ratio". The sample weight is directly applied to the gradient-Hessian statistic and indirectly affects the selection of split points, thereby giving priority to minority class anomalies and field monitoring samples, thus improving the robustness of minority class recall and overall identification.

[0067] This invention improves the problem of instability caused by multi-peak or flat probability distributions at the decision end. It proposes a dynamic threshold driven by entropy ratio / kurtosis and a Top-N trusted label output mechanism to avoid jitter and false negatives caused by a single maximum a posteriori (MAP), and significantly reduce the risk of false positives or false negatives. Attached Figure Description

[0068] Figure 1 This is a flowchart of the hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction involved in the embodiments of this application.

[0069] Figure 2 This is a visualization of the Optuna-XGBoost hyperparameter optimization iteration information involved in the embodiments of this application.

[0070] Figure 3 This is a schematic diagram illustrating the iterative optimization of XGBoost based on Bayesian hyperparameter optimization in the embodiments of this application.

[0071] Figure 4 This is the performance curve of the XGBoost model with optimal hyperparameters during the training process involved in the embodiments of this application. Figure 4 (a) is the loss value curve. Figure 4 (b) shows the accuracy curve. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0073] The following is in conjunction with the appendix Figure 1 Specific embodiments of the present invention will be described in detail;

[0074] The hyperparameter optimization method for tunnel stability prediction based on Optuna-XGBoost includes constructing a sample database of jointed and fractured rock mass tunnel stability. Based on the comprehensive stability index categories and the ratio of real to simulated data in the sample database, a training set and a validation set are divided to obtain the sample space. Then, the Optuna Bayesian optimization framework based on TPE is used to train the XGBoost classifier on the training set. During training, class weights and data source weights are constructed based on the proportion of samples at each stability level, and the proportion of real and simulated samples in the sample space, and assigned to the training samples. After multiple rounds of training, the performance is verified using a validation set to determine the optimal hyperparameters of the XGBoost classifier. The XGBoost classifier with the optimal hyperparameters is used as the final output model for multi-task intelligent prediction of tunnel stability, for risk warning and decision support in practical engineering.

[0075] Using the XGBoost model as the base classifier, a high-dimensional joint search space including both continuous and discrete parameters is defined:

[0076]

[0077] In the formula, d represents the maximum tree depth. This represents the minimum gain threshold. This represents L1 regularization, which applies an absolute value penalty to the parameters. This represents L2 regularization, used to reduce model complexity and suppress overfitting. This represents the learning rate, and s represents the subsample sampling ratio. This represents the feature sampling ratio.

[0078] The Optuna Bayesian optimization framework based on TPE introduces the Optuna hyperparameter tuning framework, treating each training iteration as an evaluation sample of hyperparameters and objective function values, and iteratively accumulating the results. It employs a Tree-structured Parzen Estimator (TPE) as a surrogate model, separately modeling subsets of hyperparameters that are "better" and "worse," and guiding the search to prioritize potentially superior regions through strategies such as likelihood ratio. This includes:

[0079] TPE, as a surrogate model, uses Gaussian kernel density estimation (KDE) to describe the distribution of hyperparameter configurations: for well-performing combinations... , A value of 0.2 assigns a higher probability to TPE and a lower probability to poor performance, thus dividing the parameter space into two regions: good performance and poor performance, and establishing probabilistic models for each. Common sampling strategies include expected improvement and likelihood ratio.

[0080]

[0081]

[0082] The above expression shows the desired improvement. It is the current known minimum value, quantifying the potential improvement that the current evaluation point may bring compared to the known optimal solution, and is suitable for low-dimensional continuous spaces. Likelihood ratio, by comparing the probability distributions of two regions, can identify the parameter ranges more likely to produce excellent results and prioritize sampling in these regions. It is suitable for high-dimensional discrete spaces and supports parallel computing environments. Considering the background of this paper's optimization of XGBoost hyperparameters using Optuna, which involves high dimensionality and discrete parameters, likelihood ratio is chosen as the sampling strategy to select the next set of hyperparameter combinations to be evaluated.

[0083] The evaluation function uses macro-average AUC and multi-class log loss. Macro-average AUC is the average area under the Receiver Operating Characteristic (ROC) curve. The formula is as follows:

[0084]

[0085] Used to measure the average performance of a model across all classes, it is less affected by class imbalance than the micro-average AUC. The closer it is to 1, the better the model's predictive performance; if it is close to 0.5, the model is no different from random guessing.

[0086] Multiclass log loss The formula is as follows:

[0087]

[0088] In the formula, N and C are the number of samples and the number of categories, respectively; and These are the actual label and predicted probability of the i-th sample in category c, respectively.

[0089] Imbalanced Classes and Field / Simulation Data Weights Embedded into Gain Calculation: Based on the proportion of samples at each stability level in the sample space, and the proportion of field monitoring and simulation samples, class weights and data source weights are constructed and assigned to training samples. In XGBoost, the weights are embedded into the calculation process of first-order and second-order gradient statistics, making the split gain calculation more sensitive to minority class samples and field samples, and improving the ability to identify critical operating conditions and disaster samples from the perspective of tree structure growth mechanism.

[0090] Based on the sample proportions of each stability level or anomaly category in the sample space The ratio of on-site monitoring to simulation samples Construct class weights With data weight And define the overall weight of each sample as:

[0091] For the Class, denoted by its sample size. The total number of samples is The number of categories is .

[0092] To alleviate class imbalance, inverse frequency weights are used:

[0093]

[0094] in For unnormalized weights, The normalized class weights ensure that the effective contributions of different classes to gradient statistics are roughly on the same order of magnitude. This explicitly improves the first and second-order gradient statistics for minority class samples. The weights in the calculation of the split gain (Gain) are used to increase the probability of the minority class split being selected.

[0095] For different data sources, the number of on-site samples is denoted as . The number of simulation samples is To avoid the field samples having too low a weight in the total gradient, the following strategy is used to weight the data source:

[0096]

[0097] And normalized to obtain

[0098]

[0099] This setting adjusts the total weight of the field samples. and the total weight of the simulation samples Within the same order of magnitude: Therefore, in the first and second gradient statistics, the influence of the field samples will not be overwhelmed by their small number.

[0100] The final overall weight for each sample is:

[0101] in The category to which sample i belongs Let i be the data source to which sample i belongs; during the decision tree growth process, for any candidate split point, the first-order and second-order gradient statistics are replaced with weighted forms:

[0102]

[0103] This yields the weighted split gain:

[0104]

[0105] Let represent the weight of the i-th sample. If The larger the sample size, the stronger its influence on node splitting decisions; It represents the second derivative of the loss function corresponding to the i-th sample with respect to the current predicted value, and is used to measure the sensitivity and stability of the update step size. This represents the first derivative of the loss function corresponding to the nth sample with respect to the current predicted value, also known as the first gradient, which reflects the direction and magnitude of the error correction for that sample under the current prediction state. , These represent the weighted sums of the first-order gradients of all samples within the left and right child nodes, respectively. , These represent the weighted sums of the second-order gradients of all samples within the left and right child nodes, respectively.

[0106] λ represents the weight of the leaf nodes. The regularization coefficient is used to suppress excessively large leaf weights and prevent overfitting. γ represents the node splitting penalty, also known as the minimum splitting loss or structural complexity penalty. A split is only worthwhile if the gain from the current split exceeds γ. Gain represents the decrease in the objective function brought about by the current candidate split, also known as the "split gain," which measures how much the model performance improves after dividing a parent node into left and right child nodes. If Gain > 0, it usually indicates that the split is valuable; if the gain is insufficient or even negative, it indicates that the split should not be performed.

[0107] When traversing features and their candidate split points, a weighted average is used. The optimal split is selected based on the principle of maximizing the split, and then... Setting a minimum gain threshold [0-10] suppresses invalid splits, thereby significantly amplifying the contribution of minority class samples and field monitoring samples to split decisions within the defined XGBoost optimization framework, and improving the ability to identify critical operating conditions and disaster samples.

[0108] During training, a maximum number of training epochs and a tolerance number of epochs are set, and early stopping is automatically triggered based on changes in the validation set evaluation metrics to avoid overfitting and invalid iterations. At the same time, a dynamic threshold adjustment algorithm is designed to adaptively adjust the classification threshold based on the predicted distribution and misclassification by comparing the credible label with the true label, thereby improving the model's fault tolerance and stability under class imbalance and multi-source data conditions.

[0109] This embodiment sets the maximum number of training rounds for XGBoost itself. The tolerance number is 500. The maximum number of training rounds is 30, at which point the training will automatically stop.

[0110] Calculate the validation set evaluation index after the t-th iteration. (e.g., macro average AUC or −Loss), and record the current best value. If continuous Wheel satisfies:

[0111]

[0112] This triggers early stopping, terminating the training process for that set of hyperparameters. This indicates the improvement tolerance or minimum effective improvement threshold, used to avoid mistaking extremely small numerical fluctuations for real performance improvements;

[0113] The probability vector of multi-class output Calculate entropy:

[0114]

[0115] And calculate the entropy ratio:

[0116]

[0117] Simultaneously calculate the kurtosis of the probability distribution. And construct the threshold function:

[0118]

[0119] When category The predicted probability satisfies When, include it in the sample Trusted tag set:

[0120]

[0121] Indicates the first The sample was predicted by the model to be the th sample. The probability of a class; Indicates the first The information entropy of the predicted probability distribution of a sample is used to measure the uncertainty of the model's classification result for that sample. Indicates the first The entropy ratio of each sample; logC represents the maximum information entropy across C categories. This is because when the probability distribution is perfectly uniform, i.e. Entropy reaches its maximum value. therefore, The value of is usually located in the interval [0,1]. Indicates the first kurtosis of the predicted probability distribution for each sample; Indicates the first An adaptive probability threshold for each sample; Let represent the set of trustworthy labels for the i-th sample.

[0122] If the real label If the prediction is successful, the threshold function parameters are iteratively updated based on the overall performance of the validation set, thereby improving the stability and fault tolerance of the model decision under imbalanced class and multi-source data conditions.

[0123] A closed-loop process for Optuna tuning and XGBoost training is constructed. In each round, the TPE model provides a new combination of hyperparameters, trains the weighted XGBoost, and outputs multi-task and multi-index validation results to Optuna. After a preset number of Bayesian optimization iterations, the hyperparameters with the best performance under a unified evaluation system and the corresponding XGBoost ensemble learning model are selected to achieve multi-task intelligent prediction of the stability of jointed and fractured rock mass tunnels.

[0124] Specifically, the closed-loop process of constructing Optuna tuning for XGBoost training is described in the first... In the round, the TPE agent model uses historical datasets. With objective function Select a new hyperparameter configuration:

[0125]

[0126] Configure candidate hyperparameters The corresponding acquisition function value; This indicates selecting the acquisition function from the hyperparameter search space. The largest candidate configuration is selected as the optimal search result for the current round.

[0127] And use this configuration to train a weighted XGBoost model. The evaluation results of multiple tasks and multiple indicators are calculated on the validation set through early stopping mechanism and dynamic threshold decision. ,Will Feedback sent to Optuna for an update. With the proxy model; after completing the preset total number of iterations. Afterwards, from The configuration with the optimal objective function value is selected.

[0128]

[0129] And the corresponding trained XGBoost ensemble learning model As the final output model for multi-task intelligent prediction of tunnel stability, it is used for risk warning and decision support in practical engineering.

[0130] In another embodiment, to fully adapt to different joint complexity and information abundance conditions, a hierarchical and information-intensity sample library construction method is adopted, and then a systematic experiment is conducted on the classification performance of deformation anomalies around the tunnel in jointed and fractured rock mass. The specific process is as follows.

[0131] Step A: Construction of a sample library based on joint complexity

[0132] First, based on the complexity of the number of joint groups and their combinations, a sample library of jointed fractured rock mass tunnels with four levels of increasing complexity was constructed, forming type I to IV libraries:

[0133] 1. Type I library: contains only a single-dip joint system, representing the tunnel working condition under a single set of joints;

[0134] 2. Type II: Contains two sets of joints with different inclinations, representing a double-set cross-joint condition;

[0135] 3. Type III: Contains three sets of complex joint combinations with different tendencies, representing multiple sets of intersecting joint systems;

[0136] 4. Type IV Library: It does not limit the number of joint groups, but directly collects actual complex joint combination conditions to simulate the real complex joint fracture rock mass tunnel environment in engineering.

[0137] Each sample library adopts a uniform data scale: the training set consists of 500 random simulation cases and 50 field cases, and the validation set consists of 100 random simulation cases and 10 field cases, thus constructing a multi-source database of "joint complexity × engineering-simulation hybrid samples" to provide basic data support for subsequent model training and comparative analysis.

[0138] Step B: Design of a hierarchical data structure based on information abundance

[0139] Based on the joint complexity classification, to analyze the impact of multi-source information completeness on the classification performance of deformable points around holes, a four-layer progressive data structure model is designed using the information abundance hierarchical approach. Four information layers—basic layer, extended layer, enhanced layer, and holographic layer—are constructed for the type I to IV sample databases, respectively.

[0140] 1. Basic layer: Contains only the geometric orientation information of each joint group (such as joint dip direction, dip angle, etc.), used to characterize the geometric features of the simplest joint;

[0141] 2. Extended layer: Based on the basic layer, geometric and statistical parameters such as the average density and trace length distribution of each group of joints are added to characterize the spatial distribution characteristics of joints;

[0142] 3. Enhancement layer: Based on the extended layer, further integrate joint structure and weathering characteristic indicators such as Jv, SR, and SCR extracted from the joint distribution and weathering degree of the face to reflect the integrity of the jointed rock mass and the weathering filling status;

[0143] 4. Holographic layer: Based on the reinforcement layer, the geostress field parameters (such as the magnitude and direction of the initial geostress) and construction support parameters (such as anchor bolts, steel arch frames, shotcrete, etc.) are integrated to form a multi-source heterogeneous "full-element information layer" for predicting the deformation and distortion points around the tunnel.

[0144] By combining "joint complexity (Type I to IV libraries) × information abundance (basic / extended / enhanced / holographic layers)," a 4×4 gridded experimental sample structure is constructed, providing a complete data architecture for the subsequent 16 joint experiments.

[0145] Step C: Training process of 4×4 gridded model based on Optuna-XGBoost

[0146] Under the four sample libraries and four information abundance conditions constructed in steps A and B, the Optuna-XGBoost hyperparameter optimization method described in Example 1 was used for model training and validation for each "library type-information layer" combination, resulting in 16 sets of cross-joint experiments:

[0147] 1. For each combination, a unified Optuna-XGBoost joint optimization process is adopted: based on the sample space and hyperparameter space, as shown in Table 1. The macro-average AUC is set as the evaluation metric and early stopping mechanism, that is, under the premise of a fixed number of training epochs, the number of tolerance epochs is used to judge whether the objective function is optimized. In this embodiment, the maximum number of training epochs of XGBoost itself is set to 500, and the number of tolerance epochs is set to 30.

[0148] Table 1. Room for optimization of XGBoost hyperparameters

[0149]

[0150] 2. Use the TPE surrogate model for Bayesian hyperparameter tuning. Figure 2 Further, multi-dimensional visualization was used to demonstrate the search trajectory during hyperparameter optimization and its impact on model performance. The results show that different combinations of hyperparameters have significantly different effects on classification performance. The L1 and L2 regularization parameters are best placed between 0 and 3 and 0 and 2, respectively, to effectively suppress overfitting; the maximum tree depth is best placed between 8 and 10 to avoid excessively shallow depth negatively impacting model performance; and the sample and feature sampling ratios are best placed between 0.75 and 0.85 and 0.85 and 1, respectively. In summary, the Bayesian joint optimization strategy, through an "exploration-exploitation" mechanism, explores the influence of model hyperparameters on the optimization objective, gradually improving the model's classification performance.

[0151] 3. The trend of the macro-average AUC value of the model in each iteration is as follows: Figure 3 As shown, with the increase of the number of iterations, the objective value gradually increases and tends to stabilize, indicating that the Bayesian optimization algorithm can effectively explore the hyperparameter space and quickly converge to a better solution.

[0152] 4. An imbalanced class structure and a "live-simulation" weighting mechanism are introduced, and a dynamic threshold algorithm is used for decision optimization. Based on the aforementioned training and AUC evaluation, this embodiment further introduces a dynamic threshold adjustment algorithm at different information abundance levels in the Type IV library to evaluate the improvement in classification accuracy and prediction stability. The original accuracy, Top 2 accuracy, dynamic accuracy, and dynamic transition rate are statistically analyzed for the base layer, extended layer, enhancement layer, and holographic layer, respectively. The results are shown in Table 2: the original accuracy gradually increased from 50.3% in the base layer to 67.8% in the holographic layer; the Top 2 accuracy increased from 64.5% to 80.5%; and the dynamic accuracy increased from 60.3% to 76.3%, indicating that the combined effect of increased information abundance and dynamic threshold adjustment significantly enhances the model's classification ability.

[0153] Table 2. Accuracy performance of the information abundance-based Type IV library under the dynamic threshold adjustment algorithm.

[0154]

[0155] 5. During training, record the changes in multi-class log loss and accuracy on both the training and validation sets to determine whether overfitting or underfitting has occurred. Taking the "Holographic Layer – Type IV Library" as an example... Figure 4 As shown, the loss values ​​of the training set and validation set gradually decrease and tend to stabilize with the increase of iterations. The accuracy of the validation set stabilizes at about 80%, and the accuracy of the training set is close to 98%, indicating that the model does not show obvious overfitting under this combination. The training process of the other 15 models shows similar performance, all exhibiting good convergence characteristics.

[0156] The embodiments described above merely illustrate specific implementation methods of this application, and while the descriptions are detailed and specific, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the technical solution of this application, and these modifications and improvements all fall within the scope of protection of this application.

Claims

1. A hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction, characterized in that, This includes constructing a stability sample database for jointed and fractured rock mass tunnels; dividing the sample space into training and validation sets based on the comprehensive stability index categories and the ratio of real and simulated data in the stability sample database; and training an XGBoost classifier using the TPE-based Optuna Bayesian optimization framework and the training set. During training, class weights and data source weights are constructed based on the proportion of samples at each stability level, the proportion of real samples and simulated samples in the sample space, and then assigned to the training samples. After multiple rounds of training, the performance effect is verified through a validation set to determine the optimal hyperparameters of the XGBoost classifier. The XGBoost classifier with the optimal hyperparameters is used as the final output model for multi-task intelligent prediction of tunnel stability, and is used for risk warning and decision support in actual engineering projects.

2. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The construction of the jointed and fractured rock mass tunnel stability sample database includes: First, based on the complexity of the number of joint groups and their combination forms, a four-level progressively more complex jointed and fractured rock mass tunnel sample database is constructed, forming type I to IV sample databases; then, a four-layer progressive data structure model is designed using the information abundance layering concept, constructing four information layers for type I to IV sample databases respectively: basic layer, extended layer, enhanced layer, and holographic layer; finally, each database type-information layer is combined into a sample database, and a total of 16 sample databases are cross-trained during training.

3. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 2, characterized in that, The type I to IV sample libraries include: Type I library: contains only a single-dip joint system, representing the tunnel working condition under a single set of joints; Type II joints: contain two sets of joints with different inclinations, representing a double-set cross-joint condition; Type III: Contains three sets of complex joint combinations with different tendencies, representing multiple sets of cross-joint systems; Type IV library: It does not limit the number of joint groups, but directly collects actual complex joint combination conditions to simulate the real complex joint fracture rock mass tunnel environment in engineering.

4. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 2, characterized in that, The base layer, extension layer, enhancement layer, and holographic layer each include: Basic layer: Contains only the geometric orientation information of each joint group, used to characterize the geometric features of the simplest joint; Extended layer: Based on the base layer, geometric and statistical parameters of each joint group are added, including average density and trace length distribution, to characterize the spatial distribution characteristics of the joints; Enhanced layer: Based on the extended layer, it further integrates joint structure and weathering characteristic indicators, including Jv, SR, and SCR, extracted from the joint distribution and weathering degree of the face, to reflect the integrity of the jointed rock mass and the weathering filling status; Holographic layer: Based on the reinforcement layer, it integrates the geostress field parameters and construction support parameters to form a multi-source heterogeneous full-element information layer for predicting the deformation and distortion points around the tunnel.

5. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The sample space is represented as follows: in, This represents the tunnel working conditions and surrounding rock characteristic vectors. The stability category label includes the tunnel perimeter convergence deformation value, the location of the maximum deformation around the tunnel, and the safety factor of the support structure.

6. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The construction category weights and data source weights include: Based on the sample proportions of each stability level or anomaly category in the sample space The ratio of real samples to simulated samples Construct class weights With data weight And define the comprehensive weight for each sample; for the first sample... Class, denoted by its sample size. The total number of samples is The number of categories is To alleviate class imbalance, inverse frequency weights are used. in, For unnormalized weights, The normalized class weights ensure that the effective contributions of different classes in gradient statistics are roughly on the same order of magnitude. For different data sources, let the number of real samples be denoted as . The number of simulation samples is Then, the data source is weighted: And normalization yields: This setting adjusts the total weight of the field samples. and the total weight of the simulation samples Within the same order of magnitude: Therefore, the influence of field samples will not be overwhelmed by their small number in the first and second gradient statistics; The final overall weight for each sample is: in, The category to which sample i belongs The data source to which sample i belongs.

7. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The evaluation function used during the training process employs macro-average AUC and multi-class log loss. ; In the formula, Let N represent the average accuracy of the i-th class of samples, and N and C be the number of samples and the number of classes, respectively. and These are the actual label and predicted probability of the i-th sample in category c, respectively.

8. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The TPE-based Optuna Bayesian optimization framework includes: introducing the Optuna hyperparameter tuning framework, treating each training iteration as an evaluation sample of hyperparameters and objective function values, and iteratively accumulating the results; employing a tree-structured Parzen estimation of TPE as a surrogate model to model subsets of hyperparameters with good and poor performance, respectively, and guiding the search to prioritize potentially superior regions through strategies such as likelihood ratio; using TPE as a surrogate model, Gaussian kernel density estimation is used to describe the distribution of hyperparameter configurations: for combinations with good performance... , Using a threshold, TPE is assigned a higher probability, while poor performance is assigned a lower probability, thus dividing the parameter space into two regions: good performance and poor performance, and establishing probability models for each region.

9. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The training process employs an early stopping mechanism and dynamic threshold decision-making, including: Set the maximum number of training rounds. With tolerance rounds Calculate the validation set evaluation index after the t-th iteration. And record the current optimal value. If continuous Wheel satisfies: This triggers early stopping, terminating the training process for that set of hyperparameters. This indicates the improvement tolerance or minimum effective improvement threshold, used to avoid mistaking extremely small numerical fluctuations for real performance improvements; The probability vector of multi-class output Calculate entropy: And calculate the entropy ratio. Simultaneously calculate the kurtosis of the probability distribution. And construct the threshold function. When category The predicted probability satisfies When, include it in the sample Trusted tag set Indicates the first The sample was predicted by the model to be the th sample. The probability of a class; Indicates the first The information entropy of the predicted probability distribution of a sample is used to measure the uncertainty of the model's classification result for that sample. Indicates the first The entropy ratio of each sample; logC represents the maximum information entropy across C categories; because when the probability distribution is perfectly uniform, i.e. Entropy reaches its maximum value. ;therefore, The value of is usually located in the interval [0,1]. Indicates the first kurtosis of the predicted probability distribution for each sample; Indicates the first An adaptive probability threshold for each sample; Let represent the set of trustworthy labels for the i-th sample; if the true label If the prediction is successful, the threshold function parameters are iteratively updated based on the overall performance of the validation set, thereby improving the stability and fault tolerance of the model decision under imbalanced class and multi-source data conditions.

10. The hyperparameter optimization method based on Optuna-XGBoost tunnel stability prediction according to claim 1, characterized in that, The multi-round training includes: In the In the round, the TPE agent model uses historical datasets. With objective function Select a new hyperparameter configuration: In the formula, Configure candidate hyperparameters The corresponding acquisition function value; This indicates selecting the acquisition function from the hyperparameter search space. The largest candidate configuration is selected as the optimal search result for the current round; And use this hyperparameter configuration to train a weighted XGBoost classifier. The evaluation results of multiple tasks and multiple indicators are calculated on the validation set through early stopping mechanism and dynamic threshold decision. ,Will Feedback sent to Optuna for an update. With the proxy model; after completing the preset total number of iterations. Afterwards, from Select the configuration with the optimal objective function value. : and the corresponding trained XGBoost classifier As the final output model for multi-task intelligent prediction of tunnel stability, it is used for risk warning and decision support in practical engineering.