Sustainable tunnel construction water inflow prediction method based on machine learning technology
Through the method based on machine learning technology, an efficient and intelligent water inrush prediction system is built, which solves the problems of insufficient accuracy and inefficiency in traditional prediction methods, and realizes high-precision and low-cost water inrush prediction in complex environments, supporting construction safety and environmental protection.
Patent Information
- Application Number
- CN202510179384.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Traditional water inrush prediction methods have problems such as insufficient accuracy, complex calculations and low efficiency, and it is difficult to accurately predict water inrush changes in tunnel construction in complex geological and hydrological environments.
Using machine learning technology-based methods, we build an efficient and intelligent prediction system, automatically process data and output results, identify key factors affecting water influx, and establish models such as random forests, XGBoost and deep neural networks to optimize water influx prediction.
It significantly improves the accuracy and efficiency of water influx forecasting, reduces calculation costs and time consumption, and provides scientific basis to help the construction team adjust the construction plan and reduce the risks of safety accidents and construction delays.
Smart Images

Figure CN120030903A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of engineering technology and environmental protection, and relates to a sustainable tunnel construction water inflow prediction method based on machine learning technology. Background Art
[0002] In recent years, my country has attached great importance to the construction of ecological civilization, and the development concept has gradually shifted from economic development to environmental protection and ecological balance. Against this background, various large-scale infrastructure construction projects, especially railway construction projects, have put forward higher requirements for environmental protection. In the construction of railway tunnels under complex geological conditions, accurate prediction of water inflow is crucial to ensure construction safety and progress and reduce environmental impact.
[0003] In the field of water inflow prediction, the commonly used methods include theoretical calculation method, numerical simulation method and non-deterministic method. The theoretical calculation method relies on simplified groundwater flow model and engineering empirical formula, which is suitable for simple geological and hydrological conditions. However, in complex geological and hydrological environments, theoretical calculation may have large errors and it is difficult to accurately reflect the changes in water inflow during tunnel construction. The numerical simulation method models the groundwater flow and solute transport process through numerical solutions, which can handle more complex geological and hydrological conditions. Although it has high accuracy, the calculation process is complicated and relies on a large amount of measured data, which is computationally intensive and time-consuming, resulting in low efficiency in practical applications. Non-deterministic methods (such as gray system models, random process models, etc.) simulate the influence of uncertain factors on water inflow through probability analysis. They have the advantages of simple concepts and the modeling process does not require detailed descriptions of geological and hydrological characteristics. Therefore, they have gradually become a research hotspot in the prediction of water inflow during tunnel construction. Traditional prediction methods generally have problems of insufficient accuracy, complex calculations and low efficiency. In recent years, machine learning technology has made significant progress in data processing and pattern recognition, showing strong prediction capabilities. Machine learning can automatically extract patterns from historical data and model complex nonlinear relationships, thus having potential advantages in water inflow prediction. Compared with theoretical calculation methods, machine learning methods can significantly improve prediction accuracy; compared with numerical simulation methods, machine learning reduces reliance on complex geological and hydrological assumptions, reduces computational costs, and improves the universality and accuracy of the model. In addition, machine learning technology can identify key factors that affect water inflow and further improve prediction efficiency. Therefore, optimizing water inflow prediction methods in combination with modern data analysis technology has become a difficult problem that needs to be solved urgently.
[0004] Studies have shown that traditional water inflow prediction models focus on physical models or statistical methods, and fail to fully utilize the advantages of machine learning technology for in-depth analysis and model optimization ( https: / / doi.org / 10.1016 / j.trgeo.2023.100978). Therefore, exploring the prediction method of tunnel construction water inflow based on machine learning will not only help improve the prediction accuracy, but also significantly improve the prediction efficiency, and has broad application prospects.
[0005] The present invention aims to construct a sustainable optimization method for predicting water inflow during tunnel construction based on machine learning technology, comprehensively consider factors such as natural geographical conditions, geological conditions, hydrogeological conditions and tunnel characteristic conditions, accurately predict water inflow during tunnel construction, provide scientific decision-making support for tunnel construction, optimize construction plans and reduce engineering risks. Summary of the invention
[0006] This paper proposes a method for predicting water inflow during tunnel construction based on machine learning technology, aiming to solve the problems of large errors and complex calculations in traditional prediction methods. By building an efficient and intelligent prediction system, automatic data processing and result output can be achieved, which improves practicality and facilitates its wide application in tunnel construction and environmental management.
[0007] The technical solution of the present invention:
[0008] A sustainable tunnel construction water inflow prediction method based on machine learning technology, the steps are as follows:
[0009] Step 1: Data collection and preprocessing;
[0010] Collect preliminary data on tunnel construction, including geological survey reports and drilling data of the tunnel site, as well as actual water inflow data and corresponding construction progress data during the construction period; organize the data to form a dataset of factors affecting the original water inflow; use interpolation and IQR methods to check for missing values and outliers, and perform descriptive statistical analysis to determine the distribution and fluctuation characteristics of each parameter;
[0011] The factors affecting the original water inflow are divided into two categories, namely engineering conditions and hydrogeological conditions; engineering conditions include construction methods (the drilling and blasting method is represented by a value of 1, and the TBM method is represented by a value of 2), the number of tunnels (the number of tunnels in the main tunnel section and the auxiliary tunnel section is 1; the number of tunnels in the main tunnel and the horizontal guide, and the TBM left and right line scenes is 2; the number of tunnels when the auxiliary pit enters the main tunnel and the horizontal guide and the large and small mileages are constructed at the same time is 4), the number of tunnel faces (the specific number needs to be determined in combination with the early construction preparation and the actual situation on site), the construction cross-sectional area (that is, the tunnel face area, represented by the maximum contour length × height, m 2When there are multiple tunnel faces, the cross-sectional area of the first tunnel shall prevail), the cumulative footage (the distance between the tunnel face and the tunnel entrance, m. When there are multiple tunnel faces, the distance between the first tunnel and the tunnel entrance shall prevail), the daily footage (the daily excavation mileage, m), the tunnel burial depth (the distance from the surface elevation to the tunnel body, m), the static water height (the distance from the stable groundwater level to the tunnel top, m); the hydrogeological conditions include the water richness (based on the borehole flow rate q, q<0.1L / s is water-poor, q=0.1~1.0L / s is weak water-rich, q=1.0~3.0L / s is moderate water-rich, q=3.0~10 .0L / s is strong water-rich; q>10.0L / s is extremely water-rich), groundwater type (mainly divided into pore water, fissure water and karst water, represented by the values 1, 2 and 3 respectively), length of crossing karst and faults (the length of the cave body passing through karst caves, faults, fractured zones and other extremely water-rich areas, m), surrounding rock grade (divided into Ⅰ to Ⅵ according to the integrity of the rock mass and the strength of the rock, represented by the values 1 to 6 respectively), permeability coefficient (unit flow rate under unit hydraulic gradient, m / d), infiltration coefficient (ratio of groundwater infiltration recharge to rainfall), rainfall (mm / d).
[0012] Furthermore, descriptive statistical analysis refers to obtaining the minimum, maximum, mean, and standard deviation of each influencing factor.
[0013] Step 2: Use multivariate statistics and machine learning methods to identify the main controlling factors of water inflow;
[0014] Firstly, the Granger causality test of rainfall and water inflow was carried out based on the vector autoregression (VAR) model in the statsmodels library using Python programming language. According to the lag time of significant causal relationship, the time interval of data resampling was determined to balance the sampling time interval of monitoring data of time-varying and non-time-varying factors.
[0015] Furthermore, the general form of the VAR model for two time series X and Y (i.e., rainfall and inflow) is as follows:
[0016]
[0017] Among them, X t and Y t represents the observed values of time series X and Y at time t; p is the maximum number of lags set; i is the lag period, and its value range is 1-p; and and are the coefficients in the model, representing the influence of X and Y at time ti on X and Y at time t; and is the error term;
[0018] Furthermore, the F-test is used for the Granger causality test. The null hypothesis is that there is no significant causal relationship between the time series X and Y. If the F-value is significant, the null hypothesis is rejected, and it is considered that there is a significant Granger causal relationship between rainfall and water inflow.
[0019] Secondly, the correlation analysis tool in SPSS 26.0 software is used to identify the influencing factors highly correlated with water inflow by using the Spearman correlation coefficient.
[0020] Finally, using the Python programming language, based on the Random Forest model algorithm, an initial water inflow prediction model is constructed using the water inflow and the data of the influencing factors of water inflow obtained in step 1. The Permutation Importance tool and the SHAP (SHapley Additive exPlanations) tool are introduced to analyze the importance ranking of each influencing factor and its marginal contribution to the water inflow prediction result, and the importance of each influencing factor is ranked according to the feature importance score and the SHAP value.
[0021] Furthermore, when constructing the initial water inflow prediction model, to ensure that the training effect of the data does not show overfitting and to ensure the repeatability of the results, 80% of the data is used as the training set, 20% of the data is used as the test set, and the random seed number is set to 42.
[0022] Furthermore, the output performance evaluation parameters of the initial water inflow prediction model include the mean absolute error (MAE), the mean square error (MSE), the root mean square error (RMSE), and the coefficient of determination (R 2 ), and the formula is:
[0023]
[0024] where n is the number of samples, y i is the measured value of each sample, is the average value of the measured values of this index, is the predicted value obtained using the model.
[0025] Furthermore, the Permutation Importance tool quantifies the impact of each influencing factor on the prediction ability of the initial water inflow prediction model by evaluating the difference between the performance of the initial water inflow prediction model on the original dataset of the influencing factors of water inflow and the performance after the influencing factors are randomly permuted. Using this Permutation Importance tool, the contribution of each influencing factor to the change in water inflow is evaluated based on the initial water inflow prediction model, and the importance of the feature values is ranked according to the Permutation Importance score.
[0026] Furthermore, the SHAP tool is based on the Shapley value in game theory, which quantifies the contribution of each influencing factor to the change in water inflow by calculating the marginal contribution of each influencing factor. The contribution and importance ranking of each influencing factor to the change in water inflow are obtained based on the bee swarm diagram obtained by the SHAP tool analysis. In the bee swarm diagram, blue indicates a positive impact on the predicted water inflow, and red indicates a negative impact on the predicted water inflow. The darker the color, the higher the degree. The SHAP value calculation formula is:
[0027]
[0028] in: is the SHAP value of the kth influencing factor; N is the set of all influencing factors; S is the feature subset of influencing factor k; f(S) is the model prediction value of subset S; f(S∪{k}) is the model prediction of the features in S plus the features
[0029] Furthermore, factors with correlation coefficient with water inflow > 0.3, feature replacement importance > 0.03, and SHAP value > 500 were screened out and determined as key influencing factors of water inflow, providing a basis for the subsequent establishment of an efficient water inflow prediction model.
[0030] Step 3: Using Python programming language, based on the key influencing factors of water inflow screened in step 2, high-efficiency water inflow prediction models are established respectively based on random forest, XGBoost and deep neural network (DNN) model algorithms, and the performance of different high-efficiency water inflow prediction models is evaluated and compared according to the evaluation parameters of the output performance of the initial water inflow prediction model in step 2. In this step, 85% of the key influencing factors and water inflow data are used as training sets, and the remaining 15% are used as test sets, and the training set and test set data of all high-efficiency water inflow prediction models are guaranteed to be the same, and the grid search method is used to adjust the parameters of the high-efficiency water inflow prediction model.
[0031] Step 4: The efficient water inflow prediction model with the closest output performance between the training set and the test set of each efficient water inflow prediction model in step 3 is determined as the optimal water inflow prediction model; a series of different hydrogeological conditions and engineering conditions scenarios are set to verify the applicability of the optimal water inflow prediction model to different engineering construction environments again to ensure its effectiveness and accuracy in practical applications.
[0032] Furthermore, the output performance evaluation parameters of the initial water inflow prediction model described in step 2 were used to evaluate the applicability of the optimal water inflow to different engineering construction environments. The mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE) and determination coefficient (R 2 ) are within an acceptable range, it is considered that the optimal water inflow prediction model is highly applicable to actual working conditions.
[0033] Step 5: Build a sustainable intelligent prediction system for water inflow during tunnel construction; use the mysql-connector-python library to achieve efficient connection between Python and MySQL to ensure fast interaction of the optimal water inflow prediction model database. The system inputs key influencing factor data into the optimal water inflow prediction model through database management to make predictions; the prediction results are returned to the database and stored for subsequent query and analysis.
[0034] Furthermore, in the sustainable intelligent prediction system for water inflow in tunnel construction, MySQL is mainly used to store and manage various parameter data of tunnel construction, historical water inflow data and the prediction results of the optimal water inflow prediction model.
[0035] Furthermore, the embedded data set of sustainable intelligent prediction of water inflow in tunnel construction will be continuously and dynamically updated, and the universality and prediction accuracy of the model will be improved by continuously adding new training samples, ensuring that the system continues to optimize the prediction performance as the amount of data increases.
[0036] Beneficial effects of the present invention:
[0037] 1. The present invention optimizes the traditional water inflow prediction method by combining multiple machine learning techniques (such as XGBoost regression, SHAP interpreter, deep neural network, etc.). Through automated data processing and model training, the interference of human factors is reduced, and the accuracy and efficiency of water inflow prediction are significantly improved, effectively reducing the computing cost and time consumption.
[0038] 2. The present invention can accurately predict the possible changes in water inflow during tunnel construction, providing a scientific basis to help the construction team adjust the construction plan in real time. This method effectively reduces the risk of safety accidents and construction delays caused by water inflow, while providing reliable data support for water resource management and groundwater protection, reducing the negative impact on the environment.
[0039] 3. The present invention not only improves the accuracy and efficiency of water inflow prediction in tunnel construction, but also promotes the application of machine learning technology in the field of groundwater prediction. By optimizing the model training and data analysis process, the present invention promotes technological innovation in water inflow prediction in the field of tunnel construction and provides new technical means for engineering practice in this field. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flow chart of a method for predicting water inflow during tunnel construction based on machine learning technology of the present invention.
[0041] Figure 2It is the Granger causality test result of rainfall and water inflow, where (a) is the drilling and blasting method and flat guide construction scenario, (b) is the drilling and blasting method and inclined shaft construction scenario, and (c) is the TBM method and left and right line construction scenario.
[0042] Figure 3 This is the Spearman correlation diagram between each influencing factor and water inflow.
[0043] Figure 4 It is an importance ranking diagram of the impact of various parameters on water inflow based on feature permutation importance (PI) calculation.
[0044] Figure 5 It is the SHAP value bee swarm diagram.
[0045] Figure 6 This is a verification diagram of the deep neural network (DNN) prediction model for the water inflow prediction of the No. 4 cross tunnel of Tunnel C.
[0046] Figure 7 This is a verification diagram of the deep neural network (DNN) prediction model for the water inflow prediction of the No. 2 inclined shaft of Tunnel F. DETAILED DESCRIPTION
[0047] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.
[0048] Example
[0049] A flow chart of a sustainable tunnel construction water inflow prediction method based on machine learning technology is shown in the figure Figure 1 As shown, the following steps are included:
[0050] Step 1: Collect preliminary data on tunnel construction, such as geological survey reports and drilling data, and collect actual water inflow data and corresponding construction progress data during the construction period to complete the establishment of a data set of water inflow data and its influencing factors. In order to establish a database with sufficient samples, a typical tunnel construction site of a high-speed railway construction was taken as the research object. The site covers drilling and blasting and TBM construction methods, different hydrogeological conditions, and a combination of main tunnels and auxiliary tunnels. The site characteristics and data period are shown in Table 1:
[0051] Table 1 Work site characteristics and data cycle
[0052]
[0053] Furthermore, the collected data were checked for missing values and outliers using interpolation and IQR methods, followed by descriptive statistical analysis. The descriptive statistics are detailed in Table 2 .
[0054] Table 2 Statistics of various index parameters
[0055] index unit Minimum Maximum average value Standard Deviation Water inflow <![CDATA[m 3 / d]]> 368.6 96775.6 15199.9 15626.1 Rainfall mm / d 0.0 205.9 10.9 23.9 Day's Progress m 0.0 15.9 3.1 3.0 Cumulative footage m 88.0 7822.6 2175.1 1309.3 Tunnel depth m 59.0 1804.0 605.7 302.9 Still water height m 0.0 1053.0 362.9 197.8 Infiltration coefficient - 0.1 1.0 0.3 0.2 Permeability coefficient m / d 0.0 0.9 0.1 0.1 Number of holes - 1.0 4.0 1.5 0.7 Number of faces - 1.0 4.0 1.7 0.9 Construction method - 1.0 2.0 1.1 0.4 Construction cross-sectional area <![CDATA[m 2 > 54.7 98.5 60.0 12.7 Surrounding rock grade - 2.0 5.0 4.2 0.6 Water richness L / s 0.2 7.0 1.5 1.6 Groundwater Type - 1.0 3.0 1.9 0.3 Length of karst and fault zones m 0.0 136.0 12.8 30.2
[0056] Step 2: Use multivariate statistics and machine learning methods to identify the main controlling factors of water inflow.
[0057] First, the Granger causality test between rainfall and water inflow is carried out. The test results are as follows: Figure 2 As shown. The B tunnel horizontal outlet, the No. 2 inclined shaft of F tunnel and the exit of K tunnel were selected as representative work points, and Granger causality analysis was performed on rainfall and water inflow. When the lag time of water inflow at the B tunnel horizontal outlet was 2 to 3 days, there was a significant causal relationship between the two (p < 0.05). There was a significant causal relationship between rainfall and water inflow in the No. 2 inclined shaft of F tunnel when the lag time was 1 to 5 days (p < 0.05), and the significance gradually weakened with the increase of lag time. There was a significant causal relationship between rainfall and water inflow at the exit of K tunnel when the lag time was 1 to 3 days (p < 0.05), but after the lag time exceeded 4 days, the causal relationship was no longer significant. The results of Granger causality analysis show that rainfall has a significant effect on water inflow only in a relatively short period of time. According to the lag time with significant causal relationship, the time interval for data resampling was determined to be 5 days.
[0058] Secondly, the correlation analysis tool in SPSS26.0 software was used to identify the influencing factors with high correlation with water inflow based on the Spearman correlation coefficient and significance test results. The analysis results are as follows: Figure 3 The factors with correlation coefficients greater than 0.3 with water inflow include permeability, construction method, water richness, groundwater type, and length of karst and faults crossed.
[0059] Finally, the Python programming language was used to build the initial random forest water inflow prediction model based on the water inflow and 15 influencing factors data obtained in step 1, and the feature permutation importance (PermutationImportance) and SHAP (SHapley Additive exPlanations) tools were introduced to analyze the importance ranking of each influencing factor and its marginal contribution to the water inflow prediction results. The analysis results are shown in Figure 2. Figure 4 and Figure 5 As shown. According to the feature permutation importance score, we can know that ( Figure 4 ), the factors with scores greater than 0.03 are fault and fracture zone length, groundwater type, static water height, construction method, rainfall and permeability coefficient. According to the SHAP value bee colony diagram ( Figure 5) shows that the factors with marginal contribution greater than 500 to water inflow are water richness (10315), length of karst, fault and fracture zone (4819), number of caves (3243), groundwater type (1652), permeability coefficient (1176), tunnel depth (650) and static water height (502).
[0060] Combining the results of Granger causal analysis, correlation analysis and importance evaluation, the construction method, number of tunnels, groundwater richness, groundwater type, fault, length of the broken zone, static water height, permeability coefficient and rainfall were finally screened out as the key influencing factors of tunnel construction water inflow. A water inflow prediction model was established based on these eight factors.
[0061] Step 3: Different efficient prediction models are established using the Python programming language pycharm library based on the key influencing factors of water inflow screened out in step 2. Before establishing the model, 15% of the sample data is randomly selected as a test set using random numbers, and the remaining 85% is used as a training set.
[0062] First, in the sklearn library, RandomForestRegressor is used for configuration and training to establish a random forest (RF) prediction model. The grid search method is used to adjust the model parameters (the number of trees in the forest, the maximum depth of the tree, the minimum number of samples per node, and the minimum number of samples per leaf) to find the optimal parameter combination to achieve a higher training set fitting effect and test set prediction effect of the model. The evaluation indicators of the random forest model when using different parameter combinations are shown in Table 3. The evaluation indicators are optimal when the model parameter combination is 300 trees in the forest, 10 maximum depth, 2 minimum number of node samples, and 1 minimum number of samples per leaf.
[0063] Table 3 Effect of different parameters on the evaluation index of random forest model
[0064]
[0065] Secondly, in the sklearn library, XGBRegressor is used for configuration and training to establish an XGBoost prediction model. The grid search method is used to adjust the model parameters (the number of decision trees, the maximum depth of the tree, the learning rate, and the sample sampling ratio) to find the optimal parameter combination to achieve a higher training set fitting effect and test set prediction effect of the model. When different parameter combinations are used, the evaluation indicators of the XGBoost model are shown in Table 4. When the model parameter combination is: the number of decision trees is 100, the maximum depth is 9, the learning rate is 0.1, and the sample sampling ratio is 1, the evaluation indicators are optimal.
[0066] Table 4 Effect of different parameters on XGBoost model evaluation indicators
[0067]
[0068] Finally, in the sklearn and Keras libraries, MLPRegressor was used for configuration and training to establish a deep neural network (DNN) prediction model. The grid search method was used to adjust the model parameters (optimizer selection, number of hidden layers, number of neurons in each layer, learning rate, activation function, batch size, and maximum number of iterations) to find the optimal parameter combination to achieve a higher training set fitting effect and test set prediction effect of the model. When different parameter combinations were used, the evaluation indicators of the deep neural network (DNN) model are shown in Table 5. The evaluation results were optimal when the model parameter combination was: number of iterations was 1000, number of hidden layers was 5, number of neurons was 128, learning rate was 0.01, and batch size was 32.
[0069] Table 5 Effect of different parameters on the evaluation index of neural network model
[0070]
[0071] Compared with the random forest model and the XGBoost model, although the deep neural network model has a slightly worse training set fitting effect, its test set prediction effect is better, and the evaluation indicators of the test set and the training set in the deep neural network model are close, indicating that the model does not have overfitting. In view of the randomness and difficulty in accurately predicting the water inflow during tunnel construction, the deep neural network water inflow prediction model has better applicability and promotion than the random forest prediction model and the XGBoost prediction model.
[0072] Step 4: Use the deep neural network (DNN) prediction model under the optimal parameter combination to predict and verify the water inflow of the No. 4 cross tunnel of C tunnel (drilling and blasting method, cross tunnel entering the main tunnel and flat guide, karst development construction scenario) and the No. 2 inclined shaft of F tunnel (drilling and blasting method, inclined shaft not entering the main tunnel and flat guide, fault zone development construction scenario). The verification period is the water volume monitoring period of the work point. The predicted water inflow obtained when the rainfall is the minimum (rainfall is zero) is the minimum water inflow, and the predicted water inflow obtained when the average of the maximum daily rainfall in the past five years is the maximum rainfall is the maximum water inflow. The verification results are as follows: Figure 6 and Figure 7 shown.
[0073] Verification results show that the deep neural network has good prediction performance. The actual water inflow at the demonstration sites of the No. 4 cross tunnel of Tunnel C and the No. 2 inclined shaft of Tunnel F is generally within the prediction range. It also has a good prediction effect on the water inflow in karst and fault areas, and can meet actual needs.
[0074] Step 5: Based on the optimal deep neural network (DNN) prediction model, an intelligent prediction system for water inflow during tunnel construction is constructed using Python and MySQL database technology. The MySQL database is responsible for efficiently storing and managing various types of data, providing basic data for water inflow prediction. Python interacts with the MySQL database through the pymysql library to extract historical data and perform cleaning and preprocessing to ensure data quality and consistency. The processed data is input into a Keras-based deep neural network (DNN) model for training. Through this model, the system can identify factors related to water inflow and make accurate predictions based on the input parameters.
[0075] The intelligent prediction system can automatically process various input data and generate real-time water inflow prediction results to assist construction site and drainage design decisions, optimize construction plans, reduce risks, and improve construction efficiency and safety.
Claims
1. A sustainable tunnel construction water inflow prediction method based on machine learning technology, characterized in that: Here are the steps: Step 1: Data collection and preprocessing; Collect preliminary data on tunnel construction, including geological survey reports and drilling data of the tunnel site, as well as actual water inflow data and corresponding construction progress data during the construction period; organize and form a dataset of factors affecting water inflow; use interpolation and IQR methods to check missing values and outliers in the data in the dataset of factors affecting water inflow, and perform descriptive statistical analysis to determine the distribution and fluctuation characteristics of each parameter; Step 2: Use multivariate statistics and machine learning methods to identify the main controlling factors of water inflow; Firstly, the Granger causality test of rainfall and water inflow was conducted based on the vector autoregression model in the statsmodels library using Python programming language. The time interval of data resampling was determined according to the lag time of significant causal relationship to balance the sampling time interval of monitoring data of time-varying and non-time-varying factors. Secondly, the correlation analysis tool in SPSS26.0 software was used to identify the influencing factors that were highly correlated with water inflow using the Spearman correlation coefficient; Finally, using Python programming language and based on the random forest model algorithm, the initial water inflow prediction model was constructed using the water inflow and water inflow influencing factor data obtained in step 1, and the feature permutation importance tool and SHAP tool were introduced to analyze the importance ranking of water inflow influencing factors and their marginal contribution to the water inflow prediction results, and the importance of water inflow influencing factors was ranked according to the feature importance score and SHAP value; Step 3: Using Python programming language, based on the key influencing factors of water inflow screened in step 2, high-efficiency water inflow prediction models are established respectively based on random forest, XGBoost and deep neural network model algorithms, and the performance of different high-efficiency water inflow prediction models is evaluated and compared according to the evaluation parameters of the output performance of the initial water inflow prediction model in step 2; Step 4: The model with the closest output performance of the training set and test set of each efficient water inflow prediction model in step 3 to the initial water inflow prediction model is determined as the optimal water inflow prediction model; different hydrogeological conditions and engineering conditions are set to verify the applicability of the optimal water inflow prediction model to different engineering construction environments again to ensure its effectiveness and accuracy in practical applications; Step 5: Build a sustainable intelligent prediction system for water inflow in tunnel construction; use the mysql-connector-python library to achieve efficient connection between Python and MySQL to ensure fast interaction of the optimal water inflow prediction model database; the sustainable intelligent prediction system for water inflow in tunnel construction inputs the key influencing factor data into the optimal water inflow prediction model through database management to make predictions; the prediction results are returned to the database and stored for subsequent query and analysis.
2. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1 is characterized in that: The specific implementation process of step one is as follows: The factors affecting the original water inflow include two categories, namely engineering conditions and hydrogeological conditions; among them, engineering conditions include construction methods, number of tunnel bodies, number of headings, construction cross-sectional area, cumulative footage, daily footage, tunnel burial depth, and static water height; hydrogeological conditions include water richness, groundwater type, length of crossing karst and faults, surrounding rock grade, permeability coefficient and infiltration coefficient.
3. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: The vector autoregression model of the two time series X and Y of rainfall and water inflow is as follows: Among them, X t and Y t represents the observed values of time series X and Y at time t; p is the maximum number of lags set; i is the lag period, and its value range is 1-p; and and are the coefficients in the model, representing the influence of X and Y at time ti on X and Y at time t; and is the error term.
4. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: The F test is used to perform the Granger causality test. The null hypothesis is that there is no significant causal relationship between time series X and Y. If the F value is significant, the null hypothesis is rejected, and it is believed that there is a significant Granger causal relationship between rainfall and water inflow.
5. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: When constructing the initial water inflow prediction model, in order to ensure that the training effect of the data does not overfit and to ensure the repeatability of the results, 80% of the data is used as the training set, 20% of the data is used as the test set, and the number of random seeds is set to 42.
6. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: The output performance evaluation parameters of the initial water inflow prediction model include mean absolute error, mean square error, root mean square error, and determination coefficient, according to the formula: Where n is the number of samples, y i is the measured value for each sample, is the average value of the measured value of this indicator. is the predicted value obtained using the model.
7. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: The feature permutation importance tool quantifies the impact of each influencing factor on the prediction ability of the initial water inflow prediction model by evaluating the difference between the performance of the initial water inflow prediction model on the original water inflow influencing factor dataset and the performance after the influencing factors are randomly permuted. Using the feature replacement importance tool, the contribution of each influencing factor to the change of water inflow is evaluated based on the initial water inflow prediction model, and the feature replacement importance is ranked according to the feature replacement importance score; The SHAP tool is based on the Shapley value in game theory. It quantifies the contribution of each influencing factor to the change in water inflow by calculating the marginal contribution of each influencing factor. The contribution and importance ranking of each influencing factor to the change in water inflow are obtained based on the bee colony diagram obtained through the SHAP tool analysis. In the bee swarm diagram, blue indicates a positive impact on the predicted water inflow, and red indicates a negative impact on the predicted water inflow. The darker the color, the higher the degree. The SHAP value calculation formula is: in: is the SHAP value of the kth influencing factor; N is the set of all influencing factors; S is the feature subset of influencing factor k; f(S) is the model prediction value of subset S; f(S∪{k}) is the model prediction of the features in S plus the features Factors with correlation coefficient with water inflow greater than 0.3, feature replacement importance greater than 0.03, and SHAP value greater than 500 were screened out and determined as key influencing factors of water inflow, providing a basis for the subsequent establishment of an efficient water inflow prediction model.
8. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: 85% of the key influencing factors and water inflow data were used as training sets, and the remaining 15% were used as test sets. The training set and test set data of all efficient water inflow prediction models were ensured to be the same, and the grid search method was used to adjust the parameters of the efficient water inflow prediction model.
9. The method for predicting water inflow in a sustainable tunnel construction based on machine learning technology according to claim 1, characterized in that: In the sustainable intelligent prediction system for water inflow in tunnel construction, MySQL is mainly used to store and manage various parameter data of tunnel construction, historical water inflow data, and the prediction results of the optimal water inflow prediction model; The embedded data set for sustainable intelligent prediction of water inflow in tunnel construction will be continuously and dynamically updated. By continuously adding new training samples, the universality and prediction accuracy of the model will be improved, ensuring that the system continues to optimize the prediction performance as the amount of data increases.
Citation Information
Patent Citations
Prediction method for bad geological type of shield tunneling based on Xgboost
CN108846521A
Water inflow prediction method and system for multi-factor adaptive section of tunnel face of water-rich tunnel
CN118862709A
Predictive Model Data Stream Prioritization
US20230123322A1
Cited By
Digital modeling method for landslide surge disaster prediction
CN120633431A
Mine pit water inflow prediction method based on artificial intelligence
CN120744335A
A method and system for predicting tunnel water inflow intervals based on physical residual adaptive constraints and seepage mechanism-guided conformal calibration.
CN122549953A
Tunnel water inflow interval prediction method and system based on physical residual adaptive constraint and seepage mechanism guided conformal calibration
CN122549953B