A TBM construction speed prediction method based on weighted random forest
By assigning weights to input parameters using a weighted random forest model, building a WRF algorithm framework, and optimizing hyperparameters, the uncertainty problem in TBM construction speed prediction was resolved, achieving high-precision construction speed prediction and abnormality warning, and ensuring the safety and efficiency of TBM construction.
Patent Information
- Application Number
- CN202210847956.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-07-19
AI Technical Summary
Existing TBM construction speed prediction models are unable to effectively identify the uncertainty of input parameters, resulting in low construction efficiency, severe tool wear, machine jams, and a lack of real-time adjustment capabilities.
A weighted random forest model was adopted to construct the WRF algorithm framework by assigning weights to different input parameters. Combined with ten-fold cross-validation to optimize hyperparameters, accurate prediction of TBM construction speed and early warning of abnormal sections were achieved.
It improves the prediction accuracy and generalization ability of TBM construction, can identify geological and mechanical uncertainties, reduce the impact of geological disasters, and ensure fast and safe construction.
Smart Images

Figure CN115115129B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of TBM tunneling performance prediction, and particularly relates to a TBM construction speed prediction method based on a weighted random forest. BACKGROUND
[0002] With the rapid economic development, in recent years, TBM has been widely popularized in many long and large tunnel excavation works. TBM tunneling mainly includes continuous excavation and tunneling at the working face, rock debris transportation and treatment, surrounding rock support and lining construction and other steps. Its efficient and intelligent mechanical construction makes the tunneling speed reach 3-10 times of the traditional drilling and blasting method, and the tunnel excavation forms a flow type construction, greatly reducing the labor intensity. The reasonable prediction and evaluation of TBM tunneling performance determine the success and benefit of TBM construction, and the effective evaluation of TBM tunneling risk is related to the construction period and cost control of infrastructure. Therefore, based on the information of the measured rock mass parameters, mechanical parameters and tunneling parameters, it is necessary to establish an accurate and effective TBM construction speed prediction and evaluation model, which has become the focus and difficulty in the field of TBM construction.
[0003] However, due to the limitations of construction equipment, construction time, experimental funds and operation errors of construction personnel, there is a certain uncertainty in the acquisition of geological conditions and tunneling parameters, and most of the point prediction models of TBM construction speed cannot identify the uncertainty of the input parameters themselves. If the tunneling parameters of TBM in the construction process cannot be adjusted in real time according to the geological conditions, it will often lead to low construction efficiency, serious tool wear problem, and even cause machine jamming accident. In view of the risk of construction process, the uncertainty of geological parameters and the uncontrollability of human factors, it is a certain challenge to carry out reasonable prediction for TBM construction speed.
[0004] Therefore, the application adopts a method of assigning different weights to optimize the hyperparameters, establishes a weighted random forest model to improve the synergistic relationship of different influencing factors, the model has a faster training speed, is not easy to overfit, and does not need complex parameter adjustment processing steps, has high precision, strong generalization performance and other advantages, and can provide scientific reference for TBM construction period estimation and construction risk early warning. SUMMARY
[0005] In view of the defects in the prior art, the application provides a TBM construction speed prediction method based on a weighted random forest, which proposes effective indexes for characterizing surrounding rock geological conditions, TBM mechanical properties and construction management factors, and selects model input parameters with strong characteristic importance according to specific engineering characteristics and prediction indexes. By assigning different weights to each input parameter, a TBM construction speed prediction model based on a weighted random forest is proposed, and reasonable weights are assigned for learning during model training. Each index can be adjusted adaptively, and the prediction result has important guiding significance for ensuring TBM rapid and safe construction.
[0006] The application provides a TBM construction speed prediction method based on a weighted random forest, which comprises the following steps:
[0007] Step 1: Constructing a TBM construction speed prediction data set considering the uncertainty of multiple source information;
[0008] Step 2: Selecting model input parameters based on geological conditions, tunneling conditions and management operation uncertainty;
[0009] Step 3: Assigning corresponding weights to different input parameters by using a weighted random forest method, and constructing a WRF algorithm framework based on the misclassification of the penalty node;
[0010] Step 4: Optimizing model hyperparameters by using ten-fold cross-validation and prediction data set, and training a TBM construction speed prediction model based on the WRF algorithm framework;
[0011] Step 5: Predicting the TBM construction speed of an unknown tunneling section based on the trained model and warning an abnormal section.
[0012] Preferably, the TBM construction speed prediction data set comprises TBM construction speeds of different geological sections and geological conditions, equipment parameters and management operation parameters.
[0013] Preferably, the selection of model input parameters comprises:
[0014] Step 2.1: Through mathematical statistical analysis of input parameters of different geological sections, invalid data at both ends of the lowest and highest thresholds is removed based on the 3σ rule;
[0015] Step 2.2: Statistics of historical TBM construction speed prediction models are performed, and comprehensive selection is performed in combination with the easiness of high-frequency use parameters and field parameters;
[0016] Step 2.3: The model input parameters are determined according to the comprehensive selection result.
[0017] Preferably, the weighted random forest method is used to assign corresponding weights to different input parameters, and a WRF algorithm framework is constructed based on the misclassification of the penalized nodes, including:
[0018] Step 3.1: Form a training set D = {X, Y} and a test set based on the prediction data set, wherein X is composed of N samples with M attributes, and Y is a unique target vector;
[0019] Step two: Bootstrap resampling method is used to resample the input training set D k times with replacement, to obtain k pseudo samples;
[0020] Step three: build a CART tree for training, randomly take m attributes as the subset of the current node, and find the optimal node separation value according to the weighted impurity G(u i ,v i );
[0021] Step four: do not prune when growing each tree, and the value of m remains unchanged during the entire training process. Stop growing when the termination condition is met;
[0022] Step five: combine the k regression trees to generate a random forest, input the test set into the trained random forest, take the average of the prediction values of a single regression tree as the output of the model, and then obtain the WRF algorithm framework.
[0023] Preferably, ten-fold cross-validation and prediction data set are used to optimize the model hyperparameters, and a TBM construction speed prediction model based on the WRF algorithm framework is trained, including:
[0024] Divide the prediction data set into 10 mutually exclusive subsets;
[0025] Select one subset as the model validation set in order, and the remaining 9 subsets as the training set. Learn the model based on the training set and test it on the validation set, and then realize 10 cross-validations of the model;
[0026] Select the optimal model parameters by comparing the average values of the evaluation indexes of the 10 groups;
[0027] Train the TBM construction speed prediction model based on the optimal model hyperparameters and the WRF algorithm framework.
[0028] Preferably, based on the trained model, the TBM construction speed of the unknown tunneling section is predicted and the abnormal section is warned, including:
[0029] Input the geotechnical parameters, TBM mechanical operation parameters and human management parameters of the unknown tunneling section into the trained model to obtain the corresponding TBM construction speed prediction value;
[0030] Based on the comparison of the predicted value and the standard value, it is determined whether there is an anomaly, and if so, a warning is given.
[0031] Preferably, the historical TBM construction speed prediction model is statistically analyzed, and the ease of obtaining high-frequency use parameters and field parameters is comprehensively screened, including:
[0032] Based on the statistical results, all historical TBM construction speed prediction models are obtained according to the time-use curve, and the use probability distribution of each historical TBM construction speed prediction model in different construction scenarios is obtained.
[0033] The historical parameter set of each historical TBM construction speed prediction model is obtained, and the superior parameters and inferior parameters in the historical parameter set are first divided, and the historical optimization factor set of the historical TBM construction speed prediction model is obtained, and the historical optimization factor set is second divided according to the optimization influence degree.
[0034] According to the number of superior units and the superior bias attribute of the first division unit in the first division result, the number of inferior units and the inferior bias attribute of the second division unit, and the grade division vector of the superior factor in the second division result, a corresponding verification mode is matched from the verification database.
[0035] Based on the verification mode, special verification elements are set to the first division unit in the first division result, and execution time verification is performed on the superior parameters of the first division unit, and supplementary verification elements are set to the second division unit in the second division result, and replacement time verification is performed on the inferior parameters of the second division unit.
[0036] According to the execution time verification result and the replacement time verification result, the to-be-referenced parameters are screened.
[0037] The same parameter analysis is performed on all to-be-referenced parameters, a same parameter appearance list is constructed, and a first parameter with a high appearance frequency is first calibrated.
[0038] According to the use probability distribution, a second parameter with a high use frequency of the same historical TBM construction speed prediction model is estimated, and the second parameter is second calibrated based on the same parameter appearance list.
[0039] Based on the first calibration result and the second calibration result, the first available parameter is screened.
[0040] The number of models corresponding to the same construction scene is obtained, and a first mapping relationship between each corresponding used model and the field parameters of the construction scene is established.
[0041] An intersection is calculated based on all the first mapping relationships to obtain the intersection times for different models and field parameters, and a second available parameter is selected according to the parameter weight of each field parameter;
[0042] A third available parameter is comprehensively selected based on the parameter adaptation degree of the first available parameter and the second available parameter.
[0043] Preferably, after the first division of the advantage parameters and the disadvantage parameters in the historical self-parameter set, the method further comprises:
[0044] A historical work log of each historical TBM construction speed prediction model in the running process is obtained, and an analysis array corresponding to the corresponding disadvantage data is obtained;
[0045] Each analysis element in the analysis array is standardized, and each two analysis elements are combined and numbered, and the maximum execution efficiency and the minimum execution efficiency of the analysis elements corresponding to the combined number are analyzed;
[0046] According to the combined number order, all the maximum execution efficiencies are constructed into a first curve and all the minimum execution efficiencies are constructed into a second curve;
[0047] The overlapping points of the first curve and the second curve are obtained, and when the execution efficiency corresponding to the overlapping points is greater than a preset efficiency, the combined number corresponding to the overlapping points is retained;
[0048] Meanwhile, the first curve is first fitted and the second curve is second fitted, the first discrete points based on the first curve and the second discrete points based on the second curve are determined and are discretely represented, and the minimum fitting interval range of the first fitting curve and the second fitting curve is obtained, and the first point in the first curve and the second curve in the minimum fitting interval range is screened, and the combined number corresponding to the first point is retained;
[0049] Based on the discrete representation result, a middle range based on the minimum fitting interval range is determined;
[0050] The first distance between each point in the middle range and each point in the minimum fitting interval range is determined;
[0051] From all the first distances, a second distance within a preset distance is screened, and the combined number corresponding to the second point matched with the second distance is retained;
[0052] Based on all the retained combined numbers, a final element is obtained as a disadvantage bias reference basis for the corresponding disadvantage data.
[0053] Preferably, in the process of predicting the TBM construction speed of the unknown tunneling section based on the trained model and warning the abnormal section, the method further comprises:
[0054] Record the rock-soil parameters of the unknown tunneling section, the TBM mechanical operation parameters and the human management parameters respectively, and the prediction thread based on the trained model;
[0055] Determine different record types, manage the key layers for the prediction thread, determine the participation function and the number of functions in the prediction process for each key layer, and calculate the prediction reliability of each key layer;
[0056]
[0057] Wherein, K1 represents the prediction reliability of the corresponding key layer; N0 represents the standard number of functions participating in the prediction of the corresponding key layer; N1 represents the actual number of functions participating in the prediction of the corresponding key layer; the value of e is 2.72; represents the actual contribution factor of the i1th participating function in the corresponding key layer; represents the standard contribution factor of the i1th participating function in the corresponding key layer; k i1 represents the prediction weight of the i1th participating function in the corresponding key layer;(k i1 ) max represents the maximum prediction weight in the N0 participating functions; represents the maximum contribution factor in the N0 participating functions; represents the prediction weight matched with the maximum contribution factor;
[0058] When the prediction reliability of all key layers is greater than the corresponding preset reliability, it is determined that the corresponding trained model is qualified;
[0059] Otherwise, trace back to the first key layer whose prediction reliability is not greater than the corresponding preset reliability, and determine the non-participating functions, and determine the first prediction space according to the ratio of the first running memory of the non-participating functions to the second running memory of the participating functions in the first key layer;
[0060] Obtain the common running memory and environment building memory of the participating functions and the non-participating functions, and determine the to-be-added space;
[0061] When the sum of the first prediction space and the to-be-added space satisfies the drivable running condition of the trained model, add a prediction space with the same sum on one side of the corresponding key layer to place the non-participating functions;
[0062] When the drivable running condition of the trained model is not satisfied, a new model is established, and the non-participating functions are placed in the new model;
[0063] Based on the processing results of the non-participating functions, the rock and soil parameters of the unknown excavation section, TBM mechanical operating parameters, and human management parameters are re-input to obtain a new construction speed, which is then compared with the original construction speed.
[0064] According to the comparison results, the corresponding alarm operation is determined.
[0065] The beneficial effects produced by the present invention are as follows:
[0066] 1. This paper proposes for the first time effective clustering indicators for characterizing surrounding rock geological conditions, TBM mechanical properties, and construction management factors. It also selects model input parameters with strong characteristic importance based on specific project characteristics and prediction indicators. This approach can, to a certain extent, characterize the uncertainties in multiple aspects, including rock and soil properties, mechanical equipment, and operation management.
[0067] 2. This paper proposes a weighted random forest method for optimizing hyperparameters by assigning different weights to each input parameter. The model has a fast training speed, is not prone to overfitting, and does not require complex parameter adjustment. It has the advantages of high precision and strong generalization ability.
[0068] 3. By analyzing the absolute error outliers of the prediction model, the present invention finds that the fluctuations of the prediction model are mutually confirmed by the results of advanced geological forecasts or the geological risks encountered during on-site excavation. The early warning based on the absolute error of the model can reduce the impact of geological disasters to a certain extent, which has important guiding significance for ensuring the rapid and safe construction of TBMs.
[0069] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0070] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0072] Figure 1 is a prediction flow chart in an embodiment of the present invention;
[0073] Figure 2 A radar chart showing the frequency of use of input parameters in an embodiment of the present invention;
[0074] Figure 3The influence of the number of CARTs k on the performance of the WRF model in the embodiment of the application;
[0075] Figure 4 The construction speed AR prediction result based on the WRF model and the absolute error in the embodiment of the application;
[0076] Figure 5 The construction geological risk map for a typical tunnel section in an embodiment application;
[0077] Figure 6 The flowchart of the TBM construction speed prediction method based on the weighted random forest in the embodiment of the application;
[0078] Figure 7 The structure diagram for determining the combination number in the embodiment of the application. DETAILED DESCRIPTION
[0079] The preferred embodiments of the application are described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the application, and are not used to limit the application.
[0080] Embodiment 1:
[0081] The application provides a TBM construction speed prediction method based on a weighted random forest, as shown in Figure 1 and 6 , including the following steps:
[0082] Step 1: Construct a TBM construction speed prediction data set considering the uncertainty of multi-source information, including the following steps:
[0083] Relying on the water conveyance tunnel project of the Lanzhou Water Source Construction Project, the following test work is carried out: on-site rock sampling, on-site testing, indoor rock physical and mechanical property testing, rock mineral composition identification, and rock wear resistance testing. Blocky rock (disturbed and damaged rock samples) in the rock muck generated by TBM under different lithological conditions is selected, and the corresponding surrounding rock coring (intact rock samples) is carried out on-site point load test to carry out on-site rock strength test. According to different geological conditions, representative lithology (covering three major lithology types of igneous rock, metamorphic rock and sedimentary rock, and taking into account hard rock and soft rock types) is selected on site, and rock samples of different lithology are taken, and the required size for testing is processed indoors. Based on the internationally recognized Cerchar abrasion method, the rock wear resistance index CAI is determined and the rock mineral composition is identified. TBM is the English Tunnel Boring Machine (Tunnel Boring Machine).
[0084] Step 2: Based on the uncertainty of geological conditions, tunneling conditions and management operations, the selection of model input parameters is carried out:
[0085] Statistics on TBM construction speed prediction models at home and abroad were conducted, and input parameter radar charts were obtained based on the usage frequency of each indicator ( Figure 2 ), based on relevant parameter statistics and on-site availability analysis, the rock mass classification system (RMR), rock abrasive coefficient (CAI), and rock mineral weighted hardness (H) were used as geological condition input parameters affecting AR to quantify the uncertainty of rock mass conditions under different geological environments. Actively controllable cutterhead thrust (TF) and rotational speed (RPM) were selected to quantify the uncertainty of mechanical equipment during construction. Other factors, such as outage (ORD), were also introduced to characterize the uncertainty of equipment and operational management. By averaging these data across different geological units, a database of TBM operating parameters under different geological conditions can be established. Descriptive statistics of the input and output parameters of the sample dataset are shown in Table 1:
[0086] Table 1 Descriptive statistics of input and output parameters of the model sample dataset
[0087]
[0088] Among them, the RMR value is approximately normally distributed, with most values ranging from 40 to 60, mainly belonging to Class III surrounding rock. Since the tunnel only involves five types of rock types, including quartz schist, diorite, granite, metamorphic andesite and sandstone, the rock abrasiveness index CAI value and the rock mineral weighted hardness value H are relatively discrete, with values ranging from 0.75 to 3.1 and 4 to 6, respectively. The cutterhead thrust presents a left-right double-skewed distribution, and most values are in the two ranges of 4000 to 5000 kN and 8000 to 9000 kN. The cutterhead speed is approximately normally distributed, mainly distributed around 6 r / min. Other downtime ORD is roughly normally distributed, with most values in the range of 30 to 40, and the measured value AR of the predicted target is also approximately normally distributed, mainly distributed in the range of 20 to 25.
[0089] Step 3: Use the weighted random forest method to assign corresponding weights to different input parameters and build the WRF algorithm framework based on penalizing node misclassification. This includes the following tasks:
[0090] First, a training input set D = {X, Y} is formed, where X is composed of N samples with M attributes and Y is a unique target vector. Then, the Bootstrap resampling method is used to perform k-times of repeated sampling with replacement on the input training sample set D to obtain k pseudo samples. The set in each pseudo sample is called in-bag data, and the data that is not sampled each time is called out-of-bag data. WRF uses out-of-bag data to perform internal error estimation to improve the generalization ability of the model. Further, a CART tree is constructed for training, and m attributes are randomly selected as a subset of the current node. According to the weighted impurity G(ui ,v i ) to find the optimal node separation value, G(u i ,v i ) value, the better the separation effect. Each tree is not pruned during split growth, and the m value remains unchanged throughout the training process. Growth stops when it meets the termination condition. Finally, k regression trees are combined to generate a random forest. The prediction set data is input into the trained random forest. The prediction values of individual regression trees are averaged as the output of the model. The final prediction result is:
[0091]
[0092] Where, f(x i ) is the final model prediction result, h i (x i ) is the result obtained by the i-th regression tree.
[0093] The quality of the separation of sub-nodes in the decision tree is generally measured by the weighted impurity G(u i ,v i ) to measure.
[0094]
[0095] Where u i and v i are a certain cut-off variable and its corresponding cut-off value, N S is the number of all training samples of the current node, N L ,N R are the number of training samples of the left child node and the right child node after segmentation, X L ,X R are the training sample sets of the left child node and the right child node after segmentation, respectively. H(X) is a function to measure the node impurity. The regression task is often expressed using the mean squared error (MSE).
[0096]
[0097] Where X is the training sample set of the current node, y i and are respectively the true value of the target variable of the current node sample and the average value of the WRF predicted value of the sample data.
[0098] Since the WRF adopts the replacement resampling method, the data of the training set samples are not the same, and the input attributes are randomly selected. The k pseudo samples are independent of each other, and the performance of the entire ensemble model can be effectively improved. The training speed is faster, and overfitting is not easy to occur, and complex parameter adjustment is not required. In addition, since the OOB data is not involved in the construction of the model regression tree, the OOB data can serve as a validation set.
[0099] Step four: using ten-fold cross-validation and prediction data set to optimize the model hyperparameters, and training the TBM construction speed prediction model based on the WRF algorithm framework:
[0100] Based on the data set of the water tunnel project in Lanzhou water source area, the corresponding optimal hyperparameters need to be searched first. For the random forest, the main hyperparameters include the regression tree k and the maximum number of features m. If the value of k is too small, the model is prone to underfitting, and if the value of k is too large, the calculation is complicated but the performance of the model cannot be significantly improved. In the WRF model, the value of k is selected in the range of [1, 800], and the mean square error of the validation set under different regression tree k is calculated iteratively. As shown in Figure 3 When the value of k is greater than 500, the MSE decreases to a stable value, and the value of k is set to 500 to reduce the complexity of the model. The value of m determines the degree of disturbance of the model, which is generally determined according to the empirical formula.
[0101] m = [log2M]
[0102] In the formula, M is the total number of features of the model input parameters, and [] represents the integer operation. In the established model, the input parameters have 6 attributes, so the value of m is estimated to be 2 according to the empirical formula.
[0103] As shown in Figure 4 , it can be found that the scatter distribution of the AR prediction value and the true value is not far apart, and the numerical values are close to each other. The average absolute error is 1.31 m / d, and the determination coefficient of the validation set reaches 0.95, indicating that the model has good prediction accuracy. Moreover, the overall trend reflected by the prediction value is similar to that of the true value, indicating that the weighted random forest algorithm can reasonably identify the geological variation during construction, and will not be inconsistent with the actual geological conditions in the excavation section;
[0104] Before performing step four, it also includes: using the same training set and validation set to establish different machine learning models, comparing and analyzing the generalization ability and robustness of the models, and selecting the best prediction model;
[0105] Using the same training set and validation set, a random forest regression model RF, a back propagation neural network model BPNN, and a support vector regression model SVR are established. The RMSE, R2 and MAPE and other related evaluation indicators are shown in Table 2.
[0106] Table 2 Comparison of evaluation indicators of different prediction models
[0107]
[0108] Among them, the improved weighted random forest model has higher prediction accuracy, and the RMSE value describing the deviation between the predicted value and the measured value of the AR prediction value of the verification sample is only 1.59, and the MAPE value is only 0.11, indicating that the prediction performance of the WRF method is the best. The R 2 values of the training and verification stages are equal to 0.96 and 0.97 respectively, and the values do not deviate much, indicating that the model does not have underfitting and overfitting phenomena. The prediction performance of the random forest model is inferior to the WRF model, but the error is smaller than that of the other two models. However, whether it is the training set or the verification set, the RMSE value and MAPE value of the SVR model are larger, and the R 2 value is significantly lower than that of other models, indicating that the sensitivity of the hyperparameters is larger, resulting in poor prediction accuracy stability of the model. In addition, the RMSE and R 2 values of the BPNN model under the training set and the verification set are quite different, indicating that the model has poor generalization ability and robustness.
[0109] In this way, the WRF related model is screened for subsequent speed prediction.
[0110] Step five: based on the trained model, the TBM construction speed of the unknown tunneling section is predicted, and the abnormal section is warned:
[0111] From the actual tunneling results, some geological intervals encountered multiple major geological risks during construction, and the prediction results of various models in the corresponding risk sections are shown in Table 3.
[0112] Table 3 Prediction results of different models in typical geological sections
[0113]
[0114] For lithological units 5-7 distributed in pile numbers T4+550~T8+737m, the tunnel passes through f2-f3 small faults ( Figure 5The prediction results of the WRF model based on the geological, mechanical, and construction management parameters collected within the section are 13.92 m / d, 11.13 m / d, and 19.95 m / d, respectively, and the absolute error of the AR prediction value of the No. 6 excavation section is 2.97 m / d, which represents an abnormal excavation condition, and can reflect that the input parameters of the model have the characteristics of representing abnormal geological and mechanical information in the TBM excavation process. The training of the model can extract effective features, and can improve the accuracy and stability of the risk division of different excavation sections.
[0115] The No. 14 geological section corresponding to the excavation stake number located at T14+622~T15+100 m has a risk of tunnel vault collapse during the excavation process, causing the TBM front shield to contact the surrounding rock (as shown in FIG. 14). Figure 5 The prediction result of the WRF model is 9.86 m / d, and the error between this value and the true value reaches 2.18 m / d, which is abnormal compared to the normal excavation section, and can reflect that the current geological section is facing certain construction risks. Compared with the other three models, the WRF model has the best prediction accuracy, which shows that the model can fully extract the information of the input parameters during the training, and give the weight of each feature vector to realize reasonable feedback.
[0116] The No. 18 and 19 geological sections have a sudden machine jam accident (as shown in FIG. 18c). Figure 5 According to the analysis of the TBM construction speed prediction results of the WRF model, the output targets of the construction speed are 6.17 m / d and 10.34 m / d, respectively, which are in a lower range, indicating that the TBM has poor performance in this geological section. Moreover, the absolute errors of the AR prediction values and the actual values are 3.3 m / d and 4.8 m / d, respectively, which are significantly higher than those of other excavation sections, indicating that the construction uncertainty is large. Therefore, the prediction results of the model can to some extent confirm the risk in the excavation process.
[0117] The No. 49 geological section has a squeezing deformation of soft surrounding rock, and the excavation stake number is located at T14+080~T14+100 m. The deformation of the top arch exceeds the gap between the tunnel wall and the shield (about 8 cm), and the surrounding rock and the shield are in extrusion contact (as shown in FIG. 49d). Figure 5 The prediction result of the WRF model is 17.25 m / d, and the absolute error of the TBM construction speed is 3.09 m / d, which is significantly abnormal. The relatively stable model has a large error fluctuation, which can reflect the changes in the excavation process.
[0118] According to the on-site excavation feedback, it is known that the geological section T19+752~T19+647m of No. 53 has serious collapse of the tunnel face and the crown, and the TBM slowly excavates through the V-class surrounding rock broken zone. Figure 5 e) The predicted value of TBM construction speed in this excavation section based on the WRF model is 12.28m / d, and the absolute error reaches 3.07m / d, and the fluctuation range is significantly higher than that of the normal excavation section. According to the absolute error of the prediction, it can be fed back that the front tunnel face is facing greater construction uncertainty, and the risk of TBM excavation obstruction or machine jamming is greater.
[0119] Embodiment 2
[0120] Based on the basis of Embodiment 1, the historical TBM construction speed prediction model is statistically analyzed, and the ease of obtaining high-frequency use parameters and field parameters is comprehensively screened, including:
[0121] Based on the statistical results, the use probability distribution of each historical TBM construction speed prediction model in different construction scenarios is obtained according to the time-use curve;
[0122] Get the historical parameter set of each historical TBM construction speed prediction model, and first divide the advantage parameters and disadvantage parameters in the historical parameter set, and at the same time, get the historical optimization factor set of the historical TBM construction speed prediction model, and according to the optimization influence degree, the second division is carried out on the historical optimization factor set;
[0123] According to the number of advantages and the advantage bias attribute of the first division unit in the first division result, the number of disadvantages and the disadvantage bias attribute of the second division unit, and the grade division vector of the advantage factor in the second division result, the corresponding verification mode is matched from the verification database;
[0124] Based on the verification mode, special verification elements are set to the first division unit in the first division result, and the execution time verification is performed on the advantage parameters of the first division unit, and at the same time, supplementary verification elements are set to the second division unit in the second division result, and the replacement time verification is performed on the disadvantage parameters of the second division unit;
[0125] According to the execution time verification result and the replacement time verification result, the to-be-referenced parameters are screened and obtained;
[0126] The same parameter analysis is performed on all to-be-referenced parameters, a same parameter occurrence list is constructed, and a first parameter with a high occurrence frequency is first calibrated;
[0127] According to the use probability distribution, a second parameter with high use frequency of the same historical TBM construction speed prediction model is estimated, and second calibration is performed based on the same parameter occurrence list;
[0128] Based on the first calibration result and the second calibration result, a first available parameter is screened;
[0129] The number of models corresponding to the use of the same construction scene is obtained, and a first mapping relationship between each corresponding used model and the field parameters of the construction scene is established;
[0130] Based on all the first mapping relationships, an intersection is obtained, and the intersection number of different models and field parameters is obtained, and according to the parameter weight of each field parameter, a second available parameter is screened;
[0131] Based on the parameter adaptation degree of the first available parameter and the second available parameter, a third available parameter is comprehensively screened.
[0132] In this embodiment, the use probability distribution is, for example, the participation probability of the model parameter when the scene 1, the scene 2, and the like are used, and finally the use probability distribution is constructed.
[0133] In this embodiment, the self-parameter set refers to various parameters of the model itself obtained by the model before each construction use, and the historical optimization factor set refers to the adjustment parameters obtained by adjusting the model according to the construction condition after each construction use of the model or the need to adjust the model.
[0134] In this embodiment, since some parameters in different parameter sets play a small role in the speed prediction process or reduce the prediction efficiency, such parameters are regarded as inferior parameters, and the remaining parameters are regarded as superior parameters.
[0135] In this embodiment, since different optimization factors have different optimization degrees, the optimization influence degree can be divided to obtain a grade division vector.
[0136] In this embodiment, the number of advantages refers to the number of superior parameters, and the number of disadvantages refers to the number of inferior parameters.
[0137] In this embodiment, the advantage bias attribute refers to the prediction advantage of the superior parameter in the prediction process, and the disadvantage bias attribute refers to the prediction disadvantage of the inferior parameter in the prediction process.
[0138] In this embodiment, the verification database includes different biases, grade division vectors, and corresponding verification modes, which facilitates matching of the verification mode to verify different division results.
[0139] In this embodiment, the special check element is related to time, mainly to set a check condition to check the execution time of the advantage parameter, and the supplementary check element is mainly to determine the replacement time required if the disadvantage parameter is replaced, so that the to-be-referenced parameter is screened according to time, and the prediction efficiency is improved to a certain extent.
[0140] In this embodiment, for example: advantage parameters 1, 2, and 3, disadvantage parameters 4 and 5, and finally the to-be-referenced parameters are 1, 2, 3, and 4.
[0141] In this embodiment, since each model has to-be-referenced parameters, the same parameter analysis is performed, and a list is constructed, the parameters with high occurrence frequency are first calibrated, and by using the probability distribution, the parameters with the frequency of use can be effectively obtained.
[0142] In this embodiment, scenario 1 corresponds to models 2 and 4, at this time, there is a first mapping relationship, and the same is true for intersection calculation, that is, the number of repeated occurrences is obtained, and combined with the weight, the second usable parameter in the scene parameter is determined.
[0143] In this embodiment, the higher the parameter adaptation degree, the more the number of third usable parameters screened.
[0144] In this embodiment, the model input parameters are: the rock mass classification system RMR value reflecting the geological conditions along the tunnel, the rock wear resistance CAI value and hardness H value, the cutter thrust TF value and the rotation speed RPM value reflecting the TBM mechanical tunneling effect, and other factors ORD quantifying human factors.
[0145] The beneficial effects of the above technical solution are: by obtaining the parameter set of the model itself and the optimization factor set, the check mode is reasonably matched to realize the check of the parameter execution time and the replacement time, ensure the prediction efficiency, and by establishing a list of parameters, it is convenient to first calibrate the high-frequency parameters, at the same time, according to the use probability distribution, it is convenient to second calibrate the parameters, two ways are adopted to realize reliable screening of the model parameters, and by determining the mapping relationship between the model and the scene, the second usable parameter is determined, finally through the adaptation degree, the rationality of parameter screening is ensured, which provides an accurate basis for subsequent model input, and it is convenient to assign accurate weights to each parameter, and the accuracy of the subsequent prediction speed is ensured.
[0146] Embodiment 3
[0147] Based on the basis of embodiment 2, after the first division of the advantage parameters and the disadvantage parameters in the historical own parameter set, the method further comprises:
[0148] Obtaining a historical working log of each historical TBM construction speed prediction model in a running process, and obtaining an analysis array corresponding to the corresponding disadvantage data;
[0149] Standardizing each analysis element in the analysis array, and combining each two analysis elements for numbering, analyzing the maximum execution efficiency and the minimum execution efficiency of the analysis element corresponding to the combined number;
[0150] According to the combined number order, all the maximum execution efficiencies are constructed into a first curve and all the minimum execution efficiencies are constructed into a second curve;
[0151] Obtaining the overlapping point of the first curve and the second curve, when the execution efficiency corresponding to the overlapping point is greater than the preset efficiency, the combined number corresponding to the overlapping point is reserved;
[0152] At the same time, the first curve is first fitted and the second curve is second fitted, the first discrete point based on the first curve and the second discrete point based on the second curve are determined and are discretely characterized, and the minimum fitting interval range of the first fitting curve and the second fitting curve is also obtained, and the first point in the first curve and the second curve in the minimum fitting interval range is screened, and the combined number corresponding to the first point is reserved;
[0153] Based on the discrete characterization result, an intermediate range based on the minimum fitting interval range is determined;
[0154] Determine the first distance between each point in the intermediate range and each point in the minimum fitting interval range;
[0155] From all the first distances, the second distance within the preset distance is screened, and the combined number corresponding to the second point matched with the second distance is reserved;
[0156] Based on all the reserved combined numbers, a final element is obtained as a disadvantage bias reference basis for the corresponding disadvantage data.
[0157] In this embodiment, the analysis element refers to the deviation of various parameters in the disadvantage data, abnormal conditions, and similar situations with the advantage data.
[0158] In this embodiment, the molecular array is various analysis elements, and the standardization refers to the standard conversion of the value of the element, which is convenient for calculating the execution efficiency, and the invalid effect brought by the analysis element is greater, and the corresponding execution efficiency is smaller.
[0159] In this embodiment, the overlapping point refers to the point where the maximum execution efficiency is equal to the minimum execution efficiency.
[0160] In this embodiment, as Figure 7As shown, A1 represents a first fitting curve, A2 represents a second fitting curve, C1 represents an intermediate range, C2 represents a minimum fitting range, and C3 represents a discrete point.
[0161] The beneficial effects of the above technical solution are: by obtaining the analysis array of disadvantage data and combining the elements two by two to determine the execution effectiveness, by constructing a curve, fitting a curve, and screening points to determine the combination number, and by constructing an intermediate range for discrete representation to screen the points again, the combination number is enriched, and finally the final element is obtained as a reference basis for disadvantage bias, which provides a basis for determining the disadvantage bias attribute and indirectly provides a basis for parameter screening, thereby indirectly improving the accuracy of the prediction speed.
[0162] Embodiment 4
[0163] Based on the basis of embodiment 1, in the process of predicting the TBM construction speed of an unknown tunneling section based on the trained model and warning the abnormal section, it further includes:
[0164] The rock-soil parameters, TBM mechanical operation parameters, and human management parameters of the unknown tunneling section are recorded based on the prediction thread of the trained model;
[0165] The different record types are determined, the prediction thread is managed by key layers, the participation function and the number of functions participating in the prediction process of each key layer are determined, and the prediction reliability of each key layer is calculated;
[0166]
[0167] Wherein, K1 represents the prediction reliability of the corresponding key layer; N0 represents the standard number of functions participating in the prediction of the corresponding key layer; N1 represents the actual number of functions participating in the prediction of the corresponding key layer; the value of e is 2.72; represents the actual contribution factor of the i1th participating function in the corresponding key layer; represents the standard contribution factor of the i1th participating function in the corresponding key layer; k i1 represents the prediction weight of the i1th participating function in the corresponding key layer;(k i1 ) max represents the maximum prediction weight of the N0 participating functions; represents the maximum contribution factor of the N0 participating functions; represents the prediction weight matched with the maximum contribution factor;
[0168] When the prediction reliability of all key layers is greater than the corresponding preset reliability, it is determined that the corresponding trained model is qualified;
[0169] Otherwise, trace the first key layer whose predicted reliability is not greater than the corresponding preset reliability, determine the non-participating functions, and determine the first prediction space according to the ratio of the first running memory of the non-participating functions to the second running memory of the participating functions in the first key layer;
[0170] Obtain the common running memory and environment building memory of the participating functions and the non-participating functions, and determine the to-be-added space;
[0171] When the sum of the first prediction space and the to-be-added space satisfies the drivable running condition of the trained model, add a prediction space with the same sum on one side of the corresponding key layer to place the non-participating functions;
[0172] When the drivable running condition of the trained model is not satisfied, establish a new model, and place the non-participating functions in the new model;
[0173] According to the processing result of the non-participating functions, the rock-soil body parameters, TBM mechanical operation parameters, and human management parameters of the unknown tunneling section are input again to obtain a new construction speed, which is compared with the original construction speed.
[0174] According to the comparison result, a corresponding alarm operation is determined to be executed.
[0175] In this embodiment, the rock-soil body parameters correspond to thread 1, and thread 1 includes key layers 2 and 8. At this time, the predicted reliabilities of key layers 2 and 8 are calculated respectively.
[0176] In this embodiment, for example, the functions that need to participate in the key layer are 10, but only 8 actually participate, so there are 2 non-participating functions left. The to-be-added space is determined by determining the memory ratio, the common running memory, and the environment building memory.
[0177] In this embodiment, the drivable running condition is set in advance, mainly for the non-participating functions.
[0178] In this embodiment, the new model is used to predict the speed in combination with the already trained model, to ensure the accuracy of the speed prediction.
[0179] In this embodiment, the alarm operation is to provide a reference for the construction personnel, to ensure the efficiency of the construction.
[0180] The beneficial effects of the above technical scheme are: in the process of predicting the speed, the management of the key layer in each thread is determined by recording each parameter thread, the prediction reliability of the key layer in the same thread is determined, a reference for predicting the speed is provided, and specifically, the relationship between the participating function and the non-participating function is traced back to add a new prediction space, the accuracy of the subsequent prediction speed is ensured, and an effective reference is provided for construction.
[0181] Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the spirit and scope of the present application. Thus, if these modifications and changes of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and changes.
Claims
1. A TBM construction speed prediction method based on weighted random forest, characterized in that, The application relates to a TBM construction speed prediction method based on a weighted random forest algorithm. The method comprises the following steps: Step 1: constructing a TBM construction speed prediction data set considering the uncertainty of multi-source information; Step 2: screening model input parameters based on geological conditions, tunneling conditions and management operation uncertainty; Step 3: assigning corresponding weights to different input parameters by using a weighted random forest method, and constructing a WRF algorithm framework based on the misclassification of a penalty node; Step 4: optimizing model hyperparameters by using ten-fold cross-validation and a prediction data set, and training a TBM construction speed prediction model based on the WRF algorithm framework; Step 5: predicting the TBM construction speed of an unknown tunneling section based on the trained model and warning abnormal sections; The screening of model input parameters comprises the following steps: Step 2.1: removing invalid data at the lowest and highest thresholds based on the 3sigma rule through mathematical statistical analysis of input parameters of different geological sections; Step 2.2: statistically analyzing historical TBM construction speed prediction models, and comprehensively screening based on the ease of acquisition of high-frequency use parameters and field parameters; Step 2.3: determining the model input parameters according to the comprehensive screening results; The comprehensive screening based on the statistical analysis of historical TBM construction speed prediction models and the ease of acquisition of high-frequency use parameters and field parameters comprises the following steps: Based on the statistical results, the use probability distribution of each historical TBM construction speed prediction model in different construction scenarios is obtained according to the time-use curve; The historical parameter set of each historical TBM construction speed prediction model is obtained, and the dominant parameters and the inferior parameters in the historical parameter set are first divided, and the historical optimization factor set of the historical TBM construction speed prediction model is obtained, and the historical optimization factor set is secondly divided according to the optimization influence degree; According to the number of dominant units and the dominant bias attribute in the first division result, the number of inferior units and the inferior bias attribute in the second division unit, and the grade division vector of the dominant factor in the second division result, a corresponding verification mode is matched from a verification database; Based on the verification mode, special verification elements are set to the first division unit in the first division result, and the execution time verification of the dominant parameters of the first division unit is performed, and supplementary verification elements are set to the second division unit in the second division result, and the replacement time verification of the inferior parameters of the second division unit is performed; According to the execution time verification result and the replacement time verification result, the to-be-referenced parameters are screened; The same parameter analysis is performed on all the to-be-referenced parameters, a same parameter appearance list is constructed, and a first parameter with a high appearance frequency is first calibrated; According to the use probability distribution, the second parameter with a high use frequency of the same historical TBM construction speed prediction model is estimated, and the second parameter is secondly calibrated based on the same parameter appearance list; Based on the first calibration result and the second calibration result, the first available parameter is screened. Obtaining the number of models corresponding to the same construction scene, and establishing a first mapping relationship between each corresponding model and the field parameters of the construction scene; Based on all the first mapping relationships, the intersection is calculated to obtain the intersection times of different models and field parameters, and according to the parameter weight of each field parameter, the second available parameter is screened; Based on the parameter adaptation degree of the first available parameter and the second available parameter, the third available parameter is comprehensively screened.
2. The TBM construction speed prediction method based on weighted random forest according to claim 1, characterized in that, The TBM construction speed prediction dataset includes: TBM construction speed and geological conditions, equipment parameters and management operation parameters in different geological sections. 3.The TBM construction speed prediction method based on weighted random forest of claim 1, wherein, The weighted random forest method is used to assign corresponding weights to different input parameters, and the WRF algorithm framework is constructed based on the misclassification of the penalty node, including: Step 3.1: Forming a training set based on the prediction dataset D ={ X , Y} and a test set, wherein X is composed of M samples with N attributes, Y is a unique target vector; Step two: Bootstrap resampling method is used to input training set D k times with replacement, get k pseudo samples; Step three: build CART tree for training, randomly take m a property as the subset of the current node, find the optimal node separation value according to weighted impurity G ( u i , v i ) Step four: no pruning is done while each tree is growing, m The value is kept constant throughout the training process and stops growing when it meets the termination condition. Step five: input the test set into the trained random forest and average the predictions of the individual regression trees to obtain the output of the model, thereby obtaining the WRF algorithm framework. k A random forest is generated by combining regression trees. The test set is input into the trained random forest, and the average of the prediction values of the individual regression trees is taken as the output of the model, thereby obtaining the WRF algorithm framework.
4. The TBM construction speed prediction method based on weighted random forest according to claim 1, characterized in that, The ten-fold cross-validation and the prediction dataset are used to optimize the model hyperparameters, and the TBM construction speed prediction model based on the WRF algorithm framework is trained, including: The prediction dataset is divided into 10 mutually exclusive subsets; Select one subset as the model validation set in order, and the remaining 9 subsets as the training set. Based on the training set, the model is learned and tested on the validation set, and then 10 cross-validations of the model are realized. The optimal model parameters are selected by comparing the average values of the evaluation indexes of the 10 groups; Based on the optimal model hyperparameters and the WRF algorithm framework, the TBM construction speed prediction model is trained.
5. The TBM construction speed prediction method based on weighted random forest according to claim 1, characterized in that, Based on the trained model, the TBM construction speed of the unknown tunneling section is predicted and the abnormal section is warned, including: The rock and soil parameters, TBM mechanical operation parameters and human management parameters of the unknown tunneling section are input into the trained model to obtain the corresponding TBM construction speed prediction value; Based on the comparison between the prediction value and the standard value, it is determined whether there is an anomaly. If there is, a warning is given.
6. The TBM construction speed prediction method based on weighted random forest of claim 1, wherein, After the first division of the advantage parameters and the disadvantage parameters in the historical self-parameter set, it further includes: Obtain the historical work log of each historical TBM construction speed prediction model during operation, and obtain the analysis array corresponding to the disadvantage data; Standardize each analysis element in the analysis array, and combine each two analysis elements to obtain the maximum execution efficiency and the minimum execution efficiency of the analysis elements corresponding to the combined number; According to the combined number order, all the maximum execution efficiencies are constructed into a first curve and all the minimum execution efficiencies are constructed into a second curve; Obtain the overlapping point of the first curve and the second curve. When the execution efficiency of the overlapping point is greater than the preset efficiency, the combined number corresponding to the overlapping point is retained; At the same time, the first curve is first fitted and the second curve is second fitted to determine the first discrete point based on the first curve and the second discrete point based on the second curve and perform discrete representation. Also, the minimum fitting interval range of the first fitting curve and the second fitting curve is obtained, and the first point in the minimum fitting interval range of the first curve and the second curve is screened, and the combined number corresponding to the first point is retained; Based on the discrete representation result, the intermediate range based on the minimum fitting interval range is determined; Determine the first distance between each point in the intermediate range and each point in the minimum fitting interval range; Filter the second distance within the preset distance from all first distances, and keep the combination number corresponding to the second point matched with the second distance; Based on all the retained combination numbers, obtain the final element as the disadvantage bias reference basis corresponding to the disadvantage data.
7. The TBM construction speed prediction method based on weighted random forest of claim 1, wherein, In the process of predicting the TBM construction speed of an unknown tunneling section based on the trained model and warning the abnormal section, the process further includes: Respectively record the rock and soil parameters, TBM mechanical operation parameters and human management parameters of the unknown tunneling section based on the prediction thread of the trained model; Respectively determine different record types, manage the key layers of the prediction thread, determine the participation function and the number of functions of each key layer in the prediction process, and calculate the prediction reliability of each key layer; K1= N0eN0+ 1 N0= K1eK1+ 1 ei1= N0ei1+ 1 ei1= N0ei1+ 1 wi1= N0wi1+ 1 wi1= N0wi1+ 1 wi1= N0wi1+ 1 wi1= N0wi1+ 1 When the prediction reliability of all key layers is greater than the corresponding preset reliability, it is determined that the corresponding trained model is qualified; Otherwise, trace back to the first key layer whose prediction reliability is not greater than the corresponding preset reliability, determine the non-participating function, and determine the first prediction space according to the ratio of the first running memory of the non-participating function to the second running memory of the participating function in the first key layer; Obtain the common running memory and environment building memory of the participating function and the non-participating function, and determine the to-be-added space; When the sum of the first prediction space and the to-be-added space meets the drivable running condition of the trained model, add a prediction space with the same sum on one side of the corresponding key layer to place the non-participating function; When the drivable running condition of the trained model is not met, a new model is established, and the non-participating function is placed in the new model; According to the processing result of the non-participating function, the rock and soil parameters, TBM mechanical operation parameters and human management parameters of the unknown tunneling section are input again to obtain a new construction speed, and the new construction speed is compared with the original construction speed; According to the comparison result, determine to perform the corresponding alarm operation.