A method and system for predicting surface settlement caused by tunnel construction based on ATD model
Through the ATD model-based method, combined with ARMA model decomposition and unsupervised learning, the problem of difficult to balance the prediction accuracy and calculation costs of surface settlement caused by tunnel construction is solved, and higher prediction accuracy and lower errors are achieved.
Patent Information
- Application Number
- CN202411711858.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The prior art has problems of difficulty in balancing accuracy and computational costs in accurately predicting surface settlement caused by tunnel construction, and the feature construction and selection methods lack universal applicability.
Using an ATD model-based approach, machine learning and deep learning algorithms are constructed to improve the accuracy of surface settlement prediction through data collection and preprocessing, ARMA model decomposition and unsupervised learning.
The accuracy of surface settlement prediction is significantly improved, the R² value is improved, the RMSE, MAE and MAPE is reduced, the 95% confidence interval of prediction error and the number of errors exceeding the critical value.
Smart Images

Figure CN119557808B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of surface settlement prediction, and in particular to a method and system for predicting surface settlement caused by tunnel construction based on an ATD model. Background Art
[0002] With the increasing popularity of subway and high-speed railway tunnels in environmentally sensitive areas such as airports, the need to control surface settlement to protect tunnels and surface structures has become more urgent. To prevent settlement from exceeding the limit, it is crucial to accurately predict the surface settlement caused by tunnel excavation and take appropriate pre-emptive measures. To accurately predict the surface settlement caused by tunnel excavation, researchers have introduced a series of analysis methods, such as fitting formulas, theoretical calculations, numerical simulations, and experimental simulations, based on mature theories such as continuum theory and random medium theory.
[0003] However, each of these methods has its own advantages and disadvantages. The implementation of fitting formulas, theoretical calculations, numerical simulations, and experimental simulations all face many challenges. The application of fitting formulas and theoretical calculations is limited by the difference between theoretical assumptions and actual conditions. At the same time, it is difficult to achieve the best balance between accuracy and computational cost when numerical simulation is applied. In addition, replicating complex construction environments through experimental simulations poses major challenges in terms of cost and technical complexity. Machine learning (ML) and deep learning (DL) have made significant progress in classification, multivariate regression prediction, and time series prediction, and have been widely used in the field of civil engineering due to their advanced data processing and pattern recognition capabilities, providing a new method for surface settlement prediction.
[0004] Based on big data optimization algorithms, maximum surface settlement (MSS) and surface settlement time series (SSTS) prediction algorithms were developed to improve the accuracy of surface settlement prediction. However, it is difficult to apply a unified optimization algorithm to the prediction needs of all tunnels due to the diverse environmental attributes, geological conditions, and construction plans of different tunnel engineering projects. These factors increase the complexity and variability of their impact on surface settlement. The SSTS dataset contains more hidden connections between parameters than the MSS dataset. The typical sample size of the MSS dataset is less than 300, while the sample size of the SSTS dataset is usually less than 300,000. These numbers are much smaller than the datasets with millions of samples that form the basis of big data optimization algorithms. Direct deployment of some optimization algorithms will inevitably lead to the inability to fully realize their potential in surface settlement prediction. Therefore, researchers have improved the effectiveness and accuracy of surface settlement prediction models through feature engineering, including data cleaning, feature construction, data enhancement, and feature selection. Similarly, the feature engineering used in MSS or SSTS also poses a challenge to the prediction algorithm. The data cleaning and feature selection processes have little effect on the accuracy. In addition, when selecting features, the sensitivity of different algorithms to the features must be considered. The same feature construction method will produce different prediction accuracy results depending on the algorithm. Therefore, it can be concluded that specific feature construction and selection methods are not universally applicable.
[0005] To solve the above problems, this study proposed a method for predicting surface settlement caused by tunnel construction based on the ATD model. Summary of the invention
[0006] The present invention is proposed in view of the above-mentioned background problems.
[0007] Therefore, the problem to be solved by the present invention is how to optimize the algorithm model to improve the accuracy of the prediction algorithm.
[0008] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0009] The first aspect of the present invention provides a method for predicting surface settlement caused by tunnel construction based on an ATD model, comprising the following specific steps:
[0010] S1. Data collection and preprocessing: Collect and process the original surface subsidence data in the monitoring area and construct the original data set;
[0011] S11, data resampling and outlier processing;
[0012] S12, data resampling: by using MATLAB resampling function to obtain effective resampling points from the original data set to remove outliers;
[0013] S13. Outlier processing: Outliers in the data are processed by using the filloutlier function and medfilt function in MATLAB; thus obtaining the corresponding data set for unsupervised learning.
[0014] S2. Constructing an ATD model based on the ARMA model: decomposing the surface settlement time series into two sub-features to enhance the learning ability of machine learning and deep learning algorithms for nonlinear surface settlement data;
[0015] S21, Stationarity verification: Calculate the autocorrelation and partial correlation coefficients to ensure that the time series meets the smoothness requirements of the ARMA model;
[0016] S22. Decompose the time series into different components based on frequency and intensity: and ;
[0017] S23, randomly dividing the above two parts into a training set and a test set according to their proportions for model training;
[0018] S24. Verify the model performance.
[0019] S3, using unsupervised learning to build corresponding networks to identify hidden patterns and internal structures;
[0020] S4, using machine learning algorithms and deep learning algorithms to independently predict sub-feature sequences on the network constructed in S3;
[0021] S5. Combining the prediction results of the sub-feature sequences obtained in S4 in a linear manner to predict surface subsidence.
[0022] Preferably, in step S1, the surface settlement data are monitored in real time using monitoring points arranged by a plurality of monitoring robots, and an original data set is constructed based on the monitoring points.
[0023] Preferably, in step S23, two one-dimensional networks are used to train and Used to reduce the amount of calculation.
[0024] Preferably, in step S24, the model performance is verified using the determination coefficient R², root mean square error RMSE, mean absolute error MAE and mean absolute percentage error MAPE.
[0025] Preferably, the machine learning algorithm in step S4 includes support vector machine SVM, random forest RF, extreme learning machine ELM and back propagation BP;
[0026] Deep learning algorithms include convolutional neural networks (CNN), long short-term memory networks (LSTM), and gated recurrent units (GRU);
[0027] Particle swarm optimization PSO and Bayesian optimization BI are used as hyperparameter optimizers for LSTM and GRU respectively; four algorithms are implemented: BI-LSTM, BI-GRU, PSO-LSTM and PSO-GRU, and the optimized hyperparameters are given.
[0028] In feature engineering-based dataset optimization, data cleaning and feature selection can generate more relevant datasets. Similarly, feature construction and data enhancement can provide additional data to help algorithms train and understand datasets more effectively. Time series decomposition is a statistical technique widely used in a variety of predictive modeling applications, including oil price forecasting and wind speed forecasting. In the most common form, time series decomposition decomposes time series into four basic components: trend term (T), seasonal variation trend (S), cyclical trend (C) and irregular variation (I). The most commonly used models are autoregression (AR), moving average (MA) and autoregressive moving average (ARMA). However, surface settlement time series data has periodic fluctuations, stages and randomness characteristics. Therefore, the present invention adopts ARMA as a data enhancement method to develop a surface settlement data decomposition method.
[0029] The ARMA model is a combination of the AR model and the MA model, so it is suitable for processing and analyzing mixed periodic and random time series. The AR model is based on two assumptions: 1) There is a strong correlation between the mark values at different time points; 2) The correlation between two time points is inversely proportional to their distance. The standard mathematical expression of the AR model is: ;in, is the time series value at time t, which is a linear combination of the previous p historical values. c is a fixed value representing the baseline or average offset of the time series data. is the ith autoregressive coefficient, is the disturbance term at time t.
[0030] The MA model is based on the assumption that the data at each time point are independent and follow the same distribution of white noise terms with constant mean and variance: ;in, is the data at the current time point (t), is the mean or expected value of the time series, is the sliding average coefficient of the i-th white noise term, which measures the influence of the corresponding white noise on the current time point. is the white noise term at the previous moment.
[0031] Therefore, the ARMA model fully considers the impact of the historical values and error terms of the time series. It combines autoregression (AR) and moving average (MA) to describe the dynamic characteristics of time series data. The ARMA model can be expressed as: ; Where p and q represent the lag orders in the AR and MA moving average models, respectively.
[0032] According to the characteristics of surface settlement time series data, this study assumes that the surface settlement Sss(t) is caused by tunnel construction. and surface settlement caused by random disturbance Based on this, the ARMA model can be used to decompose the surface subsidence time series into two parts: 1) is the historical value of the low-frequency settlement time series caused by a series of construction; 2) is the white noise value of the sedimentation time series caused by a series of random factors. Therefore, the ARMA model expression can be rewritten as: ; To simplify the model complexity, p = q is proposed. In addition, based on the assumption of the MA model, it is recommended that the average sum of the white noise terms be set to 0, that is: ; If there are n samples in total, then from t=0 to t=n and , which can form two time series respectively and : ;in, is a strongly random time series with periodicity and mean c, and \(S_t\) is a less periodic and more linear time series.
[0033] Fast Fourier Transform (FFT) is a mathematical operation that can be used to obtain the spectrum of the surface subsidence time series. Based on this, the spectrum characteristics of the surface subsidence time series can be verified. and The segmentation frequency of , thus completing the time series segmentation. The mean value of And it has a certain periodicity, and the background noise cutoff frequency can be determined according to the average background noise intensity, thereby completing data noise reduction.
[0034] The ATD model was constructed based on ARMA, and the surface settlement time series data was decomposed into two subsequences. and In order to improve the accuracy of surface subsidence prediction, the present invention introduces unsupervised learning and develops four machine learning and five deep learning algorithms based on it.
[0035] Unsupervised learning is of great significance in data mining, pattern recognition, and feature learning. The application of unsupervised learning enables researchers to gain a deeper understanding of data by elucidating the intrinsic structure and patterns of the data, thereby predicting useful information and extracting insights from the data. The introduction of unsupervised learning requires the establishment of a corresponding dataset that does not contain any labels other than the original data. In the prediction of surface subsidence time series, the prediction model usually predicts future values based on multiple historical features with time labels. In this context, the time labels can be removed to establish a corresponding dataset for unsupervised learning. The use of unsupervised learning helps to mine the hidden patterns and intrinsic structures of the data, thereby improving the prediction accuracy and generalization ability.
[0036] The surface settlement prediction method based on the ATD model starts with data collection. The automated monitoring equipment used in the present invention improves the efficiency of data collection. However, this also brings challenges such as abnormal data fluctuations, inconsistent collection intervals, and outlier anomalies. These anomalies directly affect the quality of the data set and the accuracy of prediction.
[0037] Therefore, resampling of raw data and outlier processing are indispensable stages in data processing. In addition, resampling is used to convert raw monitoring data into time series with equal intervals to meet the requirements of unsupervised learning mentioned above. In order to maintain the structure and characteristics of the data, this study established a unified data processing standard by performing spectral analysis on the data set.
[0038] After data processing, smoothness verification is performed. Autocorrelation and partial correlation coefficients are calculated to ensure that the time series meets the smoothness requirements of the ARMA model. After completing the smoothness test, the time series is decomposed into two different parts based on frequency and intensity: and .
[0039] Then, the two parts are randomly divided into training set and test set according to their proportion for model training. The input data is rearranged to form a sequence matrix X of size D×W D,W , where D is the input delay, W is the input width, and is the input dimension. As mentioned above, the input data sequence of the present invention can be expressed as X 3,W In this representation, the first and second columns represent the input feature subsequences at the current moment. and , and the last column represents the output sequence for the next prediction moment The present invention uses two one-dimensional networks instead of the original three-dimensional network to train and Then, the outputs of the two prediction models are linearly combined as This method not only reduces the amount of calculation of the activation function by at least one third, but also introduces the concept of unsupervised learning.
[0040] During the algorithm training process, the entire training set is divided into multiple small batches. In order to optimize the computing performance of the GPU and prevent fluctuations within the data batch from affecting the loss curve, it is necessary to ensure that the resampled signal contains an integer multiple of the minimum batch size. Therefore, the selected minimum batch size is a power of 2. The present invention uses a dynamic validation set so that the training set retains the validation features while maintaining a sufficient sample size. The validation is intended to prevent overfitting while ensuring the reliability of the prediction model. The number of validation iterations depends on the batch size and the simplex size.
[0041] The second aspect of the present invention provides a system for predicting surface settlement caused by tunnel construction based on an ATD model, which executes the steps of the above method, including a data acquisition robot, a data processing module, a model calculation module and an output module;
[0042] The data collection robot is used to collect surface subsidence data in a specified area and construct an original data set;
[0043] The data processing module is used to resample the original data and process outliers to improve the quality of the data set and the prediction accuracy; the processed data set is used as input, and the prediction result is obtained after calculation by the model calculation module;
[0044] The prediction results are output through the output module.
[0045] The third aspect of the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method for predicting surface settlement caused by tunnel construction based on the ATD model are implemented.
[0046] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned method for predicting surface subsidence caused by tunnel construction based on the ATD model are implemented.
[0047] The beneficial effects of the present invention are:
[0048] This paper proposes a new method based on autoregressive moving average (ARMA) time series decomposition and unsupervised learning for predicting surface settlement caused by tunnel construction. The effectiveness of the ATD model in enhancing the prediction performance of four machine learning (ML) and five deep learning (DL) algorithms is verified using the Kunming dataset. These algorithms include support vector machine (SVM), random forest (RF), extreme learning machine (ELM), back propagation (BP), convolutional neural network (CNN), particle swarm optimization (PSO), and Bayesian optimization (BI) of GRU and LSTM. The main findings are as follows:
[0049] (1) By combining data features and spectrum analysis results, the ATD model is used to divide the surface subsidence time series into two parts: 1) It is a low-frequency impact signal caused by tunnel construction; 2) It is a periodic signal composed of square and triangle waves caused by random factors. The background noise intensity and boundary frequency are 0.00983 and 1.4e-3 Hz respectively.
[0050] (2) Parabolic features represent the main structure of prediction errors in BP, CNN, and PSO-GRU. This study argues that these parabolic features represent nonlinear hidden data features of surface subsidence that are difficult to capture for both ML and DL.
[0051] (3) The ATD model divides the surface subsidence data into two subsequences with more obvious linear relationships. Therefore, the ATD model significantly enhances the learning ability of parabolic features. This may also cause the prediction models of some algorithms (such as CNN) to capture more invalid features, thereby reducing the prediction accuracy.
[0052] (4) The algorithm performance was evaluated using R², RMSE, MAE, and MAPE. Compared with the original algorithm, the application of the ATD model and unsupervised learning improved the algorithm proposed in this study by 32.62%, 19.91%, 23.37%, and 16.44%, respectively.
[0053] (5) The ATD model also reduced the 95% confidence interval of the model prediction error by 20%, and the number of samples exceeding the ±3 mm precipitation threshold was reduced by about 40% on average. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0055] Figure 1 This is a flow chart of the surface subsidence prediction method based on the ATD model;
[0056] Figure 2 It is the plan layout of the measuring points;
[0057] Figure 3 The following is a schematic diagram of the geological conditions of a typical section;
[0058] Figure 4 Schematic diagram of detection interval distribution;
[0059] Figure 5 This is the flow chart of the data processing for the monitoring point;
[0060] Figure 6 This is a schematic diagram of the time series stationarity test results;
[0061] Figure 7 It is a schematic diagram of periodic signal characteristics;
[0062] Figure 8 This is a schematic diagram of the spectrum characteristics of surface subsidence monitoring data;
[0063] Fig. 9 This is a schematic diagram of the time series decomposition of monitoring data at a typical measuring point;
[0064] Fig.10 This is a schematic diagram of improving the algorithm accuracy evaluation index;
[0065] Fig.11 It is a schematic diagram of the fit between the predicted value and the measured value;
[0066] Fig.12 is the forecast error distribution diagram;
[0067] Fig.13 This is a TD decomposition result diagram in the embodiment;
[0068] Fig.14 This is a schematic diagram of BP prediction results in the embodiment;
[0069] Fig.15 It is a schematic diagram of the BP surface subsidence prediction results in the embodiment. DETAILED DESCRIPTION
[0070] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0071] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0072] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0073] Example 1: Figure 1 The figure shows the complete flow chart of the surface subsidence prediction method based on the ATD model. After the surface subsidence data is automatically collected, data resampling and outlier processing are first performed to establish a corresponding data set for unsupervised learning. Then, the background noise intensity and boundary frequency are determined by spectrum analysis. Subsequently, the ATD model is used to decompose the time series into and Then, use the ratio divided by and The training set trains two one-dimensional networks separately.
[0074] The input data is rearranged to form a sequence matrix of size D×W X D,W , where D is the input delay, W is the input width, and is the input dimension. As mentioned above, the input data sequence of the present invention can be expressed as X 3,W In this representation, the first and second columns represent the input feature subsequences at the current moment. and , and the last column represents the output sequence for the next prediction moment The present invention uses two one-dimensional networks instead of the original three-dimensional network to train and Then, the outputs of the two prediction models are linearly combined as This method not only reduces the amount of calculation of the activation function by at least one third, but also introduces the concept of unsupervised learning.
[0075] During the algorithm training process, the entire training set is divided into multiple small batches. In order to optimize the computing performance of the GPU and prevent fluctuations within the data batch from affecting the loss curve, it is necessary to ensure that the resampled signal contains an integer multiple of the minimum batch size. Therefore, the selected minimum batch size is a power of 2. This study uses a dynamic validation set to keep the training set while retaining validation features while maintaining a sufficient sample size. Validation is designed to prevent overfitting while ensuring the reliability of the prediction model. The number of validation iterations depends on the batch size and the simplex size.
[0076] Finally, the above network is linearly integrated and the value is used as the predicted value of surface subsidence. The present invention uses four commonly used evaluation indicators to evaluate the prediction performance, specifically:
[0077] The machine learning algorithms introduced include support vector machine (SVM), random forest (RF), extreme learning machine (ELM) and back propagation (BP). Deep learning algorithms include convolutional neural network (CNN), long short-term memory network (LSTM) and gated recurrent unit (GRU). RNN-based algorithms (such as LSTM and GRU) have shown significant results in various applications of SSTS. In order to enhance the effect of the algorithm under a limited number of data samples, this study uses particle swarm optimization (PSO) and Bayesian optimization (BI) as hyperparameter optimizers for LSTM and GRU, respectively. Four algorithms were implemented: BI-LSTM, BI-GRU, PSO-LSTM and PSO-GRU, and the optimized hyperparameters are given (see Table 1 for the hyperparameters optimized using PSO and BI)
[0078] Table 1 Hyperparameters optimized using PSO and BI
[0079]
[0080] The following is a detailed description of this solution using a specific case:
[0081] A railway tunnel is 2,700 meters long and has a depth between 39 and 60 meters. The tunnel passes through the flight management area (FMA), southern work area, internal roads and external urban roads of Kunming Airport from north to south. The present invention selects a 267-meter-long section of the tunnel passing through the airport to predict surface subsidence, which contains 159 monitoring points. Figure 2 shows the distribution of monitoring points and the changes in surface elevation, with a vertical spacing of 20 meters and a horizontal spacing of 10 meters. Figure 3 depicts a typical cross-section of the tunnel, which contains three layers of rock and soil and a depth between 44 and 47 meters. To ensure measurement accuracy, seven Leica Nova MS60 3D scanning robots were arranged inside the FMA, and six TM60 precision monitoring automated measurement robots were deployed outside the FMA to monitor surface subsidence in real time.
[0082] In this example, 287,297 samples were collected from 159 monitoring points, with an average collection interval of 3.85 hours and continuous monitoring for 290 days. In other words, each monitoring point generated an average of 1,806 data points. Figure 4 The distribution of acquisition intervals at representative monitoring points and three representative SS signals are shown. Figure 5 The characteristics of the surface settlement data obtained from the construction of the Kunming high-speed railway tunnel (hereinafter referred to as the Kunming dataset) are shown. It can be seen that the acquisition frequency of the Kunming dataset is 1 time / 6 hours or 4.63e-5 Hz, and the acquisition frequency is 1 time / 2 hours or 1.38e-4 Hz when the data is abnormal.
[0083] By using the MATLAB resampling function, 1,792 (7 × 23 × 25, the minimum batch size used in this study is 25) resampled points were obtained from each monitoring point. Then, the filloutlier function and medfilt function of MATLAB were combined to deal with data outliers.
[0084] As shown in Figure 5, the resampling process does not change the intrinsic structure of the data and effectively removes outliers. The smoothness test results based on the Kunming data set are shown in Figure 6. The autocorrelation coefficient of the Kunming data set is close to 1, and most of the partial correlation coefficients are 0. Therefore, the Kunming data set meets the prerequisites for using the ARMA model.
[0085] The autocorrelation and partial correlation coefficients are calculated as follows:
[0086] Autocorrelation Coefficient and Partial Autocorrelation Coefficient are important concepts in time series analysis, which are used to describe the correlation of sequences at different lag orders.
[0087] The autocorrelation coefficient describes the correlation of a time series at different lags, that is, the degree of linear correlation between the series and its own lagged values. It measures the relationship between current values and past values. For a given time series , whose mean is Autocorrelation function at lag k Defined as: ; The partial autocorrelation function is used to measure the current value in the time series and the value of lagged k periods The pure correlation between , , …, For the lag order k, the partial autocorrelation function Defined as: ; is based on is based on , , …, of The least squares prediction value, corr represents the correlation coefficient.
[0088] The surface settlement caused by tunnel excavation is continuous and fluctuating, and can be conceptualized as a series of impact signals triggered by low-frequency construction activities. These impact signals and their spectra are shown in Figure 7. The power intensity of the square wave and the triangle wave gradually decreases, while the periodic impact wave shows a series of impact power intensities. Figure 8 shows that the spectrum signal of the surface settlement in the Kunming dataset has a significant decrease in volatility, which indicates that periodic signals are prevalent in the collected signals. The significant volatility of the signal is accompanied by a series of square waves and triangle waves, and the power intensity gradually decreases within the frequency range of 1e-3 Hz. By calculating the average intensity of these periodic extremes, the background noise intensity is obtained to be 0.00983, and the starting frequency is 8e-3 Hz. To reduce the background noise, a second-order Butterworth bandpass filter is applied, and the upper cutoff frequency is set to 8 × 10⁻³Hz.
[0089] As mentioned before, the acquisition frequencies of surface subsidence monitoring data are 1.38e-4 Hz and 4.63e-5 Hz. Therefore, it can be concluded that and The split frequency between should be equal to or higher than 1.38e-4 Hz. In addition, according to the second part of the paper, this study made the following refinements: 1 It is a low-frequency impact signal; 2) is a periodic signal composed of square and triangle waves. However, Figure 8 shows that and The initial strength is the same. and The boundary of is set at 1.4e-3 Hz, which is about 10 times the acquisition frequency of 1.38e-4 Hz and is located between the first and second cycles. Similarly, by setting the upper cutoff frequency to 1.4e-3 Hz, the time series data can be decomposed into and Two parts, as shown in Figure 9. The time series data is divided into two different parts.
[0090] The linear combination process is as follows:
[0091] The time series decomposition model generally includes an additive model and a multiplicative model. The ATD modeling of the present invention adopts an additive model, and the theoretical formula is: ; Among them, after decomposition and Sequence as Fig.13 As shown; respectively and The training set and the test set are divided into two sets according to the ratio of 0.8. Predictive models and Prediction model, prediction results such as Fig.14 As shown; the above prediction result sequence and Direct addition, that is: ; Get the predicted value of surface settlement, such as Fig.15 shown.
[0092] The algorithm mentioned in Section 3 uses a delayed one-dimensional input structure for model training. During this process, validation is performed every 6400 iterations.
[0093] The most commonly used metrics for evaluating predictive model performance include the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE).
[0094] R² values range from 0 to 1, with values closer to 1 indicating a better model fit.
[0095] ; RMSE is a common indicator to measure the difference between the observed value and the true value. The lower the RMSE value, the smaller the deviation between the observed value and the true value.
[0096] ; MAE and MAPE calculate the average and percentage of the absolute values of the differences between the predicted values and the actual values, respectively. The smaller these values, the better.
[0097] ; ; Table 2 shows the four evaluation indicators of the surface settlement prediction value of each algorithm, including the original algorithm and the improved algorithm. As shown in Table 2, the average R² of each original algorithm is 0.9938, the root mean square error is 0.2162 mm, the average error is 0.1449 mm, and the average error percentage is 8.536%. Table 2 Algorithm accuracy evaluation indicators detailed table and Figure 10 show that, except for the convolutional neural network (CNN), the ATD model improves the R² of all algorithms by an average of 32.62%, and the RMSE, MAE and MAPE by an average of 19.91%, 23.37% and 16.44% respectively. Among the algorithms, the improvements of back propagation (BP) and bidirectional gated recurrent unit (BI-GRU) are the most significant, with R², RMSE, MAE and MAPE increased by 68.83%, 44.30%, 44.27% and 17.38% respectively.
[0098] Table 2 Detailed table of algorithm accuracy evaluation indicators
[0099]
[0100] To calculate the evaluation metrics, use y i The value of is used as the denominator to ensure that any small probability prediction error can be averaged by the evaluation index. In addition, when the error is significant and the actual value is close to zero, the calculation of the indicator will be abnormal, resulting in distortion. Therefore, in Figure 11, the fit between the predicted value and the measured value, the horizontal and vertical axes represent the true value and the predicted value, respectively, for further error analysis. Figure 11, the fit between the predicted value and the measured value, shows the parabolic relationship between the surface subsidence values predicted by the original algorithm and the measured values, indicating that there is an unidentified correlation between the two data sets. In particular, the parabolic features represent the main structure of the predicted data in the BP, CNN and particle swarm optimization (PSO) models.
[0101] This embodiment believes that the parabolic features represent nonlinear hidden data features of surface subsidence, which are difficult to identify through BP, CNN and PSO optimization. In response to this challenge, the ATD model significantly enhances the learning ability of parabolic features. The algorithms mentioned in this study all showed enhanced model performance and higher prediction accuracy when used in conjunction with the ATD model. It is worth noting that the average accuracy improvement of machine learning (ML) is greater than that of deep learning (DL) algorithms. This is because the ATD model divides the surface subsidence data into two subsequences with more obvious linear relationships. ML algorithms are better at extracting hidden patterns and structures of data from linear relationships, and therefore show better prediction accuracy. DL is able to capture nonlinearity and long-term dependencies, which may cause the prediction models of some algorithms (such as CNN) to capture more invalid features, thereby reducing prediction accuracy.
[0102] Figure 12 Prediction error distribution plot shows the dispersion of prediction errors of each algorithm relative to the measured values. The consistency of the prediction performance improvement of each algorithm by ATD was verified by analyzing the distribution of prediction errors in the test set (except for CNN). In order to quantitatively analyze the consistency of predictions, a settlement cutoff of ±3 mm was set. For the results of extreme learning machine (ELM), support vector machine (SVM), BI-GRU and BI-long short-term memory network (LSTM), the number of residual errors exceeding the threshold was the least, and there were no cases that exceeded or fell below the threshold. Overall, the residual error levels were similar across the entire range of actual settlement, demonstrating the consistency of predictions of the ELM, SVM, BI-GRU and BI-LSTM models. In contrast, there were 272 residual errors exceeding the threshold for the results of random forest (RF), 2 for CNN, 19 for BI-GRU and 32 for BI-LSTM. After applying ATD, the number of residual errors exceeding the critical value for BP, BI-GRU and BI-LSTM decreased to 0, 2 and 6, respectively. In contrast, it increased to 5 and 42 for RF and CNN, respectively.
[0103] Table 3 shows the statistical analysis of the prediction errors, indicating that the prediction values of all algorithms are smaller than the actual measured values and more accurate. After applying ATD, the average value of the prediction error decreased from -0.074 to -0.052, and the average error value decreased by 50.19%. The average length of the 95% confidence interval is 0.0030, while the average length of ATD is -0.0024. In contrast, ATD can reduce the 95% confidence interval by 20%.
[0104] In summary, the ATD model combined with unsupervised learning significantly improves the performance of the algorithms proposed in this study, including SVM, RF, ELM, BP, CNN, PSO optimized GRU and LSTM. The ATD model not only improves the prediction accuracy, but also reduces the confidence interval of the prediction error and the number of errors exceeding the critical value.
[0105] Table 3 Algorithm statistical error analysis table
[0106]
[0107] In summary, this embodiment shows that the method for predicting surface settlement caused by tunnel construction based on the ATD model has good accuracy. Compared with the original algorithm, the algorithm performance is improved and the algorithm error is greatly reduced.
[0108] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for predicting surface settlement caused by tunnel construction based on the ATD model, characterized by: The specific steps include: S1. Data collection and preprocessing: Collect and process the original surface subsidence data in the monitoring area and construct the original data set; S2. Constructing an ATD model based on the ARMA model: decomposing the surface settlement time series into two sub-features to enhance the learning ability of machine learning and deep learning algorithms for nonlinear surface settlement data; S3, using unsupervised learning to build corresponding networks to identify hidden patterns and internal structures; S4, using machine learning algorithms and deep learning algorithms to independently predict sub-feature sequences on the network constructed in S3; S5, combining the prediction results of the sub-feature sequences obtained in S4 in a linear manner to predict surface subsidence; The construction of the ATD model in step S2 includes the following steps: S21, Stationarity verification: Calculate the autocorrelation and partial correlation coefficients to ensure that the time series meets the smoothness requirements of the ARMA model; S22. Decompose the time series into different components based on frequency and intensity: and ; S23, randomly dividing the above two parts into a training set and a test set according to their proportions for model training; S24, verifying the model performance; Based on this, the ARMA model can be applied to decompose the surface subsidence time series into two parts: 1) is the historical value of the low-frequency settlement time series caused by a series of construction; 2) is the white noise value of the sedimentation time series caused by a series of random factors, so the ARMA model expression can be rewritten as: ; In order to simplify the complexity of the model, p = q is proposed. In addition, based on the assumption of the MA model, it is recommended that the average sum of the white noise terms be set to 0, that is: ; If there are n samples in total, then the number of samples from t=0 to t=n is and , which can form two time series respectively and : ; in, is a strongly random time series with periodicity and mean c, is a less periodic and more linear time series.
2. The method for predicting surface subsidence caused by tunnel construction based on the ATD model according to claim 1 is characterized by: In step S1, the surface settlement data is monitored in real time using the monitoring points formed by the arrangement of multiple monitoring robots, and an original data set is constructed based on this.
3. The method for predicting surface subsidence caused by tunnel construction based on the ATD model according to claim 1 is characterized by: The process of processing the data set in step S1 includes: S11, data resampling and outlier processing; S12, data resampling: by using MATLAB resampling function to obtain effective resampling points from the original data set to remove outliers; S13. Outlier processing: Outliers in the data are processed by using the filloutlier function and medfilt function in MATLAB; thus obtaining the corresponding data set for unsupervised learning.
4. The method for predicting surface subsidence caused by tunnel construction based on the ATD model according to claim 3 is characterized by: In step S23, two one-dimensional networks are used to train and Used to reduce the amount of calculation.
5. The method for predicting surface subsidence caused by tunnel construction based on the ATD model according to claim 4 is characterized by: In step S24, the model performance is verified using the determination coefficient R², root mean square error RMSE, mean absolute error MAE and mean absolute percentage error MAPE.
6. The method for predicting surface settlement caused by tunnel construction based on the ATD model according to claim 1, characterized in that: In step S4, the machine learning algorithms include support vector machine SVM, random forest RF, extreme learning machine ELM and back propagation BP; Deep learning algorithms include convolutional neural networks (CNN), long short-term memory networks (LSTM), and gated recurrent units (GRU); Particle swarm optimization PSO and Bayesian optimization BI are used as hyperparameter optimizers for LSTM and GRU respectively; four algorithms are implemented: BI-LSTM, BI-GRU, PSO-LSTM and PSO-GRU, and the optimized hyperparameters are given.
7. A system for predicting surface settlement caused by tunnel construction based on an ATD model, executing the steps of the method according to any one of claims 1 to 6, characterized in that: It includes a data acquisition robot, a data processing module, a model calculation module and an output module; The data collection robot is used to collect surface subsidence data in a specified area and construct an original data set; The data processing module is used to resample the original data and process outliers to improve the quality of the data set and the prediction accuracy; the processed data set is used as input, and the prediction result is obtained after calculation by the model calculation module; The prediction results are output through the output module.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for predicting surface subsidence caused by tunnel construction based on the ATD model as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting surface subsidence caused by tunnel construction based on the ATD model as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
EMD-SVR-based ground surface settlement amount prediction method
CN107092744A
CNN-LSTM model-based method for predicting and early warning land subsidence of region along railway
CN113886917A