Method for selecting redundant observations of tunnel plane network based on simulation and LOOCV verification
By using simulation and LOOCV verification, an accuracy fitting model for redundant observations and total observations was constructed, which solved the problem of lack of criteria for selecting redundant observations in tunnel construction surveying, and achieved optimization of tunnel construction surveying accuracy and reduction of engineering risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA RAILWAY FIRST GROUP CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
In tunnel construction surveying, the lack of a unified standard for selecting redundant observations leads to inflated accuracy in the calculation of lateral breakthrough error in tunnel construction surveying accuracy assessment, or failure to meet the standard requirements, which affects the quality of tunnel engineering and construction safety.
A simulation-based and LOOCV-validated approach was adopted. The minimum number of simulations was determined through confidence interval analysis, and a precision fitting model of redundant observations and total observations was constructed. The stability and reliability of the model were verified using the leave-one-out cross-validation (LOOCV) method, and the selection of redundant observations was optimized.
It provides a scientific basis for selecting redundant observations in tunnel construction surveying, improves the accuracy of tunnel construction surveying accuracy assessment and engineering construction efficiency, reduces engineering risks, and is applicable to the optimization of construction surveying schemes for tunnels of different lengths.
Smart Images

Figure CN122365877A_ABST
Abstract
Description
Technical Field
[0001] This invention pertains to tunnel construction surveying technology, and particularly relates to a method for selecting redundant observations in a tunnel plane network based on simulation and LOOCV verification. Background Technology
[0002] With the implementation of the national strategy to build a strong transportation nation in the new era and the continuous development of railway tunnel engineering technology, tunnel construction surveying has become a core technical link in tunnel construction, and its accuracy has become one of the core indicators for evaluating the quality and safety of tunnel engineering. In the current railway tunnel construction surveying process, limited by the harsh surveying environment and extremely poor observation conditions inside the tunnel, high redundancy and high reliability design schemes are often pursued to ensure the accuracy of construction surveying. However, there is no unified standard for determining redundancy. In the accuracy evaluation of tunnel construction surveying, if the redundancy ratio of redundant observations to the total number of observations is too high, it will lead to an overestimation of the accuracy of the lateral breakthrough error, causing the estimated lateral breakthrough error value to deviate from actual application; if the redundancy ratio of redundant observations to the total number of observations is too low, it will cause the estimated lateral breakthrough error value to fail to meet the accuracy requirements in the specifications. International research also faces the same problem in the selection of redundant observations. Traditional tunnel construction surveying usually uses optimized redundant observations to control tunnel breakthrough error. During the construction of immersed tunnels in the ocean, the reliability of the observation scheme can be evaluated by calculating the average redundant observation components. Based on this, the layout of measuring points and the frequency of observations can be optimized, effectively reducing positioning errors during the lowering of the immersed tube and increasing the number of redundant observations. Then, by analyzing the number of redundant observations, the optimal observation path can be determined to minimize the impact of the marine environment on the measurement results. In general railway tunnels, to explore the impact of different observation schemes on the accuracy of the control network, a method combining free-station measurements with gyro-oriented measurements is often used to increase the number of redundant observations, thereby improving the uncertainty of the control network and mitigating the problem of decreased measurement accuracy in long tunnels. In summary, it is necessary to use virtual simulation measurement technology, based on the cross-traverse control network within the tunnel, to deeply explore the profound relationship between tunnel construction measurement accuracy and the number of redundant observations. The aim is to provide a new method for evaluating the accuracy of tunnel construction measurements and to establish a universal principle for selecting redundant observations. Summary of the Invention
[0003] To address the above problems, this invention provides a method for selecting redundant observations in a tunnel plane network based on simulation and LOOCV verification.
[0004] This invention discloses a method for selecting redundant observations in a tunnel plane network based on simulation and LOOCV verification. The method calculates the minimum number of simulations required to meet the target accuracy requirements using confidence interval analysis, constructs an accuracy fitting model based on the redundant observations and the total number of observations, and verifies the model's stability and reliability using leave-one-out cross-validation (LOOCV). Specifically:
[0005] Determining the number of simulations.
[0006] Using simulation measurement technology, a simulation experiment was designed to analyze the lateral breakthrough error inside a railway tunnel, and the probability that the lateral breakthrough error is distributed at twice the lateral breakthrough error was obtained.
[0007] The mean probability of the lateral penetration error distribution obtained from statistics is used as the theoretical probability. The half-width of the confidence interval is calculated using the formula for calculating the half-width of the confidence interval. The half-width of the confidence interval is then added to the theoretical probability to obtain the probability interval of the lateral penetration error distribution obtained from 1000 simulations at a 95% confidence level of the normal distribution. The specific calculation formula is as follows:
[0008] (1)
[0009] in, The confidence interval is half-width. For probability distribution, This is the value corresponding to the 95% confidence level. The number of simulations.
[0010] Using the maximum difference of 1.84% between the actual simulated probability and the theoretical probability as the maximum allowable error range, and employing the effect size-sample size inverse formula, the minimum number of simulations required to achieve the target accuracy is calculated using the actual simulated probability values. The specific calculation formula is as follows:
[0011] (2)
[0012] in, To minimize the number of simulations, This is the value corresponding to the 95% confidence level. For probability distribution, This represents the acceptable error range.
[0013] Establishment of the fitting model.
[0014] The second-order measurement accuracy for tunnels was uniformly adopted, and N simulation experiments were conducted on tunnels of different lengths. The experimental data were statistically analyzed based on the theory of average redundant observation components.
[0015] Based on the obtained statistical data, the average redundant observation component value is set as The original average redundant observation component value is set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. At the same time, the excess observations are set as The total number of redundant observations is set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. The specific calculation formula is as follows:
[0016] (3)
[0017] in, , denoted as undetermined coefficients of a first-order polynomial, and n represents the number of tunnels of different lengths.
[0018] (4)
[0019] in, For all Sum of mean For each point, based on the fitted line Predicted value, SSE is The sum of squares of the deviations between the actual values and the average, SST is... The sum of squares of the deviations between the actual values and the predicted values of the regression model.
[0020] Coefficient of determination It is an important indicator in statistics used to evaluate the goodness of fit of regression models, measuring the proportion of the variation of the dependent variable that can be explained by the independent variable.
[0021] Using the above formula, the first-order polynomial obtained by data fitting based on the average redundant observation component value and the original average redundant observation component value is: =1.6993 -0.5417, =89.43%; Based on the redundant observations and the total number of observations, the first-order polynomial is obtained by fitting: =0.3115 -8.2195, =99.96%.
[0022] Therefore, the fitted model obtained using the redundant observations and the total number of observations was chosen.
[0023] Validation of the generalization of the fitted model.
[0024] The fitting model built based on the redundant observations and the total number of observations consists of 18 fitting points. These fitting points together form a small sample dataset, which is suitable for leave-one-out cross-validation (LOOCV) method.
[0025] During the validation process, each fitted point is used as the validation set, and the remaining 17 points are used as the training set to validate the fitted model. Specifically, the 17 points in the validation set are used as the fitted points to perform least squares first-order polynomial fitting, and the two core parameters of the fitted formula are obtained statistically and the coefficient of determination is calculated. The prediction error of the validation points in the validation set is calculated. The two core parameters are the slope and the intercept.
[0026] After 18 leave-one-out cross-validations, the mean values of the slope and intercept are obtained and substituted into the fitted equation to evaluate the stability of the fitted result.
[0027] To further evaluate the model's generalization ability, its generalization gap needs to be examined. This involves calculating the ratio of the average validation error to the average training error, which represents the consistency of the model's performance on the training set and the unseen validation set. The calculation formula is as follows:
[0028] (5)
[0029] in, To average the verification error, The average training error, To verify the number of times, For the number of training sessions.
[0030] The average validation error refers to the average of the "single-sample validation error" generated in each validation step during the entire cross-validation process. Its core is to quantify the model's prediction bias on "single samples not involved in training" by iteratively excluding individual samples as the validation set and using the remaining samples as the training set. The average value ultimately reflects the overall validation performance of the model. The average training error refers to the statistical average of the "single training error" of all training samples when the model makes predictions on the "training set". Its core is used to quantify the model's fit to the "trained data" and to serve as a key benchmark for subsequent comparison of generalization.
[0031] Evaluate the model's generalization and overfitting risk.
[0032] The beneficial technical effects of this invention compared to the prior art are as follows:
[0033] This invention uses confidence interval analysis to calculate the minimum number of simulations required to meet the target accuracy requirements, and explores the influence of measurement level on lateral breakthrough error. Based on this, an accuracy fitting model based on redundant observations and total observations is constructed. Leave-one-out cross-validation (LOOCV) is used to verify the model's stability and reliability. The results show that the model possesses excellent fitting accuracy and strong generalization ability, providing quantitative support for redundancy optimization. To further verify the model's applicability in engineering practice, empirical verification is conducted based on actual tunnel construction survey data. Statistical analysis results show that the tunnel lateral breakthrough error calculated using this fitting model better matches the probability distribution of twice the tunnel's lateral mean square error. This invention not only provides a solid theoretical basis for the scientific selection of redundant observations in tunnel construction surveying, but also innovatively proposes a new approach to tunnel construction survey accuracy assessment. It has significant practical implications for optimizing tunnel construction surveying schemes, reducing engineering risks, and improving construction efficiency, and can provide a reference for subsequent similar projects. Attached Figure Description
[0034] Figure 1 The fitting formula and scatter plot are established using the average redundant observation components.
[0035] Figure 2 The fitted formula and scatter plot were created using the extra observations. Detailed Implementation
[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0037] With the continuous advancement of railway tunnel construction technology, tunnel construction surveying, as a core link in ensuring project quality and construction safety, is becoming increasingly critical in terms of accuracy control within complex tunnel environments. Current engineering designs often ensure measurement accuracy by increasing observation redundancy; however, the lack of unified standards for redundancy selection leads to problems such as unreasonable resource allocation and imbalances in accuracy control.
[0038] This invention discloses a method for selecting redundant observations in a tunnel plane network based on simulation and LOOCV verification. The method calculates the minimum number of simulations required to meet the target accuracy requirements using confidence interval analysis, constructs an accuracy fitting model based on the redundant observations and the total number of observations, and verifies the model's stability and reliability using leave-one-out cross-validation (LOOCV). Specifically:
[0039] Determining the number of simulations.
[0040] In the current calculation of the accuracy of tunnel construction measurement, there is a problem that the calculation accuracy of the transverse connection error is too high, which leads to the transverse connection error not meeting the error distribution law. In order to solve the above problem, through simulation measurement technology, 5000 simulation experiments were designed to analyze the transverse connection error of a 10km long railway tunnel. The probability of the transverse connection error distribution being twice the transverse connection error is shown in Table 1.
[0041] Table 1. Probability that the lateral breakthrough error of a 10km long tunnel is within twice the lateral breakthrough mean error
[0042]
[0043] As shown in Table 1, the mean probability of the lateral penetration error distribution obtained from statistics is used as the theoretical probability. The half-width of the confidence interval is calculated using the formula for calculating the half-width of the confidence interval. Then, the theoretical probability is added to obtain the probability interval of the lateral penetration error distribution obtained from 1000 simulations at a 95% confidence level of the normal distribution. The specific calculation formula is as follows:
[0044] (1)
[0045] in, The confidence interval is half-width. For probability distribution, This is the value corresponding to the 95% confidence level (1.96). The number of simulations.
[0046] The confidence interval half-width of the probability distribution of the lateral penetration error after 1000 simulations is 2.08%, and the confidence interval is [84.98%, 89.14%]. The probability interval obtained from the actual simulation is [85.60%, 88.90%]. This shows that the probability distribution of the lateral penetration error obtained from the above simulation experiment meets the distribution requirements of the confidence interval.
[0047] To further confirm the number of simulations required for subsequent experiments, the maximum allowable error range was set at 1.84%, the maximum difference between the actual simulated probability and the theoretical probability. Using the effect size-sample size inverse formula, the minimum number of simulations required to achieve the target accuracy was calculated based on the actual simulated probability values. The specific calculation formula is as follows:
[0048] (2)
[0049] in, To minimize the number of simulations, This is the value corresponding to the 95% confidence level (1.96). For probability distribution, This represents the acceptable error range.
[0050] Calculated using the formula, with an allowable error range of 1.84%, the minimum number of simulations required to achieve the target accuracy is calculated to be 1279 using the effect size-sample size inverse formula.
[0051] The impact of measurement levels in simulation;
[0052] The "Specifications for Surveying and Mapping in High-Speed Railway Engineering" stipulates that tunnels with a length of 3-6km should use the third-class horizontal surveying level, those with a length of 6-9km should use the second-class horizontal surveying level, and those with a length of over 9km should also use the second-class horizontal surveying level. To facilitate standardized simulation experiments and data analysis, it is necessary to investigate the impact of different surveying levels on the probability distribution of lateral tunnel breakthrough errors.
[0053] To ensure the reliability of the experiment, the cross-traverse control network was used as the unified measurement network type, and three typical tunnels of different lengths were selected: 4km tunnel (matching the third-order plane measurement level), 8km tunnel (matching the second-order tunnel measurement level), and 12km tunnel (matching the second-order plane measurement level). For each type of tunnel, simulation measurement experiments were carried out using the above three measurement levels. Each group of experiments was repeated 1279 times, and the probability distribution of lateral breakthrough error of tunnels of different lengths under different measurement levels was statistically analyzed based on the experimental results. The results are shown in Table 2.
[0054] Table 2. Probability distribution of lateral breakthrough error for each tunnel under different measurement levels
[0055]
[0056] Table 2 shows that the differences in the lateral breakthrough error distribution probabilities obtained from the three different tunnel lengths under different measurement levels are not significant. To determine whether the differences between the data are significant, we can calculate whether the confidence interval of the data difference includes zero. Taking the simulation experimental data of an 8km tunnel as an example, the specific steps are as follows:
[0057] (1) Calculate the differences between the three sets of data.
[0058] The experimental data in Table 2 show that the differences between the 8km tunnels are 0.16%, 0.23%, and 0.07%, respectively.
[0059] (2) Calculate the standard error
[0060] The formula for calculating the standard error is as follows:
[0061] ;
[0062] in, For standard error, , For probability distribution, The number of simulations.
[0063] Based on the probability distribution of lateral breakthrough error of an 8km-long tunnel under different measurement levels, three sets of standard error values were calculated using the standard error formula. , , The values are 0.01283, 0.01281, and 0.01273, respectively.
[0064] (3) Determine the critical value of the confidence interval
[0065] The value corresponding to the confidence level based on the 95% normal distribution is used as the critical value of the confidence interval, i.e. =1.96.
[0066] (4) Calculate the confidence interval for the difference data
[0067] The calculation formula is as follows:
[0068] ;
[0069] The calculated confidence intervals for the differences in each data point are [-2.36%, 2.67%], [-2.28%, 2.74%], and [-2.42%, 2.56%].
[0070] Following the same method described above, the simulation experimental data of the 4km tunnel and the 12km tunnel were calculated and analyzed respectively. The confidence intervals for the differences in data for the 4km tunnel were [-2.24%, 2.70%], [-2.43%, 2.48%], and [-1.99%, 2.93%]; and the confidence intervals for the differences in data for the 12km tunnel were [-2.16%, 2.94%], [-2.34%, 2.80%], and [-1.94%, 3.18%].
[0071] Since each confidence interval contains a zero point, the differences between the data groups can be considered statistically insignificant; that is, when conducting simulation experiments on different lengths under different measurement levels, the measurement level can be regarded as having no significant impact on the probability distribution of the transverse penetration error.
[0072] Establishment of the fitting model.
[0073] As shown in the above calculations, the measurement level has no significant impact on the probability distribution of lateral breakthrough error in the simulation experiment. To address the issue of how to select redundant observations when the lateral breakthrough error of tunnels of different lengths meets the error distribution law, a uniform second-order measurement accuracy for tunnels was adopted. 1279 simulation experiments were conducted on the following tunnels of different lengths, and the experimental data were statistically analyzed based on the theory of average redundant observation components. The results are shown in Table 3.
[0074] Table 3. Redundant observation conditions for tunnels of different lengths that satisfy the distribution law of transverse breakthrough error.
[0075]
[0076] Based on the statistical data obtained in Table 3, the average redundant observation component value is set as The original average redundant observation component value was set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. At the same time, the excess observations are set as The total number of redundant observations is set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. The specific calculation formula is as follows:
[0077] (3)
[0078] in, , denoted as undetermined coefficients of a first-order polynomial, and n represents the number of tunnels of different lengths.
[0079] (4)
[0080] in, For all Sum of mean For each point, based on the fitted line Predicted value, SSE is The sum of squares of the deviations between the actual values and the average, SST is... The sum of squares of the deviations between the actual values and the predicted values of the regression model.
[0081] Coefficient of determination It is an important indicator in statistics used to evaluate the goodness of fit of regression models, measuring the proportion of the variation of the dependent variable that can be explained by the independent variable.
[0082] Using the above formula, the first-order polynomial obtained by data fitting based on the average redundant observation component value and the original average redundant observation component value is: =1.6993 -0.5417, =89.43%; Based on the redundant observations and the total number of observations, the first-order polynomial is obtained by fitting: =0.3115 -8.2195, =99.96%. The magnitude of the coefficient of determination shows that in the former fit, only 89.43% of the variance in the dependent variable can be explained by the independent variable, far less than the latter's 99.96%. Therefore, the latter has a better fit. To further analyze the differences between the two fits, the fit points and fit curves of both are plotted, as shown below. Figure 1 , Figure 2 As shown. Figure 1 As shown, in the fitting formula obtained based on the average redundant observation component values and the original average redundant observation component values, the fitting points are randomly distributed and most of them deviate significantly from the fitted line, resulting in poor fitting performance. Figure 2 As shown, in the fitting formula obtained based on the redundant observations and the total number of observations, the fitting points are evenly distributed and most of them are close to both sides of the fitting line, indicating a good fitting effect.
[0083] Therefore, the fitting model obtained by using the redundant observations and the total number of observations has a better fitting effect. The generalization ability will be verified in the next step, and the applicability of the model will be verified using actual tunnel data.
[0084] Validation of the generalization of the fitted model.
[0085] Leave-One-Out Cross-Validation (LOOCV) was originally a method for evaluating the performance of machine learning models, particularly suitable for small datasets (≤100 samples). In this method, with N data points, the model undergoes N training and validation cycles, leaving one data point as the validation set each time, and using the remaining N-1 data points as the training set. Each data point has the opportunity to be tested once as a validation set, resulting in N test results for the model. These N results are then averaged to obtain the final model performance. Furthermore, LOCV can validate some of the model's hyperparameters, obtaining robust performance evaluation results across the entire dataset without the need for additional validation set partitioning. This characteristic gives LOCV a significant advantage in small-sample scenarios, making full use of limited data resources while providing a relatively reliable estimate of the model's generalization ability.
[0086] The fitting model built based on the redundant observations and the total number of observations consists of 18 fitting points. These fitting points together form a small sample dataset, which is suitable for leave-one-out cross-validation (LOOCV) method.
[0087] During the validation process, each fitted point was used as the validation set and the remaining 17 points were used as the training set to validate the fitted model. Specifically, the 17 points in the validation set were used as the fitted points to perform least squares first-order polynomial fitting. The two core parameters of the fitted formula (slope and intercept) were obtained statistically, and the coefficient of determination was calculated. The prediction error of the validation points in the validation set was calculated, and the results are shown in Table 4.
[0088] Table 4. Calculation results of leave-one-out cross-validation method
[0089]
[0090] As shown in Table 4, after 18 leave-one-out cross-validations, the mean values of the slope and intercept are substituted into the fitted equation. =0.3115, -8.2086, =99.96%, compared to the fitting formula obtained by directly fitting 18 fitting points. =0.3115 -8.2195, =99.96%, the two core parameters are basically consistent, and the coefficients of determination are completely equal. This result indicates that the model fits stably, has low sensitivity to different subsets of data, and performs well in cross-validation evaluation.
[0091] To further evaluate the model's generalization ability, its generalization gap needs to be examined. This involves calculating the ratio of the average validation error to the average training error, which represents the consistency of the model's performance on the training set and the unseen validation set. The calculation formula is as follows:
[0092] (5)
[0093] in, To average the verification error, The average training error, To verify the number of times, For the number of training sessions.
[0094] The average validation error refers to the average of the "single-sample validation error" generated in each validation step during the entire cross-validation process. Its core is to quantify the model's prediction bias on "single samples not involved in training" by iteratively excluding individual samples as the validation set and using the remaining samples as the training set. The average value ultimately reflects the overall validation performance of the model. The average training error refers to the statistical average of the "single training error" of all training samples when the model makes predictions on the "training set". Its core is used to quantify the model's fit to the "trained data" and to serve as a key benchmark for subsequent comparison of generalization. The results calculated by the above formula are shown in Table 5.
[0095] Table 5 Calculation results of validation error and training error
[0096]
[0097] Table 5 shows that the calculated average validation error is 2.1197, the average training error is 1.7731, and the ratio of average validation error to average training error is 1.1196. Based on practical experience in machine learning, a ratio close to 1 indicates a small generalization gap and no overfitting. Based on the above validation process, it can be considered that the ratio of average validation error to average training error meets practical standards, the model has good generalization ability, there is no risk of overfitting, and it can be directly applied.
[0098] Applicability verification of the fitted model.
[0099] To clarify the applicability of the fitting model obtained by fitting the redundant observations to the total number of observations in the actual tunnel construction measurement accuracy evaluation, various observation data obtained during the actual construction measurement of three different tunnels were selected for statistical analysis and calculation. The calculation results are shown in Table 6.
[0100] Table 6 Statistical Results and Calculations of Measured Tunnel Data
[0101]
[0102] As shown in Table 6, the lateral connection errors of the three tunnels all meet the requirements of the "High-Speed Railway Engineering Surveying Specification". However, the lateral connection error value does not meet the requirement of being within twice the lateral connection error value. To solve this problem, the previously verified fitting model was used to update the selected redundant observations and the calculation was re-performed. The specific results are shown in Table 7.
[0103] Table 7 Results of recalculation using the fitted model based on measured tunnel data
[0104]
[0105] As shown in Table 7, the new lateral breakthrough errors are 72.33 mm, 41.15 mm, and 29.02 mm, respectively. At this point, the lateral breakthrough errors of all three tunnels are within twice the lateral breakthrough error, satisfying the condition based on a normal distribution level of 2. in principle.
[0106] The above calculation results confirm that the optimal fitting model obtained from tunnel simulation measurement experimental data not only has excellent fitting effect, but also has good applicability in actual tunnel engineering and can be applied to the accuracy evaluation of actual tunnel construction measurement.
[0107] This invention addresses the problem of selecting redundant observations when the lateral breakthrough error of tunnels of different lengths meets a distributional pattern. It employs simulation experiments on tunnels of varying lengths using second-order measurement accuracy. By comparing the performance of two fitting models—one based on the average redundant observation component and the other on the original average redundant observation component, and the other based on the number of redundant observations and the total number of redundant observations—it was found that the linear fitting model based on the number of redundant observations and the total number of redundant observations achieved a determination coefficient of 99.96%, significantly better than the former's 89.43%. Furthermore, its scatter points were closely distributed on both sides of the fitted line, indicating that this model has a superior fitting effect. Therefore, the fitting relationship between the number of redundant observations and the total number of redundant observations was selected as the basis for determining the required number of redundant observations for different tunnel lengths, and its validity was verified.
[0108] To verify the generalization ability and applicability of the fitted model established based on redundant observations and total observations, this invention employs leave-one-out cross-validation for evaluation. The validation results show that the model's core parameters (slope and intercept) are stable, the coefficient of determination remains at 99.96%, and the ratio of average validation error to average training error is close to 1 (1.1196), indicating good generalization ability and no risk of overfitting.
[0109] This invention applies the fitted optimization model to the accuracy evaluation of construction surveys for three actual tunnels. The results show that when using the original redundant observations, the probability distribution of the tunnel lateral breakthrough error does not satisfy the normal distribution level of 2. Based on this principle, the tunnel lateral breakthrough error value calculated using this fitting model better matches the distribution pattern within twice the lateral breakthrough mean error. Verification shows that this fitting model has good applicability and reliability in the accuracy assessment of actual tunnel construction measurements.
Claims
1. A method for selecting redundant observations in a tunnel plane network based on simulation and LOOCV verification, characterized in that, The minimum number of simulations required to meet the target accuracy requirements was calculated using confidence interval analysis. An accuracy fitting model based on redundant observations and the total number of observations was constructed. The stability and reliability of the model were verified using leave-one-out cross-validation (LOOCV). Determining the number of simulations; Using simulation measurement technology, a simulation experiment was designed to analyze the lateral breakthrough error inside a railway tunnel, and the probability that the lateral breakthrough error is distributed at twice the lateral breakthrough error was obtained. The mean probability of the lateral penetration error distribution obtained from statistics is used as the theoretical probability. The half-width of the confidence interval is calculated using the formula for calculating the half-width of the confidence interval. The half-width of the confidence interval is then added to the theoretical probability to obtain the probability interval of the lateral penetration error distribution obtained from 1000 simulations at a 95% confidence level of the normal distribution. The specific calculation formula is as follows: (1) in, The confidence interval is half-width. For probability distribution, This is the value corresponding to the 95% confidence level. The number of simulations; Using the maximum difference of 1.84% between the actual simulated probability and the theoretical probability as the maximum allowable error range, and employing the effect size-sample size inverse formula, the minimum number of simulations required to achieve the target accuracy is calculated using the actual simulated probability values. The specific calculation formula is as follows: (2) in, To minimize the number of simulations, This is the value corresponding to the 95% confidence level. For probability distribution, This represents the acceptable error range; Establishing a fitting model; The second-order measurement accuracy for tunnels was uniformly adopted, and N simulation experiments were conducted on tunnels of different lengths. The experimental data were statistically analyzed based on the theory of average redundant observation components. Based on the obtained statistical data, the average redundant observation component value is set as The original average redundant observation component value was set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. At the same time, the excess observations are set as The total number of redundant observations is set as Calculate the least squares first-order fitting polynomial and its coefficient of determination. The specific calculation formula is as follows: (3) in, , represents the undetermined coefficients of a first-order polynomial, and n represents the number of tunnels of different lengths. (4) in, For all Sum of mean For each point, based on the fitted line Predicted value, SSE is The sum of squares of the deviations between the actual values and the average, SST is... The sum of squares of the deviations between the actual values and the predicted values of the regression model; Coefficient of determination It is an important indicator in statistics used to evaluate the goodness of fit of regression models, measuring the proportion of the variation of the dependent variable that can be explained by the independent variable. Using the above formula, the first-order polynomial obtained by data fitting based on the average redundant observation component value and the original average redundant observation component value is: =1.6993 -0.5417, =89.43%; Based on the redundant observations and the total number of observations, the first-order polynomial is obtained by fitting: =0.3115 -8.2195, =99.96%; Therefore, the fitted model obtained by using the redundant observations and the total number of observations is chosen; Validation of the generalization ability of the fitted model; The fitting model built based on the redundant observations and the total number of observations consists of 18 fitting points. These fitting points together form a small sample dataset, which is suitable for leave-one-out cross-validation (LOOCV) method. During the validation process, each fitted point is used as the validation set, and the remaining 17 points are used as the training set to validate the fitted model. Specifically, the 17 points in the validation set are used as the fitted points to perform least squares first-order polynomial fitting, and the two core parameters of the fitted formula are obtained and the coefficient of determination is calculated. The prediction error of the validation points in the validation set is calculated. The two core parameters are the slope and the intercept. After 18 leave-one-out cross-validations, the mean values of the slope and intercept are obtained and substituted into the fitted formula to evaluate the stability of the fitted results. To further evaluate the model's generalization ability, its generalization gap needs to be examined. This involves calculating the ratio of the average validation error to the average training error, which represents the consistency of the model's performance on the training set and the unseen validation set. The calculation formula is as follows: (5) in, To average the verification error, The average training error, To verify the number of times, For the number of training iterations; The average validation error refers to the average of the "single-sample validation error" generated in each validation step during the entire cross-validation process. Its core is to quantify the model's prediction bias on "single samples not involved in training" by iteratively excluding individual samples as the validation set and using the remaining samples as the training set. The average value ultimately reflects the overall validation performance of the model. The average training error refers to the statistical average of the "single training error" of all training samples when the model makes predictions on the "training set". Its core is used to quantify the model's fit to the "trained data" and is a key benchmark for subsequent comparison of generalization. Evaluate the model's generalization and overfitting risk.