Data-driven Automatic Posture Correction Method and System

By building a digital twin model and a data-driven prediction model, the path deviation problem of tunnel boring machines in complex environments is solved, real-time and accurate correction of tunnel boring machines is achieved, the accuracy and efficiency of tunnel boring are improved, and the stability and safety of tunnel boring are ensured.

CN119918160BActive Publication Date: 2025-07-04GUIZHOU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510411220.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing tunnel boring machine attitude control system is difficult to achieve real-time and accurate path correction under complex working conditions. Traditional methods rely on manual judgment and manual intervention and cannot meet the needs of high precision and efficiency.

Method used

A digital twin model of TBM and surrounding rock environment is constructed, and data is collected in real time through attitude and position sensors, combined with the GRU model and Q-learning algorithm, the thrust required by the oil cylinder is predicted, automatic deviation correction is achieved, and the tunnel boring process is optimized.

Benefits of technology

It improves the accuracy and efficiency of tunnel boring, reduces equipment losses, ensures the stability and safety of tunnel boring, and realizes high-precision automated control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918160B_ABST
    Figure CN119918160B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data-driven modeling and prediction, and discloses a data-driven automatic attitude correction method and system, including: by constructing a digital twin model of the TBM and the surrounding rock environment, it realizes the omni-directional and real-time monitoring of the tunnel boring process, provides high-precision basic data, and optimizes the visualization and prediction of the actual boring. Using attitude and position sensors to collect the actual position and operation data of the TBM in real time, and synchronizing these data to the digital twin model, ensuring a high degree of consistency between the model and the actual progress, and greatly improving the real-time and accuracy of the data. By comparing the actual trajectory with the designed trajectory, the real deviation of the trajectory is calculated; after each deviation correction, the required thrust is predicted in combination with the surrounding rock parameters and the cylinder state, ensuring the continuous optimization of the boring process. It improves the degree of automation, reduces equipment wear, and by dynamically adjusting the correction trajectory, it enhances the precision, efficiency and safety in the tunnel boring process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data-driven modeling and prediction, and in particular to a method and system for automatic posture correction based on data-driven. Background Art

[0002] With the continuous development of tunnel excavation technology, tunnel boring machines (TBMs) have been widely used in underground projects as an efficient tunnel construction equipment. TBMs have achieved large-scale and efficient tunnel excavation through mechanization and automation, greatly improving construction efficiency and reducing manual operations. However, in actual operations, tunnel boring machines face complex underground environments, such as geological changes, equipment wear, and the accuracy of construction paths. These factors often cause the actual path of the tunnel boring machine to deviate from the designed path. In order to ensure that the tunnel boring machine advances according to the predetermined design trajectory, precise posture control has become a core technology in tunnel construction. In recent years, with the continuous advancement of digitalization and automation technology, the automated control system based on sensor data and digital twin technology has gradually become the main research direction for the posture adjustment and correction of tunnel boring machines. By real-time monitoring of the operating status of the tunnel boring machine and combining it with a data-driven prediction model, the operating posture of the tunnel boring machine can be effectively predicted and adjusted to improve work accuracy.

[0003] Although the existing technology has made some progress in the attitude control and correction of tunnel boring machines, there are still some technical difficulties. Traditional attitude correction methods often rely on simple feedback control algorithms, which are difficult to adapt to complex underground environments and changeable construction conditions, resulting in low efficiency and accuracy of attitude adjustment. In addition, the existing correction system usually uses rule-based algorithms for control, ignoring the possibility of in-depth mining of real-time data and intelligent optimization. Based on these problems, the existing technology has great limitations in the real-time and accuracy of path deviation prediction, pressure application and adjustment, and cannot achieve dynamic adjustment and intelligent optimization. Existing methods often fail to fully utilize the advantages of sensor data and digital twin models, and it is difficult to ensure the efficient operation of tunnel boring machines when dealing with path deviations in complex environments. At the same time, although digital twin technology has been applied in many fields, there is still a lack of mature systems that combine real-time data and prediction models in the precise control of tunnel boring machines. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is that in the existing traditional attitude control system of tunnel boring machines, the path deviation cannot be corrected in real time and accurately. Especially under complex working conditions, manual adjustment and traditional deviation correction methods are difficult to meet the requirements of high precision and high efficiency. By constructing a digital twin model and a data-driven prediction model, automatic deviation correction of the tunnel boring machine attitude is realized, overcoming the deficiencies of relying on manual judgment and manual intervention in traditional methods, improving the deviation correction accuracy and automation level, and ensuring the stability and safety of tunnel boring operations.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: A data-driven automatic attitude deviation correction method, including:

[0007] Construct a digital twin model of the TBM and the surrounding rock environment;

[0008] Real-time collect the actual position and real-time operation data of the TBM through attitude and position sensors, determine the actual tunneling trajectory of the TBM, and synchronize the actual position and real-time operation data of the TBM to the digital twin model;

[0009] Compare the actual trajectory with the designed trajectory to calculate the true trajectory deviation;

[0010] When the true trajectory deviation exceeds the deviation correction threshold, plan a deviation correction trajectory and calculate the expected deviation correction parameters of the actual trajectory and the deviation correction trajectory;

[0011] Taking the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders as input parameters, predict the thrust required for the four driving cylinders through a data-driven model in the digital twin model;

[0012] Apply the predicted thrust to the cylinders with corresponding numbers, carry out tunneling in the next cycle step, measure the actual position and attitude information of the TBM in the next cycle, determine the actual trajectory, and then calculate the deviation correction parameters for the next cycle;

[0013] Assemble the deviation correction parameters, surrounding rock parameters, and cylinder thrust into historical time series data for enhancing the training of the data-driven model.

[0014] As a preferred solution of the data-driven automatic attitude deviation correction method described in the present invention, wherein: the real-time operation data includes: attitude data, position data, trajectory deviation, cylinder pressure data, vibration acceleration data, propulsion force data, and surrounding rock parameters;

[0015] Remove the noise and outliers in the real-time operation data, and normalize the real-time operation data with different dimensions;

[0016] Synchronizing the attitude data to the digital twin model includes transmitting the real-time collected data from the TBM to the digital twin platform through MQTT;

[0017] Synchronizing the time of the real-time operation data through the GPS clock and adding a timestamp after each piece of real-time operation data;

[0018] When MQTT is delayed or interrupted, ensure data is not lost through a caching mechanism and synchronize when MQTT resumes;

[0019] When an exception occurs in the caching mechanism, perform data backup through redundant storage to ensure data integrity;

[0020] Synchronize the real-time operation data to the digital twin model in real time and update the status of the digital twin model.

[0021] As a preferred solution of the data-driven attitude automatic correction method described in the present invention, calculating the deviation amount between the correction trajectory and the actual trajectory includes discretizing the correction trajectory and the actual trajectory into data points according to the tunneling progress step and calculating the deviation based on the data points;

[0022]

[0023]

[0024] Wherein, represents the deviation between the actual trajectory and the designed trajectory, d represents the correction threshold, which is taken as 0.02 meters in the straight section and 0.005*R in the turning section, and R represents the radius of the turning section, represents the coordinate of the TBM on the X-axis at the actual position in the current time step, represents the X-axis coordinate of the designed trajectory in the current time step, represents the coordinate of the TBM on the Y-axis at the actual position in the current time step; represents the Y-axis coordinate of the designed trajectory in the current time step; represents the TBM at the actual position in the current time step on axis coordinate; represents the Z-axis coordinate of the designed trajectory in the current time step;

[0025] Assume that the tunneling progress step each time is 0.01 meters; the time step is 0.5 seconds;

[0026] The position information of the TBM needs to be updated every time a tunneling progress step is completed;

[0027] The designed trajectory is the discrete point coordinates of the ideal trajectory designed in the early project planning;

[0028] The formula for calculating the expected deviation parameter is expressed as:

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035] Among them, represents the component of the expected error in the X direction, and X represents the coordinate of the X-axis in three-dimensional space. represents the coordinate of the X-axis of the actual position of the TBM at the current time step. represents the coordinate of the X-axis of the current deviation correction trajectory. represents the path deviation of the TBM in the Y-axis direction. represents the coordinate of the actual position of the TBM at the current time step on the Y-axis; represents the coordinate of the Y-axis of the current deviation correction trajectory; represents the path deviation of the TBM in the Z-axis direction. represents the coordinate of the actual position of the TBM at the current time step on the axis. represents the coordinate of the Z-axis of the current deviation correction trajectory; represents the actual angle of the TBM and the expected angle between the deviations; represents the actual angle at the current time step and the expected angle between the deviations; represents the actual angle at the current time step and the expected angle between the deviations.

[0036] As a preferred solution of the data-driven attitude automatic deviation correction method described in the present invention, wherein: based on the GRU model, predict the pressure values that need to be applied to the four cylinders ; the expected deviation correction parameters include , , , , and ;

[0037] Integrate the input parameters at the current moment with the data of historical time steps to form a time series data matrix, and perform preprocessing, which is expressed by the formula:

[0038]

[0039] Select the data of the past time steps as the input sequence, and the historical time series data matrix is expressed by the formula:

[0040]

[0041] Among them, represents the surrounding rock parameter vector at the current moment , represents the cylinder state parameter at the current moment ;

[0042] The preprocessing includes using Kalman filtering to remove noise in the time series data, identifying and processing outliers in the time series data, ensuring that all time series data are synchronized in time, and normalizing the time series data with different dimensions;

[0043] The input time window size of the GRU model is the time series data of the past 10 time steps;

[0044] Divide the time series data into training set, validation set and test set in the ratio of 70%, 15% and 15% in chronological order;

[0045] The hyperparameters of the GRU model include the number of GRU units, time step, learning rate, batch size and number of training epochs;

[0046] The number of GRU units is 512 hidden units in the GRU layer;

[0047] The time step is 25 time steps, representing the length of the entire input sequence;

[0048] The learning rate is 0.0005;

[0049] The batch size is 64 samples used in one training;

[0050] The number of training epochs is 100 epochs;

[0051] Adjust the hyperparameters of the GRU model by grid search method;

[0052] The GRU model includes an input layer, a hidden layer and an output layer;

[0053] The input layer receives the preprocessed time series data;

[0054] The hidden layer includes multiple stacked GRU units for extracting temporal features;

[0055] The formula representation of the hidden layer is:

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] Among them, represents the update gate at the current time ; represents the Sigmoid activation function; represents the weight matrix input to the update gate, with a dimension of where is the dimension of the hidden state, is the dimension of the input; represents the input data vector at the current time ; represents the hidden state at the previous time for the weight matrix of the update gate, with a dimension of ; ; represents the hidden state vector at the previous time ; represents the bias term of the update gate, with a dimension of ; represents the hyperparameter for adjusting the gain of the activation function; represents the reset gate at the current time to control the reset of the current hidden state; represents the weight matrix input to the reset gate, with a dimension of ; represents the weight matrix of the hidden state at the previous time for the reset gate, with a dimension of ; represents the dimension of the bias term of the reset gate as ; represents the decay factor in the integral function when controlling the reset gate; represents the candidate hidden state at the current time to generate the candidate hidden state at the current time; represents the hyperbolic tangent activation function; The weight matrix representing the input to the candidate hidden state, with dimensions of ; Represents the hidden state at the previous time step The weight matrix of the hidden state at the previous time step for the candidate hidden state, with dimensions of ; Represents the reset gate The weighted operation on the hidden state at the previous time step ; Represents element-wise multiplication; Represents the bias term of the candidate hidden state, with dimensions of ; Represents the hyperparameter for controlling the weighting of historical information; Represents the final hidden state at the current time step ; Represents the historical time step index; Represents the number of historical time steps; Represents the hidden state at the previous time step vector; Represents the sum of time steps; Represents the time step index, Represents the hidden state at time step

[0062] The output layer includes four output nodes, corresponding respectively to the thrust prediction values, and the formula is expressed as:

[0063] Wherein, is the weight matrix; is the bias term, representing the bias of each output; is the thrust prediction values of the four hydraulic cylinders, with a dimension of 4, representing the thrust values of the four hydraulic cylinders the thrust prediction values.

[0064] As a preferred solution of the data-driven attitude automatic correction method described in the present invention, wherein: the data-driven model further includes defining the state space, action space, and reward function of the Q-learning algorithm;

[0065] The state space includes the current path deviation amount, the pose data of the current tunnel boring machine, and the current hydraulic cylinder pressure state;

[0066] The action space includes different pressure adjustment schemes for the four hydraulic cylinders;

[0067] The reward function includes giving a positive reward when the path deviation amount decreases; giving a negative reward when the path deviation amount increases;

[0068] When the deviation decreases after the oil cylinder applies pressure, a positive reward is given; when the deviation increases after the oil cylinder applies pressure, a negative reward is given;

[0069] Initialize the Q-table to record the Q-values of each state-action pair, and set the initial value of Q to 0;

[0070] According to the current state, select an action using the ε-greedy strategy to balance exploration and exploitation;

[0071] Apply the selected pressure adjustment plan to the oil cylinder to adjust the attitude of the tunnel boring machine;

[0072] Calculate the reward value based on the new path deviation and the oil cylinder pressure state, and update the Q-table;

[0073] The formula for updating the Q-table is expressed as:

[0074]

[0075] Among them, represents the new state after executing the action, represents the best action in the new state, represents the immediate reward of the current action; represents the learning rate, represents the discount factor;

[0076] The iteration process continues until the Q-table converges or reaches the preset 100 rounds of training times, and then the training stops.

[0077] As a preferred solution of the data-driven attitude automatic correction method described in the present invention, wherein: the data-driven model further includes using a genetic algorithm to optimize the learning rate, discount factor, and exploration rate in the Q-learning algorithm;

[0078] The value ranges of the learning rate α, discount factor γ, and exploration rate ε are respectively defined as (0, 1);

[0079] Evenly divide the defined range of each parameter (α, γ, ε) into 10 intervals;

[0080] Randomly select a value from each interval of each parameter as a sample point;

[0081] Combine the sample points of each parameter to generate several groups parameter combinations to form an initial population;

[0082] The evaluation criteria of the fitness function include the convergence speed of the cumulative reward sum;

[0083] The cumulative reward sum includes that the higher the cumulative reward obtained during the training process, the higher the fitness;

[0084] The convergence speed includes The faster the Q-table converges with the parameter combination, the higher the fitness;

[0085] Using each set of parameters Train the Q-learning model and record the cumulative reward and the number of convergence rounds during the training process;

[0086] The formula of the fitness function is expressed as:

[0087] Where represents the fitness function; represents the parameter combination, represents the cumulative reward weight, represents the standard deviation, represents the mean; CR(θ) is the cumulative reward; represents the convergence speed weight, represents the convergence speed function, CS) represents the maximum value of the convergence speed, represents the time step;

[0088] The value range of F(θ) is from 0 to 1. The closer F(θ) is to 1, the higher the fitness of the parameter combination (α, γ, ε);

[0089] Adopt the roulette wheel selection method to select the parameter combination (α, γ, ε) with high fitness to enter the crossover and mutation operations;

[0090] Execute the single-point crossover method to generate new individuals;

[0091] Set the mutation probability to 5% to determine whether to mutate each set of parameter combinations (α, γ, ε);

[0092] Randomly fine-tune the parameter values of a selected set of parameter combinations (α, γ, ε) to generate a new parameter combination;

[0093] Select the parameter combinations with the top 30% fitness from the newly generated parameter combinations and perform local search using the hill climbing algorithm;

[0094] Local search includes generating the fine-tuned parameter combinations as the neighborhood solutions for each selected parameter combination, and the formula is expressed as:

[0095]

[0096] Where represents the optimized parameter combination obtained by local search, represents the neighborhood solution set of the current parameter combination, and argmax represents among Find the solution that maximizes the objective function among the neighborhood solutions, representing the cumulative reward value of this solution, representing the convergence error of this solution;

[0097] Evaluate the fitness of the neighborhood solutions, and select the neighborhood solution with the highest fitness as the new solution; replace the original parameter combination with the better parameter combination after local search ; ;

[0098] Repeat the execution of selection, crossover, mutation and local search to generate a new population; calculate the fitness values of each group of parameter combinations and retain the individual with the best performance;

[0099] When the population fitness no longer improves significantly, stop the iteration.

[0100] A data-driven attitude automatic rectification system, wherein:

[0101] A digital twin model module for constructing a digital twin model of the TBM and the surrounding rock environment;

[0102] A synchronization module that collects the actual position and real-time operation data of the TBM in real time through attitude and position sensors, determines the actual tunneling trajectory of the TBM, and synchronizes the actual position and real-time operation data of the TBM to the digital twin model;

[0103] A comparison module that compares the actual trajectory with the designed trajectory and calculates the true deviation of the trajectory;

[0104] A planning module that, when the true deviation of the trajectory exceeds the rectification threshold, plans a rectification trajectory and calculates the expected rectification parameters of the actual trajectory and the rectification trajectory;

[0105] A prediction module that uses the expected rectification parameters, the surrounding rock parameters at the current position, and the current state of the cylinders as input parameters, and predicts the thrust required for the four driving cylinders through a data-driven model in the digital twin model;

[0106] A calculation module that applies the predicted thrust to the cylinders with corresponding numbers, performs tunneling in the next cycle step, measures the actual position and attitude information of the TBM in the next cycle, determines the actual trajectory, and then calculates the rectification parameters for the next cycle;

[0107] An enhancement module that assembles the rectification parameters, the surrounding rock parameters, and the cylinder thrust into historical time series data for enhancing the training of the data-driven model.

[0108] A computer device, comprising: a memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of any one of the methods of the present invention.

[0109] A computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, the steps of any one of the methods described in the present invention are implemented.

[0110] Advantages of the present invention: By constructing a digital twin model of the TBM and the surrounding rock environment, all-round and real-time monitoring of the tunnel boring process is realized, high-precision basic data is provided, and the visualization and prediction of actual boring are optimized. The actual position and operation data of the TBM are collected in real time by attitude and position sensors, and these data are synchronized to the digital twin model, ensuring a high degree of consistency between the model and the actual progress, and greatly improving the timeliness and accuracy of the data. This synchronization process provides a reliable basis for subsequent trajectory comparison and deviation correction, making it possible to detect and correct trajectory deviations in real time. By comparing the actual trajectory with the designed trajectory, the true trajectory deviation e is calculated, and when the deviation exceeds the threshold, a deviation correction trajectory is planned, further improving the accuracy and timeliness of deviation correction. After each deviation correction, the required thrust is predicted by combining the surrounding rock parameters and the cylinder state, ensuring the continuous optimization of the boring process. This data-driven closed-loop control mechanism can significantly reduce manual intervention, improve the degree of automation, avoid error accumulation, reduce equipment wear, and improve the accuracy, efficiency and safety in the tunnel boring process by dynamically adjusting the deviation correction trajectory. Description of the Drawings

[0111] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0112] Figure 1 It is the overall flowchart of the data-driven attitude automatic deviation correction method provided by the first embodiment of the present invention. Specific Embodiments

[0113] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0114] Example 1, referring to Figure 1 , which is an embodiment of the present invention, provides a data-driven attitude automatic deviation correction method, including:

[0115] S1: Build a digital twin model of the TBM and the surrounding rock environment.

[0116] Collect detailed hardware data of the tunnel boring machine and use CAD to build the geometric model, kinematic model, and dynamic model of the TBM.

[0117] The geometric model describes the shape of the tunnel boring machine and the spatial position relationship of each component. It includes the overall dimensions of the tunnel boring machine, the shapes of each component, the motion connections and constraint conditions; the geometric characteristics of the cutter head, support structure, cylinders, hydraulic system, etc. of the tunneling machine, as well as their relative positions and cooperation methods.

[0118] The kinematic model describes the relative motion laws of each component of the tunnel boring machine, especially the motion trajectories, speeds, accelerations, etc. of each cylinder and other moving components. This model takes into account the motion relationships and coordination of each component during the operation of the tunneling machine and provides the motion states, angle changes, and constraint conditions of each component during the operation.

[0119] The dynamic model describes the mechanical actions and their responses of the tunnel boring machine during operation, mainly involving the dynamic behaviors such as the thrust, torque, acceleration, and vibration of the tunneling machine. It includes the external forces, frictional forces, and hydraulic pressures received by each component during operation, as well as their impacts on the overall motion and attitude of the machine. This model can simulate the dynamic response of the tunnel boring machine and provide a theoretical basis for subsequent attitude correction control.

[0120] Establish the behavior models of the TBM under different working conditions (such as forward, turning, undulating, etc.).

[0121] Install sensors at the key parts of the tunnel boring machine (such as cylinders, power systems, attitude sensors, accelerometers, gyroscopes, etc.) to collect the pose data, path deviation data, cylinder pressure data, vibration acceleration data, thrust data, and soil and tunnel environment data of the TBM in real time.

[0122] Furthermore, by establishing the digital twin model of the tunnel boring machine (TBM) and combining the accurate descriptions of its geometric, kinematic, and dynamic models, it provides a reliable physical and behavioral basis for subsequent attitude control and path correction. By installing sensors at the key parts of the TBM and collecting various operation data (such as pose, path deviation, cylinder pressure, vibration acceleration, thrust, etc.) in real time, it can realize the dynamic monitoring and precise analysis of the machine state. These data provide the necessary inputs for the subsequent real-time attitude correction of the TBM through data-driven control algorithms, thus realizing more efficient and precise tunnel boring operations and optimizing its working performance and stability.

[0123] S2: Collect the actual position and attitude data of the TBM in real time through the attitude and position sensors, determine the actual tunneling trajectory of the TBM, and synchronize the actual position and attitude data of the TBM to the digital twin model.

[0124] The attitude data includes pose data, path deviation data, cylinder pressure data, vibration acceleration data, thrust data, and soil and tunnel environment data.

[0125] Remove the noise and outliers from the real-time operation data, and normalize the real-time operation data with different dimensions.

[0126] Synchronizing the attitude data to the digital twin model includes transmitting the real-time collected data from the tunnel boring machine to the digital twin platform through MQTT.

[0127] Synchronize the time of the real-time operation data through the GPS clock, and add a timestamp after each piece of real-time operation data.

[0128] When the MQTT is delayed or interrupted, ensure that the data is not lost through the cache mechanism, and synchronize it when the MQTT resumes.

[0129] When an exception occurs in the cache mechanism, perform data backup through redundant storage to ensure the integrity of the data.

[0130] Synchronize the real-time operation data to the digital twin model in real time, and update the state of the digital twin model.

[0131] Furthermore, by collecting various operation data of the tunnel boring machine in real time (such as pose data, path deviation data, cylinder pressure data, etc.) and synchronizing this data to the digital twin model, it is ensured that the model can reflect the actual working state of the tunnel boring machine in real time. By removing noise and outliers and normalizing the real-time data, the accuracy and consistency of the data are guaranteed. Using the MQTT protocol for data transmission and time synchronization ensures the efficiency and timeliness of data transmission. In addition, through the cache mechanism and redundant storage technology, the stability and integrity of the data transmission process are guaranteed, preventing data loss caused by delays or interruptions, thus ensuring the accuracy and real-time update of the digital twin model. This process ensures the integrity, real-time nature, and efficiency of the data, providing a reliable data basis for subsequent attitude correction and control strategy optimization.

[0132] S3: Compare the actual path with the preset design path and calculate the path deviation amount.

[0133] Discretize the deviation correction trajectory and the actual trajectory into data points according to the tunneling step.

[0134]

[0135]

[0136] Among them, represents the deviation between the actual trajectory and the designed trajectory, d represents the deviation correction threshold, the threshold is taken as 0.02 m in the straight section and 0.005*R in the turning section, and R represents the radius of the turning section. represents the coordinate on the X-axis of the actual position of the TBM at the current time step. represents the coordinate of the X-axis of the designed trajectory at the current time step. represents the coordinate on the Y-axis of the actual position of the TBM at the current time step; represents the coordinate of the Y-axis of the designed trajectory at the current time step; represents the actual position of the TBM at the current time step on axis coordinate. represents the coordinate of the Z-axis of the designed trajectory at the current time step.

[0137] Assume that the advancement step length each time is 0.01 m; the time step length is 0.5 s.

[0138] The position information of the TBM needs to be updated every time an advancement step length is completed.

[0139] Furthermore, by comparing the actual path of the tunnel boring machine with the preset designed path, the path deviation amount can be accurately calculated, so as to realize the precise positioning and attitude adjustment of the tunnel boring machine. By discretizing the path into a series of points and updating the position of the tunnel boring machine within each time step, the position change of the boring machine during real-time operation can be carefully tracked and compared with the preset path. The calculated path deviation amount (including the deviations in the X, Y, and Z directions) provides a quantitative basis for subsequent attitude deviation correction. This step ensures that the system can monitor and correct the attitude of the boring machine in real time, avoid the decline of construction accuracy caused by path deviation, and thus improve the safety and efficiency of the boring operation.

[0140] S4: When the true deviation of the trajectory exceeds the deviation correction threshold, plan the deviation correction trajectory and calculate the expected deviation correction parameters of the actual trajectory and the deviation correction trajectory.

[0141] The designed trajectory is the coordinate of the discrete points of the ideal trajectory designed in the early project planning.

[0142] The formula for calculating the expected deviation amount is expressed as:

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149] Among them, represents the component of the expected error in the X direction, and X represents the coordinate of the X-axis in three-dimensional space. represents the coordinate on the X-axis of the actual position of the TBM at the current time step. represents the coordinate of the X-axis of the current deviation correction trajectory. represents the path deviation of the TBM in the Y-axis direction. represents the coordinate on the Y-axis of the actual position of the TBM at the current time step; represents the coordinate of the Y-axis of the current deviation correction trajectory; represents the path deviation of the TBM in the Z-axis direction. represents the coordinate of the actual position of the TBM at the current time step on the axis. represents the coordinate of the Z-axis of the current deviation correction trajectory; represents the actual angle of the TBM and the expected angle between the deviations; represents the actual angle at the current time step and the expected angle between the deviations; represents the actual angle at the current time step and the expected angle between the deviations.

[0150] Furthermore, when the deviation between the actual trajectory of the tunnel boring machine (TBM) and the preset design trajectory exceeds the set deviation correction threshold, the expected deviation correction parameters between the actual trajectory and the deviation correction trajectory are calculated to plan a scheme for adjusting the path. Specifically, this process includes calculating the path deviation in the X, Y, and Z directions and the angular deviation between the actual attitude and the expected attitude to ensure that the tunnel boring machine can be adjusted according to the predetermined trajectory. Through the accurate calculation of these deviation amounts, a clear basis can be provided for the application of the cylinder thrust, thereby achieving accurate trajectory correction, ensuring that the TBM advances smoothly along the ideal trajectory, and avoiding engineering risks and time delays caused by deviating from the design trajectory.

[0151] S5: Using the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders as input parameters, predict the required thrust of the four driving cylinders through a data-driven model in the digital twin model.

[0152] The data-driven model includes using a GRU model based on the real-time path deviation to predict the pressure values that need to be applied to the four cylinders. .

[0153] According to the real-time TBM state and the prediction results of the GRU, use the Q-learning algorithm to select the optimal pressure adjustment plan, output the results adjusted by the plan, and use them as the actual pressure values of the corresponding cylinders.

[0154] Predict the pressure values that need to be applied to the four cylinders through the GRU model .

[0155] The input parameters include the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders.

[0156] The expected deviation correction parameters include , , , , and .

[0157] Integrate the input parameters at the current moment with the data of historical time steps to form time series data and perform preprocessing.

[0158] The preprocessing includes using Kalman filtering to remove the noise in the time series data, identifying and processing the outliers in the time series data, ensuring that all time series data are synchronized in time, and normalizing the time series data with different dimensions.

[0159] Using Kalman filtering to remove the noise in the time series data includes establishing a state space model for the operation of the TBM, including position, attitude, and cylinder state variables. Kalman filtering first needs to construct a state space model for the time series data. It includes a state transition equation: describing how the system transfers from the previous state to the current state in one time step. The formula is expressed as:

[0160]

[0161] where, is the system state at the current moment, is the state transition matrix, is the process noise, indicating the uncertainty of the system during the transfer. Observation equation: describing how to generate the observed value from the current system state.

[0162] The observation equation formula is expressed as:

[0163]

[0164] where, is the observed value, is the observation matrix, is the observation noise.

[0165] Initial state Typically, the first observation of the time series data is used as the initial estimate.

[0166] Initial error covariance matrix Represents the uncertainty of the initial estimate and is usually set to a relatively large value.

[0167] At each time step, the Kalman filter predicts the state at the current moment by using the state transition equation and updates the estimated covariance matrix.

[0168] Predict the current state:

[0169]

[0170] Predict the error covariance matrix:

[0171]

[0172] where, is the process noise covariance matrix. In the update step, after receiving new observation data, the Kalman filter updates the predicted state estimate through the observation equation and calculates the new error covariance matrix, calculates the Kalman gain , representing the weight of the correction to the predicted value at the current moment:

[0173]

[0174] where, is the observation noise covariance matrix.

[0175] Update the estimated state:

[0176]

[0177] where, is the residual between the observation value and the predicted value. Update the error covariance matrix:

[0178]

[0179] where, is the identity matrix.

[0180] The Kalman filter repeats the prediction and update steps recursively. At each time step, the algorithm optimizes the state estimate based on the previous estimate and the new observation data, and continuously reduces the impact of noise on the data.

[0181] The input time window size of the GRU model is the time series data of the past 10 time steps.

[0182] The time series data is divided into a training set of 70%, a validation set of 15%, and a test set of 15% in chronological order.

[0183] The hyperparameters of the GRU model include the number of GRU units, the time step, the learning rate, the batch size, and the number of training epochs.

[0184] The number of GRU units is 512 hidden units in the GRU layer.

[0185] The time step is 25 time steps, representing the length of the entire input sequence.

[0186] The learning rate is 0.0005.

[0187] The batch size is 64 samples used in one training.

[0188] The number of training epochs is 100 epochs.

[0189] The hyperparameters of the GRU model are adjusted by the grid search method.

[0190] List all the hyperparameters to be adjusted and their candidate values.

[0191] Learning rate: {0.0001, 0.0005, 0.001}

[0192] Batch size: {32, 64, 128}

[0193] Number of hidden layer units: {128, 256, 512}

[0194] Number of training epochs: {50, 100, 150}

[0195] According to the set candidate values, generate all possible combinations of hyperparameters. For the candidate values given above, the grid search will generate the following combinations:

[0196] (0.0001, 32, 128, 50)

[0197] (0.0001, 32, 128, 100)

[0198] (0.0001, 32, 128, 150)

[0199] (0.0001, 64, 128, 50)

[0200] (0.0001, 64, 128, 100)

[0201] (0.0001, 64, 128, 150)

[0202] (0.0001, 128, 128, 50)

[0203] (0.0001, 128, 128, 100)

[0204] (0.0001, 128, 128, 150) ...

[0206] (All other possible combinations)

[0207] For each set of hyperparameter combinations, train the model and evaluate the model's performance on the validation set. Use the mean squared error as the evaluation metric and select the current hyperparameter combination;

[0208] Train the GRU model on the training set.

[0209] Evaluate the performance of the trained model on the validation set.

[0210] Record the evaluation results. After training and evaluating all hyperparameter combinations, select the hyperparameter combination that performs best on the validation set, i.e., select those hyperparameter combinations that can minimize the loss function, maximize the accuracy, or other objectives.

[0211] Use the selected best hyperparameter combination to retrain the GRU model on the entire training set. After training, use this model to evaluate its generalization ability on the test set and confirm its performance.

[0212] Analyze the evaluation results of all hyperparameter combinations and determine which parameters have the greatest impact on the model performance.

[0213] The GRU model includes an input layer, a hidden layer, and an output layer.

[0214] The input layer receives the preprocessed time series data.

[0215] The hidden layer includes multiple stacked GRU units for extracting temporal features.

[0216] The formula representation of the hidden layer is:

[0217]

[0218]

[0219]

[0220]

[0221]

[0222] Among them, represents the update gate at the current time ; represents the Sigmoid activation function; represents the weight matrix input to the update gate, with dimensions , where is the dimension of the hidden state, is the dimension of the input; represents the input data vector at the current time ; represents the hidden state at the previous time ; represents the weight matrix of the hidden state at the previous time to the update gate, with dimensions ; represents the hidden state vector at the previous time represents the bias term of the update gate, with dimensions ; represents the hyperparameter for adjusting the gain of the activation function; represents the reset gate at the current time , controlling the reset of the current hidden state; represents the weight matrix input to the reset gate, with dimensions ; represents the weight matrix of the hidden state at the previous time to the reset gate, with dimensions ; represents the dimension of the bias term of the reset gate ; represents the decay factor in the integral function when controlling the reset gate; represents the candidate hidden state at the current time , generating the candidate hidden state at the current time; represents the hyperbolic tangent activation function; represents the weight matrix input to the candidate hidden state, with dimensions ; represents the weight matrix of the hidden state at the previous time to the candidate hidden state, with dimensions ; represents the weighted operation of the reset gate on the hidden state at the previous time ; represents element-wise multiplication; represents the bias term of the candidate hidden state, with dimensions ; represents the hyperparameter for controlling the weighting of historical information; represents the current time The final hidden state. Indicates the historical time step index; Indicates the number of historical time steps; Indicates the previous moment The hidden state vector; Indicates the sum of time steps; Indicates the time step index, Indicates The hidden state at the moment.

[0223] The output layer includes four output nodes, corresponding respectively to The predicted thrust value, expressed by the formula:

[0224] Where, Is the weight matrix; Is the bias term, indicating the bias of each output; Is the predicted thrust value of the four cylinders, with a dimension of 4, indicating the thrust values of the four cylinders The predicted thrust value.

[0225] Define the state space, action space and reward function of the Q-learning algorithm.

[0226] The state space includes the current path deviation, the pose data of the current tunnel boring machine, and the current cylinder pressure state.

[0227] The action space includes different pressure adjustment schemes for the four cylinders.

[0228] The reward function includes giving a positive reward when the path deviation decreases; giving a negative reward when the path deviation increases.

[0229] When the deviation decreases after the cylinder applies pressure, give a positive reward; when the deviation increases after the cylinder applies pressure, give a negative reward.

[0230] Initialize the Q-table to record the Q-values of each state-action pair, and set the initial value of Q to 0. Select an action according to the current state using the ε-greedy strategy to balance exploration and exploitation.

[0231] Apply the selected pressure adjustment scheme to the cylinders to adjust the attitude of the tunnel boring machine.

[0232] Calculate the reward value according to the new path deviation and cylinder pressure state, and update the Q-table.

[0233] The formula for updating the Q-table is expressed as:

[0234]

[0235] Where, Represents the new state after performing an action, Represents the best action in the new state, Represents the immediate reward for the current action; Represents the learning rate, Represents the discount factor.

[0236] The iterative process continues until the Q - table converges or reaches the preset 100 - round training times, at which point the training stops.

[0237] The data - driven model further includes optimizing the learning rate, discount factor, and exploration rate in the Q - learning algorithm using a genetic algorithm.

[0238] The value ranges of the learning rate α, discount factor γ, and exploration rate ε are respectively defined as (0, 1).

[0239] The defined range of each parameter (α, γ, ε) is evenly divided into 10 intervals.

[0240] Randomly select a value from each interval of each parameter as a sample point.

[0241] Combine the sample points of each parameter to generate several groups Parameter combinations to form an initial population.

[0242] The evaluation criteria of the fitness function include the convergence speed of the cumulative reward sum.

[0243] The cumulative reward sum includes that the higher the cumulative reward obtained during the training process, the higher the fitness.

[0244] The convergence speed includes, The faster the parameter combination makes the Q - table converge, the higher the fitness.

[0245] Use each group of parameters Train the Q - learning model and record the cumulative reward and convergence rounds during the training process.

[0246] The formula of the fitness function is expressed as:

[0247] Where, Represents the fitness function; Represents the parameter combination, Represents the cumulative reward weight, Represents the standard deviation, Represents the mean; CR(θ) is the cumulative reward; Represents the convergence speed weight, Represents the convergence speed function, CS) represents the maximum value of the convergence speed, Indicates the time step. The value range of is from 0 to 1. The closer F(θ) is to 1, the higher the fitness of the parameter combination (α, γ, ε).

[0248] Adopt the roulette wheel selection method to select the parameter combination (α, γ, ε) with high fitness to enter the crossover and mutation operations.

[0249] Calculate the total fitness of all candidate parameter combinations . Suppose there are m individuals in the population, and the fitness of each individual is , then .

[0250] For each individual , calculate its selection probability :

[0251]

[0252] Individuals with high fitness will have a higher selection probability.

[0253] According to the selection probabilities of each individual, construct a cumulative probability distribution , where

[0254] The -th individual's cumulative probability is:

[0255]

[0256] Among them, represents the total fitness of all individuals in the population. represents the selection probability of the -th individual. represents the cumulative probability of the s-th individual. represents a random number between 0 and 1, which is used to determine which individual to select. represents the fitness value of the s-th individual. represents the number of individuals in the population.

[0257] Generate a random number between 0 and 1 . Find the individual s that satisfies in the cumulative probability distribution and take it as the selected individual. Execute the above steps until the required number of individuals is selected.

[0258] Execute the single-point crossover method to generate new individuals.

[0259] According to the roulette wheel selection method, select two individuals with higher fitness as the parents and .

[0260] Randomly select a position in the dimension of the parameter vector as the crossover point. For example, assume the parameter combination; copy the parameters before the crossover point from to offspring 1, and copy the parameters after the crossover point from to offspring 1. Perform the reverse operation on another offspring 2:

[0261] Offspring 1: ; Offspring 2:

[0262] Two new individuals are obtained as offspring through the above operations.

[0263] According to the mutation probability, perform mutation operations on some individuals and randomly change the genes of the individuals.

[0264] Select the top 30% of the individuals with the highest fitness from the newly generated individuals for local search.

[0265] Iteratively optimize through operations such as selection, crossover, mutation, and local search until the fitness no longer improves significantly.

[0266] Set the mutation probability to 5% and determine whether to mutate each group of parameter combinations (α, γ, ε).

[0267] Randomly fine-tune the parameter values of a selected group of parameter combinations (α, γ, ε) to generate a new parameter combination.

[0268] Select the top 30% of the parameter combinations with the highest fitness from the newly generated parameter combinations and perform local search using the hill-climbing algorithm.

[0269] Local search includes generating the fine-tuned parameter combinations for each selected parameter combination as the neighborhood solutions, which is expressed by the formula:

[0270]

[0271] where, represents the optimized parameter combination obtained through local search, represents the set of neighborhood solutions of the current parameter combination, argmax represents finding the solution that maximizes the objective function among the neighborhood solutions, represents the cumulative reward value of this solution, represents the convergence error of this solution.

[0272] Evaluate the fitness of the neighborhood solutions, select the neighborhood solution with the highest fitness as the new solution; replace the original parameter combination with the better parameter combination after local search .

[0273] ​Repeat the operations of selection, crossover, mutation, and local search to generate a new population; calculate the fitness values of each group of parameter combinations and retain the individual with the best performance.

[0274] Stop the iteration when the population fitness no longer improves significantly.

[0275] Furthermore, ensure a balance between efficiency and effectiveness in the optimization process of the genetic algorithm. When the population fitness no longer improves significantly, it means that the current parameter combination is close to the optimal solution. Further iteration may not bring substantial performance improvement and may even lead to waste of resources. Therefore, by setting the termination condition that the fitness no longer improves significantly, it can effectively prevent the algorithm from running excessively, save computing resources, and avoid overfitting problems caused by over-optimization. In addition, this step ensures that the algorithm converges to a high-quality parameter combination within a reasonable time, thereby improving the prediction accuracy of the data-driven model and the overall control performance of the system. This design not only improves the efficiency of the optimization process but also enhances the robustness and reliability of the model in practical applications, ensuring that the tunnel boring machine can move forward stably and accurately along the designed trajectory.

[0276] S6: Apply the predicted thrust to the cylinders with corresponding numbers, conduct tunneling for the next cycle step, measure the actual position and attitude information of the TBM in the next cycle, determine the actual trajectory, and then calculate the deviation correction parameters for the next cycle;

[0277] S7: Assemble the deviation correction parameters, surrounding rock parameters, and cylinder thrust into historical time series data to enhance the training of the data-driven model.

[0278] Furthermore, by applying the predicted thrust to the cylinders with corresponding numbers, conduct the tunneling process for the next cycle, and simultaneously measure the actual position and attitude information of the tunnel boring machine (TBM) in real time to obtain the actual trajectory. By comparing the actual trajectory with the preset designed trajectory, calculate the new trajectory deviation, and thus obtain the deviation correction parameters for the next cycle. These deviation correction parameters, surrounding rock parameters, and cylinder thrust are aggregated into historical time series data to enhance the performance of the training data-driven model and further optimize the prediction ability. The core of this step is to continuously update and adjust the model, accurately correct the future thrust prediction based on real-time data, so that the TBM can efficiently and precisely adjust its attitude and trajectory in a complex environment and achieve more accurate automatic control.

[0279] Embodiment 2, an embodiment of the present invention, provides a data-driven attitude automatic deviation correction system, including:

[0280] A digital twin model module that constructs a digital twin model of the TBM and the surrounding rock environment.

[0281] The synchronization module collects the actual position and attitude data of the TBM in real time through the attitude and position sensors, determines the actual tunneling trajectory of the TBM, and synchronizes the actual position and attitude data of the TBM to the digital twin model.

[0282] The comparison module compares the actual trajectory with the designed trajectory and calculates the true trajectory deviation.

[0283] The planning module plans a deviation correction trajectory and calculates the expected deviation correction parameters of the actual trajectory and the deviation correction trajectory when the true trajectory deviation exceeds the deviation correction threshold.

[0284] The prediction module takes the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders as input parameters, and predicts the thrust required for the four driving cylinders through a data-driven model in the digital twin model.

[0285] The calculation module applies the predicted thrust to the cylinders with corresponding numbers, conducts tunneling in the next cycle step, measures the actual position and attitude information of the TBM in the next cycle, determines the actual trajectory, and then calculates the deviation correction parameters for the next cycle.

[0286] The enhancement module assembles the deviation correction parameters, the surrounding rock parameters, and the cylinder thrust into historical time series data for enhancing the training of the data-driven model.

[0287] Embodiment 3, an embodiment of the present invention, which is different from the previous two embodiments in that:

[0288] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0289] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0290] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0291] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0292] Embodiment 4, an embodiment of the present invention, provides a data-driven automatic attitude correction method and system. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.

[0293] Select a TBM equipped with high-precision attitude and position sensors, and install attitude sensors, position sensors, cylinder pressure sensors, vibration acceleration sensors, and thrust sensors at its key parts to collect the attitude data of the TBM in real time. These data are transmitted to the digital twin platform in real time through the MQTT protocol and synchronized with the GPS clock to ensure that each piece of data is attached with an accurate timestamp.

[0294] During the test, the TBM began to tunnel according to the preset design trajectory, and at the same time, it collected and transmitted real-time data such as its position, attitude, cylinder pressure, vibration acceleration, thrust force, and surrounding rock parameters. Through the digital twin model to synchronize and process the real-time data, the system can immediately determine the actual tunneling trajectory of the TBM, compare it with the design trajectory, and calculate the real trajectory deviation. When the trajectory deviation exceeds the set deviation correction threshold (the deviation threshold for the straight section is 0.02 meters, and the deviation threshold for the turning section is 0.005 times the turning radius R), the system automatically starts the deviation correction algorithm. The deviation correction algorithm first plans the deviation correction trajectory and calculates the expected deviation correction parameters between the actual trajectory and the deviation correction trajectory. Then, based on the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders, the data-driven model predicts the thrust force required to be applied to the four driving cylinders through the GRU model. The predicted thrust force value is then applied to the cylinders with corresponding numbers to adjust the attitude of the TBM and carry out the tunneling of the next cycle step. During this process, the system continuously measures the actual position and attitude information of the TBM in the next cycle, determines the new actual trajectory, and recalculates the deviation correction parameters for the next cycle. All deviation correction parameters, surrounding rock parameters, and cylinder thrust forces are assembled into historical time series data to further enhance the prediction accuracy of the training data-driven model. To ensure data integrity, when the MQTT transmission is delayed or interrupted, the system temporarily stores the data through the cache mechanism and synchronizes it after the transmission resumes; if an exception occurs in the cache mechanism, data backup is performed through redundant storage to ensure that the data is not lost.

[0295] During the implementation process, the GRU model uses the time series data of historical time steps, combines the current trajectory deviation and surrounding rock parameters, and accurately predicts the thrust force required by the cylinders. The Q-learning algorithm then selects the optimal pressure adjustment scheme according to the real-time state and the prediction results of the GRU to achieve dynamic and intelligent trajectory deviation correction. By optimizing the hyperparameters (learning rate, discount factor, exploration rate) of Q-learning through the genetic algorithm, the convergence speed and prediction accuracy of the algorithm are further improved. The whole process forms a closed-loop control system to ensure that the TBM can tunnel along the design trajectory efficiently and accurately in a complex surrounding rock environment, significantly reducing the trajectory deviation and improving the tunneling efficiency and safety.

[0296] In this experiment, 100 tunneling cycle steps were tested, each cycle step was 0.01 meters long, and the time step was 0.5 seconds. The following are the actual data records of some key time steps:

[0297] In the initial stage (steps 1 - 10), the TBM tunnels along the designed trajectory with a small trajectory deviation, and the maximum deviation is 0.01 meters. As the tunneling process progresses (step 20), due to the uneven distribution of the surrounding rock strength, the TBM shows a slight trajectory deviation. The deviation in the X direction is 0.015 meters, the deviation in the Y direction is -0.012 meters, and the deviation in the Z direction is 0.018 meters. When entering step 30, the deviation further increases. The deviation in the X direction reaches 0.025 meters, exceeding the deviation correction threshold of 0.02 meters for the straight section, and the system automatically activates the deviation correction algorithm. The GRU model predicts the thrust values of the cylinders to be 16.50 MPa, 15.80 MPa, 14.90 MPa, and 16.20 MPa respectively. After applying the thrust, the trajectory deviation of the TBM significantly decreases in step 40. The deviation in the X direction drops to 0.005 meters, the deviation in the Y direction drops to -0.004 meters, and the deviation in the Z direction drops to 0.007 meters. After 50 steps of tunneling, the trajectory deviation further stabilizes. The deviation in the X direction remains at 0.003 meters, the deviation in the Y direction is -0.002 meters, and the deviation in the Z direction is 0.004 meters. During the entire tunneling process, the GRU model gradually optimizes the thrust prediction by continuously learning historical data and real-time feedback, achieving precise trajectory deviation correction.

[0298] Through the analysis of the experimental data, it can be clearly seen the significant advantages of the data-driven automatic attitude deviation correction method during the TBM tunneling process. First of all, in the initial stage when the deviation correction algorithm is not activated, the trajectory deviation of the TBM gradually increases with the increase of the tunneling step length, indicating that the traditional manual adjustment method is difficult to keep the TBM tunneling stably along the designed trajectory in the face of a complex surrounding rock environment. However, once the trajectory deviation exceeds the preset threshold, the automatically activated deviation correction algorithm of the system responds quickly. Through the thrust prediction of the GRU model and the strategy optimization of Q-learning, the thrust adjustment of the four cylinders significantly reduces the trajectory deviation. During this process, relying on its deep learning ability for time series data, the GRU model can accurately predict the thrust values required for each cylinder, thus achieving precise attitude adjustment.

[0299] Compared with the traditional method, the automatic attitude deviation correction algorithm of the present invention demonstrates its innovation and superiority in multiple aspects. The traditional method usually relies on the experience of the operator or preset static rules and is difficult to cope with the real-time changing surrounding rock conditions and dynamic trajectory deviations. While the present invention realizes dynamic and real-time trajectory deviation correction through the real-time data synchronization of the digital twin model, the precise thrust prediction of the GRU model, and the intelligent decision-making of the Q-learning algorithm. The experimental data shows that after implementing the deviation correction algorithm of the present invention, the trajectory deviation of the TBM significantly decreases in each tunneling cycle step. Especially in the key deviation correction steps, the trajectory deviation drops below 0.005 meters, fully demonstrating the high efficiency and accuracy of the algorithm.

[0300] In addition, the hyperparameters of Q-learning optimized by the genetic algorithm enable the deviation correction algorithm to have a faster convergence rate and higher adaptability during the training process. The genetic algorithm improves the performance of the Q-learning algorithm in different tunneling environments by optimizing the learning rate, discount factor, and exploration rate, making the adjustment of the cylinder thrust more reasonable and efficient. In the experiment, Q-learning optimized by the genetic algorithm showed a higher cumulative reward and a faster convergence rate after 100 rounds of training, further verifying the application potential and technological advancement of the present invention in complex tunneling environments.

[0301] In summary, this embodiment fully demonstrates the innovation and advantages of the data-driven automatic attitude correction method during the TBM tunneling process through detailed data recording and analysis. This algorithm can not only adjust the attitude and trajectory of the TBM in real time and accurately, reducing the deviation during tunneling, but also improve the tunneling efficiency and safety through intelligent thrust prediction and strategy optimization, overcoming the deficiencies of traditional methods in dynamic environments, and reflecting significant technological progress and application value.

[0302] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A data-driven automatic attitude correction method, characterized in that Including: Constructing a digital twin model of the TBM and the surrounding rock environment; Real-time collecting the actual position and real-time operation data of the TBM through attitude and position sensors, determining the actual tunneling trajectory of the TBM, and synchronizing the actual position and real-time operation data of the TBM to the digital twin model; Comparing the actual trajectory with the designed trajectory to calculate the true trajectory deviation; When the true trajectory deviation exceeds the deviation correction threshold, planning a deviation correction trajectory and calculating the expected deviation correction parameters of the actual trajectory and the deviation correction trajectory; Taking the expected deviation correction parameters, the surrounding rock parameters at the current position, and the current pressure of the cylinders as input parameters, predicting the required thrust of the four driving cylinders in the next time step through a data-driven model in the digital twin model; the data-driven model includes using a GRU model based on real-time path deviation parameters to predict the pressure values that need to be applied by the four cylinders; According to the real-time TBM state and the prediction results of the GRU, using the Q-learning algorithm to select the optimal pressure adjustment scheme, outputting the result adjusted by the scheme and using it as the actual pressure value of the corresponding cylinder; Applying the actual pressure value of the cylinder to the corresponding numbered cylinder, carrying out tunneling in the next cycle step, measuring the actual position and attitude information of the TBM in the next cycle, determining the actual trajectory, and then calculating the expected deviation correction parameters of the next cycle; Assembling the expected deviation correction parameters, surrounding rock parameters, and cylinder thrust of the next cycle into historical time series data to enhance the training of the data-driven model.

2. The data-driven automatic attitude correction method according to claim 1, characterized in that: The real-time operation data includes: attitude data, position data, trajectory deviation, cylinder pressure data, vibration acceleration data, propulsion force data, and surrounding rock parameters; Removing the noise and outliers of the real-time operation data and normalizing the real-time operation data with different dimensions; Synchronizing the attitude data to the digital twin model includes transmitting the real-time collected data from the TBM to the digital twin platform through MQTT; Synchronizing the time of the real-time operation data through the GPS clock and adding a timestamp after each piece of real-time operation data; When MQTT is delayed or interrupted, ensuring that the data is not lost through a caching mechanism and synchronizing when MQTT resumes; When an exception occurs in the caching mechanism, backing up the data through redundant storage to ensure the integrity of the data; Real-time synchronizing the real-time operation data to the digital twin model and updating the state of the digital twin model.

3. The data-driven automatic attitude correction method according to claim 1, wherein: Calculating the deviation amount between the deviation correction trajectory and the actual trajectory includes discretizing the deviation correction trajectory and the actual trajectory into data points according to the tunneling step length and calculating the deviation according to the data points; ; ; Among them, represents the deviation between the actual trajectory and the designed trajectory, d represents the deviation correction threshold, the threshold is 0.02 meters in the straight section and 0.005*R in the turning section, and R represents the radius of the turning section. represents the coordinate on the X-axis of the actual position of the TBM at the current time step. represents the X-axis coordinate of the designed trajectory at the current time step. represents the coordinate on the Y-axis of the actual position of the TBM at the current time step; represents the Y-axis coordinate of the designed trajectory at the current time step; represents the actual position of the TBM at the current time step on axis coordinate; represents the Z-axis coordinate of the designed trajectory at the current time step; Assuming that the tunneling step length each time is 0.01 meters; the time step length is 0.5 seconds; Updating the position information of the TBM every time a tunneling step length is completed; The designed trajectory is the discrete point coordinates of the ideal trajectory designed in the early project planning; The formula for calculating the expected deviation parameter is expressed as: ; ; ; ; ; ; Among them, represents the component of the expected error in the X direction, where X represents the coordinate of the X-axis in three-dimensional space, represents the coordinate of the X-axis of the actual position of the TBM at the current time step, represents the coordinate of the X-axis of the current deviation correction trajectory, represents the path deviation of the TBM in the Y-axis direction, represents the coordinate of the actual position of the TBM at the current time step on the Y-axis; represents the coordinate of the Y-axis of the current deviation correction trajectory; represents the path deviation of the TBM in the Z-axis direction, represents the actual position of the TBM at the current time step on the axis coordinate, represents the coordinate of the Z-axis of the current deviation correction trajectory; represents the actual angle of the TBM and the expected angle the deviation between; represents the deviation between the actual angle and the expected angle at the current time step; represents the deviation between the actual angle and the expected angle at the current time step.

4. The data-driven automatic attitude correction method according to claim 3, wherein: Based on the GRU model, predict the pressure values that need to be applied to the four oil cylinders ; The described expected rectification parameters include , , , , and ; Integrating the input parameters at the current moment with the data of the historical time steps to form a time series data matrix and performing preprocessing, the formula is expressed as: ; Select the data of the past time steps as the input sequence, and the historical time series data matrix is expressed by the formula: ; Among them, represents the surrounding rock parameter vector at the current moment, represents the cylinder state parameter at the current moment; The preprocessing includes using Kalman filtering to remove noise in the time series data, identifying and processing outliers in the time series data, ensuring that all time series data are synchronized in time, and normalizing time series data with different dimensions; The input time window size of the GRU model is the time series data of the past 10 time steps; The time series data are divided in the ratio of 70% training set, 15% validation set, and 15% test set in chronological order; The hyperparameters of the GRU model include the number of GRU units, time step, learning rate, batch size, and number of training epochs; The number of GRU units is 512 hidden units in the GRU layer; The time step is 25 time steps, representing the length of the entire input sequence; The learning rate is 0.0005; The batch size is 64 samples used in one training; The number of training epochs is 100; The hyperparameters of the GRU model are adjusted by the grid search method; The GRU model includes an input layer, a hidden layer, and an output layer; The input layer receives the preprocessed time series data; The hidden layer includes multiple stacked GRU units for extracting temporal features; The formula of the hidden layer is expressed as: ; ; ; ; ; ; Among them, represents the update gate at the current moment ; represents the Sigmoid activation function; represents the weight matrix input to the update gate, with a dimension of , where is the dimension of the hidden state, is the dimension of the input; represents the input data vector at the current moment ; represents the hidden state at the previous moment ; represents the weight matrix of the hidden state at the previous moment for the update gate, with a dimension of ; represents the hidden state vector at the previous moment represents the bias term of the update gate, with a dimension of ; represents the hyperparameter used to adjust the gain of the activation function; represents the reset gate at the current moment , controlling the reset of the current hidden state; represents the weight matrix input to the reset gate, with a dimension of ; represents the weight matrix of the hidden state at the previous moment for the reset gate, with a dimension of ; represents the dimension of the bias term of the reset gate ; represents the decay factor in the integral function when controlling the reset gate; represents the candidate hidden state at the current moment , generating the candidate hidden state at the current moment; represents the hyperbolic tangent activation function; represents the weight matrix input to the candidate hidden state, with a dimension of ; represents the weight matrix of the hidden state at the previous moment for the candidate hidden state, with a dimension of ; represents the weighted operation of the reset gate on the hidden state at the previous moment ; represents element-wise multiplication; represents the bias term of the candidate hidden state, with a dimension of ; represents the hyperparameter for controlling the weighting of historical information; represents the final hidden state at the current moment ; Indicates the historical time step index; Indicates the number of historical time steps; Indicates the previous moment of the hidden state vector; Indicates the sum of time steps; Indicates the time step index, Indicates the hidden state at the moment; The output layer includes four output nodes, corresponding respectively to the predicted thrust values, which are expressed by the formula as: ; Among them, is the weight matrix; is the bias term, representing the bias of each output; is the predicted thrust value of the four oil cylinders, with a dimension of 4, representing the thrust values of the four oil cylinders predicted thrust value.

5. The data-driven automatic attitude correction method according to claim 4, wherein: The data-driven model further includes defining the state space, action space, and reward function of the Q-learning algorithm; The state space includes the current path deviation, the pose data of the current tunnel boring machine, and the current cylinder pressure state; The action space includes different pressure adjustment schemes for four cylinders; The reward function includes giving a positive reward when the path deviation decreases and a negative reward when the path deviation increases; When the deviation decreases after the cylinder applies pressure, a positive reward is given; when the deviation increases after the cylinder applies pressure, a negative reward is given; Initialize the Q-table to record the Q-values of each state-action pair, and set the initial value of Q to 0; Select an action according to the current state using the ε-greedy strategy to balance exploration and exploitation; Apply the selected pressure adjustment scheme to the cylinder to adjust the attitude of the tunnel boring machine; Calculate the reward value according to the new path deviation and cylinder pressure state, and update the Q-table; The formula for updating the Q-table is expressed as: ; Among them, represents the new state after performing the action, represents the best action in the new state, represents the immediate reward of the current action; represents the learning rate, represents the discount factor; The iterative process continues until the Q-table converges or reaches the preset 100 training times, and then the training stops.

6. The data-driven automatic attitude correction method according to claim 5, wherein: The data-driven model further includes optimizing the learning rate, discount factor, and exploration rate in the Q-learning algorithm using the genetic algorithm; The value ranges of the learning rate α, discount factor γ, and exploration rate ε are respectively defined as (0, 1); The defined range of each parameter (α, γ, ε) is evenly divided into 10 intervals; Randomly select a value from each parameter interval as a sample point; Combine the sample points of each parameter to generate several groups of parameter combinations to form an initial population; The evaluation criteria of the fitness function include the convergence speed of the cumulative reward sum; The cumulative reward sum includes that the higher the cumulative reward obtained during the training process, the higher the fitness; The convergence speed includes, The faster the Q-table converges with the parameter combination, the higher the fitness; Use each set of parameters Train the Q-learning model and record the cumulative reward and the number of convergence rounds during the training process; The formula of the fitness function is expressed as: ; Among them, represents the fitness function; represents the parameter combination, represents the cumulative reward weight, represents the standard deviation, represents the mean; CR(θ) is the cumulative reward; represents the convergence speed weight, represents the convergence speed function, CS) represents the maximum value of the convergence speed, represents the time step; The value range of F(θ) is from 0 to 1. The closer F(θ) is to 1, the higher the fitness of the parameter combination (α, γ, ε); Use the roulette wheel selection method to select the parameter combination (α, γ, ε) with high fitness to enter the crossover and mutation operations; Execute the single-point crossover method to generate new individuals; Set the mutation probability to 5% to determine whether to mutate each set of parameter combinations (α, γ, ε); Randomly fine-tune the parameter values of a selected set of parameter combinations (α, γ, ε) to generate a new set of parameter combinations; Select the parameter combinations with the top 30% fitness rankings from the newly generated parameter combinations and perform local search using the hill-climbing algorithm; Local search includes generating fine-tuned parameter combinations for each selected parameter combination As a neighborhood solution, it is expressed by the formula: ; Among them, represents the optimized parameter combination obtained by local search, represents the set of neighborhood solutions of the current parameter combination, and argmax represents finding the solution that maximizes the objective function among the neighborhood solutions, represents the cumulative reward value of this solution, represents the convergence error of this solution; Evaluate the fitness of the neighborhood solutions, and select the neighborhood solution with the highest fitness as the new solution; replace the original parameter combination with the better parameter combination after local search Replace the original parameter combination ; Repeat the selection, crossover, mutation, and local search operations to generate a new population; calculate the fitness values of each set of parameter combinations and retain the individual with the best performance; Stop the iteration when the population fitness no longer improves significantly.

7. A data-driven automatic attitude correction system using the method according to any one of claims 1-6, characterized in that: A digital twin model module that constructs a digital twin model of the TBM and the surrounding rock environment; A synchronization module that uses attitude and position sensors to collect the actual position and real-time operation data of the TBM in real time, determines the actual tunneling trajectory of the TBM, and synchronizes the actual position and real-time operation data of the TBM to the digital twin model; A comparison module that compares the actual trajectory with the designed trajectory and calculates the true trajectory deviation; A planning module that, when the true trajectory deviation exceeds the correction threshold, plans a correction trajectory and calculates the expected correction parameters for the actual trajectory and the correction trajectory; A prediction module that uses the expected correction parameters, the surrounding rock parameters at the current position, and the current state of the cylinders as input parameters, and predicts the thrust required for the four driving cylinders through a data-driven model in the digital twin model; A calculation module that applies the predicted thrust to the cylinders with corresponding numbers, performs tunneling in the next cycle step, measures the actual position and attitude information of the TBM in the next cycle, determines the actual trajectory, and then calculates the correction parameters for the next cycle; An enhancement module that assembles the correction parameters, the surrounding rock parameters, and the cylinder thrust into historical time series data for enhancing the training of the data-driven model.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the data-driven automatic attitude correction method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the data-driven automatic attitude correction method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Shield attitude automatic control method and system based on data driving

    CN114075980A

  • Shield tunneling attitude control method and system based on digital twinning

    CN119466838A