A shield machine posture self-adaptive control method and system under shield construction conditions

By constructing a virtual construction environment and reinforcement learning model using the DGDPG method, the tunneling posture parameters of the TBM are optimized, solving the problem of low accuracy and stability in the posture control of the tunnel boring machine. This achieves high-precision and stable adaptive control, supporting intelligent, autonomous, and unmanned construction of the tunnel boring machine.

CN116220713BActive Publication Date: 2025-11-11HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310186685.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-11-11
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing methods for controlling the attitude of tunnel boring machines (TBMs) rely on human experience, resulting in poor control accuracy, low stability, and an inability to achieve adaptive control. In particular, it is difficult to optimize the combination of TBM parameters in complex underground environments.

Method used

A shield tunneling machine attitude adaptive control system based on deep gated recurrent neural network combined with deterministic policy gradient method (DGDPG) is adopted. By constructing a virtual construction environment and reinforcement learning model, the TBM tunneling attitude parameters are optimized to achieve high-precision and stable adaptive control.

Benefits of technology

The accuracy and stability of the shield machine's attitude control were improved, with pitch angle control accuracy increased by 60.14% and stability increased by 90.21%. Vertical deviation control accuracy and stability were increased by 76.52% and 62.13% respectively, realizing intelligent, autonomous, and unmanned construction of the shield machine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116220713B_ABST
    Figure CN116220713B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of shield tunneling construction control technology, and specifically discloses a method and system for adaptive control of shield machine attitude under shield tunneling conditions. The method includes: constructing a virtual construction environment for TBM attitude adaptive control; determining the state and reward of a reinforcement learning model during TBM tunneling, wherein the TBM state is updated according to constitutive relations, and the reward is used for forward speed and TBM attitude; constructing a DGDPG model for the shield tunneling process based on the above virtual construction environment, state, and reward, and optimizing and training the DGDPG model; and evaluating the TBM tunneling process based on the optimized DGDPG model. This invention addresses the characteristics of traditional shield tunneling processes, such as non-repeatability, poor control accuracy, and low level of intelligence, by achieving high-precision adaptive control of shield machine tunneling attitude based on deep reinforcement learning methods, thereby overcoming the shortcomings of current shield machine attitude prediction and optimization methods in engineering applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of shield tunneling construction control technology, and relates to a shield machine attitude adaptive control method and system under shield tunneling construction conditions. More specifically, it relates to a shield machine attitude adaptive control method (DGDPG) under shield tunneling construction conditions based on a gated recurrent neural network combined with a deep deterministic policy gradient method. Background Technology

[0002] Tunnel boring machines (TBMs) are widely recognized as a common method for urban underground construction due to their environmental friendliness, efficiency, and safety. The attitude of the TBM is particularly important during the tunneling process, and maintaining a good attitude is a prerequisite for high-quality tunnel construction. However, during TBM tunneling, the adjustment of shield tunneling parameters still relies on the operator's experience, which not only consumes human resources but also reduces the safety and stability of the TBM. Exploring autonomous TBM tunneling technology is an urgent requirement. It is worth noting that the implementation of autonomous TBM tunneling depends on the accurate prediction of TBM parameters, allowing for appropriate adjustments to the upcoming tunneling process based on the prediction results. Previous studies on TBM parameter prediction have relied on empirical and semi-empirical methods, typically based on project-specific data. Therefore, in the shield tunneling process of rail transit, it is necessary to adopt effective methods for adaptive attitude control optimization of the shield tunneling.

[0003] While deep learning-based TBM parameter prediction has achieved a certain level of accuracy, it still falls short of adaptive control, especially adaptive attitude control. On one hand, existing TBM parameter prediction methods often require a large number of features to form a training set, enabling the trained model to achieve higher accuracy. However, this contradicts engineering practice. During tunnel excavation, only a few active parameters (such as hydraulic cylinder pressure) can be modified; passive parameters are collected by various sensors and calculated after changes in active parameters. Although using passive parameters for TBM attitude prediction can improve accuracy, these passive parameters are usually not directly controllable during excavation to adjust the TBM attitude in a timely manner. Therefore, accurately controlling the TBM attitude using only a small number of operable active parameters is one of the challenges of autonomous tunneling technology. On the other hand, based on the predicted TBM parameters, only limited evaluation and improvement of the excavation process can be made. After establishing the prediction model, inputting different parameters will only yield corresponding results. To obtain the optimal combination of parameters, some research has begun to focus on multi-objective optimization of TBM parameters. Currently, commonly used optimization algorithms include Particle Swarm Optimization (PSO), evolutionary algorithms, and their improved versions. These optimization methods are based on known information to find the optimal combination of tunnel parameters within a limited range. While the results did prove effective, these offline optimization methods could not adapt to the constantly changing environment before it was even excavated. In other words, in complex underground environments, optimization methods could not achieve adaptive control. Furthermore, the actual TBM tunneling process is non-repeatable and cannot be canceled once completed. If the action strategy is not optimal, the resulting errors cannot be corrected. This is another significant challenge in achieving autonomous TBM tunneling.

[0004] Therefore, how to use intelligent algorithms to process tunnel boring machine (TBM) construction data to control the TBM's attitude and guide its autonomous, unmanned construction has become a crucial issue. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for adaptive control of tunnel boring machine (TBM) attitude under TBM construction conditions. Addressing the characteristics of traditional TBM tunneling processes, such as non-repeatability, poor control accuracy, and low level of intelligence, this invention achieves high-precision adaptive control of TBM tunneling attitude based on deep reinforcement learning methods, thereby overcoming the shortcomings of current TBM attitude prediction and optimization methods in engineering applications.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for adaptive control of the attitude of a tunnel boring machine under shield tunneling construction conditions is proposed, comprising the following steps:

[0007] S1 constructs a virtual construction environment for TBM attitude adaptive control;

[0008] S2 determines the state and reward of the reinforcement learning model during the TBM tunneling process, where the state of the TBM is updated according to the constitutive relation, and the reward is used for the forward speed and TBM attitude.

[0009] Based on the aforementioned virtual construction environment, status, and rewards, S3 constructs a DGDPG model for the shield tunneling process and optimizes and trains the DGDPG model.

[0010] S4 evaluates and describes the results of the TBM tunneling process based on the optimized DGDPG model.

[0011] As a further preferred option, step S1 specifically includes the following steps: by setting a reasonable soil constitutive model, the feedback data between the cutter and the soil can be obtained, the optimal combination of hydraulic cylinder thrust of the TBM can be calculated, and basic parameters such as the minimum acquisition interval of the sensor are considered, thereby constructing a virtual construction environment that can be used for TBM attitude adaptive control.

[0012] As a further preferred option, in step S2, updating the state of the TBM according to the constitutive relation specifically includes the following steps: First, based on the current pitch angle, the net travel and vertical deviation are obtained by advancing the tunnel. At the same time, the random settlement caused by the gravity of the TBM in the vertical direction is added, and the range of the settlement value is determined according to the project profile. Then, the pitch angle is updated according to the force difference between the upper and lower thrust cylinder groups.

[0013] As a further preferred embodiment, the computational model for updating the state of TBM according to the constitutive relation in step S2 is as follows:

[0014]

[0015] Where (x,y) and (x * ,y * ) represent the positions before and after the update, respectively, y s This indicates random settlement caused by the gravity of the TBM. θ and θ represent the force of each thrust cylinder during adaptive tunneling. * These represent the pitch angles before and after the update, respectively, and Δ(·) represents the difference in input parameters. This refers to the pressure of the upper hydraulic cylinder assembly. For the pressure of the lower cylinder assembly, k (e) is the formation elastic coefficient.

[0016] As a further preferred option, in step S2, the reward for forward speed and TBM attitude specifically includes: for forward speed, the TBM records the original tunnel data at equal time intervals, with the net travel distance of each record being stable at around 20mm; for TBM attitude, the deviation is expected to be 0 under ideal conditions, and during dynamic adjustment, the TBM state is restricted to a reliable range. When the TBM attitude triggers a boundary, a penalty value is given in the reward to restrict the TBM operation to an allowable range.

[0017] As a further preferred option, in step S2, the order of magnitude of the propulsion speed and the TBM attitude are balanced according to equations (2) and (3):

[0018]

[0019]

[0020] in, Th represents the force of each thrust cylinder during adaptive tunneling, ω represents the net stroke threshold for each tunneling step, and Status represents the coefficient matrix. (a) This represents the TBM attitude at the current step, where θ is the pitch angle before the update.

[0021] As a further preferred embodiment, in step S3, the DGDPG model includes an action module for interacting with the virtual underground environment, an evaluation module, and an experience storage module, wherein...

[0022] The action module output network first obtains the current TBM attitude and outputs the thrust cylinder force and the maximum Q value in the current state. Then, the thrust cylinder force and the current TBM attitude are input into the evaluation module output network to obtain a new Q value. Based on the new Q value, the action module output network is updated accordingly using the gradient ascent method. The thrust cylinder force from the action module output network will change the virtual underground environment, and the TBM attitude will be updated. The new TBM attitude is passed to the action module target network to obtain the thrust cylinder force at the next moment. The new attitude and the new thrust cylinder force are input into the evaluation module target network to obtain a new target Q value. This value is calculated together with the reward obtained from the virtual underground environment to obtain the corresponding time difference error. The evaluation module output network is updated by minimizing the error.

[0023] As a further preferred option, step S4 specifically includes: evaluating the TBM attitude and its stability using the mean and standard deviation; more specifically, using the area enclosed by the attitude curve and the design axis to represent the attitude error, and evaluating the extent of improvement of adaptive control relative to manual control based on this metric.

[0024] As a further preferred embodiment, the calculation model for the improvement magnitude is as follows:

[0025]

[0026] Among them, Area MC and Area AC These represent the areas enclosed by the attitude curve and the design axes for manual and adaptive control, respectively.

[0027] According to another aspect of the present invention, an adaptive control system for the attitude of a tunnel boring machine under tunneling conditions is also provided for implementing the above-described method.

[0028] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages:

[0029] 1. This invention proposes a method based on a deep gated recurrent neural network combined with a deterministic policy gradient method (DGDPG). By optimizing the range of key shield tunneling attitude control parameters such as vertical attitude deviation and pitch angle, the goal of adaptive attitude control optimization in shield tunneling engineering is achieved.

[0030] 2. This invention, using real-world case data samples, employs a reinforcement learning-based improved DGDPG algorithm to establish a nonlinear mapping relationship between the geological environment interaction during shield tunneling construction and the shield's attitude parameters, thereby improving the accuracy of adaptive attitude control optimization for shield tunneling projects. Compared to manual control, adaptive attitude control shows significant improvements in both accuracy and stability. The adaptive strategy increases rewards by 93.31% and stability by 90.72%. The improvement in pitch angle control is mainly reflected in the increased control stability, with an increment of 90.21%. Pitch angle control accuracy is improved by 60.14%. Therefore, the DGDPG method exhibits good soil adaptability and adaptive attitude control accuracy and stability. This invention provides a new approach and implementation method for intelligent, autonomous, and unmanned shield tunneling construction. Attached Figure Description

[0031] Figure 1 This is a flowchart of a preferred embodiment of the present invention regarding an adaptive control method for the attitude of a tunnel boring machine under tunneling conditions;

[0032] Figure 2 This is a schematic diagram of the overall structure of the DGDPG model according to a preferred embodiment of the present invention;

[0033] Figure 3 This is a detailed network structure diagram of the action module and evaluation module according to a preferred embodiment of the present invention;

[0034] Figure 4 (a) in the diagram is a schematic of the average reward value for each generation during the training process; Figure 4(b) in the diagram is a schematic of the total reward value for each generation during the training process;

[0035] Figure 5 (a) in the diagram is a schematic diagram showing the change in the net tunneling distance of the tunnel boring machine's posture during the training process. Figure 5 (b) in the diagram is a schematic diagram of the pitch angle change of the tunnel boring machine during the training process. Figure 5 (c) in the diagram is a schematic diagram of the vertical deviation of the tunnel boring machine's attitude during the training process;

[0036] Figure 6 It is a cumulative distribution curve of the net travel distance for each segment during the excavation of a single tunnel.

[0037] Figure 7 (a) in the figure is a comparison of the pitch angle results of manual control and adaptive control in a ring shield tunnel. Figure 7 (b) in the figure is a comparison of the vertical deviation results of manual control and adaptive control in a ring shield tunnel. Figure 7 (c) in the figure is a comparison of the reward values ​​for each step of manual control and adaptive control in the ring shield tunneling machine;

[0038] Figure 8 (a) in the figure is a comparison of the vertical deviation of the ring shield tunnel trajectory under manual control and adaptive control. Figure 8 (b) in the figure is a comparison of the pitch angles of the ring shield trajectory under manual control and adaptive control. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0040] like Figure 1 As shown, this embodiment provides an adaptive attitude control method for shield tunneling construction based on DGDPG, including the following steps:

[0041] Step 1: Construct a virtual tunnel boring machine (TBM) construction environment and quantitatively obtain the correlation between the cutter and the soil, and between the environment and the thrust. At the same time, assume the basic parameters of key components such as soil and data acquisition sensors based on actual working conditions to ensure that the parameterized virtual environment information can be accurately used for adaptive control.

[0042] Step 2: Determine the state and reward of the reinforcement learning model during the TBM tunneling process, specifically including:

[0043] Step 2.1: The state of TBM is updated according to the constitutive relation.

[0044] First, based on the current pitch angle, the net travel and vertical deviation are obtained by advancing the tunnel. To make the tunneling process similar to a real underground construction scenario, this example adds random settlement due to the vertical gravity of the TBM; the range of settlement values ​​is determined based on the project profile. Random settlement is applied to the vertical direction for each tunneling step. Then, the pitch angle is updated based on the force difference between the upper and lower thrust cylinder groups. The specific calculation process is shown in Equation (1).

[0045]

[0046] Where (x,y) and (x * ,y * The ') indicates the position before and after the update. y s This indicates random settlement caused by the gravity of the TBM. This represents the force of each thrust cylinder during adaptive tunneling. θ and θ' * These represent the pitch angles before and after the update. Δ(·) represents the difference in input parameters.

[0047] Step 2.2: Determining the reward for the reinforcement learning model during TBM tunneling.

[0048] The reward calculation is divided into two parts: one for advance speed and the other for TBM attitude. For advance speed, the TBM records raw tunnel data at equal time intervals, with each recorded net travel distance consistently around 20mm. Excessively high advance speeds can cause the attitude to deteriorate rapidly within a time step, while excessively low advance speeds can lead to rapid settlement accumulation in the surrounding area, which is also detrimental to TBM construction. Therefore, in this example, the target travel distance for the intelligent agent within a time step is 20mm.

[0049] Ideally, the deviation from the TBM attitude should be zero. However, in underground environments, settlement and elasticity coefficients change, and the TBM attitude is constantly undergoing dynamic adjustment. During this dynamic adjustment, the TBM state should be constrained within a reliable range to improve the convergence speed of the training process and provide attitude warnings in engineering applications. When the TBM attitude triggers a boundary, a penalty value is given in the reward to limit the TBM operation within the allowable range. Specific boundaries and penalty values ​​should be set according to construction specifications. In this example, [-20°, 20°] and [-50mm, 50mm] are used as the boundaries for the pitch angle and vertical deviation, respectively, and -10° is used as the boundary. 9 As a penalty value.

[0050] Furthermore, different construction projects place varying degrees of emphasis on vertical deviation and pitch angle. This example adds a state weighting coefficient to the reward function to control the reward value for different projects. This example calculates the reward by balancing the order of magnitude of the advance speed and TBM attitude according to equations (2) and (3). In the ideal state, the advance speed is 20 mm / step, the TBM attitude deviation is 0, and the reward reaches its maximum value of 0. In other cases, the reward is less than 0.

[0051]

[0052]

[0053] Action (a) This represents the thrust cylinder force used in adaptive control during tunneling. Th represents the net travel threshold for each tunneling step. ω represents the coefficient matrix, which is 0.1 for pitch angle and 1 for vertical deviation. (a) This indicates the TBM pose at the current step.

[0054] Step 3: Construction and training of the DGDPG model for the shield tunneling construction process, specifically including:

[0055] Step 3.1, the basic structure of the DGDPG model.

[0056] The DGDPG model consists of three core components: an action module, an evaluation module, and an experience storage module. These three modules interact with the virtual underground environment to adaptively control the tunnel excavation process and achieve optimal excavation posture. The overall structure of the DGDPG model is as follows: Figure 2 As shown.

[0057] Both the action module and the evaluation module contain two identical networks: an output network and a target network. The action module's output network first obtains the current TBM attitude and outputs the thrust cylinder force and the maximum Q-value in the current state. Then, the thrust cylinder force and the current TBM attitude are input into the evaluation module's output network to obtain a new Q-value. Based on the new Q-value, the action module's output network is updated using gradient ascent. The thrust cylinder force from the action module's output network changes the virtual underground environment, and the TBM attitude is updated accordingly. Passing the new TBM attitude to the action module's target network yields the thrust cylinder force for the next time step. The TBM attitude and the new thrust cylinder force are input into the evaluation module's target network to obtain a new target Q value. This value is calculated together with the reward obtained from the virtual underground environment to obtain the corresponding time difference error, and the evaluation module's output network is updated by minimizing the error. At this point, both the action module's target network and the evaluation module's target network undergo soft updates. In the above process, the TBM attitude of the current step, the thrust cylinder force of the current step, the TBM attitude of the next step, and the obtained reward are all stored in the experience storage module. When the model needs to be updated, it is passed back to the corresponding model by randomly sampling from the experience storage module. The detailed network structure description of the action module and the evaluation module is as follows: Figure 3 As shown.

[0058] The action module network's role is to output the action with the highest expected value based on the input TBM state. Therefore, its input layer accepts two variables: pitch angle and vertical deviation. Due to the stochastic nature of the underground environment, a batch normalization layer is introduced to process the input data, preventing extreme cases in the state data from causing the network to become insensitive to other data. Furthermore, batch normalization can significantly accelerate deep networks, which can improve the convergence speed of DGDPG and benefit its application in TBM construction. In this invention, the batch normalization calculation process is as follows:

[0059] Layer input: {x 1,2,…,n =(θ,Δy) 1,2,…,n}

[0060] Initialize trainable parameters: α, β

[0061] The first step is to calculate the average value of the small batch:

[0062] The second step is to calculate the variance of the small batch:

[0063] The third step is to standardize small batches: Here, ∈ represents a very small value to prevent the variance from being 0.

[0064] Layer output: {y 1,2,…,n=(F top ,F down ) 1,2,…,n}, y=αx * +β

[0065] In this invention, the nonlinear expression of the model is enhanced by five fully connected layers, each containing 600, 500, 400, 300, and 2. It is particularly important to note that the TBM parameters need to be adjusted according to the pressures of the upper and lower cylinders; therefore, the last layer contains two parameters. Assuming the final output value is the actual cylinder pressure, the result of the dense layer is then normalized to the interval [0,1] using the sigmoid function (Equation (4)), and the result is reduced to the actual interval range of the output layer using Equation (5).

[0066]

[0067] t real =Sigmoid(σ)*(t) u -t l )+t l (5)

[0068] After obtaining the TBM action corresponding to the current state, the evaluation module network evaluates the action module network's result by incorporating the impact of the current action on the environment. The input to the evaluation module network is the current thrust output by the action network and the thrust in the new environment after responding to this action. Considering that the thrust of each cylinder is not independent but mutually influential during adjustment, three GRU layers are added to the critical network to compute the feature sequence. The number of neurons in the three GRU layers are 500, 400, and 300, respectively. Finally, two dense layers are added to improve the nonlinear expression, and the Q-value is finally output, completing the construction of the DGDPG model.

[0069] Step 3.2, Training the DGDPG model.

[0070] After constructing the virtual underground environment and determining the relevant states and rewards during the tunnel excavation process, the DGDPG model was trained and applied under the aforementioned conditions. The DGDPG hyperparameters are shown in Table 1, as are detailed information about the training process.

[0071] Table 1. Hyperparameters of the DGDPG training process

[0072]

[0073] During training, the DGDPG agent gains probing ability by adding Gaussian noise to the output actions. The Gaussian noise is centered on the output action, with an initial standard deviation of 1000. To facilitate model convergence, the standard deviation is halved every 5000 sets. When the agent terminates an event and reaches an endpoint or satisfies boundary constraints, it restarts the next event until the model converges. In this case, only two strict constraints are set: one for the pitch angle (-20°, 20°) and the other for the vertical deviation (-50mm, 50mm). When the agent's attitude during tunneling exceeds the constraints, the current tunneling process stops, and a larger penalty is applied to the corresponding reward; in this example, -10 is chosen. 9 .

[0074] The average reward for each step and the total reward for each generation are as follows: Figure 4 As shown. In this example, the confidence interval (CI) is set to the mean ± 3std. Since the standard deviation of Gaussian noise is large in the early stages, it can be assumed that the intelligent agent is in a stochastic exploration phase and struggles to complete a single tunnel excavation. The reward is -10 for steps exceeding the limit. 9 The penalty value is calculated, which is much smaller than the reward before the penalty, so the reward for the early segments is distributed in -10. 9 Around 5000 tunnel excavations. While the reward remained relatively stable over a long period, the number of steps in each segment varied significantly, resulting in large differences in the average reward per step. This continued until the exploration range was halved for the first time, at which point the average reward increased over a small range before decreasing to a very small range to plateau. This situation was considered trapped in a local optimum, and the intelligent agent behaved significantly differently from its expected behavior. This continued until the exploration range was halved for the third time, at which point it was 1 / 8 of its initial value. After more than 15000 tunnel excavations, the intelligent agent sampled and learned enough steps to finally break free of the local optimum. The total reward value no longer approached -10. 9 The average reward increases significantly. Afterward, neither the total nor the average reward changes significantly during the tunnel excavation process, and the training process is considered convergent. It should be noted that the convergence of the training process cannot reach the theoretical state where constraint penalties are not triggered at all, because this example simulates the actual process of TBM gravity settlement by consistently adding vertical random negative values ​​and Gaussian noise.

[0075] Changes in TBM posture during training, such as Figure 5As shown, the overall trend is consistent with the average reward trend. It should also be noted that this example marks the cumulative deviation range of manual excavation, with the green dashed line representing the upper limit and the red dashed line representing the lower limit. Since the vertical deviation is weighted 10 times the pitch angle in the reward function, the intelligent agent will prioritize reducing vertical deviation, which is inferred to be a significant reason why the intelligent agent gets stuck in local optima. After episode 15000, the intelligent agent experiences a small phase with large vertical deviation and small pitch angle, before eventually converging to a global optimum with a pitch angle slightly greater than 0 and both pitch angle and vertical deviation being small.

[0076] Figure 6 The cumulative distribution curve of the net travel for each set during a single tunnel excavation is shown. The x-axis represents the net travel, and the y-axis represents the cumulative count of the corresponding net travel. In this example, the training set was divided into five groups to demonstrate convergence. The second group performed the worst, starting slightly greater than 0, rapidly increasing around 100 mm, and then leveling off around 400 mm. Very few sets reached 1400 mm. The starting points of the other four groups gradually shifted backward; at positions with smaller net travel, the cumulative count decreased, and the number of sets capable of completing a 1400 mm tunnel excavation increased.

[0077] Step 4: Evaluation and Result Description of TBM Tunneling Process Based on DGDPG Model, specifically including:

[0078] After obtaining the actual DGDPG model of the optimal tunneling strategy, the results of actions taken according to the strategy need to be evaluated using a uniform standard. In this example, the TBM attitude and its stability are evaluated using the mean and standard deviation. It is important to note that the average value of the original attitude offset does not reflect its variation pattern and is small when the attitude value is distributed above and below 0. Therefore, this example suggests using the area enclosed by the attitude curve and the design axis to represent the attitude error and to evaluate the magnitude of improvement (IM) of adaptive control relative to manual control based on this metric, as shown in Equation (6).

[0079]

[0080]

[0081] Area MC and Area AC These represent the areas enclosed by the attitude curve and the design axes for manual and adaptive control, respectively. Att i Indicates the i-th th The posture of stepping. Δstep i Indicates the steps to maintain the posture.

[0082] Applying the corresponding evaluation methods to the trained DGDPG model reveals that, compared to manual attitude control, DGDPG-based adaptive TBM attitude control significantly improves the accuracy and stability of the tunneling process. After determining the applicable underground environment for the trained DGDPG model, this example selected 10 manual tunneling intervals as a comparative experimental scenario. The comparative experimental results are as follows: Figure 7 As shown. The DGDPG model concentrates the pitch angle at around 4.42°, while the manually controlled pitch angle is more dispersed, fluctuating from 7.9° to 13.3°. Both control methods allow the TBM to maintain a "lifted" state during tunneling to adapt to settlement. Although most of the vertical deviation data in the manually controlled model is near 0, a large number of data are also near the maximum value, indicating a lack of finesse in manual control, specifically manifested in delayed operations and excessive attitude adjustments. In contrast, the vertical deviation and reward of the adaptive control are more concentrated and stable, ranging from -10.92mm to 7.93mm. The tunneling trajectory is as follows... Figure 8 As shown, the solid line is the vertical trajectory of the TBM, which can be considered as the central axis of the tunnel formed in the linear tunneling interval. The dashed line represents the pitch angle at each position, and the background color represents the elastic coefficient for each 20mm interval. The action strategy of the adaptive control process is usually consistent with the manual strategy, with the difference appearing in the later stages of the tunneling process. At this time, manual control cannot correct the increasing trend of vertical deviation in the "eye-level" state, while the adaptive strategy can keep the vertical deviation of the TBM at around 0. In addition, in the early stage, due to the larger pitch angle of manual control, the resistance to TBM settlement is better. The comprehensive comparison results of the entire tunneling process are shown in Table 2. Table 2.1 Comparison of manual control and adaptive control results in 10 shield tunnel sections.

[0083] In summary, compared with manual control, adaptive attitude control has significant advantages in both accuracy and stability.

[0084]

[0085] Significant improvements were achieved. The adaptive strategy improved reward by 93.31% and stability by 90.72%. The improvement in pitch angle control was mainly reflected in increased control stability, with an increment of 90.21%. Pitch angle control accuracy improved by 60.14%. While the improvement in vertical deviation control was not as good as the reward improvement, it still achieved a 76.52% improvement in accuracy and a 62.13% improvement in stability. Therefore, based on a series of results, it can be concluded that the DGDPG method has good soil adaptability and adaptive attitude control accuracy and stability.

[0086] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for adaptive attitude control of a tunnel boring machine under shield tunneling construction conditions, characterized in that, Includes the following steps: S1 constructs a virtual construction environment for TBM attitude adaptive control; S2 determines the state and reward of the reinforcement learning model during the TBM tunneling process, where the state of the TBM is updated according to the constitutive relation, and the reward is used for the forward speed and TBM attitude. In step S2, the computational model for updating the state of TBM according to the constitutive relation is as follows: Where (x,y) and (x * ,y * ) represent the positions before and after the update, respectively, y s This indicates random settlement caused by the gravity of the TBM. θ and θ represent the force of each thrust cylinder during adaptive tunneling. * These represent the pitch angles before and after the update, respectively, and Δ(·) represents the difference in input parameters. This refers to the pressure of the upper hydraulic cylinder assembly. For the pressure of the lower cylinder assembly, k (e) The elastic coefficient of the formation; In step S2, the order of magnitude of the propulsion speed and TBM attitude are balanced according to equations (2) and (3): in, Th represents the force of each thrust cylinder during adaptive tunneling, ω represents the net stroke threshold for each tunneling step, and Status represents the coefficient matrix. (a) This indicates the TBM attitude at the current step, where θ is the pitch angle before the update. Based on the aforementioned virtual construction environment, status, and rewards, S3 constructs a DGDPG model for the shield tunneling process and optimizes and trains the DGDPG model. S4 evaluates the TBM tunneling process based on the optimized DGDPG model.

2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: by setting a reasonable soil constitutive model, obtaining feedback data between the cutter and the soil, calculating the optimal combination of hydraulic cylinder thrust for the TBM, and considering the minimum acquisition interval of the sensors, a virtual construction environment for TBM attitude adaptive control is constructed.

3. The method according to claim 1, characterized in that, In step S2, updating the state of the TBM according to the constitutive relation specifically includes the following steps: First, based on the current pitch angle, the net travel and vertical deviation are obtained by advancing the tunnel. At the same time, the random settlement caused by the TBM's gravity in the vertical direction is added, and the range of the settlement value is determined according to the project profile. Then, the pitch angle is updated based on the force difference between the upper and lower thrust cylinder groups.

4. The method according to claim 1, characterized in that, In step S2, the rewards for forward speed and TBM attitude specifically include: for forward speed, the TBM records the original tunnel data at equal time intervals, with the net travel distance of each record being stable at around 20mm; for TBM attitude, the deviation is expected to be 0 under ideal conditions. During dynamic adjustment, the TBM state is restricted to a reliable range. When the TBM attitude triggers a boundary, a penalty value is given in the rewards to restrict the TBM operation to an allowable range.

5. The method according to any one of claims 1-4, characterized in that, In step S3, the DGDPG model includes an action module for interacting with the virtual underground environment, an evaluation module, and an experience storage module, wherein... The action module output network first obtains the current TBM attitude and outputs the thrust cylinder force and the maximum Q value in the current state. Then, the thrust cylinder force and the current TBM attitude are input into the evaluation module output network to obtain a new Q value. Based on the new Q value, the action module output network is updated accordingly using the gradient ascent method. The thrust cylinder force from the action module output network will change the virtual underground environment, and the TBM attitude will be updated. The new TBM attitude is passed to the action module target network to obtain the thrust cylinder force at the next moment. The new attitude and the new thrust cylinder force are input into the evaluation module target network to obtain a new target Q value. This value is calculated together with the reward obtained from the virtual underground environment to obtain the corresponding time difference error. The evaluation module output network is updated by minimizing the error.

6. The method according to claim 5, characterized in that, Step S4 specifically includes: evaluating the TBM attitude and its stability using the mean and standard deviation; more specifically, using the area enclosed by the attitude curve and the design axis to represent the attitude error, and evaluating the improvement of adaptive control over manual control based on this metric.

7. The method according to claim 6, characterized in that, The calculation model for the magnitude of the improvement is as follows: Among them, Area MC and Area AC These represent the areas enclosed by the attitude curve and the design axes for manual and adaptive control, respectively.

8. A shield machine attitude adaptive control system under shield tunneling construction conditions, characterized in that, Used to implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method for predicting soil displacement caused by shield tunneling in soil-rock composite stratum

    CN114961751A

  • Deviation rectification control method and apparatus for shield tunneling attitude

    WO2022179266A1