Shield tunneling machine earth pressure balance autonomous control method based on warm start DDPG

Through the temperature-start DDPG method, combined with data pre-training and reinforcement learning, the soil pressure balance control of the shield machine is optimized, which solves the real-time interaction problem of the shield machine when soil quality changes, improves control accuracy and adaptability, and reduces training costs and trial and error times.

CN120507972APending Publication Date: 2025-08-19NORTHEASTERN UNIV CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510605222.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The deep learning model of existing shield machines is difficult to achieve real-time interaction when soil quality changes, resulting in slow decision making, affecting safety and adaptability. In addition, the training of reinforcement learning agents takes a long time, many trial and error times, and poor control accuracy.

Method used

The shield machine soil pressure balance autonomous control method based on temperature-start DDPG is adopted, and key features and rules are learned through data temperature-start, combined with whale optimization algorithm to optimize the Actor network hyperparameters, and interact with the actual construction environment in reinforcement learning, optimize control strategies, and build a closed-loop control chain of state perception-decision optimization-execution feedback.

Benefits of technology

It significantly reduces the cost of invalid trial and error in the early stage of the training of the agent, improves the response speed and accuracy of decision-making, and realizes intelligent decision-making and stable control of shield machine soil pressure balance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120507972A_ABST
    Figure CN120507972A_ABST
Patent Text Reader

Abstract

The invention provides a shield tunneling machine earth pressure balance autonomous control method based on warm start DDPG, and relates to the technical field of shield tunneling machines. Comprising data warm start and reinforcement learning training; warm start is carried out on the Actor network based on a large amount of construction data, and hidden key features and rules in the data are learned; on the basis of warm start, a deep reinforcement learning method is introduced, a control strategy is continuously optimized through interaction with an actual construction environment, and finally intelligent decision-making of earth pressure balance of the shield tunneling machine is achieved. Through data pre-training, the Actor network can preliminarily adapt to a complex construction environment, the number of trial and error times is reduced, and the invalid trial and error cost at the initial stage of intelligent agent training is remarkably reduced. And reinforcement learning training forms a closed-loop control chain of state perception-decision optimization-execution feedback through real-time interaction with a dynamic construction environment, so that not only is gradual fine adjustment of a control strategy in the tunneling process realized, but also the decision response speed and precision are improved, and finally, intelligent decision making of earth pressure balance of the shield tunneling machine is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of shield machines, and in particular to a shield machine earth pressure balance autonomous control method based on a warm start DDPG (Deep Deterministic Policy Gradient) algorithm. Background Art

[0002] With the rise of artificial intelligence algorithms, improving the intelligence of shield machines has become a research hotspot in the shield tunneling field. Furthermore, to address surface subsidence caused by pressure imbalance, various data-driven optimization and control methods have been adopted, significantly improving tunneling safety. Existing models often rely on large amounts of data for training, but the timeliness of construction data limits their real-time performance. Furthermore, due to the complexity and variability of the tunneling environment, deep learning models struggle to interact in real time with soil changes, making them unable to adapt to environmental changes. This limitation is particularly prominent in safe tunneling, especially when encountering sudden geological changes, where the model's slow response can affect decision-making accuracy. Therefore, while deep learning has shown promise in autonomous intelligent decision-making, its limitations in real-time performance and adaptability still need to be addressed.

[0003] Deep reinforcement learning effectively overcomes the limitations of deep learning in intelligent decision-making. With its powerful self-learning and environmental perception capabilities, it has been widely adopted in the field of intelligent decision-making. Deep reinforcement learning has achieved significant application progress in various aspects of shield machines. Pin Zhan et al. proposed a hybrid model combining DQN-PSO and ELM, optimizing the weights of the ELM using DQN-PSO to improve prediction accuracy. Ya-kun ZHANG et al. integrated mechanical models and construction data to establish a hybrid model, promoting the partial autonomy of shield machines. X. Liu et al. established a hierarchical control model based on reinforcement learning and a CPS system, improving control accuracy. Soranzo Enrico et al. used a DQN model to determine tunnel face support pressure to reduce subsequent settlement. J. Xu et al. used multiple reinforcement learning agents to correct shield machine posture deviations. Elbaz Khalid et al. used DQN and ELM models to predict thrust and cutterhead torque of shield machines. Xu J. et al. integrated iterative deepening with a soft actor-critic algorithm to construct a posture correction model that achieves adaptive posture adjustment. X. Liu et al., based on DDPG, implemented autonomous adjustment of shield machine delays and multiple tunneling parameters, enhancing the autonomous capabilities of an earth pressure balance shield machine. These methods improved the autonomous capabilities of the shield machine to varying degrees, but they also suffered from issues such as lengthy training of the reinforcement learning agent, a large search space, and numerous trial-and-error cycles. Furthermore, some reinforcement learning agents exhibited poor performance, control accuracy, and an inability to adapt to continuous action space tasks, limiting their autonomous control. A single reinforcement learning agent requires extensive trial-and-error to explore the optimal strategy, making learning efficiency and timeliness key considerations. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a shield machine earth pressure balance autonomous control method based on warm start DDPG. Under the premise of ensuring control accuracy, it can realize adaptive adjustment of the sealed cabin pressure, maintain the stable state of the system, and avoid the influence of small disturbances.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A warm-start DDPG-based autonomous control method for shield machine earth pressure balance (EPB) consists of two parts: data warm-start and reinforcement learning training. First, the actor network is warm-started based on a large amount of construction data to learn the key features and patterns hidden in the data. Then, based on the warm-start, a deep reinforcement learning method is introduced to continuously optimize the control strategy through interaction with the actual construction environment, ultimately achieving intelligent decision-making for the shield machine's EPB.

[0007] During the data warm start, a nonlinear relationship model between the tunneling parameters of the shield machine and the sealed chamber pressure during excavation is first established, and the model is decomposed into a multi-objective dynamic decision-making problem. At the same time, the input and output of the Actor network are determined based on the operating characteristics of the shield machine and the control task, and the corresponding tunneling parameters are selected to provide dependent data support for subsequent training. Secondly, the three hyperparameters of the Actor network (number of iterations, learning rate, and number of neurons) are optimized based on the whale optimization algorithm (WOA), and the Actor network is trained based on the optimal hyperparameters. Finally, the decision-making effect of the Actor network is verified based on construction data under different geological conditions.

[0008] In the reinforcement learning training, first, the state space, action space, and reward function of the DDPG agent are defined based on the working characteristics of the shield machine and the control task characteristics; second, the DDPG agent is designed and relevant hyperparameters are selected; finally, the agent is trained through interaction with the virtual environment, realizing intelligent decision-making for the shield machine earth pressure balance based on the pre-trained deep reinforcement learning agent.

[0009] Furthermore, during the data warm start, the process of establishing the nonlinear relationship model between the tunneling parameters of the shield machine during excavation and the sealed cabin pressure is as follows:

[0010] When the shield machine advances forward at a certain speed, the cutterhead rotates to cut the soil in front of the excavation face, breaking it into debris and entering the sealed cabin. The amount of soil entering the sealed cabin per unit time, Q1, is shown in formula (1). As the soil in the sealed cabin accumulates, the pressure in the cabin gradually increases. At this time, the soil is discharged from the sealed cabin through the rotation of the screw conveyor to adjust the cabin pressure. The amount of soil discharged per unit time, Q2, by the screw conveyor is shown in formula (2).

[0011] Q1=πR 2 V (1)

[0012] Q2=ηπATn0 (2)

[0013] Where R is the cutter head radius; V is the propulsion speed; η is the soil discharge efficiency; A is the effective cross-sectional area of the screw conveyor; T is the blade pitch; n0 is the screw conveyor speed;

[0014] Based on the principle of balance between the amount of slag in and out, the continuity equation of the slag in the sealed cabin is derived, as shown in formula (3); the nonlinear relationship between the sealed cabin pressure and the propulsion speed and the screw conveyor speed is further mapped into formula (4).

[0015]

[0016]

[0017] Among them, C ep is the external leakage coefficient of the sealed cabin; P is the sealed cabin pressure; is the first-order derivative of P; P0 is the leakage pressure of the sealed cabin; V e is the sealed cabin capacity; β e is the compressibility coefficient of the material in the cabin;

[0018] Based on the Laplace transform result of formula (4) and the characteristics of intelligent decision-making, the input of the Actor network is set as: real-time sealed cabin pressure, optimal sealed cabin pressure under current soil conditions, and the absolute value of the error between the real-time pressure and the optimal pressure; the output of the Actor network is set as: propulsion speed and screw conveyor speed.

[0019] Furthermore, in the data warm start, the corresponding excavation parameters are screened as follows:

[0020] Actual construction data was collected and filtered according to the input and output of the Actor network. A total of 20,000 sets of data were screened to form a pre-training dataset. This dataset covers the construction status and transition process of nine types of soil, including excavation parameters such as sealing chamber soil pressure, optimal sealing chamber soil pressure, sealing chamber pressure error, screw conveyor speed, and propulsion speed.

[0021] Furthermore, in the data warm start, the Actor network is a neural network based on a multi-layer perceptron structure, including an input layer, a first hidden layer, a second hidden layer and an output layer.

[0022] Furthermore, during the data warm start, the hyperparameter optimization process is performed using a global optimization method based on the whale optimization algorithm. The specific process is as follows:

[0023] Step 1.1: Initialize the number of whales M and the maximum number of iterations T;

[0024] Step 1.2: Initialize the population: X i ,i=1,2,...,M;M represents the population size;set X * is the current optimal position;

[0025] Step 1.3: Check if there are whales exceeding the search space and make modifications;

[0026] Step 1.4: Calculate the fitness value of each position and find the position vector X of the current best individual * ;

[0027] Step 1.5: Determine whether the predation probability p is less than the set threshold; if so, proceed directly to step 1.6; otherwise, enter the bubble net predation stage, based on the formula D = |X * -X(t)| and X(t+1) = D·ebl ·cos(2πl) + X * Update the individual position information; where D represents the distance between the current search individual and the current best individual, X(t) is the position vector of the current individual in the bubble net predation stage, X(t + 1) represents the position vector of the best individual in the next iteration after update, t is the number of iterations, X * is the position vector of the current best individual, l is a random number uniformly distributed in the interval [-1, 1]; b is a constant; e bl is a parameter that changes with the number of iterations and is used to adjust the search range and convergence;

[0028] Step 1.6: Determine whether the coefficient vector |B| is less than 1; if so, enter the stage of hunting prey, and use D′ = |C·X * - x′(t)| and X′(t + 1) = X * - B·D′ to update the position information; otherwise, enter the stage of searching for prey, based on D″ = |C·X rand - X″(t)| and X″(t + 1) = X rand - B·D″ to update the position information; where D′ is the distance between the current individual and the best individual in the stage of hunting prey, X′(t) is the position vector of the current individual in the stage of hunting prey, X′(t + 1) is the position vector of the best individual in the next iteration in the stage of hunting prey, B and C represent coefficient vectors; X″(t) is the position vector of the current individual in the stage of searching for prey, X″(t + 1) represents the position vector of the best individual in the next iteration in the stage of searching for prey, C″ is the relationship between the current search individual and the random individual in the stage of searching for prey, X rand is the position vector of the random individual in the stage of searching for prey;

[0029] Step 1.7: After the position information update is completed, calculate the fitness value and compare it with the fitness value of the initial best position information, and select the best individual position with a smaller fitness value;

[0030] Step 1.8: Increment the number of iterations t by 1. If t < T, return to Step 1.3; otherwise, end the optimization process and output the best hyperparameters of the Actor network.

[0031] Furthermore, the state space, action space, and reward function of the DDPG agent are defined as follows:

[0032] State space: It is the input source for the agent to perceive the environment and make decisions, and is used to represent the environmental state during the construction of the shield machine to ensure that the DDPG agent can timely perceive the dynamic changes of the construction environment; the defined state variables are shown in formula (5):

[0033] S t = [P, P′, P″] (5)

[0034] Among them, P, P′, and P″ are the real-time sealed cabin pressure, the optimal sealed cabin pressure, and the absolute value of the error between the real-time pressure and the optimal pressure, respectively;

[0035] Action space: The action space is defined as shown in formula (6):

[0036] a t =[V,n0] (6)

[0037] Among them, n0 and V are the screw conveyor speed and propulsion speed respectively;

[0038] Reward function: It uses a combination of positive rewards, negative rewards, and boundary rewards. The specific reward function is as follows:

[0039]

[0040] Positive rewards are used to encourage excellent performance of the intelligent agent. When the sealed cabin pressure is consistent with the optimal value, positive rewards are given to enhance its goal-orientedness towards the correct strategy. Negative rewards are used to help the intelligent agent identify adverse behaviors. When the pressure deviates significantly from the optimal value, negative rewards are given to correct invalid operations in a timely manner. Boundary rewards are used to accelerate the intelligent agent's avoidance of actions that do not meet the requirements. When the sealed cabin pressure reaches a dangerous value, boundary rewards are given, thereby improving the stability of training and avoiding operational errors.

[0041] Furthermore, the DDPG agent adopts an Actor-Critic architecture, which is composed of an Actor network and a Critic network; the Actor network is responsible for generating the optimal strategy, and the Critic network performs value evaluation on the actions taken; the Critic network includes an input layer, a hidden layer 1, a hidden layer 2, and an output layer; the Critic network performs value evaluation based on the current state and action combination, and outputs a corresponding value function.

[0042] Furthermore, the training process of the agent is as follows:

[0043] First, the weights of the Actor network and the Critic network are initialized, and an experience replay buffer is created to store the agent's interaction experience. At each time step, the agent selects an action based on the current state and applies the action to the environment, receiving a reward and the next state. At this point, the agent stores the experience in the experience replay buffer. The stored experience includes state, action, reward, and next state.

[0044] Secondly, small batches of samples are randomly sampled from the buffer for training. The Critic network uses the Bellman equation to calculate the target Q value and updates its weights by minimizing the mean squared error loss function. The Actor network uses the policy gradient method to update based on the Critic's Q value to optimize its policy. In addition, the Target network used to supervise the training of the Actor network periodically adopts a soft update strategy to ensure the stability of the training process.

[0045] The entire training process is repeated until the agent's performance reaches the expected level;

[0046] The specific training process is as follows:

[0047] Step 2.1: Initialize the sealed cabin earth pressure environment model and the Actor network after warm start;

[0048] Random noise is represented by ξ, and random network parameters θ are used μ and θ Q Initialize the Actor network μ(s|θ μ ) and Critic network Q(s|θ Q );

[0049] Copy the parameter values θ of the Actor network separately μ →θ μ′ and the parameter value θ of the Critic network Q →θ Q′ , initialize the Target Actor network μ′(s|θ μ′ ) and Target Critic network Q′(s|θ Q′ ), where θ μ ,θ Q are the parameters of the Actor network and the Critic network respectively; θ μ′ ,θ Q′ These are the parameters of the Target Actor network and the Target Critic network respectively;

[0050] Step 2.2: Initialize the entire experience replay pool R;

[0051] Step 2.4: Initialize random noise ξ for motion exploration; obtain the initial state s1 from the sealed cabin soil pressure environment;

[0052] Step 2.5: For each time step t, perform the following operations:

[0053] Select action a based on the model's current strategy and noise t =μ(s|θ μ )+ξ, where μ(s|θ μ) represents the strategy output by the Actor network;

[0054] Execute action a t Observe the error between the sealed cabin soil pressure value and the standard pressure value, and obtain the reward r t , the environment state becomes s t+1 ;

[0055] (s t ,a t ,r t ,s t+1 ) Stored in the experience replay pool R;

[0056] Randomly sampling N tuples {(s i ,a i ,r i ,s′ i )} i=1,…,N ;

[0057] For each tuple, use the Target network to calculate y = r i +γQ′(s i+1 ,a i+1 ,θ Q′ ); where y is the cumulative expected value, γ is the discount factor, Q′(s i+1 ,a i+1 ,θ Q′ ) is the evaluation value of the next state action;

[0058] Minimize the target loss: To update the current Critic network;

[0059] Calculate the sampled policy gradient as follows and update the current Actor network:

[0060]

[0061] in, is the gradient of the policy objective function; is the gradient of the Critic network to the action; Output the gradient of the policy to the parameters for the Actor;

[0062] Update the Target network as follows:

[0063] θ μ′ =τθ μ +(1-τ)θ μ′ ;

[0064]

[0065] Among them, τ represents the learning rate, and its value range is (0,1);

[0066] Step 2.6: Repeat steps 2.3 to 2.5 until the agent's cumulative reward is maximized, that is, the sealed cabin pressure value output by the agent strategy converges to the set pressure reference value.

[0067] The beneficial effects of adopting the above technical solution are as follows: the shield machine earth pressure balance autonomous control method based on warm start DDPG provided by the present invention introduces a deep reinforcement learning method on the basis of warm start of the Actor network, and continuously optimizes the control strategy through interaction with the actual construction environment. Data pre-training enables the Actor network to initially adapt to the complex construction environment, reduce the number of trial and error, and significantly reduce the cost of ineffective trial and error in the early stage of intelligent agent training. Reinforcement learning training forms a closed-loop control chain of state perception-decision optimization-execution feedback through real-time interaction with the dynamic construction environment, which not only realizes the gradual and refined adjustment of the control strategy during the excavation process, but also improves the response speed and accuracy of the decision, and ultimately realizes intelligent decision-making on the earth pressure balance of the shield machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 A schematic diagram of a shield machine earth pressure balance autonomous control method based on warm start DDPG provided by an embodiment of the present invention;

[0069] Figure 2 Actor network structure diagram provided for an embodiment of the present invention;

[0070] Figure 3 This is a diagram of the Critic network structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0071] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0072] like Figure 1 As shown, the method of this embodiment consists of two major parts: data warm-start and reinforcement learning training. First, the actor network is warm-started based on a large amount of construction data to learn the key features and patterns hidden in the data, laying a solid foundation for subsequent intelligent decision-making. Then, based on the warm-start, deep reinforcement learning methods are introduced to continuously optimize the control strategy through interaction with the actual construction environment.

[0073] For the warm start phase, a nonlinear relationship model between the tunneling parameters and the sealed chamber pressure during the shield machine excavation process was first established. This model was decomposed into a multi-objective dynamic decision-making problem. The inputs and outputs of the actor network were determined based on the shield machine's operating characteristics and control tasks, and the corresponding tunneling parameters were selected to provide dependent data support for subsequent training. Secondly, the actor network's three hyperparameters (number of iterations, learning rate, and number of neurons) were optimized using WOA, and the actor network was established based on the optimal hyperparameters. Finally, the decision-making effectiveness was verified using construction data under different geological conditions.

[0074] During shield machine excavation, the screw conveyor's rotational speed and propulsion speed are key parameters for regulating the amount of soil in and out. This example, based on the principle of balanced soil in and out, deeply analyzes the complex relationship between these two excavation parameters and the sealed chamber pressure. This derivation does not account for soil loss during shield machine advancement.

[0075] As the shield machine advances at a certain speed, the cutterhead rotates and cuts the soil in front of the excavation face, breaking it into debris and entering the sealed chamber. The amount of soil entering the sealed chamber per unit time is shown in formula (1). As the soil accumulates in the sealed chamber, the pressure inside the chamber gradually increases. At this time, the soil is discharged from the sealed chamber through the rotation of the screw conveyor to regulate the chamber pressure. The amount of soil discharged by the screw conveyor per unit time is shown in formula (2).

[0076] Q1=πR 2 V (1)

[0077] Q2=ηπATn0 (2)

[0078] Among them, R is the cutter head radius; V is the propulsion speed; η is the soil discharge efficiency; A is the effective cross-sectional area of the screw conveyor; T is the blade pitch; and n0 is the screw conveyor speed.

[0079] Based on the principle of soil inflow and outflow balance, the soil continuity equation within the sealed cabin can be derived, as shown in Equation (3). Because the sealed cabin has good sealing properties, the leakage pressure can be approximated to 0 MPa. Through the above analysis, the nonlinear relationship between the sealed cabin pressure and the propulsion speed and screw conveyor speed can be further mapped into Equation (4).

[0080]

[0081] Among them, C ep is the external leakage coefficient of the sealed cabin; P is the sealed cabin pressure; is the first-order derivative of P; P0 is the leakage pressure of the sealed cabin; V e is the sealed cabin capacity; β e is the compressibility coefficient of the material in the cabin.

[0082] Based on the Laplace transform results of formula (4) and the characteristics of intelligent decision-making, the inputs of the Actor network are set as: real-time sealed chamber pressure, optimal sealed chamber pressure under current soil conditions, and the error between the real-time pressure and the optimal pressure. The outputs of the Actor network are set as: propulsion speed and screw conveyor speed. Proper adjustment of these two will ensure safe and efficient construction. The subsequent reinforcement learning training environment is also built based on this mechanism model.

[0083] This embodiment is based on the construction background of Metro Line 10 in a certain place. The engineering geology is a complex sand and gravel stratum. Therefore, the stratum structure of this project is loose, complex and changeable, and has poor self-stability. The data used for warm start-up are actual construction data from a certain section of the project. The database of this project includes a total of 227 excavation parameters and 23,316 samples, including parameters such as propulsion speed, screw conveyor speed, optimal sealing cabin pressure value, and sealing cabin pressure at each monitoring point. This embodiment screens the database according to the input and output of the Actor network, and screens out a total of 20,000 sets of data to form a pre-training data set. The data set covers the construction status and transition process of 9 types of soil, such as silt, sand and sandstone. The specific data information is shown in Table 1.

[0084] Table 1 Data information

[0085] Input parameters unit Maximum Minimum average value Sealed cabin earth pressure MPa 0.34 0.13 0.1805 Optimal sealed compartment earth pressure MPa 0.32 0.15 0.2850 Sealed cabin pressure error MPa 0.13 0 0.1095 Screw conveyor speed rpm 11.8 2.3 7.7437 Advance speed mm / min 64 33 49.7793

[0086] The Actor network is a neural network based on a multi-layer perceptron (MLP) structure. In order to minimize the complexity of the network while ensuring model performance, a network structure with two hidden layers was selected. Specifically, the Actor network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. Figure 2 As shown in the figure, P, P′, and P″ are the real-time sealed chamber pressure, the optimal sealed chamber pressure, and the absolute value of the error between the real-time pressure and the optimal pressure, respectively. n0 and V are the screw conveyor speed and propulsion speed, respectively. This design balances complexity and computational efficiency, capturing the nonlinear characteristics of the data while avoiding the risk of overfitting or increased computational cost caused by an overly complex model structure.

[0087] To optimize the performance of the Actor Network, this example uses a global optimization method based on the Whale Optimization Algorithm (WOA) to optimize four key hyperparameters: the learning rate, the number of iterations, and the number of neurons in the first and second hidden layers. The initial search range for each hyperparameter is set as follows: the learning rate range is (0.0001, 0.1), the number of iterations range is (1, 100), the number of neurons in the first hidden layer range is (1, 100), and the number of neurons in the second hidden layer range is (1, 100). These ranges are based on previous experience and experimental data and can well cover the potential optimal values of the model parameters. The optimal hyperparameters for the Actor Network are shown in Table 2.

[0088] Table 2 Optimal hyperparameters for the Actor network

[0089] Number of iterations 98 Learning rate 0.007 Number of neurons in hidden layer 1 57 Number of neurons in hidden layer 2 98

[0090] During the initial stage of model training, the actor network is warm-started based on actual construction data. First, the WOA algorithm globally optimizes the search space to find more suitable parameter combinations and gradually optimize network performance. Second, the network is trained using construction data to provide it with a certain amount of prior knowledge, thereby accelerating convergence during subsequent training. This process effectively narrows the agent's action search range, reduces the number of trial and error times during training, and ultimately improves learning efficiency and decision accuracy. The warm-start process is as follows:

[0091] Step 1.1: Initialize the number of whales N and the maximum number of iterations T.

[0092] Step 1.2: Initialize the population: X i ,i=1,2,...,M;M represents the population size;set X * is the current optimal position.

[0093] Step 1.3: Check if there are whales exceeding the search space and modify it.

[0094] Step 1.4: Calculate the fitness value of each position and find the current optimal position X * .

[0095] Step 1.5: Determine whether the predation probability p is less than the set threshold; if so, proceed directly to step 1.6; otherwise, enter the bubble net predation stage, based on the formula D = |X * -X(t)| and X(t+1) = D·e bl ·cos(2πl)+X *Update the individual position information; where D represents the distance between the current search individual and the current best individual, X(t) is the position vector of the current individual in the bubble net predation stage, X(t + 1) represents the position vector of the best individual in the next iteration after update, t is the iteration number, X * is the position vector of the current best individual, l is a random number uniformly distributed in the interval [-1, 1]; b is a constant; e bl is a parameter that changes with the iteration number and is used to adjust the search range and convergence;

[0096] Step 1.6: Determine whether the coefficient vector |B| is less than 1; if so, enter the stage of hunting prey, and use D′ = |C·X * - X′(t)| and X′(t + 1) = X * - B·D′ to update the position information; otherwise, enter the stage of searching for prey, based on D″ = |C·X rand - X″(t)| and X″(t + 1) = X rand - B·D″ to update the position information; where D′ is the distance between the current individual and the best individual in the stage of hunting prey, X′(t) is the position vector of the current individual in the stage of hunting prey, X′(t + 1) is the position vector of the best individual in the next iteration in the stage of hunting prey, B and C represent coefficient vectors; X″(t) is the position vector of the current individual in the stage of searching for prey, X″(t + 1) represents the position vector of the best individual in the next iteration in the stage of searching for prey, D″ is the relationship between the current search individual and the random individual in the stage of searching for prey, X rand is the position vector of the random individual in the stage of searching for prey;

[0097] Step 1.7: After the position information update is completed, calculate the fitness value and compare it with the fitness value of the initial best position information, and select the best individual position with a smaller fitness value.

[0098] Step 1.8: Increment the iteration number t by 1. If t < T, return to Step 1.3; otherwise, end the optimization process and output the best hyperparameters of the Actor network.

[0099] Step 1.9: Divide the construction data into a training set and a test set according to 8.5:1.5.

[0100] Step 1.10: Continuously update the weights and biases to train the Actor network based on the training set.

[0101] Step 1.11: Examine the decision-making effect of the Actor network based on the test set.

[0102] For reinforcement learning training, the state space, action space, and reward function of the DDPG agent are first defined based on the operating characteristics of the shield machine and the control task. Next, the DDPG agent is designed and its hyperparameters are selected. Finally, the agent is trained through interaction with the virtual environment, achieving intelligent decision-making for shield machine earth pressure balance based on the pre-trained deep reinforcement learning agent.

[0103] The reinforcement learning training component designs the agent framework and training strategies, systematically building an intelligent decision-making model for complex and dynamic environments. Through interactive training with the environment, the agent gradually improves its adaptability and task-solving capabilities, achieving keen perception and dynamic response to uncertain environments, ensuring accurate and efficient decision-making in diverse real-world scenarios.

[0104] Reinforcement learning agents involve three key elements: state space, action space, and reward function. By carefully designing the reward function, action space, and state space, DDPG agents can better understand their environment and make appropriate decisions, thereby improving their overall performance.

[0105] The state space is the input source used by the agent to perceive the environment and make decisions. A well-designed state space ensures that the agent fully grasps the key information that influences decision-making during the construction process, thereby enhancing its adaptability to the environment and improving the accuracy and intelligence of overall decision-making. Formula (5) represents the defined state variables.

[0106] S t =[P, P′, P″] (5)

[0107] The state space is used to characterize the environmental state during the shield machine construction process, ensuring that the DDPG agent can perceive the dynamic changes of the construction environment in a timely manner.

[0108] Reasonable action space design enables the agent to explore different control strategies and thus optimize the specific actions it performs. Formula (6) represents the defined action space.

[0109] a t =[V,n0] (6)

[0110] The agent optimizes the pressure inside the sealed chamber during excavation by selecting different combinations of propulsion speed and conveyor speed, ensuring dynamic equilibrium with the water and soil pressure at the excavation surface. This multidimensional action space design allows the agent to fully explore the complexities of shield machine operations, thereby finding a more optimal excavation strategy. This multidimensional action space ensures that the agent fully explores the complex relationships between shield machine excavation parameters, thereby better optimizing the control strategy and ultimately finding a more efficient and stable excavation solution.

[0111] The reward function is the core mechanism for evaluating the performance of an agent's behavior, helping it gradually adjust its strategy through quantitative feedback. A precisely defined reward function encourages agents to make decisions that promote construction efficiency and safety, while avoiding potentially risky actions. In intelligent decision-making solutions and systems, a trinity design combining positive rewards, negative rewards, and boundary rewards is employed to provide agents with clearer optimization goals. The specific reward function is shown below:

[0112]

[0113] Positive rewards are used to encourage the agent's outstanding performance. When the chamber pressure is consistent with the optimal value, positive rewards are given, reinforcing its goal-oriented approach to correct strategies. Negative rewards help the agent identify adverse behaviors. When the pressure deviates significantly from the optimal value, negative rewards are given to promptly correct ineffective actions. Boundary rewards accelerate the agent's avoidance of substandard actions. When the chamber pressure reaches a dangerous level, boundary rewards are given, thereby improving training stability and preventing operational errors.

[0114] DDPG adopts the Actor-Critic architecture, which is mainly composed of the Actor network and the Critic network. The Actor network is responsible for generating the optimal strategy, while the Critic network evaluates the value of the actions taken. By introducing the experience replay and target network mechanisms, the stability and efficiency of DDPG in the learning process are significantly improved, especially for decision-making tasks in high-dimensional continuous action spaces. The Actor network structure has been explained in detail in the data warm start section. The Critic network consists of four parts: input layer, hidden layer 1, hidden layer 2, and output layer. The number of neurons in the hidden layer is 50. The specific structure is as follows Figure 3 shown.

[0115] The Critic network evaluates the value of the current state and action combination, outputting a corresponding value function. This evaluation result provides a basis for subsequent strategy updates. Simultaneously, by calculating the value of actions, the Critic network promptly transmits information to the Actor network, guiding its optimization strategy and ensuring that the shield machine maintains optimal operating conditions in various complex situations.

[0116] The reinforcement learning training process can be summarized as follows: First, the weights of the Actor network and the Critic network are initialized, and an experience replay buffer is created to store the agent's interaction experience. At each time step, the agent selects an action based on the current state, applies the action to the environment, and receives rewards and the next state. At this time, the agent stores the experience (state, action, reward, next state) in the experience replay buffer. Secondly, a small batch of samples are randomly sampled from the buffer for training. The Critic network uses the Bellman equation to calculate the target Q value and updates its weights by minimizing the mean squared error loss function. The Actor network is updated based on the Critic's Q value using the policy gradient method to optimize its policy. In addition, the Target network used to supervise the Actor network training periodically adopts a soft update strategy to ensure the stability of the training process. The entire training process is repeated until the performance of the agent reaches the expected level. The specific training process is as follows:

[0117] Step 2.1: Initialize the sealed cabin earth pressure environment model and the Actor network after warm start.

[0118] Random noise is represented by ξ, and random network parameters θ are used μ and θ Q Initialize the Actor network μ(s|θ μ ) and Critic network Q(s|θ Q ).

[0119] Copy the parameter values θ of the Actor network and the Critic network respectively μ →θ μ′ and θ Q →θ Q′ , initialize the TargetActor network μ′(s|θ μ′ ) and Target Critic network Q′(s|θ Q′ ), where θ μ ,θ Q are the parameters of the Actor and Critic networks respectively. μ′ ,θ Q′ are the parameters of the Target Actor and Target Critic networks respectively.

[0120] Step 2.2: Initialize the entire experience replay pool R.

[0121] Step 2.4: Initialize random noise ξ for motion exploration; obtain the initial state s1 from the soil pressure environment of the sealed cabin.

[0122] Step 2.5: For each time step t, perform the following operations:

[0123] Select action a based on the model's current strategy and noise t =μ(s|θ μ )+ξ;where μ(s|θ μ ) represents the strategy output by the Actor network.

[0124] Execute action a t Observe the error between the sealed cabin soil pressure value and the standard pressure value, and obtain the reward r t , the environment state becomes s t+1 .

[0125] (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool R.

[0126] Randomly sampling N tuples {(s i ,a i ,r i ,s′ i )} i=1,…,N .

[0127] For each tuple, use the Target network to calculate y = r i +γQ′(s i+1 ,a i+1 ,θ Q′ ); where y is the cumulative expected value, γ is the discount factor, Q′(s i+1 ,a i+1 ,θ Q′ ) is the evaluation value of the next state action.

[0128] Minimize the target loss: To update the current Critic network.

[0129] Calculate the sampled policy gradient as follows and update the current Actor network:

[0130]

[0131] in, is the gradient of the policy objective function; is the gradient of the Critic network to the action; Outputs the gradient of the policy with respect to the parameters for the Actor.

[0132] Update the Target network as follows:

[0133] θ μ′ =τθ μ +(1-τ)θ μ′ ;

[0134]

[0135] Among them, τ represents the learning rate, and its value range is (0,1); θ μ Represents the Actor network parameters; θ Q Represents the Critic network parameters; θ μ′ Represents the Target Actor network parameters; θ Q′ Represents the Target Critic network parameters.

[0136] Step 2.6: Repeat steps 2.3 to 2.5 until the agent's cumulative reward is maximized, that is, the sealed cabin pressure value output by the agent strategy converges to the set pressure reference value.

[0137] The DDPG agent achieves earth pressure balance control by constructing a closed-loop decision-making system for dynamic interaction between the shield machine and the geological environment. The target pressure in the sealed chamber, the real-time monitored pressure, and their deviation form a multidimensional state space input. Based on this state information feedback, the actor network generates optimal coordinated adjustment commands for propulsion speed and screw conveyor speed. The shield machine executes these control commands, regulating the soil flow rate within the sealed chamber to achieve earth pressure balance control. This creates a continuous decision-making process of "state perception → value assessment → action generation → pressure feedback." This mechanism dynamically decouples the tightly coupled relationships among the shield machine's multiple parameters, autonomously maintaining stable sealed chamber pressure under time-varying geological conditions.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A shield machine earth pressure balance autonomous control method based on warm start DDPG, characterized by: The proposed shield machine earth pressure balance autonomous control method includes two parts: data warm start and reinforcement learning training. First, the actor network is warm started based on a large amount of construction data to learn the key features and patterns hidden in the data. Subsequently, based on the warm start, a deep reinforcement learning method was introduced to continuously optimize the control strategy through interaction with the actual construction environment, ultimately achieving intelligent decision-making for the earth pressure balance of the shield machine; During the data warm start, a nonlinear relationship model between the tunneling parameters of the shield machine and the sealed chamber pressure during excavation is first established, and the model is decomposed into a multi-objective dynamic decision-making problem. At the same time, the input and output of the Actor network are determined based on the shield machine's operating characteristics and control tasks, and the corresponding tunneling parameters are selected to provide dependent data support for subsequent training. Secondly, the Whale Optimization Algorithm (WOA) was used to optimize the three hyperparameters of the Actor network: the number of iterations, the learning rate, and the number of neurons. The Actor network was then trained based on the optimal hyperparameters. Finally, the decision-making effect of the Actor network was tested based on construction data under different geological conditions. In the reinforcement learning training, first, the state space, action space, and reward function of the DDPG agent are defined based on the working characteristics of the shield machine and the control task characteristics; second, the DDPG agent is designed and relevant hyperparameters are selected; finally, the agent is trained through interaction with the virtual environment, realizing intelligent decision-making for the shield machine earth pressure balance based on the pre-trained deep reinforcement learning agent.

2. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 1 is characterized by: During the data warm start, the process of establishing the nonlinear relationship model between the tunneling parameters of the shield machine and the sealed cabin pressure during the excavation process is as follows: When the shield machine advances forward at a certain speed, the cutterhead rotates to cut the soil in front of the excavation face, breaking it into debris and entering the sealed cabin. The amount of soil entering the sealed cabin per unit time, Q1, is shown in formula (1). As the soil in the sealed cabin accumulates, the pressure in the cabin gradually increases. At this time, the soil is discharged from the sealed cabin through the rotation of the screw conveyor to adjust the cabin pressure. The amount of soil discharged per unit time, Q2, by the screw conveyor is shown in formula (2). Q1=πR 2 V (1) Q2=ηπATn0 (2) Where R is the cutter head radius; V is the propulsion speed; η is the soil discharge efficiency; A is the effective cross-sectional area of the screw conveyor; T is the blade pitch; n0 is the screw conveyor speed; Based on the principle of balance between the amount of slag in and out, the continuity equation of the slag in the sealed cabin is derived, as shown in formula (3); the nonlinear relationship between the sealed cabin pressure and the propulsion speed and the screw conveyor speed is further mapped into formula (4); Among them, C ep is the external leakage coefficient of the sealed cabin; P is the sealed cabin pressure; is the first-order derivative of P; P0 is the leakage pressure of the sealed cabin; V e is the sealed cabin capacity; β e is the compressibility coefficient of the material in the cabin; Based on the Laplace transform result of formula (4) and the characteristics of intelligent decision-making, the input of the Actor network is set as: real-time sealed cabin pressure, optimal sealed cabin pressure under current soil conditions, and the absolute value of the error between the real-time pressure and the optimal pressure; the output of the Actor network is set as: propulsion speed and screw conveyor speed.

3. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 1 is characterized by: In the data warm start, the corresponding excavation parameters are selected as follows: Collect the actual construction data, screen the collected data according to the input and output of the Actor network, and a total of 20,000 groups of data are screened out to form a pre-training data set. This data set covers the construction states and transition processes of 9 types of soil, and the tunneling parameters included are the soil pressure in the sealed cabin, the optimal soil pressure in the sealed cabin, the pressure error in the sealed cabin, the speed of the screw conveyor, and the propulsion speed.

4. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 1 is characterized by: In the warm start of the data, the Actor network is a neural network based on the multi-layer perceptron structure, including an input layer, a first hidden layer, a second hidden layer, and an output layer.

5. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 1 is characterized by: In the warm start of the data, the hyperparameters are optimized using a global optimization method based on the whale optimization algorithm. The specific process is as follows: Step 1.1: Initialize the number of whales M and the maximum number of iterations T; Step 1.2: Initialize the population: X i ,i=1,2,...,M;M represents the population size;set X * is the current optimal position; Step 1.3: Check whether there are whales exceeding the search space and make modifications; Step 1.4: Calculate the fitness value of each position and find the position vector X of the current best individual * ; Step 1.5: Determine whether the predation probability p is less than the set threshold; if so, proceed directly to step 1.6; otherwise, enter the bubble net predation stage, based on the formula D = |X * -X(t)| and X(t+1) = D·e bl ·cos(2πl)+X * Update individual position information; where D represents the distance between the current search individual and the current best individual, X(t) is the position vector of the current individual in the bubble net predation phase, X(t+1) represents the position vector of the best individual in the next iteration after the update, t is the number of iterations, and X * is the position vector of the current best individual, l is a random number uniformly distributed in the interval [-1,1]; b is a constant; e bl It is a parameter that changes with the number of iterations and is used to adjust the search range and convergence; Step 1.6: Determine whether the coefficient vector |B| is less than 1; if so, enter the prey hunting phase and use D′=|C·X * -X′(t)| and X′(t+1) = X * -B·D′ updates the position information; otherwise, it enters the prey search phase, based on D″=|C·X rand -X″(t)| and X″(t+1) = X rand -B·D″ updates the position information; where D′ is the distance between the current individual and the best individual in the prey-hunting stage, X′(t) is the position vector of the current individual in the prey-hunting stage, X′(t+1) is the position vector of the best individual in the next iteration of the prey-hunting stage, B and C represent coefficient vectors; X″(t) is the position vector of the current individual in the prey-searching stage, X″(t+1) represents the position vector of the best individual in the next iteration of the prey-searching stage, D″ is the relationship between the current search individual and the random individual in the prey-hunting stage, X rand is the position vector of a random individual during the prey search phase; Step 1.7: When the update of the position information ends, calculate the fitness value and compare it with the fitness value of the initial best position information, and select the best individual position with a smaller fitness value; Step 1.8: Increment the iteration number t by 1. If t < T, return to Step 1.3; otherwise, end the optimization process and output the best hyperparameters of the Actor network.

6. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 1 is characterized by: The state space, action space, and reward function of the DDPG agent are defined as follows: State space: It is the input source for the agent to perceive the environment and make decisions, and is used to characterize the environmental state during the construction of the shield machine to ensure that the DDPG agent can timely perceive the dynamic changes of the construction environment; the defined state variables are shown in formula (5): S t =[P,P′,P″] (5) Among them, P, P′, and P″ are the real-time pressure in the sealed cabin, the optimal pressure in the sealed cabin, and the absolute value of the error between the real-time pressure and the optimal pressure, respectively; Action space: The defined action space is shown in formula (6): a t =[V,n0] (6) Among them, n0 and V are the rotation speed of the screw conveyor and the propulsion speed, respectively; Reward function: It adopts a combination of positive rewards, negative rewards, and boundary rewards. The specific reward function is as follows: Positive rewards are used to encourage the excellent performance of the agent. When the pressure in the sealed cabin is consistent with the optimal value, a positive reward is given to enhance its goal orientation for the correct strategy; negative rewards are used to help the agent identify adverse behaviors. When the pressure deviates greatly from the optimal value, a negative reward is given to timely correct ineffective operations; boundary rewards are used to accelerate the agent's avoidance of actions that do not meet the requirements. When the pressure in the sealed cabin reaches a dangerous value, a boundary reward is given, thereby improving the stability of training and avoiding operation errors.

7. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 6 is characterized by: The DDPG agent adopts an Actor-Critic architecture, which consists of an Actor network and a Critic network; the Actor network is responsible for generating the optimal strategy, and the Critic network evaluates the value of the actions taken; the Critic network includes an input layer, a first hidden layer, a second hidden layer, and an output layer; the Critic network conducts value evaluation based on the current state and action combination and outputs a corresponding value function.

8. The shield machine earth pressure balance autonomous control method based on warm start DDPG according to claim 7 is characterized by: The training process of the agent is specifically as follows: First, the weights of the Actor network and the Critic network are initialized, and an experience replay buffer is created to store the agent's interaction experience. At each time step, the agent selects an action based on the current state and applies the action to the environment, receiving a reward and the next state. At this point, the agent stores the experience in the experience replay buffer. The stored experience includes state, action, reward, and next state. Secondly, small batches of samples are randomly sampled from the buffer for training; the Critic network uses the Bellman equation to calculate the target Q value and updates its weights by minimizing the mean squared error loss function; the Actor network uses the policy gradient method to update based on the Critic's Q value to optimize its policy; In addition, the Target network used to supervise the training of the Actor network regularly adopts a soft update strategy to ensure the stability of the training process; The entire training process is repeated until the agent's performance reaches the expected level; The specific training process is as follows: Step 2.1: Initialize the sealed cabin earth pressure environment model and the Actor network after warm start; Random noise is represented by ξ, and random network parameters θ are used μ and θ Q Initialize the Actor network μ(s|θ μ ) and Critic network Q(s|θ Q ); Copy the parameter values θ of the Actor network separately μ →θ μ′ and the parameter value θ of the Critic network Q →θ Q′ , initialize the TargetActor network μ′(s|θ μ′ ) and Target Critic network Q′(s|θ Q′ ), where θ μ ,θ Q are the parameters of the Actor network and the Critic network respectively; θ μ′ ,θ Q′ These are the parameters of the Target Actor network and the Target Critic network respectively; Step 2.2: Initialize the entire experience replay pool R; Step 2.4: Initialize random noise ξ for motion exploration; obtain the initial state s1 from the soil pressure environment of the sealed cabin; Step 2.5: For each time step t, perform the following operations: Select action a based on the model's current strategy and noise t =μ(s|θ μ )+ξ, where μ(s|θ μ ) represents the strategy output by the Actor network; Execute action a t Observe the error between the sealed cabin soil pressure value and the standard pressure value, and obtain the reward r t , the environment state becomes s t+1 ; (s t ,a t ,r t ,s t+1 ) Stored in the experience replay pool R; Randomly sampling N tuples {(s i ,a i ,t i ,s′ i )} i=1,…,N ; For each tuple, use the Target network to calculate y = r i +γQ′(s i+1 ,a i+1 ,θ Q′ ); where y is the cumulative expected value, γ is the discount factor, Q′(s i+1 ,a i+1 ,θ Q′ ) is the evaluation value of the next state action; Minimize the target loss: To update the current Critic network; Calculate the sampled policy gradient as follows and update the current Actor network: in, is the gradient of the policy objective function; is the gradient of the Critic network to the action; Output the gradient of the policy to the parameters for the Actor; Update the Target network as follows: i μ′ =tθ μ +(1-τ)θ μ′ ; Among them, τ represents the learning rate, and its value range is (0,1); Step 2.6: Repeat steps 2.3 to 2.5 until the agent's cumulative reward is maximized, that is, the sealed cabin pressure value output by the agent strategy converges to the set pressure reference value.

Citation Information

Cited By

  • Rail transit full-scene intelligent construction cooperative control method and system

    CN120951834A

  • Rail transit full-scene intelligent construction collaborative control method and system

    CN120951834B