A method and device for tunnel boring machine control based on physical information reinforcement learning
By combining earth pressure balance theory with deep neural networks, a tunneling environment for the tunnel boring machine (TBM) was constructed, and physical laws were added to the reinforcement learning algorithm. This solved the problem of lag in manual operation of the TBM and enabled efficient, safe, and automated excavation of the TBM.
Patent Information
- Authority / Receiving Office
- CN ยท China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2023-07-28
- Publication Date
- 2026-07-17
AI Technical Summary
The manual operation of existing tunnel boring machines (TBMs) during tunnel excavation is lagging, leading to project delays and cost overruns. Traditional machine learning methods lack physical interpretability and cannot optimize TBM parameters over long time steps, affecting operational automation.
By combining earth pressure balance theory with deep neural networks, a simulated tunneling environment for a tunnel boring machine (TBM) is constructed. By incorporating physical laws into the reward function and penalty function using a dual-delay depth deterministic algorithm, a TBM tunneling control method based on physical information reinforcement learning is developed. This method allows for real-time adjustment of TBM parameters to achieve tunneling speed and earth pressure balance.
It significantly improved the tunneling speed and earth pressure balance of the tunnel boring machine, realizing efficient, safe and automated excavation of the tunnel boring machine. The tunneling speed increased by 65.9%, the earth pressure balance increased by 72.7%, and the overall improvement was 69.3%.
Smart Images

Figure CN117145503B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of tunnel boring machine (TBM) construction technology, and more specifically, relates to a TBM tunneling control method and device based on physical information reinforcement learning. Background Technology
[0002] Earth pressure balance (EPB) tunnel boring machines (TBMs) are specialized equipment for mechanized excavation of urban tunnels, revolutionizing the field of underground construction. However, the complex geological conditions encountered during tunnel excavation limit the limitations of traditional TBM operation methods that rely on experience-based rules. One major reason is the inherent lag in manual TBM operation, which can lead to project delays and cost overruns. Therefore, optimizing EPB TBM performance through effective methods is a crucial step in achieving automated operation during tunnel excavation.
[0003] Currently, optimization efforts for EPB TBMs primarily focus on improving tunneling efficiency, enhancing safety, and reducing costs by predicting TBM parameters. Common parameter optimization methods utilize machine learning (ML) techniques to capture complex relationships within datasets, thereby enabling TBM parameter prediction. However, standard machine learning methods often lack physical interpretability and are unreliable in engineering applications. To address this challenge, researchers have begun to focus on Physical Information Machine Learning (PIML) methods, which integrate prior knowledge of fundamental physical processes into machine learning algorithms to develop more accurate and interpretable models. For TBM operation problems, embedding the physical laws of TBM-soil interaction into machine learning algorithms allows for the use of virtual ML models to describe the real performance of EPB TBMs operating in soil.
[0004] Furthermore, to address the issue that traditional TBM parameter optimization methods cannot consider every behavior over longer time steps and are unsuitable for continuous tunnel excavation operations, reinforcement learning (RL) has been gradually introduced into TBM operations. This involves training an agent capable of dynamically adjusting TBM parameters in real time to achieve the required tunneling speed (AS) and maintain earth pressure balance (EPB) during excavation. However, the success of agent training is highly dependent on the defined simulation environment. Therefore, how to utilize rich TBM operation datasets to train an RL agent capable of real-time TBM parameter optimization, and ensure that the learned strategy meets the physical constraints and principles of TBM operation, has become one of the urgent problems to be solved in achieving TBM operation automation. Summary of the Invention
[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a shield tunneling control method and device based on physical information reinforcement learning. Firstly, it integrates earth pressure balance theory with deep neural networks (DNNs) to construct a simulated environment for the response of the air pressure balance (AS) and soil pressure (CP) to TBM operations. Based on this, it adds physical laws to the reward function and penalty of the dual-delay depth deterministic algorithm (TD3) to significantly improve TBM performance. By integrating the physical laws of the EPB TBM working mechanism into the simulation environment and reward function, this invention not only makes it possible to simulate complex systems composed of TBMs and soil using virtual environments trained with machine learning techniques, but also helps to improve the tunneling speed and earth pressure balance of TBMs through improved RL methods, ultimately achieving efficient, safe, and automated TBM excavation.
[0006] To achieve the above objectives, according to one aspect of the present invention, a tunnel boring machine (TBM) tunneling control method based on physical information reinforcement learning is proposed, comprising the following steps:
[0007] Based on TBM operation data, S1 embeds the earth pressure balance theory into a model with DNN as the neural network architecture to construct an environmental network model of the TBM during tunnel construction, and simulates the physical information environment of the TBM during tunnel construction based on this environmental network model.
[0008] Based on the physical information environment, S2 constructs a physics-based dual-delay depth deterministic algorithm model by considering the physical laws and constraints of earth pressure balance inside and outside the tunnel boring machine, non-negativity of tunneling speed, and the pressure in the middle earth chamber being between the pressure in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm. SHAP is used to evaluate and interpret the above-mentioned environmental network model, and at the same time, the dual-delay depth deterministic algorithm model is evaluated.
[0009] Based on the dual-delay depth deterministic algorithm model, S3 dynamically adjusts the TBM parameters in real time to achieve the required tunneling speed and maintain the balance of earth pressure during the excavation process.
[0010] As a further preferred embodiment, the process also includes the acquisition and preprocessing of the TBM operating data;
[0011] The TBM operating data includes:
[0012] Tunnel depth h, total thrust TF, screw conveyor pressure SCP, tunneling speed AS, and soil chamber pressure at the top, middle and bottom of the tunnel.
[0013] As a further preferred option, step S1 includes the following steps:
[0014] S11 constructs a theoretical model of earth pressure balance to maintain the balance between earth pressure and earth pressure in the earth chamber of the tunnel boring machine during tunnel excavation.
[0015] S12 constructs a differential equation for the tunnel depth using the earth pressure balance theoretical model, and reconstructs the differential equation based on the earth pressure at different heights measured in real time to obtain the physical loss.
[0016] S13 adds this physical loss to the observation loss, which is the squared error between the measured observation and the predicted value, to construct the loss function used to train the environment network model;
[0017] S14 trains the environmental network model based on TBM operation data and the loss function, and simulates the physical information environment of TBM during tunnel construction based on the environmental network model.
[0018] As a further preferred embodiment, in step S12, the reconstruction of the differential equation based on the real-time measured soil pressure at different heights includes:
[0019]
[0020] In step S13, the new loss function used to train the PDNN model includes:
[0021]
[0022] In the formula, It is a weight used to measure physical loss; These are the top, middle, and bottom soil pressures predicted by the DNN model, respectively. It is the predicted tunneling speed; ๐ฃ, ๐ ๐ , ๐ ๐ and ๐ ๐ต These are the observed design output values.
[0023] As a further preferred embodiment, in step S2, the consideration of the physical laws and constraints of the earth pressure balance inside and outside the tunnel boring machine, the non-negativity of the tunneling speed, and the fact that the pressure in the middle earth chamber is between the pressures in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm includes:
[0024] The following formula is used to normalize the rewards for tunneling speed and soil pressure:
[0025]
[0026] In the formula, ๐ฃ ๐๐๐ฅ and ๐ฃ ๐๐๐ These are the maximum and minimum tunneling speeds in the dataset, respectively; โ is the weight. These are the top, middle, and bottom soil pressures predicted by the DNN model.
[0027] As a further preferred embodiment, in step S2, the loss function of the dual-delay deep deterministic algorithm model is defined by the TD error, which is the root mean square error of the Q-value at the current time step based on the Bellman equation:
[0028]
[0029] In the formula, ๐ต is the batch sampled from the response buffer; ๐ and ๐ are the current state and behavior; It represents the next state after taking the current action; ๐ is the corresponding reward. Record whether the time step has been completed; ๐พ is the loss rate;
[0030] To reduce the correlation between consecutive updates, the dual-delay deep deterministic algorithm model employs delayed updates to stabilize the learning process. Furthermore, it reduces the likelihood of the policy getting trapped in local optima through target policy smoothing. The target behavior after policy smoothing includes:
[0031]
[0032] In the formula, ๐ is the error added to the behavior, limited to the range (โ๐, ๐), ๐ ๐ฟ๐๐ค and ๐ ๐ป๐๐โ These are the lower and upper limits of the behavioral space.
[0033] As a further preferred embodiment, in step S3, the evaluation and interpretation of the above-mentioned environmental network model using SHAP includes:
[0034] (1) Evaluation of the environmental network model: RMSE is used to measure the deviation between the predicted value and the true value, and a20_index is used with a 20% tolerance to measure the reliability of the model. 2 Used to assess the degree of agreement between predicted and actual values.
[0035] (2) Interpretability of the environmental network model, using Shapley values โโto measure the importance of input features:
[0036]
[0037] In the formula, ๐ is the set containing all features, ๐ represents all non-zero entries, K is the dimension of the feature, and ๐ ๐ฅ (๐)=๐ธ[๐(๐ฅ)|๐ฅ ๐ ] is the expectation of the model's output value.
[0038] According to another aspect of the present invention, a tunnel boring machine control system based on physical information reinforcement learning is also provided, comprising:
[0039] The first main module: Based on TBM operation data, the earth pressure balance theory is embedded into a model with DNN as the neural network architecture to construct an environmental network model of TBM during tunnel construction, and the physical information environment of TBM during tunnel construction is simulated based on this environmental network model.
[0040] The second main module: Based on the physical information environment, by considering the physical laws and constraints of the earth pressure balance inside and outside the tunnel boring machine, the non-negativity of the tunneling speed, and the fact that the pressure in the middle earth chamber is between the pressure in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm, a physical-based dual-delay depth deterministic algorithm model is constructed. SHAP is used to evaluate and interpret the above-mentioned environmental network model, and at the same time, the dual-delay depth deterministic algorithm model is evaluated.
[0041] The third main module: Based on the aforementioned dual-delay depth deterministic algorithm model, the TBM parameters are dynamically adjusted in real time to achieve the required tunneling speed and maintain the balance of earth pressure during the excavation process.
[0042] According to another aspect of the invention, an electronic device is also provided, comprising:
[0043] At least one processor, at least one memory, and a communication interface; wherein,
[0044] The processor, memory, and communication interface communicate with each other;
[0045] The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the methods involved in any of the above embodiments.
[0046] According to another aspect of the present invention, a non-transitory computer-readable storage medium is also provided, the non-transitory computer-readable storage medium storing computer instructions that cause the computer to perform the methods involved in any of the above embodiments.
[0047] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages:
[0048] 1. This invention integrates the physical laws of the EPB TBM working mechanism into the simulation environment and reward function, which not only makes it possible to simulate complex systems composed of TBM and soil using virtual environments trained by machine learning technology, but also helps to improve the tunneling speed and earth pressure balance of TBM through an improved RL method, ultimately achieving efficient, safe and automated excavation of TBM.
[0049] 2. Based on data samples from engineering examples, this invention verifies the effectiveness of the proposed method. Data collected during tunnel excavation, including soil pressure, earth pressure (replaced by the pressure in the middle soil chamber), total thrust, screw conveyor pressure, tunneling speed, and tunnel depth, were used for model training and evaluation. Results show that the pTD3 algorithm proposed in this invention increases the tunneling speed of the tunnel boring machine (TBM) by 65.9% and improves the earth pressure balance during excavation by 72.7%, resulting in an overall improvement of 69.3%. This demonstrates that by incorporating physical laws into the simulation environment and reward function, the tunneling speed and earth pressure balance of the TBM can be significantly improved, strongly proving that the PIRL method can simultaneously improve the construction efficiency and stability of tunnel excavation. This invention provides a novel approach and method for optimizing EPB TBM performance and achieving automated operation during EPB TBM excavation. Attached Figure Description
[0050] Figure 1 This is a flowchart of a tunnel boring machine tunneling control method based on physical information reinforcement learning, which is an embodiment of the present invention.
[0051] Figure 2 This is a diagram illustrating the working mechanism of EPB TBM according to an embodiment of the present invention;
[0052] Figure 3 This is a flowchart of the TD3 algorithm involved in the embodiments of the present invention;
[0053] Figure 4 This is a schematic diagram of a PDNN structure that integrates physical laws, as described in an embodiment of the present invention.
[0054] Figure 5 (a) in the figure represents the PDNN prediction of ๐ ๐ , ๐ ๐ and ๐ ๐ต ; Figure 5 (b) in the figure represents the prediction made by the PDNN. Distribution; Figure 5 (c) in the figure represents the ฮฑ predicted by the DNN. ๐ , ๐ ๐ and ๐ ๐ต ; Figure 5 (d) in the figure represents the prediction of the DNN. Distribution;
[0055] Figure 6 (a) in the figure represents the SHAP analysis results using a model trained with PDNN. Figure 6 (b) in the figure shows the SHAP analysis results using a DNN-trained model.
[0056] Figure 7 (a) in the figure represents the cumulative AS reward as construction progresses. Figure 7 (b) represents the total cumulative reward as construction progresses. Figure 7 (c) in the figure represents the cumulative CP reward as construction progresses;
[0057] Figure 8 (a) shows a comparison of the AS performance of TBM using manual operation and pTD3. Figure 7 (b) shows the earth pressure balance of the TBM using manual operation and pTD3. Comparison;
[0058] Figure 9 (a) shows the TF and SCP designed using the pTD3 algorithm. Figure 9 (b) in the diagram represents the manually designed TF and SCP;
[0059] Figure 10 In the diagram, (a) shows the TF and SCP distributions designed using the pTD3 algorithm. Figure 10 (b) in the diagram represents the manually designed TF and SCP distributions. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0061] like Figure 1 , Figure 2 , Figure 3 and Figure 4 As shown in the figure, the present invention provides a shield tunneling control method based on reinforcement learning, namely a novel physical information reinforcement learning (PIRL) method to control the performance of EPB TBM, aiming to improve the tunneling speed and stability of TBM and solve the problem of low automation of TBM operation in engineering applications.
[0062] The core idea of โโthis method is to first construct a simulated response environment of Earth Pressure Balance (AS) and Compressed Pressure (CP) to TBM operations by integrating Earth Pressure Balance (EPB) theory with deep neural networks (DNNs). Based on this, physical laws are added to the reward function and penalty of the dual-delay deep deterministic algorithm (TD3) model to significantly improve TBM performance. This method includes four main steps: data acquisition and preprocessing, development of a physically-based reinforcement learning model, model evaluation, and interpretation. The overall process is described in [reference needed]. Figure 1As shown, the specific steps are as follows:
[0063] 1. Data Acquisition and Preprocessing
[0064] (1) Data Acquisition. Maintaining a balance between external earth pressure and internal chamber pressure is the fundamental principle behind the operation of an EPB TBM, ensuring stability and preventing ground collapse. Therefore, before tunnel construction, two symmetrical pressure sensors are often installed at the top, middle, and bottom of the EPB TBM to monitor and collect the TBM's pressure pressure (CP) in real time during tunnel construction. Figure 2 As shown. Furthermore, during the tunnel construction phase, the total thrust (TF) of the EPB TBM, the screw conveyor pressure (SCP), and the AS were also recorded. According to the earth pressure balance theory in soil mechanics, earth pressure is related to depth. Therefore, the tunnel depth (h) also needs to be determined and collected based on the design drawings.
[0065] (2) Data Preprocessing. Using the data collected by the sensors, the average CP of the top, middle, and bottom sensors can be calculated. It should be noted that due to a lack of detailed information about the geological conditions, the earth pressure (P) is... ๐๐๐๐๐ The intermediate CP is used to approximate the data. Considering that significant differences in data scale may lead to difficulties in model training convergence, and that repeated normalization consumes considerable computational resources during training, increasing the data reduction problem, it is necessary to reconstruct the feature parameters by transforming their units to the same scale. Specifically, the tunnel depth is in meters (m), the earth chamber pressure is in bars (bar), and the total thrust is in ร10โปยนโฐ units. 3 kN, the pressure unit of the screw conveyor belt is MPa, and the unit of tunneling speed is m / s.
[0066] 2. Development of Physical Information-Based Reinforcement Learning Models
[0067] The basic idea behind the PIRL model development is as follows: First, earth pressure balance theory is embedded as physical knowledge into a model with a DNN (Digital Neural Network) architecture to obtain the PIML model. The main task of this model is to establish a physical information environment for reasonably and reliably simulating the behavior of an EPB (Electronic Power Shield) TBM during tunnel construction, enabling it to receive behavior and simulate the generation of the next state. Then, based on the environment trained with a Physical Information Deep Neural Network (PDNN), a physical-based dual-delay depth deterministic (pTD3) algorithm is developed by considering physical laws and constraints such as earth pressure balance inside and outside the tunnel boring machine, the non-negativity of tunneling speed, and the fact that the pressure in the middle earth chamber is between the pressures in the top and bottom earth chambers, in the reward function and penalty of the dual-delay depth deterministic algorithm (TD3). This algorithm is used to judge the performance of the behavior. Specific steps mainly include: clarifying the physical laws in the EPB TBM, using the dual-delay depth deterministic algorithm for strategy optimization, developing a physical information-based simulation environment, and designing a physical information-based reward function and constraints, such as... Figure 1 As shown.
[0068] (1) Clarify the physical laws in EPB TBM
[0069] Maintaining a balance between earth pressure (EP) and the earth chamber pressure within the tunnel boring machine (TBM) during tunnel excavation is the fundamental principle behind the operation of an EPB TBM. A conceptual diagram of the EPB TBM system is shown below. Figure 2 As shown. From Figure 2 As can be seen, the cutterhead, cavity, screw, and conveyor all play important roles in maintaining earth pressure balance. According to soil mechanics, the definitions for different situations are given in formula (1):
[0070]
[0071] In summary, when developing the PIRL model in the later stages, it is necessary to consider the AS, TF, SCP, h of the TBM, and the soil pressure ( ) are used as input parameters for the model.
[0072] (2) Use a dual-delay deep deterministic algorithm for strategy optimization
[0073] TD3 is a method for finding optimal behavior in Actor-Critic networks based on Q-values. Its model structure is as follows: Figure 3 As shown. By Figure 3 It can be seen that this group of networks includes the current Actor network. The model consists of a target Actor network, a current Critic network, and a target Critic network. The Critic network is used to estimate the Q-value, and the Actor network is used to select the action that maximizes the Q-value. It should be noted that when estimating the Q-value using two Critic networks, the smaller Q-value is selected for estimating the target value. The dataset used to train the Q-function of the current Critic network is a response buffer sample that collects previous state and operational experience. Furthermore, the loss function of this model is defined by the TD error, which is the root mean square error (RMSE) of the Q-value at the current time step based on the Bellman equation, as shown in Equation (2):
[0074] (2)
[0075] In the formula, ๐ต is the batch sampled from the response buffer; ๐ and ๐ are the current state and behavior; It represents the next state after taking the current action; ๐ represents the corresponding reward. Record whether the time step has been completed. It's the loss rate.
[0076] Furthermore, to reduce the correlation between consecutive updates, the TD3 algorithm employs delayed updates to stabilize the learning process, and also reduces the possibility of the policy getting trapped in local optima through target policy smoothing. The target behavior after policy smoothing is shown in formula (3):
[0077]
[0078] In the formula, ๐ is the error added to the behavior, limited to the range (โ๐, ๐); ๐ ๐ฟ๐๐ค and ๐ ๐ป๐๐โ These are the lower and upper limits of the behavioral space.
[0079] (3) Develop a simulation environment
[0080] To more accurately reflect the relationship between operation (behavior), TBM performance (state), and construction site conditions (environment), earth pressure balance theory is introduced into the development of the environmental network, such as... Figure 4 As shown. By Figure 4 It can be seen that the specific process of integrating physical laws into the PDNN structure can be described as follows: First, by solving... The partial derivatives yield an ordinary differential equation, as shown in formula (4):
[0081]
[0082] Secondly, by using the soil pressure (CP) measured by sensors at different heights to reconstruct equation (4), we can obtain:
[0083]
[0084] In the formula, ฯ, ฯ, and ฯ represent the average CP measured by the top, middle, and bottom sensors, respectively; d is the distance between the top and bottom sensors. Obviously, the ordinary differential equation (5) eliminates all uncontrollable parameters, retaining only the measured soil pressure and tunnel depth. The former can be acquired by sensors, and the latter can be determined according to the planning and design before construction. According to the working mechanism of EPB TBM, the model simulating the excavation process should also follow formula (5), and its simulation error can be regarded as physical loss. By adding this physical loss to the observation loss of the square error between the measured observation and the predicted value, the new loss function used to train the PDNN model can be rewritten as:
[0085]
[0086] It is a weight used to measure physical loss; These are the top, middle, and bottom soil pressures predicted by the DNN model, respectively. It is the predicted tunneling speed. These are the observed design output values.
[0087] (4) Design reward function and constraints based on physical information
[0088] Based on the simulated AS and CP response environment established in Part (3), it can be seen that the reward for the new state consists of two parts. One part evaluates the tunneling speed, which should be as large as possible; the other part evaluates the stability, which should be as close as possible to the working mechanism of the TBM. In order to reasonably measure the above two criteria, it is necessary to consider not only the balance between the internal CP and external EP of the tunnel boring machine, but also the scale difference between CP and AS. Because during the training process, larger scale criteria often generate higher rewards and attract more attention to the algorithm. To solve the above problems, formula (7) is used to normalize the rewards of AS and CP.
[0089]
[0090] In the formula, ๐ฃ๐๐๐ฅ and ๐ฃ๐๐๐ are the maximum and minimum tunneling speeds in the dataset, respectively; ๐ is the weight of the two criteria in the reward equation.
[0091] To make the algorithm more sensitive to design requirements, certain constraints should be set for both AS and CP standards. For example, AS should not be negative during excavation. To maintain the stability of the TBM, CP should always satisfy ฯ.๐ <0 ๐ <0 ๐ต If the next state does not satisfy the two constraints mentioned above, the simulation at the current time step will terminate, and the total reward will be penalized by -1. Furthermore, the earth pressure inside and outside the earth chamber should remain balanced. However, considering the small difference between EP and CP in actual projects, the design constraints of this standard are not as strict as the previous two standards. The total reward will only be penalized when the difference between CP and EP exceeds 10%.
[0092] 3. Model Evaluation and Interpretability
[0093] As mentioned earlier, the construction of a simulated environment is of great significance for deep reinforcement learning. Therefore, before evaluating the physical information-based dual-delay deep deterministic (pTD3) model developed in this invention, the constructed environment network model should be evaluated and SHAP analyzed in advance.
[0094] (1) Evaluation of the environmental network model. To test the model's performance, three criteria were selected to evaluate the accuracy of the prediction results, including root mean square error (RMSE), a20_index, and ฮฑ. 2 Among them, RMSE is used to measure the deviation between the predicted value and the true value; a20_index measures the reliability of the model with a 20% tolerance; and ๐ 2 Used to evaluate the degree of agreement between predicted and actual values. In summary, the three metrics follow different logics, and using them simultaneously provides a more comprehensive understanding of the model's performance. Used to calculate RMSE, a20_index, and ๐ . 2 The mathematical expressions are formulas (8)-(10):
[0095]
[0096] In the formula, is the average of all samples, and a20 is the number of samples with a residual error not exceeding ยฑ20% of the sample value. Clearly, by training the model more accurately, the RMSE will approach 0, and a20_index and a2 will approach 1.
[0097] (2) Interpretability of the environmental network model. To explore the relationship between tunnel depth and intermediate CP, this study uses Shapley Additive Interpretation (SHAP). The basic idea of โโthe SHAP method is to use Shapley values โโto measure the importance of input features, and its mathematical expression is shown in formula (11):
[0098]
[0099] In the formula, ๐ is the set containing all features, ๐ collects all non-zero entries, K is the dimension of the feature, and ๐ ๐ฅ (k) = ๐ธ[k(k)|k ๐ The output is derived from the model to be interpreted.
[0100] According to the principles of soil mechanics, there are two rules that should be followed between tunnel depth and soil pressure: (1) Soil pressure increases with increasing depth; (2) Since tunnel depth is directly related to soil weight, the Shapley value of tunnel depth should also increase with its own increase.
[0101] (3) Evaluation of the pTD3 model
[0102] To evaluate whether the proposed pTD3 algorithm can significantly improve the excavation efficiency and safety of TBMs, the AS, EPB, CP, TF, and SCP obtained by the pTD3 algorithm are compared with those obtained manually. If the comparative analysis shows that the AS obtained based on pTD3 is larger, the EPB is smaller, the CP is more balanced, and the standard deviations of TF and SCP are smaller, then the developed pTD3 algorithm is proven to be effective and can significantly improve the performance of TBMs.
[0103] According to another aspect of the present invention, a shield tunneling control system based on physical information reinforcement learning is also provided, comprising: a first main module: based on TBM operation data, embedding earth pressure balance theory into a model with DNN as the neural network architecture to construct an environmental network model of TBM during tunnel construction, and simulating the physical information environment of TBM during tunnel construction based on the environmental network model;
[0104] The second main module: Based on the physical information environment, by considering the physical laws and constraints of the earth pressure balance inside and outside the tunnel boring machine, the non-negativity of the tunneling speed, and the fact that the pressure in the middle earth chamber is between the pressure in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm, a physical-based dual-delay depth deterministic algorithm model is constructed. SHAP is used to evaluate and interpret the above-mentioned environmental network model, and at the same time, the dual-delay depth deterministic algorithm model is evaluated.
[0105] The third main module: Based on the aforementioned dual-delay depth deterministic algorithm model, the TBM parameters are dynamically adjusted in real time to achieve the required tunneling speed and maintain the balance of earth pressure during the excavation process.
[0106] The methods in the embodiments of the present invention are implemented using electronic devices; therefore, it is necessary to describe the relevant electronic devices. For this purpose, embodiments of the present invention provide an electronic device comprising: at least one processor, a communication interface, at least one memory, and a communication bus, wherein the at least one processor, the communication interface, and the at least one memory communicate with each other via the communication bus. The at least one processor can invoke logical instructions in the at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.
[0107] Furthermore, when the logical instructions in at least one of the aforementioned memories can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various method embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments. Specific Implementation Example 1
[0111] Reference Figure 1 As shown, this embodiment provides a novel physical information reinforcement learning method to improve the performance of TBMs in terms of both tunneling speed and earth pressure balance. The specific steps are as follows:
[0112] Step 1: Data Acquisition and Preprocessing. This specifically includes:
[0113] Step 1.1 Data Acquisition. This invention selects measurement data from the inner track of the EPBTBM used in the C885 tunnel of the Singapore MRT Circle Line Phase 6 project as a case study. This tunnel track comprises a total of 696 loops. The recorded data for each loop mainly includes: tunnel depth (h), total thrust (TF), screw conveyor pressure (SCP), tunneling speed (AS), and top, middle, and bottom CP values โโcollected by six pressure sensors installed on the EPB TBM. As mentioned earlier, the middle CP can be used to approximately reflect the earth pressure (S). ๐๐๐๐๐ ).
[0114] Step 1.2 Data Preprocessing. Considering the six deep learning networks in the RL algorithm and the significant differences in data scale, the training model may struggle to converge. Furthermore, the numerous training steps and repeated normalization require more computational resources, leading to further data reduction issues. Therefore, each feature is reconstructed by transforming the units to the same scale; details are shown in Table 1.
[0115] Table 1 Features considered in PDNN model training
[0116]
[0117]
[0118] Approximate the pressure of the central earthwork in the previous time step.
[0119] Step 2, Model Development:
[0120] (1) Establish a simulation environment
[0121] As mentioned earlier, the collected dataset contained 696 data points. To check the generality of the proposed method and avoid data leakage, the first 550 data points of the lining rings were selected for training the environment model, and these 550 data points were randomly divided into training and test sets in a 4:1 ratio. The physical information simulation environment was built based on a basic deep neural network (DNN) with three hidden layers, which served to consider the underlying principles of EPB TBM operation and simulate the TBM's response to specific operations under different geological conditions. Each hidden layer in the developed model contained 32 neurons, which were activated using the tanh function. The weights and biases were optimized using the Adam optimizer with a learning rate of 0.001. The working mechanism of the EPB TBM was incorporated into the loss function with a weight of 0.5. The specific architecture was determined through repeated trials.
[0122] (2) Training the pTD3 model
[0123] During the construction of the final 146 lining rings, the RL algorithm will be responsible for automatically adjusting the TBM's TF and SCP to optimize tunneling speed and earth pressure balance performance. To ensure the actor and critic networks' ability to describe the relationship between state, behavior, and Q-value, a DNN with 3 hidden layers and 128 neurons per hidden layer is used to train the six networks in the RL algorithm. The activation function for all neurons is the Sigmoid function, and the model parameters are optimized using the Adam optimizer with a learning rate of 0.001. In training the critic network, 10,000 randomly generated samples are used to fill a replay buffer with a batch size of 256. The discount factor ฯ is 0.98, and the soft update factor ฯ of the network is 0.02. The noise added to the behavior follows the formula ฯโฝฯ(0, 0.1โฯ). ๐ฅ The normal distribution of ) where ? ๐ฅ The standard deviations (STDs) of TF and SCP are calculated based on samples from the training environment network. The limits for TF and SCP are defined as [0, 0] and [20, 10], respectively. At each time step, the algorithm terminates if AS and CP violate the rules in Algorithm 2, or if all remaining lining rings are successfully constructed.
[0124] Step 3, Model Evaluation and Interpretation; specifically including:
[0125] Step 3.1: Evaluation and Interpretation of the Constructed Physical Information Deep Neural Network Model
[0126] (1) Evaluation of PDNN model
[0127] To verify the superiority of the developed PDNN model, its prediction performance on TBM (including AS, S) was tested. ๐ , ๐ ๐ต ,S ๐ RMSE, a20_index, and ๐ 2 The values โโare compared with the training results of the basic DNN model, as shown in Table 2. It can be clearly observed that there is almost no difference between using PDNN and DNN in terms of accuracy in predicting training data. For example, in predicting AS and S... ๐ , ๐ ๐ต and ๐ ๐ At that time, 2 The values โโwere 0.95, 0.98, 0.94, and 0.97, respectively. However, the DNN model performed poorly in predicting the AS and ฮฑ values โโfor the test data. ๐ , ๐ ๐ต and ๐ ๐ At that time, 2 The initial values โโwere 0.78, 0.95, 0.81, and 0.91, respectively. After using the model that considers physics knowledge, these numbers increased to 0.81, 0.95, 0.88, and 0.91, respectively. This indicates that overfitting is less common when training the model using PDNN, and the generalization performance when predicting test data is better. Furthermore, to evaluate the model's adherence to physics knowledge, the predictions of the two methods were compared... The distributions of these organisms were compared, and the results are as follows: Figure 5 As shown. From Figure 5 As can be seen, the average value of PDNN is 0.106 and the standard deviation is 0.099, which is significantly better than DNN.
[0128] Table 2. Model performance with and without integrated physical laws
[0129]
[0130]
[0131] (2) Interpretability of PDNN models
[0132] As mentioned earlier, to analyze the relationship between tunnel depth and intermediate CP, a comparative analysis of Shapley values โโrelated to tunnel depth was conducted using both integrated and unintegrated physical law methods. The results are as follows: Figure 6 As shown. From Figure 6 As can be seen, the model trained with PDNN performs better than the model trained with DNN. The latter has more clusters and fewer exceptions in the graph, and the relationship between CP and depth, as well as the Shapley value of depth, perfectly match the two physical rules mentioned above.
[0133] Step 3.2: Evaluation and Interpretation of the Constructed Physical Information Reinforcement Learning Model
[0134] To evaluate whether the optimization results of the proposed pTD3 algorithm are superior to manual operation, Figure 7 The cumulative rewards obtained from AS and EPB, and the total reward of the final model, are presented for both the pTD3 and manual operation methods. Table 3 summarizes the performance of TBM on AS and EPB when using the manual operation and pTD3 algorithms. Figure 8 The differences between using pTD3 and manually operated AS and CP in each lining ring are shown. Figure 9 The differences between TF and SCP when constructing each lining ring using the pTD3 algorithm and when operating manually were plotted. Figure 10 The statistics summarize the TF and SCP operations using the pTD3 algorithm and manual operation. Detailed information is shown in Table 4.
[0135] Based on the above results, it can be found that (1) the pTD3 algorithm proposed in this invention significantly improves the performance of TBM. Figure 8 As can be seen, in the construction of almost all lining rings, the AS optimized by pTD3 exceeded that of manual operation. Moreover, in a considerable number of lining rings, the CP obtained by manual operation was obviously unbalanced, while the pTD3 algorithm can control it within a very small range. Specifically, the pTD3 algorithm improved AS by 65.9% and EPB by 72.7%, and considering these two factors, the overall improvement was 69.3%. This strongly demonstrates that the present invention has important significance for improving the efficiency and stability of tunnel construction. (2) The operation implemented by the pTD3 algorithm is more stable than that of manual operation. Figure 9 This indicates that TF and SCP can be controlled within a certain range throughout the construction process. In the TF-SCP coordinate system, the data points of the pTD3 algorithm are more densely distributed. Based on numerical evaluation of the TBM's TF and SCP using pTD3 and manual operation, it can be found that the average values โโof TF and SCP are almost different between the two methods. The average TF using pTD3 and manual operation is 12.84 ร 10โปโถ. 3 ๐๐ and 12.92ร10 3 The SCP values โโusing pTD3 and manual operation were 4.12 MPa and 4.98 MPa, respectively. However, after using pTD3, the standard deviation of TF decreased from 0.90 ร 10โปโถ. 3 The value decreased to 0.69 ร 10โปโถ. 3 At 0.63 MPa, the standard deviation of SCP decreased to 0.47 MPa. The narrowing of the operating parameter range is beneficial to the stability of TBM operation. Therefore, the application of pTD3 enhances the safety of tunnel construction.
[0136] Table 3 Comparison of AS and CP between manual operation and pTD3 algorithm
[0137]
[0138] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A shield tunneling control method based on physical information reinforcement learning, characterized in that, Includes the following steps: Based on TBM operation data, S1 embeds the earth pressure balance theory into a model with DNN as the neural network architecture to construct an environmental network model of the TBM during tunnel construction, and simulates the physical information environment of the TBM during tunnel construction based on this environmental network model. Based on the physical information environment, S2 constructs a physics-based dual-delay depth deterministic algorithm model by considering the physical laws and constraints of earth pressure balance inside and outside the tunnel boring machine, non-negativity of tunneling speed, and the pressure in the middle earth chamber being between the pressure in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm. SHAP is used to evaluate and interpret the above-mentioned environmental network model, and at the same time, the dual-delay depth deterministic algorithm model is evaluated. Based on the dual-delay depth deterministic algorithm model, S3 dynamically adjusts the TBM parameters in real time to achieve the required tunneling speed and maintain the balance of earth pressure during the excavation process. It also includes the acquisition and preprocessing of the TBM's operational data; The TBM operating data includes: Tunnel depth h, total thrust TF, screw conveyor pressure SCP, tunneling speed AS, and soil chamber pressure at the top, middle and bottom of the tunnel; Step S1 includes the following steps: S11 constructs a theoretical model of earth pressure balance to maintain the balance between earth pressure and earth pressure in the earth chamber of the tunnel boring machine during tunnel excavation. S12 constructs a differential equation for the tunnel depth using the earth pressure balance theoretical model, and reconstructs the differential equation based on the earth pressure at different heights measured in real time to obtain the physical loss. S13 adds this physical loss to the observation loss, which is the squared error between the measured observation and the predicted value, to construct the loss function used to train the environment network model; S14 trains the environmental network model based on TBM operation data and the loss function, and simulates the physical information environment of TBM during tunnel construction based on the environmental network model. In step S12, the reconstruction of the differential equation based on the real-time measured soil pressure at different heights includes: ๏ผ In the formula, ๐ ๐ , ๐ ๐ and ๐ ๐ต These represent the average soil pressure values โโmeasured by sensors at the top, middle, and bottom of the tunnel cross-section, respectively, where d is the vertical distance between the top and bottom sensors. It is the pressure in the central earth chamber of the tunnel boring machine. It is the change in depth; In step S13, the new loss function used to train the PDNN model includes: ๏ผ In the formula, It is a weight used to measure physical loss; These are the top, middle, and bottom soil pressures predicted by the DNN model, respectively. is the predicted tunneling speed; h is the depth.
2. The shield tunneling control method based on physical information reinforcement learning according to claim 1, characterized in that, In step S2, the consideration of the earth pressure balance inside and outside the tunnel boring machine, the non-negativity of the tunneling speed, and the physical laws and constraints of the middle earth chamber pressure being between the top and bottom earth chamber pressures in the reward function and penalty of the dual-delay depth deterministic algorithm includes: The following formula is used to normalize the rewards for tunneling speed and soil pressure: ๏ผ In the formula, is the weight coefficient, Sigmoid() is the Sigmoid activation function, and r is the corresponding system normalized output reward.
3. The shield tunneling control method based on physical information reinforcement learning according to claim 2, characterized in that, In step S2, the loss function of the dual-delay deep deterministic algorithm model is defined by the TD error, which is the root mean square error of the Q-value at the current time step based on the Bellman equation: ๏ผ In the formula, B is the batch sampled from the response buffer; s and a are the current state and behavior; It is the next state after taking the current action. Record whether the time step has been completed; ๐พ is the loss rate; It is a Q-value function; The dual-delay deep deterministic algorithm model employs delayed updates to stabilize the learning process and also reduces the possibility of the policy getting trapped in local optima through target policy smoothing. The target behavior after policy smoothing includes: ๏ผ In the formula, ๐ is the error added to the behavior, limited to the range (โ๐, ๐), ๐ ๐ฟ๐๐ค and ๐ ๐ป๐๐โ These are the lower and upper limits of the behavioral space.
4. The shield tunneling control method based on physical information reinforcement learning according to claim 1, characterized in that, In step S2, the evaluation and interpretation of the above-mentioned environmental network model using SHAP includes: (1) Evaluation of the environmental network model: RMSE is used to measure the deviation between the predicted value and the true value, and a20_index is used with a 20% tolerance to measure the reliability of the model. 2 Used to assess the degree of agreement between predicted and actual values; (2) Interpretability of the environmental network model, using Shapley values โโto measure the importance of input features: ๏ผ In the formula, ๐ is the set containing all features, ๐ represents all non-zero entries, K is the dimension of the feature, and ๐ ๐ฅ (๐) = ๐ธ[๐(๐ฅ)|๐ฅ๐] is the expectation of the model output value.
5. A shield tunneling control system based on physical information reinforcement learning, used to implement the shield tunneling control method based on physical information reinforcement learning as described in any one of claims 1-4, characterized in that, include: The first main module: Based on TBM operation data, the earth pressure balance theory is embedded into a model with DNN as the neural network architecture to construct an environmental network model of TBM during tunnel construction, and the physical information environment of TBM during tunnel construction is simulated based on this environmental network model. The second main module: Based on the physical information environment, by considering the physical laws and constraints of the earth pressure balance inside and outside the tunnel boring machine, the non-negativity of the tunneling speed, and the fact that the pressure in the middle earth chamber is between the pressure in the top and bottom earth chambers in the reward function and penalty of the dual-delay depth deterministic algorithm, a physical-based dual-delay depth deterministic algorithm model is constructed. SHAP is used to evaluate and interpret the above-mentioned environmental network model, and at the same time, the dual-delay depth deterministic algorithm model is evaluated. The third main module: Based on the aforementioned dual-delay depth deterministic algorithm model, the TBM parameters are dynamically adjusted in real time to achieve the required tunneling speed and maintain the balance of earth pressure during the excavation process.
6. An electronic device, characterized in that, include: At least one processor, at least one memory, and a communication interface; wherein, The processor, memory, and communication interface communicate with each other; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to perform the method described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the method described in any one of claims 1 to 4.