Shield tunneling control system and method based on agent collaborative decision

By using a shield tunneling control system based on intelligent agent collaborative decision-making, the problem of collaborative intelligence among multiple subsystems in shield construction has been solved, thereby improving the stability and efficiency of the shield tunneling process, enhancing adaptability under complex geological conditions, and improving the competitiveness of the shield machine.

CN121680192APending Publication Date: 2026-03-17CHINA RAILWAY TUNNEL GROUP CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies lack a global collaborative intelligent decision-making mechanism for multiple subsystems in tunnel boring machines (TBMs), making it difficult to balance stability, efficiency, and economy in TBM construction under complex geological conditions. In particular, there are gaps in the reliance on imports for high-end core components of large TBMs and in terms of intelligent technology.

Method used

A shield tunneling control system based on agent-based collaborative decision-making is adopted, including a parameter collaborative optimization module, a collaborative control module, an attention mechanism module, a decomposition module, and a real-time decision-making and execution module. The collaborative control and decision-making of each subsystem are realized through multi-agent reinforcement learning technology.

Benefits of technology

It has improved the construction quality and safety of tunnel boring machines, enhanced their adaptability to complex geological conditions, created core technologies with independent intellectual property rights, and improved the competitiveness of tunnel boring machines in the high-end market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121680192A_ABST
    Figure CN121680192A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of shield tunneling control, and discloses a shield tunneling control system based on agent collaborative decision, which comprises a parameter collaborative optimization module, a collaborative control module, an attention mechanism module, a decomposition module and a real-time decision and execution module, according to the method, industry technology transformation is promoted, and shield control is upgraded from a traditional mode depending on artificial experience to an intelligent decision-making mode driven by data and algorithms. Through a multi-agent reinforcement learning technology, the problem of data islands is attempted to be solved, and cooperative intelligence among subsystems in the tunneling process is realized. Through precise control of the thrust vector and multi-parameter collaborative optimization, precise control of the shield tunneling track and attitude is realized, and the axis deviation is expected to be obviously reduced. The intelligent cooperative control can reduce the quantity fluctuation caused by human factors, improves the stability and reliability of tunnel construction, and guarantees the construction safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of shield tunneling control technology, and particularly relates to a shield tunneling control system and method based on intelligent agent collaborative decision-making. Background Technology

[0002] While my country has achieved approximately 90% domestic production of tunnel boring machine (TBM) components, large TBMs with a diameter of 12 meters or more still rely on imports for high-end core components such as main bearings. Furthermore, there is a gap in intelligent systems compared to foreign "cloud-based TBM management systems," and the lack of intelligent core modules puts my country at a disadvantage in the high-end market. The construction process still relies heavily on human experience and operator feel, leading to fluctuations in the accuracy of tunneling posture control and segment assembly quality due to human error, making it difficult to guarantee the stability of construction quality. Simultaneously, the large amount of data generated during TBM construction is scattered between the owner and the contractor, forming "data silos." Inconsistent data standards and difficulties in cross-platform sharing hinder the development of intelligent decision-making systems.

[0003] Existing publicly available technologies, such as US Patent US 2021 / 0287459 A1, "Shield Tunneling Deviation Rectifying Intelligent Decision-making Method Based on Reinforcement Learning," focus on intelligent correction of tunneling deviations, but fail to resolve the conflicting objectives among multiple subsystems during shield tunneling. The four major subsystems—propulsion, cutterhead, screw conveyor, and grouting—exist an inherent contradiction between efficiency, stability, and cost, forming a "coordination paradox," where locally optimal decisions lead to global imbalances. Taking propulsion and cutterhead as an example, the propulsion system prioritizes tunneling efficiency, while the cutterhead system needs to maintain stable torque. Existing independent control logic easily leads to torque overload and shutdown, reducing overall efficiency.

[0004] Furthermore, a conflict exists between the screw conveyor and the soil chamber pressure: a conflict of "slag discharge versus pressure stabilization." Traditional control methods often employ fixed ratios or threshold triggers, failing to establish a closed-loop coordination mechanism with the propulsion system. When the formation moisture content changes drastically, the imbalance between the slag discharge rate and the tunneling rate causes a sudden drop in soil chamber pressure, leading to decreased face stability and excessive surface settlement. A conflict of "quantity-efficiency versus cost" also exists between the grouting system and surface settlement control. Fixed grouting volume or lag correction modes cannot achieve dynamic balance, easily resulting in material waste, increased formation reinforcement costs, and project delays.

[0005] In summary, existing technologies generally lack a global collaborative intelligent decision-making mechanism for multiple subsystems in tunnel boring machines (TBMs), failing to achieve dynamic balance and real-time coupling of "thrust-torque," "slag removal-pressure stabilization," and "sinking control-cost." This makes it difficult to balance stability, efficiency, and economy in TBM construction under complex geological conditions, becoming a key bottleneck restricting the intelligent development of TBMs in my country. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a shield tunneling control system and method based on intelligent agent collaborative decision-making.

[0007] This invention is implemented as follows: a shield tunneling control system based on intelligent agent collaborative decision-making, comprising sequential data interaction:

[0008] The module includes a parameter collaborative optimization module, a collaborative control module, an attention mechanism module, a decomposition module, and a real-time decision-making and execution module.

[0009] The parameter co-optimization module is used to integrate the precise control of the thrust vector with key parameters such as cutterhead torque and tunneling speed to achieve synchronous optimization of tunneling attitude and efficiency.

[0010] The collaborative control module is used to establish control interfaces for the propulsion system, cutterhead system, screw conveyor system and grouting system as independent intelligent agents;

[0011] The attention mechanism module is used to allocate agent weights according to different construction stages and dynamically focus on key information.

[0012] The decomposition module is used to distribute the global reward signal to each agent to solve the credit allocation problem;

[0013] The real-time decision-making and execution module predicts future stratigraphic trends through a long short-term memory network, and generates control commands through a reinforcement learning model before sending them to the execution mechanism.

[0014] Furthermore, the parameter collaborative optimization module:

[0015] (1) Data acquisition and preprocessing;

[0016] (2) Establishment of parameter fusion model;

[0017] (3) Attitude control and tunneling efficiency are optimized simultaneously.

[0018] Furthermore, the data acquisition and preprocessing include:

[0019] Thrust vector data acquisition: Real-time acquisition of the magnitude and direction data of the thrust vector through high-precision sensors installed on the propulsion system;

[0020] Data acquisition of cutterhead torque and tunneling speed: Torque sensors and speed sensors are installed on the cutterhead system and tunneling drive device to collect cutterhead torque and tunneling speed data respectively;

[0021] Data preprocessing: The collected raw data is filtered and denoised to eliminate noise and outliers.

[0022] Furthermore, the parameter fusion model is established as follows:

[0023] Establish a mathematical model: Based on the physical relationship between thrust vector, cutterhead torque, and tunneling speed, establish a mathematical model that integrates parameters; for example, consider the influence of thrust vector on cutterhead torque and tunneling speed, as well as the interaction between cutterhead torque and tunneling speed;

[0024] Model training: The established mathematical model is trained using historical construction data, and the model parameters are adjusted through optimization algorithms.

[0025] Furthermore, the attitude control and tunneling efficiency are optimized simultaneously:

[0026] Attitude control target setting: Based on the tunnel design axis and construction requirements, set the control targets for the tunneling attitude, including the tunneling direction and pitch angle;

[0027] Setting tunneling efficiency targets: Taking into account project progress and cost factors, set tunneling efficiency targets, including daily tunneling length and tunneling speed range;

[0028] Optimization algorithm application: Multi-objective optimization algorithms, including genetic algorithms and particle swarm optimization, are used to optimize the parameter fusion model.

[0029] Furthermore, the attention mechanism module:

[0030] 1) Extraction of key information

[0031] Construction phase division: Based on the characteristics of tunnel construction, the construction process is divided into different phases, including the excavation phase, the support phase, and the grouting phase;

[0032] Key information definition: Define key information for each construction stage; for example, in the tunneling stage, key information may include ground hardness, cutterhead torque, and tunneling speed; in the support stage, key information may include support pressure and support deformation.

[0033] 2) Attention weight calculation

[0034] Attention model establishment: An attention mechanism model is adopted, including a self-attention model and a multi-head attention model, to calculate the attention weights of key information and agents at different construction stages;

[0035] Weight calculation method: Based on the importance of key information and the agent to the global goal, the attention weight is calculated through the attention model; the larger the weight, the more important the information or agent is in the current construction stage.

[0036] Given an input sequence , where xi represents the feature vector of the i-th key information or agent, with dimension d; the calculation process of the self-attention model is as follows:

[0037] Q=XW Q K=XW K V=XW V

[0038] Among them W Q W K W V It is a learnable weight matrix with dimensions d×d. k , d×d k , d×d v ; usually d k =d v ;

[0039] Calculate attention score S

[0040] S=QK T

[0041] The attention score S is an n×n matrix, where Sij represents the similarity between the i-th query vector and the j-th key vector;

[0042] Calculate attention weight α

[0043] Normalizing the attention score S yields the attention weight α:

[0044]

[0045] Where the softmax function is α is also an n×n matrix, and αij represents the attention weight of the i-th key information or agent to the j-th key information or agent.

[0046] 3) Improved accuracy of decision-making

[0047] Information fusion: Based on the calculated attention weights, key information and agents at different construction stages are fused to highlight the role of important information and agents;

[0048] Decision generation: Based on the fused information, more accurate decision instructions are generated to guide the various intelligent agents to carry out collaborative control;

[0049] The decomposition module:

[0050] (a) Definition of global reward signal

[0051] Global target quantification: The global targets of tunneling efficiency, energy consumption, axis control, and surface settlement are quantified and transformed into measurable indicators;

[0052] Global reward function establishment: Based on the quantified global goal, a global reward function is established; the value of the reward function reflects the degree to which the current construction status achieves the global goal, and the larger the value, the better the achievement.

[0053] (b) Value function decomposition

[0054] Decomposition method selection: Value function decomposition technology is adopted, including linear decomposition and nonlinear decomposition methods, to reasonably distribute the global reward signal to each subsystem agent;

[0055] Decomposition principle formulation: Formulate decomposition principles to ensure that the reward signals allocated to each agent accurately reflect its contribution to the global goal; for example, allocate more reward signals to agents that contribute more to tunneling efficiency.

[0056] (c) Solving the credit allocation problem

[0057] Credit assessment: The credit of each agent is assessed based on the reward signals assigned to each agent; the credit assessment results reflect the performance and contribution of each agent in the collaborative process.

[0058] Ensuring the effectiveness of collaboration: By decomposing the value function and evaluating credit, the problem of credit allocation in multi-agent systems is solved, ensuring that each agent can work actively and collaboratively to achieve the optimal global goal;

[0059] Another objective of this invention is to provide a shield tunneling control method based on multi-agent reinforcement learning, comprising:

[0060] Step 1: The precise control of thrust vector is deeply integrated with key parameters such as cutterhead torque and tunneling speed through the parameter co-optimization module; thus achieving synchronous optimization of attitude control and tunneling efficiency.

[0061] Step 2: The propulsion system, cutterhead system, screw conveyor system, and grouting system are treated as independent intelligent agents through the collaborative control module; each intelligent agent has its own local objectives, and the global objectives such as tunneling efficiency, energy consumption, axis control, and surface settlement are jointly optimized through the collaborative mechanism.

[0062] Step 3: By employing the attention mechanism module, the system can dynamically focus on key information and agents at different construction stages, thereby improving the accuracy of decision-making.

[0063] Step 4: By using the value function decomposition technique through the decomposition module, the global reward signal is reasonably distributed to each subsystem agent, solving the credit allocation problem in multi-agent systems and ensuring the effectiveness of collaboration.

[0064] Step 5: The real-time decision-making and execution module collects sensor data each time; the LSTM network predicts the formation change trend in the next 5 seconds; the MARL model generates optimization instructions, which are then sent to the execution mechanism after security verification.

[0065] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the shield tunneling control method based on multi-agent reinforcement learning.

[0066] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the shield tunneling control method based on multi-agent reinforcement learning.

[0067] Another objective of this invention is to provide an information data processing terminal for implementing the shield tunneling control system based on intelligent agent collaborative decision-making.

[0068] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0069] (1) Promoting technological transformation in the industry: Upgrading shield tunneling control from the traditional model that relies on human experience to an intelligent decision-making model driven by data and algorithms. By using multi-agent reinforcement learning technology, we attempt to solve the problem of data silos and achieve collaborative intelligence among various subsystems in the tunneling process.

[0070] (2) Improve construction quality and safety: Through precise control of thrust vector and multi-parameter collaborative optimization, the tunneling trajectory and attitude of the shield can be precisely controlled, which is expected to significantly reduce axis deviation. Intelligent collaborative control can reduce quality fluctuations caused by human factors, improve the stability and reliability of tunnel construction, and ensure construction safety.

[0071] (3) Enhance adaptability to special working conditions: Through deep reinforcement learning algorithms, the shield tunneling system is equipped with adaptive optimization capabilities under complex geological conditions (such as the limit radius of a narrow space), effectively responding to changes in strata.

[0072] (4) Building core competitiveness: The multi-agent reinforcement learning collaborative control system involved in this patent directly addresses the pain points of intelligentization in the industry, which helps to build core technologies with independent intellectual property rights and enhance the competitiveness of my country's tunnel boring machines in the high-end market.

[0073] The technical solution of this invention solves a technical problem that people have long desired to solve but have never been able to achieve: the multi-agent reinforcement learning collaborative control system involved in this patent directly addresses the pain points of intelligentization in the industry, helps to create core technologies with independent intellectual property rights, and enhances the competitiveness of my country's tunnel boring machines in the high-end market. Attached Figure Description

[0074] Figure 1 This is a block diagram of a shield tunneling control system based on intelligent agent collaborative decision-making provided in an embodiment of the present invention.

[0075] Figure 2 This is a flowchart of the parameter collaborative optimization module method provided in the embodiment of the present invention.

[0076] Figure 3 This is a flowchart of the attention mechanism module method provided in an embodiment of the present invention.

[0077] Figure 4 This is a flowchart of a shield tunneling control method based on multi-agent reinforcement learning provided in an embodiment of the present invention.

[0078] Figure 1 The module consists of: 1. Parameter collaborative optimization module; 2. Collaborative control module; 3. Attention mechanism module; 4. Decomposition module; and 5. Real-time decision-making and execution module. Detailed Implementation

[0079] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0080] like Figure 1 As shown, an embodiment of the present invention provides a shield tunneling control system based on intelligent agent collaborative decision-making, comprising:

[0081] Parameter collaborative optimization module 1, collaborative control module 2, attention mechanism module 3, decomposition module 4, real-time decision-making and execution module 5;

[0082] The parameter co-optimization module 1, connected to the co-control module 2, is used to deeply integrate the precise control of the thrust vector with key parameters such as cutterhead torque and tunneling speed; thereby achieving synchronous optimization of attitude control and tunneling efficiency.

[0083] The collaborative control module 2, connected to the attention mechanism module 3 and the decomposition module 4, is used to treat the propulsion system, cutterhead system, screw conveyor system and grouting system as independent intelligent agents; each intelligent agent has its own local goal, and through the collaborative mechanism, they jointly optimize global goals such as tunneling efficiency, energy consumption, axis control and surface settlement.

[0084] Attention mechanism module 3, connected to collaborative control module 2, is used to employ an attention mechanism to enable the system to dynamically focus on key information and agents at different construction stages, thereby improving the accuracy of decision-making.

[0085] The decomposition module 4, connected to the collaborative control module 2, is used to rationally distribute the global reward signal to each subsystem agent through value function decomposition technology, thereby solving the credit allocation problem in multi-agent systems and ensuring the effectiveness of collaboration.

[0086] The real-time decision-making and execution module 5, connected to the attention mechanism module 3 and the decomposition module 4, is used to predict the formation change trend in the next 5 seconds after each sensor data acquisition; the MARL model generates optimization instructions, which are then sent to the execution mechanism after security verification.

[0087] The shield tunneling control system based on intelligent agent collaborative decision-making provided in this embodiment of the invention is implemented as follows: During shield tunneling, the system continuously collects operating status and ground response data through a high-frequency sensor array (including earth pressure, cutterhead torque, propulsion cylinder attitude, screw conveyor load, motor power, synchronous grouting pressure, and surface settlement monitoring point feedback, etc.) installed on the shield machine, and inputs this data into parameter collaborative optimization module 1. Parameter collaborative optimization module 1 adopts multi-parameter coupled modeling, linking the shield machine thrust vector control with cutterhead torque and tunneling speed for solution. Instead of local optimization based on a single indicator (such as attitude deviation or propulsion speed), it establishes a joint objective function of "attitude control-tunneling efficiency," collaboratively optimizing key parameters such as leading-edge earth pressure stability, head attitude (pitch angle, yaw angle), cutterhead cutting load, and propulsion rate, enabling the shield machine to maintain efficient tunneling while satisfying axis control and controlled ground disturbance. The resulting collaborative parameter objective is then sent in real-time to collaborative control module 2 as a global scheduling reference.

[0088] The collaborative control module 2 models the shield propulsion system, cutterhead system, screw conveyor system, and synchronous grouting system as four intelligent agents with decision-making capabilities. Each agent performs self-learning and self-adjustment based on its own local objectives: the propulsion system's main objectives are attitude correction, axis control, and earth pressure balance; the cutterhead system's main objectives are cutterhead cutting torque stability and tool wear suppression; the screw conveyor system's main objective is the stable matching of slag discharge volume and chamber pressure; and the grouting system's main objective is the suppression of surface settlement by the synchronous grouting volume-grouting pressure curve. Through the multi-agent collaborative mechanism, these subsystems are not controlled independently, but rather undergo joint strategy updates under global objectives (tunneling efficiency, energy consumption constraints, surface settlement control, and attitude error constraints). To avoid mutual interference between subsystems in complex strata or curved routes, the collaborative control module 2 introduces an attention mechanism module 3 and a decomposition module 4. The attention mechanism module 3 dynamically weights key information at different time periods. For example, in soft soil high settlement sensitive areas, priority is given to the grouting-earth pressure coupling state, and in hard interlayers or gravel strata, priority is given to the cutterhead impact torque and propulsion resistance, thereby improving the targeting of parameter tuning. The decomposition module 4 adopts a global reward allocation strategy based on value function decomposition, which allocates the overall system reward (such as "keeping settlement ≤ threshold and propulsion rate ≥ target value") to each agent according to the interpretable contribution, solving the credit allocation problem of "who caused the problem cannot be attributed" in traditional multi-agent control, thereby ensuring that each subsystem does not pursue local optima on its own, but rather collaborates towards a unified tunneling quality goal.

[0089] At the execution level, the real-time decision-making and execution module 5 uses the acquisition cycle as the control cycle and performs inference processing on the latest complete sensor data frame. First, the long short-term memory (LSTM) network built into module 5 predicts the stratum disturbance trend on the timescale of 5 seconds in front of the propulsion head, including the soil plastic flow trend, chamber pressure fluctuation trend, and attitude deviation trend, predicting the evolution of the working condition in advance from the perspective of "what will happen". Subsequently, based on the above prediction results, the trained multi-agent reinforcement learning (MARL) policy network generates a set of joint control instructions, including: thrust correction amount and thrust distribution matrix for each propulsion cylinder, cutterhead speed and torque upper limit adjustment amount, slag discharge rate target curve of screw conveyor, and dynamic set value of synchronous grouting volume / pressure, etc. Before the instruction set is issued, it will pass through the safety verification submodule. The verification content includes: whether it exceeds the allowable working range of the equipment, whether it causes overshoot risk to attitude control, and whether it may cause sudden pressure drop or backflow due to excessive grouting, thereby preventing unsafe operations from being directly written into the actuator. Only verified instructions are written into the servo control loop of each execution unit in real time, realizing online, feedforward, and collaborative shield tunneling control to maintain a stable, precise, and low-disturbance tunneling process in urban sections with complex geology, settlement limits, and strict spatial attitude requirements.

[0090] like Figure 2 As shown, the parameter collaborative optimization module provided in this embodiment of the invention:

[0091] S101, Data Acquisition and Preprocessing;

[0092] S102, Establishment of parameter fusion model;

[0093] S103, attitude control and tunneling efficiency are optimized simultaneously.

[0094] Data acquisition and preprocessing provided in this embodiment of the invention:

[0095] Thrust vector data acquisition: Real-time acquisition of the magnitude and direction data of the thrust vector through high-precision sensors installed on the propulsion system;

[0096] Data acquisition of cutterhead torque and tunneling speed: Torque sensors and speed sensors are installed on the cutterhead system and tunneling drive device to collect cutterhead torque and tunneling speed data respectively;

[0097] Data preprocessing: The collected raw data is filtered and denoised to eliminate noise and outliers.

[0098] The parameter fusion model established according to the embodiments of the present invention:

[0099] Establish a mathematical model: Based on the physical relationship between thrust vector, cutterhead torque, and tunneling speed, establish a mathematical model that integrates parameters; for example, consider the influence of thrust vector on cutterhead torque and tunneling speed, as well as the interaction between cutterhead torque and tunneling speed;

[0100] Model training: The established mathematical model is trained using historical construction data, and the model parameters are adjusted through optimization algorithms.

[0101] The embodiments of this invention provide simultaneous optimization of attitude control and tunneling efficiency:

[0102] Attitude control target setting: Based on the tunnel design axis and construction requirements, set the control targets for the tunneling attitude, including the tunneling direction and pitch angle;

[0103] Setting tunneling efficiency targets: Taking into account project progress and cost factors, set tunneling efficiency targets, including daily tunneling length and tunneling speed range;

[0104] Optimization algorithm application: Multi-objective optimization algorithms, including genetic algorithms and particle swarm optimization, are used to optimize the parameter fusion model.

[0105] like Figure 3 As shown, the attention mechanism module provided in this embodiment of the invention:

[0106] S201, Key Information Extraction

[0107] Construction phase division: Based on the characteristics of tunnel construction, the construction process is divided into different phases, including the excavation phase, the support phase, and the grouting phase;

[0108] Key information definition: Define key information for each construction stage; for example, in the tunneling stage, key information may include ground hardness, cutterhead torque, and tunneling speed; in the support stage, key information may include support pressure and support deformation.

[0109] S202, Attention Weight Calculation

[0110] Attention model establishment: An attention mechanism model is adopted, including a self-attention model and a multi-head attention model, to calculate the attention weights of key information and agents at different construction stages;

[0111] Weight calculation method: Based on the importance of key information and the agent to the global goal, the attention weight is calculated through the attention model; the larger the weight, the more important the information or agent is in the current construction stage.

[0112] Given an input sequence , where xi represents the feature vector of the i-th key information or agent, with dimension d; the calculation process of the self-attention model is as follows:

[0113] Q=XW Q K=XW K V=XW V

[0114] Among them W Q W K W V It is a learnable weight matrix with dimensions d×d. k , d×d k , d×d v ; usually d k =d v ;

[0115] Calculate attention score S

[0116] S=QK T

[0117] The attention score S is an n×n matrix, where Sij represents the similarity between the i-th query vector and the j-th key vector;

[0118] Calculate attention weight α

[0119] Normalizing the attention score S yields the attention weight α:

[0120]

[0121] Where the softmax function is α is also an n×n matrix, and αij represents the attention weight of the i-th key information or agent to the j-th key information or agent.

[0122] S203, Improved Decision-Making Accuracy

[0123] Information fusion: Based on the calculated attention weights, key information and agents at different construction stages are fused to highlight the role of important information and agents;

[0124] Decision generation: Based on the fused information, more accurate decision instructions are generated to guide the various intelligent agents to carry out collaborative control;

[0125] The decomposition module:

[0126] (a) Definition of global reward signal

[0127] Global target quantification: The global targets of tunneling efficiency, energy consumption, axis control, and surface settlement are quantified and transformed into measurable indicators;

[0128] Global reward function establishment: Based on the quantified global goal, a global reward function is established; the value of the reward function reflects the degree to which the current construction status achieves the global goal, and the larger the value, the better the achievement.

[0129] (b) Value function decomposition

[0130] Decomposition method selection: Value function decomposition technology is adopted, including linear decomposition and nonlinear decomposition methods, to reasonably distribute the global reward signal to each subsystem agent;

[0131] Decomposition principle formulation: Formulate decomposition principles to ensure that the reward signals allocated to each agent accurately reflect its contribution to the global goal; for example, allocate more reward signals to agents that contribute more to tunneling efficiency.

[0132] (c) Solving the credit allocation problem

[0133] Credit assessment: The credit of each agent is assessed based on the reward signals assigned to each agent; the credit assessment results reflect the performance and contribution of each agent in the collaborative process.

[0134] Ensuring the effectiveness of collaboration: By decomposing the value function and evaluating credit, the problem of credit allocation in multi-agent systems is solved, ensuring that each agent can work actively and collaboratively to achieve the optimal global goal;

[0135] like Figure 4As shown in the figure, the shield tunneling control method based on multi-agent reinforcement learning provided by this embodiment of the invention includes:

[0136] The S301 integrates precise thrust vector control with key parameters such as cutterhead torque and tunneling speed through a parameter co-optimization module, achieving synchronous optimization of attitude control and tunneling efficiency.

[0137] S302 treats the propulsion system, cutterhead system, screw conveyor system, and grouting system as independent intelligent agents through a collaborative control module; each intelligent agent has its own local objectives, and together they optimize global objectives such as tunneling efficiency, energy consumption, axis control, and surface settlement through a collaborative mechanism.

[0138] S303 employs an attention mechanism through its attention mechanism module, enabling the system to dynamically monitor key information and agents at different construction stages, thereby improving the accuracy of decision-making.

[0139] S304, through the decomposition module and the use of value function decomposition technology, rationally distributes the global reward signal to each subsystem agent, solves the credit allocation problem in multi-agent systems, and ensures the effectiveness of collaboration;

[0140] The S305 collects sensor data every time through the real-time decision-making and execution module; predicts the formation change trend in the next 5 seconds through the LSTM network; generates optimization instructions through the MARL model, and sends them to the execution mechanism after security verification.

[0141] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the shield tunneling control method based on multi-agent reinforcement learning.

[0142] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the shield tunneling control method based on multi-agent reinforcement learning.

[0143] Another objective of this invention is to provide an information data processing terminal for implementing the shield tunneling control system based on intelligent agent collaborative decision-making.

[0144] Specific implementation of the present invention:

[0145] (I) Example 1: Urban subway tunnel construction

[0146] Scenario: A section of Beijing Metro Line 12, with a stratum consisting of alternating layers of silty clay and gravel, and a radius of curvature of 400m.

[0147] Implementation steps:

[0148] Deploy lidar and fiber optic sensors with a sampling frequency of 50Hz;

[0149] Initialize MARL model parameters: learning rate 0.001, discount factor 0.95;

[0150] A conservative strategy was adopted for the first 100 meters of construction (with a safety weight of 0.7 in the reward function), and then the efficiency-first strategy was switched to (with a safety weight of 0.3).

[0151] Effect:

[0152] The average daily tunneling progress is 18 rings (compared to 12 rings using the traditional method).

[0153] The maximum surface settlement is 4.2 mm (design requirement <10 mm).

[0154] (II) Example 2: Construction of cross-river and cross-sea tunnels

[0155] Scene: A crossing of the Yangtze River, water pressure 0.85 MPa, stratum is highly permeable sand.

[0156] Key technologies:

[0157] The intelligent agent of the grouting system dynamically adjusts the ratio of the two liquid grouts, reducing the gelation time from 45s to 18s.

[0158] Adopt a risk-sensitive reward function:

[0159] , where ΔP is the mud pressure fluctuation.

[0160] Effect:

[0161] The incidence of water inrush accidents has been reduced to 0;

[0162] Leakage at the segment joints is less than 0.2 L / min.

[0163] I. Specific application areas or related products of this invention.

[0164] 1. Application Areas

[0165] Urban rail transit: Suitable for subway tunnels in densely populated urban areas, reducing the impact on surrounding buildings;

[0166] Railway and highway tunnels: dealing with adverse geological conditions such as fault fracture zones and karst in mountain tunnels;

[0167] Underwater tunnels: Solving the problem of attitude control during high water pressure and long-distance tunneling.

[0168] 2. Related Products

[0169] Intelligent tunnel boring machine: A tunnel boring machine that integrates this system (such as China Railway Construction's "Dream") can achieve autonomous tunneling and parameter self-optimization;

[0170] Construction monitoring platform: It connects with CCCC Tianhe's "Shield Tunneling Cloud" platform to provide real-time decision support;

[0171] Aftermarket services: Predictive maintenance models trained on construction data extend equipment lifespan.

[0172] Example 1: Hardware and software integration implementation of a multi-agent reinforcement learning system

[0173] A central control server is deployed as the master control node on a typical tunnel boring machine (TBM) platform, employing an NVIDIA Jetson AGX industrial-grade computing module to perform real-time inference of the reinforcement learning model. The system communicates with the PLC controller via a CAN bus. Each agent corresponds to the propulsion cylinder group, cutterhead servo drive, screw conveyor, and grouting pump group, respectively. Each agent node has a built-in ARM Cortex-A72 microprocessor to perform local decision-making tasks. The master control node and agent nodes establish a data synchronization mechanism via TSN industrial Ethernet, ensuring that the sampling and command transmission delay does not exceed 10 milliseconds, thereby achieving distributed multi-agent collaborative control.

[0174] Example 2: Model Training and Deployment Scheme for Parameter Co-optimization

[0175] Historical construction data under typical geological conditions was collected in a laboratory environment, including multi-dimensional time-series samples such as thrust vector, cutterhead torque, tunneling speed, and ground resistance. A nonlinear multi-objective optimization model was constructed using Python and the PyTorch framework, with attitude error and energy consumption function as dual-objective optimization inputs. The fusion network parameters were trained using a particle swarm optimization algorithm, and the model was deployed to the central server on-site in ONNX format after training. During runtime, the module performs forward inference every 0.5 seconds to output the thrust vector correction, achieving real-time control with attitude deviation less than 0.1 degrees.

[0176] Example 3: Implementation of Attention Mechanism and Multi-Stage Information Fusion

[0177] The system defines different key feature sets for each of the tunneling, support, and grouting stages. During tunneling, it focuses on cutterhead torque, ground hardness, and advance speed; during support, it focuses on the shield tail gap and grouting pressure; and during grouting, it focuses on grout density and surface settlement rate. A self-attention network is used to calculate the weight of each feature to the global objective, and a multi-head mechanism is employed to fuse feature information. When switching between construction stages, the system automatically resets the attention parameters based on the stage discrimination signal, ensuring the model always focuses on the most influential variables, thus improving the adaptability and accuracy of decision-making.

[0178] Example 4: Implementation of Value Function Decomposition and Credit Allocation Algorithm

[0179] The system employs a nonlinear value decomposition architecture (QMIX-type structure), where the global value function is formed by a weighted combination of multiple local value functions. Each agent performs actions in its local environment and receives immediate rewards. The central coordinator calculates the global reward and decomposes it into sub-reward signals using a neural network. A credit evaluation network adaptively adjusts the reward coefficients of each agent based on its historical contribution, achieving differentiated incentives for the propulsion, cutterhead, auger conveyor, and grouting subsystems, thereby improving the overall convergence speed and the stability of the tunneling process.

[0180] Example 5: Implementation of Real-time Decision-Making and Secure Execution Mechanism

[0181] The system deploys a safety isolation module at the on-site control layer, ensuring the safety of control commands through a dual-channel decision verification mechanism. After collecting sensor data once per cycle, the Long Short-Term Memory (LSTM) network predicts the hardness and resistance trends of the strata in the next 5 seconds. The MARL model outputs multi-dimensional control commands, including thrust correction, cutterhead speed, screw conveyor speed, and grouting pressure adjustment. The safety module verifies the commands through threshold judgment, anomaly detection, and logic interlocking before writing them into the PLC for execution. If a sudden change in strata or an excessive attitude deviation is detected, the system automatically switches to safety mode, executing tunneling deceleration and pressure stabilization control to ensure the safety of equipment and tunnel structure.

[0182] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0183] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A shield tunneling control system based on collaborative decision of agents, characterized in that, The parameter collaborative optimization module, the collaborative control module, the attention mechanism module, the value decomposition module, and the real-time decision and execution module are sequentially performed for data interaction. The parameter collaborative optimization module is used for fusing precise control of the thrust vector with key parameters of the cutterhead torque and the tunneling speed, so as to realize synchronous optimization of the tunneling posture and efficiency. The collaborative control module is used for establishing control interfaces of the propulsion system, the cutterhead system, the screw conveyor system, and the grouting system as independent agents respectively. The attention mechanism module is used for dynamically focusing on key information according to the agent weight distribution in different construction stages. The value decomposition module is used for distributing the global reward signal to each agent to solve the credit distribution problem. The real-time decision and execution module predicts the trend of the stratum parameter change through a long short-term memory network, generates a control instruction through a reinforcement learning model, and then issues the control instruction to an execution mechanism, so as to realize adaptive adjustment of the shield tunneling parameters.

2. The control system of claim 1, wherein, The software part of the system is deployed on a central computing unit, and the central computing unit includes a data acquisition submodule, a feature extraction submodule, a reinforcement learning training submodule, and a safety decision submodule. The hardware part of the system includes a propulsion cylinder thrust sensor, a cutterhead torque sensor, a screw conveyor flowmeter, a grouting pressure sensor, and a posture measurement unit. The central computing unit and each execution device realize millisecond-level data synchronization through an industrial Ethernet.

3. The control system of claim 1, wherein, The parameter collaborative optimization module includes a data acquisition unit, a data preprocessing unit, and a parameter fusion unit. The parameter fusion unit establishes a multivariate nonlinear fusion model based on the coupling relationship between the thrust vector and the cutterhead torque. The fusion model is trained through a gradient descent optimization algorithm of historical data, so as to realize synchronous minimization of the posture deviation and the propulsion energy consumption.

4. The control system of claim 1, wherein, The collaborative control module adopts a multi-agent reinforcement learning structure. Each agent updates the strategy through a shared global state vector. A central learner maintains a global value function, and each agent executes a local action strategy. The convergence and stability of multi-objective reinforcement learning are realized through a policy gradient method.

5. The control system of claim 4, wherein, The system adopts a weighted summation type multi-objective reward function to balance multiple objectives of the tunneling efficiency, energy consumption optimization, axis deviation control, and ground settlement control. The weight coefficients of each sub-objective sum to 1, and are dynamically adjusted by the central learner according to the construction stage. In the gradient calculation process, a multi-objective constraint term is introduced to balance the influence of the advantage function and the constraint function, so as to realize stable strategy updating.

6. A shield tunneling control method based on collaborative decision of agents, characterized in that, The system includes the following steps: (1) collecting the thrust vector, the cutterhead torque, the tunneling speed, and the stratum parameters; (2) jointly optimizing the posture and the tunneling efficiency based on the parameter collaborative optimization module; (3) under the multi-agent framework, each subsystem independently decides and optimizes the tunneling performance through a global coordination mechanism; (4) the attention mechanism dynamically adjusts the agent weight according to the construction stage; (5) the value decomposition module distributes the global reward signal to each agent; (6) the prediction module predicts the trend of the stratum parameter change based on a long short-term memory network; (7) a reinforcement learning model generates a control instruction and issues the control instruction to an execution mechanism, so as to realize real-time tunneling control.

7. The method of claim 6, wherein, The reinforcement learning model adopts an adaptive discount factor mechanism to dynamically balance long-term rewards and short-term responses; When the amplitude of stratum disturbance increases, the discount factor decreases to improve response speed; When the excavation process is stable, the discount factor increases to enhance long-term excavation quality.

8. The method of claim 6, wherein, The attention mechanism module is implemented through a feature weighting network; The network input is a key parameter matrix of the current construction stage, and the output is the attention weight of each parameter; The similarity between the query vector, key vector and value vector is calculated to obtain the weighting coefficient, thereby adaptively focusing on the most relevant feature information.

9. The method of claim 6, wherein, The value decomposition module adopts a nonlinear value function decomposition network to represent the global reward function as a weighted combination of local value functions of each agent; The weighting coefficient is updated in real time through a credit evaluation network to dynamically adjust the contribution among multiple agents.

10. The method of claim 6, wherein, The real-time decision and execution process includes a sensor interface layer, a prediction layer, a decision layer and an execution layer; The prediction layer predicts the trend of stratum parameter changes within the next 5 seconds; The decision layer outputs control actions based on reinforcement learning; The execution layer outputs to the hydraulic system, cutter drive system, screw conveyor and grouting system after safety verification of the control instructions, thereby realizing closed-loop control of shield posture and propulsion parameters.

Citation Information

Patent Citations

  • Digital twin systems and methods for transportation systems

    US20210287459A1