Special-shaped tunnel structure intelligent agent and intelligent agent control method

By employing an active control method for intelligent agents in irregularly shaped tunnel structures, and utilizing sensing units and reinforcement learning models to optimize the stress and deformation of tunnel structures, the control challenges of traditional reinforcement technologies under complex stress environments have been solved, thereby improving construction safety and efficiency.

CN120867779APending Publication Date: 2025-10-31TONGJI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511008180.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional reinforcement techniques fail to proactively adjust and control the complex stress changes of tunnel structures, making it difficult to adapt to the combined effects of excavation sequence and external water and soil pressure. Furthermore, intelligent tunnel construction systems lack structural stiffness matching and coordinated stress adaptability during dynamic construction processes.

Method used

An intelligent agent with an irregular tunnel structure is adopted, including a main structure, a sensing unit and a control unit. It achieves adaptive optimization through a reinforcement learning model, and combines a servo actuator and a PLC controller to achieve active control and stress deformation optimization.

Benefits of technology

Active stress control of irregular tunnel structures has been achieved, which has improved construction safety and efficiency, reduced construction costs, and reduced the risk of structural deformation and stress concentration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120867779A_ABST
    Figure CN120867779A_ABST
Patent Text Reader

Abstract

The invention discloses a special-shaped tunnel structure intelligent agent and an intelligent agent control method, and belongs to the field of tunnel engineering intelligent construction.The special-shaped tunnel structure intelligent agent comprises a tunnel structure intelligent agent composed of a shield segment structure, an intelligent sensing unit and a self-adaptive servo control unit; and an action strategy is automatically generated to carry out stress control, so that the stress deformation self-adaption reaches an optimized integrated structure system. The sensing unit is a laser range finder, an MEMS inclinometer and other deformation and mechanical sensing sensors arranged in a pipe piece and a servo control and inner supporting rod piece. The control unit is composed of active control equipment such as a permanent or temporary support servo actuator arranged in a stress concentration area. The stress deformation self-adaptive control of the whole structural intelligent body is completed by sensing the environmental state of the stress deformation of the structural system through the sensing unit and carrying out the autonomous generation and active application of the action strategy of the control unit through reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent construction of tunnel engineering, and particularly relates to an intelligent agent for irregular tunnel structures and an intelligent agent control method. Background Technology

[0002] With the rapid advancement of urbanization and urban renewal, more and more irregularly shaped tunnel structures will emerge during the reconstruction and expansion of existing underground structures at the merging and diverging sections of underground road ramps. However, these structures have complex stress characteristics, and conventional methods of underground excavation pose high risks and significant environmental impacts, especially in soft soil areas where they are prone to excessive ground settlement and collapse. Regarding the safety control of these irregularly shaped structures, passive temporary supports are often used to resist external soil and water loads, but these are costly and have low reliability. Meanwhile, with the gradual rise of intelligent sensing and control in underground engineering, intelligent methods such as deep learning models can be used for real-time feedback, improving the reliability of tunnel structure reinforcement.

[0003] The current problem is:

[0004] Traditional reinforcement techniques are temporary support structures that are not permanently integrated with the tunnel design. They are also passive supports, meaning they only respond to external forces and lack the ability to actively adjust and control. This makes them difficult to adapt to the complex dynamic changes in stress caused by the combined effects of excavation steps and external water and soil pressures.

[0005] Currently, most intelligent tunnel construction systems focus solely on monitoring, sensing, and information feedback, without being linked with control systems. This results in insufficient consideration of structural stiffness matching, structural coordination, and system stress adaptability during dynamic construction processes. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides an intelligent agent for irregularly shaped tunnel structures, comprising:

[0007] The main structure, composed of shield tunnel segments, forms the basic load-bearing framework for irregularly shaped tunnels.

[0008] A sensing unit, installed on the main structure, is used to sense the stress and deformation state of the tunnel structure and output sensing data;

[0009] The control unit is located in the stress concentration area or mechanical transmission path of the main structure and is used to adjust the mechanical properties of the tunnel structure according to the input control commands.

[0010] The control system, connected to the sensing unit and the control unit, is used to receive the output data of the sensing unit, generate control commands based on the reinforcement learning model, and send the control commands to the control unit to complete the adaptive optimization of the stress deformation of the tunnel structure.

[0011] Furthermore, the main structure includes a main tunnel segment structure, a ramp tunnel segment structure, and a connecting pipe curtain steel-concrete structure. The main tunnel segment structure and the ramp tunnel segment structure are connected by the connecting pipe curtain steel-concrete structure to form an irregular cross-section tunnel structure.

[0012] Furthermore, the sensing unit includes a laser rangefinder, a MEMS inclinometer, and an axial force gauge. The laser rangefinder is used to measure the radial convergence of the tube segment, the MEMS inclinometer is used to measure the inclination angle of the tube segment, and the axial force gauge is used to measure the axial force of the servo actuator.

[0013] Furthermore, the control unit includes a servo actuator, which is located in the stress concentration area or mechanical transmission path of the main structure. The servo actuator is connected to the segment by high-strength bolts and reinforced by steel bolts in the bolt holes at the segment joints.

[0014] Furthermore, the control system includes a PLC controller and a reinforcement learning module. The PLC controller is used to receive the output data of the sensing unit and adjust the displacement and thrust of the servo actuator according to the control instructions generated by the reinforcement learning module.

[0015] Furthermore, the main structure can be integrated with the control unit and control system to form a permanent and temporary combined structural system. Some servo actuators can be integrated into the segment structure as permanent components. The permanent components can be locked after control is in place, and the stress state can be adjusted when necessary. Some servo actuators can be used as a temporary servo support system. The temporary servo support system is used to enhance structural stability during construction and is removed after construction is completed.

[0016] On the other hand, the present invention also provides a method for controlling an intelligent agent, comprising:

[0017] Sensing units are installed on the main structure to sense the stress and deformation state of the tunnel structure.

[0018] Install control units in stress concentration areas or mechanical transmission paths of the main structure;

[0019] The stress and deformation state data of the tunnel structure are obtained through the sensing unit, and the obtained stress and deformation state data are transmitted to the control system.

[0020] The control system calculates the stress-deformation state data based on a reinforcement learning model and generates control commands.

[0021] Control commands are sent to the control unit, which then adjusts the mechanical properties of the tunnel structure according to the commands to achieve adaptive optimization of stress deformation.

[0022] Preferably, the step of generating control commands based on a reinforcement learning model in the control system includes:

[0023] Construct a reinforcement learning model based on a value network, a target value network, a policy network, and a target policy network;

[0024] The force and deformation state data acquired by the sensing unit is used as input to the reinforcement learning model to generate the action strategy of the control unit;

[0025] The effectiveness of the current action strategy is evaluated based on the reward function, and the network parameters are updated accordingly.

[0026] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0027] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.

[0028] Compared with the prior art, the present invention has the following advantages and technical effects:

[0029] (1) By integrating traditional structural systems with active control devices such as servo, the problem of active control of stress deformation of irregularly shaped and gradually changing cross-section tunnel structures is solved, and the adaptiveness to complex stress environment is realized.

[0030] (2) The tunnel structure intelligent agent has the ability to perceive the environment and take autonomous actions. By perceiving the deformation and stress state of the force system, the control unit automatically generates action parameters to achieve active control.

[0031] (3) Structural intelligent bodies can cope with complex stress environments such as irregular cross sections, reduce structural deformation and stress concentration, improve the safety and efficiency of construction, and reduce costs and construction risks. Attached Figure Description

[0032] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0033] Figure 1 This is a schematic diagram illustrating the composition and principle of the intelligent agent for the irregular tunnel structure in an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram illustrating the composition and control unit setup of the irregular-shaped tunnel structure system for underground road diversion sections according to an embodiment of the present invention.

[0035] Figure 3 This is a schematic diagram of the intelligent agent sensing unit setup for a tunnel structure according to an embodiment of the present invention;

[0036] Figure 4 This is a schematic diagram of the intelligent agent control method and system composition according to an embodiment of the present invention;

[0037] Figure 5 This is a schematic diagram of the implementation path of an embodiment of the present invention;

[0038] Figure 6 This is a schematic diagram of the relationship between the reinforcement learning network models in an embodiment of the present invention;

[0039] Figure 7 This is a schematic diagram of the neural network architecture of the reinforcement learning agent according to an embodiment of the present invention;

[0040] The components include: 1. Main tunnel segment structure; 2. Ramp tunnel segment structure; 3. Connecting pipe curtain steel-concrete structure; 4. Soil between main tunnel and ramp tunnel; 5. Permanently installed servo control device; 6. Temporary servo support system; 7. Inclinometer; 8. Laser rangefinder; 9. Axis force gauge; 10. Laser emitter; and 11. Target. Detailed Implementation

[0041] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0042] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0043] Example 1

[0044] This embodiment provides an intelligent agent for irregularly shaped tunnel structures and an intelligent agent control method, including:

[0045] (1) Main structure:

[0046] The main structure adopts a traditional steel structure, steel-concrete composite structure, or concrete structure to form the basic load-bearing frame. In the intelligent body of the irregular tunnel structure in the merging and diverging section of the underground road ramp, the structural cross-section includes the main tunnel segment structure 1, the ramp tunnel segment structure 2, and the connecting pipe curtain steel-concrete structure 3. At the junction of the main tunnel segment structure 1 and the connecting pipe curtain steel-concrete structure 3, a permanently installed servo control device 5 is set up, and a temporary servo support system 6 is added inside the tunnel if necessary. The segments between the two pipe curtains are gradually removed during construction excavation, and finally connected to form the irregular cross-section tunnel structure.

[0047] (2) Control unit:

[0048] The control unit consists of a permanently installed servo control device 5 along the structural stress concentration and mechanical transmission path. In this structure, stress concentration and potential damage can easily occur at the angle between the pipe curtain and the shield tunnel. By installing oblique servo actuators in conjunction with the sensing unit, and based on the action strategy generated by reinforcement learning, the PLC adjusts the top force to make the stress and deformation of the entire structural system more balanced. Once the control is in place, the component is locked. During operation, it can be opened for stress state adjustment when necessary. During construction, to further enhance control capabilities, a temporary servo support system can be added to provide overall structural support and ensure structural stability.

[0049] The servo actuator, acting as the control unit of the structural intelligent entity, connects the servo head to the tunnel segments via high-strength bolts, further reinforced by steel bolts in the bolt holes at the segment joints. For non-circular, irregularly shaped tunnel segments, the servo head is positioned at critical stress areas such as segment bends and joints. The control system primarily consists of a servo support system and a PLC controller. It allows for real-time adjustment of servo parameters based on displacement and pressure observations. By regulating the flow and pressure of hydraulic oil, the servo head of the jack provides precise support force, achieving synchronized and optimized control of the tunnel structure's stress and deformation.

[0050] (3) Sensory unit:

[0051] The sensing unit is used to sense the dynamic changes in the structural stress state caused by changes in the stress environment, providing data support for the subsequent generation of active control strategies. It mainly uses deformation and mechanical sensing sensors such as inclinometer 7, laser rangefinder 8, and axial force gauge 9 to acquire data such as servo axial force, actuator distance, segment tilt angle, and convergence, to sense the structural stress and deformation state and serve as input data for the control unit, contributing to the generation of control strategies and actions. The laser rangefinder 8 also includes a laser emitter 10 and a target 11.

[0052] The placement of sensing units needs to be optimized based on the stress characteristics of the tunnel structure and construction requirements. Typically, sensors are placed at segment joints, critical nodes of the servo support system, and areas subject to high stress. These locations are the most sensitive to changes in structural stress, effectively capturing the structure's dynamic response. As permanent servo components, sensors must be embedded in the structure, integrated with it, to ensure stability during construction and operation.

[0053] To enable real-time data acquisition and transmission, the sensing unit is connected to the control system via a dedicated signal line. The strain, displacement, and axial force data acquired by the sensors are transmitted to the control system for analysis and processing.

[0054] Based on the cooperative control principle of agent structure-perception-control, the relevant control methods and systems include the following characteristics:

[0055] (1) Perception of environment and stress state:

[0056] By utilizing displacement and pressure sensors deployed at the servo head locations, the strain, displacement, and axial force parameters of the support system and tunnel segment structure are monitored in real time, and dynamic response data such as tunnel segment inclination and convergence are obtained, thereby achieving real-time sensing of the tunnel segment structure.

[0057] (2) Control strategy learning:

[0058] Based on the structural stress state perception parameters (input state space S of the network model), the PLC controller compares the preset thresholds (maximum allowable displacement threshold, axial force safety threshold) with real-time data, and generates servo parameters (output motion space A of the network model) and control commands through algorithms such as reinforcement learning according to the coordination relationship between axial force and displacement, thereby determining the displacement and thrust adjustment of the servo actuator and other control units.

[0059] The reinforcement learning model comprises four networks: a value network, a target value network, a policy network, and a target policy network. The value network receives the state information output from the perception units, evaluates the current control unit's action based on the target value, and provides a reward function value for the structural response. It can use both single-action rewards and dynamic rewards across actions to calculate the maximum total reward. Based on the obtained reward values, the value network and the target value network calculate gradients to update the value network, and simultaneously pass these gradients to the policy network for updates. The target network partially replicates the parameters of the policy and value networks for soft updates.

[0060] The value assessment of environmental states and agent actions is achieved through state value functions and action value functions. The calculation methods for the state value function V(s) and action value function Q(s,a) are as follows:

[0061]

[0062] In the formula: t represents the discretization time step, corresponding to the PLC control cycle; γ represents the reduction factor, γ t The product of the reduction factors at time step t; r tdenoted by t, the reward value at time t is represented; a∈A represents the control action, corresponding to the servo parameter vector of the agent's structural control unit, including the supporting force; a0 represents the initial control action, corresponding to the first set of output servo parameter vectors; a represents the control action at any time, corresponding to the servo parameter vector calculated under state s; s∈S represents the state, corresponding to the real-time monitoring parameter vector of the agent's structural perception unit, including the strain, displacement, and axial force of the support system and the segment structure; s0 represents the initial environmental state, corresponding to the initial monitoring parameter vector of the agent's perception unit; s represents the state at any time, corresponding to the monitoring parameter vector of the agent's perception unit at time step t.

[0063] Specifically, model training can employ the Deep Q-Network (DQN) algorithm, incorporating experience replay and a target network. The neural network model architecture of the agent is as follows: Figure 7 As shown, the environmental state, based on the structural stress state response sensing parameters, constitutes a multidimensional state space (Δ) consisting of 3n elements. r1 N h1 ,δ r1 ,...,Δ rn N hn ,δ rn The input layer is used as the action layer; the actions, based on the servo parameters of the control unit, constitute a multi-dimensional action space (Δ) consisting of 2n elements. r1 ',N h1 ',...,Δ rn ',N hn As the output layer, the servo parameters also form part of the input layer for the next state parameters.

[0064] In the formula:

[0065] Δ r1 ...Δ rn N represents the displacement of the servo head at the control unit. h1 ...N hn Indicates the axial force of the servo head at the control unit; δ r1 ...δ rn This indicates radial convergence at the observation point of the tunnel segment, calculated from data monitored by a laser rangefinder and tilt sensor; Δ r1 '...Δ rn ' indicates the displacement of the servo head at the next state control unit; N h1 '...N hn 'Indicates the axial force of the servo head at the next state control unit;

[0066] Among them, the state transition samples (s) generated by the interaction between the agent and the environment t ,r t ,a t ,s t+1The data is stored in the experience pool, and batch data is randomly sampled during training to eliminate temporal correlation. The selection of actions introduces noise, and an ε-greedy strategy is adopted to select actions with probability ε and select argmaxQ(s,a') with probability 1-ε.

[0067] In the formula: s t a represents the complete environmental state perceived by the agent at time step t; t The agent indicates that at time step t, the agent determines the time step based on s. t The control action performed; r t This indicates that the agent executes control action a at time step t. t Immediate rewards from post-environmental feedback; s t+1 This indicates that the agent performs a control action. t The new state after the environment transitions, i.e. the environment state at time step t+1; Q(s,a') represents the long-term value prediction of the action a' performed in state s; argmaxQ(s,a') represents the optimal action that maximizes the Q value from all possible control actions a∈A; ε represents a probability parameter with a value range of 0 < ε < 1, and a decay strategy is adopted.

[0068] The value network updates the gradient by minimizing the temporal difference error, and the calculation method is as follows:

[0069]

[0070] In the formula: This represents the gradient operator, which is the partial derivative with respect to the network parameters θ; θ represents the current network parameters; θ - represents the target network parameters; s' represents the next state; a' represents the next control action; s represents the current state; a represents the current control action; γ represents the discount factor, whose value range is 0 < γ < 1; Q represents the action value function of the target network in the next state s' and control action a'; θ (s,a) represents the action value function of the current network under the current state s and control action a; This means finding the maximum value of all possible actions a'.

[0071] The reward function is designed as follows:

[0072]

[0073] In the formula: R t This represents the reward function setpoint; n represents the number of observation points; ω0 represents the weighting coefficient of the servo structure state parameter; ω1...ω n Indicates the weighting coefficient of the structural state parameter of the tunnel segment; Δ r1 ...Δ rnThis represents the displacement of the servo head at the control unit at n observation points; |Δ r1 |...|Δ rn | represents absolute value operation; Δ max Indicates the maximum allowable displacement of the servo head; N h1 ...N hn δ represents the axial force of the servo head at the control unit at n observation points; r1 ...δ rn This represents the radial convergence at the segment observation points at n observation points, calculated using data from a laser rangefinder and tilt sensor; N safe This indicates the axial force safety threshold of the servo head.

[0074] In this embodiment, the reinforcement learning network model utilizes information about the servo load applied by the servo structure; it simulates the state of the tunnel environment through numerical simulation or physical experiments to obtain the environmental state observation inputs for the reinforcement learning evaluation network and policy network models; based on the observation inputs and preset policy gradients, it trains and updates the evaluation network and policy network models until the models meet the preset training conditions, including the displacement and constraint conditions of the structural system; finally, it continuously receives reward and penalty values ​​from environmental feedback to update and optimize the parameters of the policy network until the reward value reaches the expected value or converges, and outputs the final servo parameters generated by the policy network.

[0075] (3) Applying control actions:

[0076] Based on the control commands output by the PLC, the displacement and thrust of the servo head are precisely controlled by adjusting the oil pressure of the hydraulic cylinder of the jack. The dynamic adjustment of the servo head acts on the segment joints and support nodes, actively compensating for the additional stress caused by stratum deformation by applying force, maintaining the balance of internal and external forces, and suppressing defects such as segment misalignment and joint opening caused by localized stress concentration.

[0077] In practical applications, as tunnel excavation progresses, the structural condition is continuously monitored, and strain, displacement, and axial force parameters are collected. The servo control parameters are dynamically adjusted according to the force changes at each stage of construction to achieve the coordinated control goal of uniform structural force distribution and minimizing deformation throughout the entire construction process.

[0078] Example 2

[0079] like Figure 1 , Figure 4 and Figure 5 As shown, this embodiment provides an intelligent agent for irregularly shaped tunnel structures and an intelligent agent control method, including:

[0080] (1) Determine the structural layout and construction steps of the irregularly shaped intelligent agent:

[0081] Depending on the specific circumstances of the project, taking the merging and diverging sections of an underground road as an example, such as... Figure 2 and Figure 3 As shown, the structural layout and construction steps of the irregular tunnel are determined. In this scheme, the structural cross-section includes the main tunnel segment structure 1, the ramp tunnel segment structure 2, and the connecting pipe curtain steel-concrete structure 3. At the junction of the main tunnel segment structure 1 and the connecting pipe curtain steel-concrete structure 3, a permanent servo control device 5 is set up. If necessary, a temporary servo support system 6 is added inside the tunnel.

[0082] In terms of construction steps, the main tunnel segment structure 1 and the ramp tunnel segment structure 2 are constructed first using the shield tunneling method, and then the transverse steel-concrete pipe curtain structure 3 is constructed, thus forming a closed irregular tunnel cross-section.

[0083] (2) Construction of irregular structures, installation of control units and sensing units:

[0084] According to the construction steps in step 1, after the construction of the irregular structure is completed, a permanent servo control device 5 is installed. The servo head and the pipe segment are connected by high-strength bolts, and the connection is reinforced by steel bolts in the bolt holes at the pipe segment joints. If necessary, a temporary servo support system 6 can be added to further enhance the control capability.

[0085] After the control unit is installed, inclinometer 7, laser rangefinder 8, axial force gauge 9 and other deformation and mechanical sensing sensors are installed on the tube structure.

[0086] Ultimately, a physical structure system of an intelligent agent is formed, which includes the main tube structure, control unit and sensing unit.

[0087] (3) Construction work under complex conditions such as soil excavation between the main line and ramp tunnels:

[0088] The construction is stuck in a complex construction condition, which causes the irregular structure system to be stressed and deformed. In this scheme, the excavation of the soil 4 between the main line and the ramp tunnel will cause the segment structure to tilt and the tunnel to converge. At the same time, the initial force is applied to the permanently installed servo control device 5 and the temporary servo support system 6, resulting in the initial actuator elongation.

[0089] (4) Agent's perception of environment and structural state:

[0090] Inclinometers 7 installed on the tunnel lining structure, and deformation and mechanical sensing sensors such as laser rangefinders 8 and axial force gauges 9 installed on the control unit, are used to sense tunnel convergence and the stress and deformation state of the structural system. The laser rangefinders also include a laser emitter 10 and a target 11.

[0091] (5) Generation of action policies through reinforcement learning for structural agents:

[0092] Establishing a reinforcement learning model is a core control method for structural intelligent agents. It is based on the structural response parameters of the sensing units (the input state space S of the network model) and the servo parameters of the control units (the output action space A of the network model). The network model comprises four networks, such as... Figure 6 As shown, the network consists of a value network, a target value network, a policy network, and a target policy network. The value network evaluates the current control unit's actions based on the target value, providing a reward function value for the structural response. It can use both single-action rewards and dynamic rewards for each action to calculate the maximum total reward. The value network and the target value network calculate gradients based on the obtained reward values ​​to update the value network, and these gradients are then passed to the policy network for simultaneous updates. The target network partially replicates the parameters of the policy network and the value network for soft updates.

[0093] The value assessment of environmental states and agent actions is achieved through state value functions and action value functions. The calculation methods for the state value function V(s) and action value function Q(s,a) are as follows:

[0094]

[0095] In the formula: t represents the discretization time step, corresponding to the PLC control cycle; γ represents the reduction factor, γ t r represents the cumulative product of the reduction factors at time t; t denoted by , represents the reward value at time t; a∈A represents the control action, corresponding to the servo parameter vector of the agent's structural control unit, including the supporting force; a0 represents the initial control action, corresponding to the first set of output servo parameter vectors; a represents the control action at any time, corresponding to the servo parameter vector calculated under state s. s∈S represents the state; corresponding to the real-time monitoring parameter vector of the agent's structural sensing unit, including the strain, displacement, and axial force of the support system and the segment structure; s0 represents the initial environmental state, corresponding to the initial monitoring parameter vector of the agent's sensing unit; s represents the state at any time, corresponding to the monitoring parameter vector of the agent's sensing unit at time step t.

[0096] Specifically, model training can employ the Deep Q-Network (DQN) algorithm, incorporating experience replay and a target network. The neural network model architecture of the agent is as follows: Figure 7 As shown, the environmental state, based on the structural stress state response sensing parameters, constitutes a multidimensional state space (Δ) consisting of 3n elements. r1 N h1 ,δ r1 ,...,Δ rn N hn ,δ rn The input layer is used as the action layer; the actions, based on the servo parameters of the control unit, constitute a multi-dimensional action space (Δ) consisting of 2n elements. r1 ',Nh1 ',...,Δ rn ',N hn As the output layer, the servo parameters also form part of the input layer for the next state parameters.

[0097] In the formula: Δ r1 ...Δ rn N represents the displacement of the servo head at the control unit. h1 ...N hn Indicates the axial force of the servo head at the control unit; δ r1 ...δ rn This indicates radial convergence at the observation point of the tunnel segment, calculated from data monitored by a laser rangefinder and tilt sensor; Δ r1 '...Δ rn ' indicates the displacement of the servo head at the next state control unit; N h1 '...N hn 'Indicates the axial force of the servo head at the next state control unit;

[0098] Among them, the state transition samples (s) generated by the interaction between the agent and the environment t ,r t ,a t ,s t+1 The data is stored in the experience pool, and batch data is randomly sampled during training to eliminate temporal correlation. The selection of actions introduces noise, and an ε-greedy strategy is adopted to select actions with probability ε and select argmaxQ(s,a') with probability 1-ε.

[0099] In the formula: s t a represents the complete environmental state perceived by the agent at time step t; t The agent indicates that at time step t, the agent determines the time step based on s. t The control action performed; r t This indicates that the agent executes control action a at time step t. t Immediate rewards from post-environmental feedback; s t+1 This indicates that the agent performs a control action. t The new state after the environment transitions, i.e. the environment state at time step t+1; Q(s,a') is the long-term value prediction of the action a' performed in state s; argmaxQ(s,a') represents the optimal action that maximizes the Q value from all possible actions a∈A; ε represents a probability parameter with a value range of 0 < ε < 1, and a decay strategy is adopted.

[0100] The value network updates the gradient by minimizing the temporal difference error, and the calculation method is as follows:

[0101]

[0102] In the formula: This represents the gradient operator, which is the partial derivative with respect to the network parameters θ; θ represents the current network parameters; θ - represents the target network parameters; s' represents the next state; a' represents the next control action; s represents the current state; a represents the current control action; γ represents the discount factor, whose value range is 0 < γ < 1; Q represents the action value function of the target network in the next state s' and action a'; θ (s,a) represents the action value function of the current network under the current state s and action a; This means finding the maximum value of all possible actions a'.

[0103] The reward function is designed as follows:

[0104]

[0105] In the formula: R t This represents the reward function setpoint; n represents the number of observation points; ω0 represents the weighting coefficient of the servo structure state parameter; ω1...ω n Indicates the weighting coefficient of the structural state parameter of the tunnel segment; Δ r1 ...Δ rn Δ represents the displacement of the servo head at the control unit at n observation points; max Indicates the maximum allowable displacement of the servo head; N h1 ...N hn δ represents the axial force of the servo head at the control unit at n observation points; r1 ...δ rn This represents the radial convergence at the segment observation points at n observation points, calculated using data from a laser rangefinder and tilt sensor; N safe This indicates the axial force safety threshold of the servo head.

[0106] In this embodiment, reinforcement learning is used to generate servo actions for the control unit. First, information on the servo load applied to the servo structure is collected. The state of the tunnel environment is simulated through numerical simulation or physical experiments to obtain the environmental state observation input for the reinforcement learning model. Based on the observation input and the preset policy gradient, the evaluation network and policy network models are trained and updated until the models meet the preset training conditions, including the displacement and constraint conditions of the structural system. Finally, the reward and penalty values ​​from the environmental feedback are continuously received and updated to optimize the parameters of the policy network until the reward value reaches the expected value or converges. The final servo parameters generated by the policy network are then output.

[0107] (6) Servo control unit PLC active control:

[0108] Based on the control commands output by the PLC, the displacement and jacking force of the servo head are precisely controlled by adjusting the hydraulic pressure of the jack's hydraulic cylinder. The dynamic adjustment of the servo head acts on the segment joints and support nodes, actively compensating for the additional stress caused by ground deformation and maintaining internal and external force balance. In practical applications, as tunnel excavation progresses, the structural condition is continuously monitored, and strain, displacement, and axial force parameters are collected. The servo control parameters are dynamically adjusted according to the force changes at each stage of construction, achieving the coordinated control goal of uniform structural force distribution and minimal deformation throughout the entire construction process.

[0109] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An intelligent agent with an irregular tunnel structure, characterized in that, include: The main structure, composed of shield tunnel segments, forms the basic load-bearing framework for irregularly shaped tunnels. A sensing unit, installed on the main structure, is used to sense the stress and deformation state of the tunnel structure and output sensing data; The control unit is located in the stress concentration area or mechanical transmission path of the main structure and is used to adjust the mechanical properties of the tunnel structure according to the input control commands. The control system, connected to the sensing unit and the control unit, is used to receive the output data of the sensing unit, generate control commands based on the reinforcement learning model, and send the control commands to the control unit to complete the adaptive optimization of the stress deformation of the tunnel structure.

2. The intelligent agent according to claim 1, characterized in that, The main structure includes a main tunnel segment structure, a ramp tunnel segment structure, and a connecting pipe curtain steel-concrete structure. The main tunnel segment structure and the ramp tunnel segment structure are connected by the connecting pipe curtain steel-concrete structure to form an irregular cross-section tunnel structure.

3. The intelligent agent according to claim 1, characterized in that, The sensing unit includes a laser rangefinder, a MEMS inclinometer, and an axial force gauge. The laser rangefinder is used to measure the radial convergence of the tube segment, the MEMS inclinometer is used to measure the inclination angle of the tube segment, and the axial force gauge is used to measure the axial force of the servo actuator.

4. The intelligent agent according to claim 1, characterized in that, The control unit includes a servo actuator, which is located in the stress concentration area or mechanical transmission path of the main structure. The servo actuator is connected to the segment by high-strength bolts and reinforced by steel bolts in the bolt holes at the segment joints.

5. The intelligent agent according to claim 1, characterized in that, The control system includes a PLC controller and a reinforcement learning module. The PLC controller is used to receive the output data of the sensing unit and adjust the displacement and thrust of the servo actuator according to the control instructions generated by the reinforcement learning module.

6. The intelligent agent according to claim 1, characterized in that, The main structure, control unit, and control system form an integrated structural system combining permanent and temporary components. Some servo actuators are integrated into the segment structure as permanent components. The permanent components are locked after being controlled and the stress state adjustment is activated when necessary. Some servo actuators serve as a temporary support system, which is used to enhance structural stability during construction and is removed after construction is completed.

7. The control method for an intelligent agent according to any one of claims 1-6, characterized in that, include: Sensing units are installed on the main structure to sense the stress and deformation state of the tunnel structure. Install control units in stress concentration areas or mechanical transmission paths of the main structure; The stress and deformation state data of the tunnel structure are obtained through the sensing unit, and the obtained stress and deformation state data are transmitted to the control system. The control system calculates the stress-deformation state data based on a reinforcement learning model and generates control commands. Control commands are sent to the control unit, which then adjusts the mechanical properties of the tunnel structure according to the commands to achieve adaptive optimization of stress deformation.

8. The method according to claim 1, characterized in that, The steps for the control system to generate control commands based on a reinforcement learning model include: Construct a reinforcement learning model based on a value network, a target value network, a policy network, and a target policy network; The force and deformation state data acquired by the sensing unit is used as input to the reinforcement learning model to generate the action strategy of the control unit; The effectiveness of the current action strategy is evaluated based on the reward function, and the network parameters are updated accordingly.

9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 7-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 7-8.

Citation Information

Patent Citations

  • Shield tunnel deformation active regulation and control method

    CN116066154A

  • Tunnel shield segment with self-sensing performance and intelligent water seepage early warning method thereof

    CN118462221A

  • Tunnel servo support system based on reinforcement learning and adaptive control method

    CN118605181A

  • Intelligent tunnel special-shaped section duct piece structure adopting servo control

    CN118757179A

  • Assembling type shield tunnel self-adaptive reinforcement and repair servo supporting system, construction method and computer system

    CN119844126A