Dynamic decision-making method, device and system for long and short term benefit collaborative optimization of construction organization
By constructing a high-fidelity virtual simulation environment for tunnel construction and the SoftActor-Critic intelligent decision-making model, the problems of insufficient simulation fidelity and low efficiency of heuristic algorithms in tunnel construction are solved, realizing dynamic optimization and real-time decision-making in tunnel construction and meeting the complex needs of extra-long tunnel construction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-03-13
- Publication Date
- 2026-05-05
AI Technical Summary
Existing tunnel construction organization design and schedule optimization technologies suffer from insufficient simulation fidelity, low computational efficiency of heuristic algorithms, and a lack of long-term value assessment capabilities, making it difficult to meet the dynamic interaction and real-time response requirements of modern extra-long tunnel construction.
A high-fidelity virtual simulation environment for tunnel construction was constructed. A smart decision-making model based on SoftActor-Critic was used for virtual training and real-time reasoning. The reinforcement learning mechanism was used to achieve synergistic optimization of the long-term and short-term interests of the construction organization.
It enables high-fidelity simulation and dynamic decision-making within milliseconds, allowing for real-time response to unexpected changes during construction, optimization of construction progress, and a balance between short-term and long-term interests.
Smart Images

Figure CN121980962A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent construction technology in civil engineering, and specifically relates to a dynamic decision-making method, device, and system for synergistic optimization of long-term and short-term interests in construction organization. Background Technology
[0002] In modern transportation infrastructure construction, the construction of extra-long tunnels (such as railway or highway tunnels exceeding 20 kilometers) is an extremely complex system engineering project. In order to shorten the construction period, the "long tunnel, short construction" strategy is usually adopted, which means advancing the tunnel through a parallel pilot tunnel (referred to as a pilot tunnel) and opening a transverse connecting passage (referred to as a transverse passage) at a specific location, thereby adding new tunneling faces in the main tunnel and realizing parallel operation of multiple working faces.
[0003] However, existing tunnel construction organization design and schedule optimization technologies face severe challenges and are unable to meet the needs of actual engineering projects: Existing simulation technologies lack sufficient fidelity and dynamic coupling mechanisms. Traditional project scheduling software (such as P6 and Project) is based on static network diagrams (CPM / PERT) and cannot simulate the dynamic interactions during construction. For example, when a sudden shortage of materials (such as concrete and steel bars) occurs on the construction site, traditional models cannot automatically calculate how this shortage non-linearly affects the tunneling speed of all working faces. In addition, for the dynamic evolution of the topology over time, such as "the horizontal tunneling reaching a specific position triggers the opening of a cross passage, and the opening of the cross passage triggers a new working face for the main tunnel," existing simulations lack an endogenous modeling mechanism and often require manual adjustments to the schedule.
[0004] The short-sightedness and inefficiency of existing optimization algorithms: While commonly used heuristic search algorithms such as Genetic Algorithm (GA) and Particle Swarm Optimization (PSO) can find optimal solutions to some extent, they suffer from two fundamental flaws: The computational efficiency is low and the decision-making is open-loop: Faced with a massive number of decision combinations (dozens of horizontal channels, each channel's activation time, direction, and construction method), the search space is huge, and convergence takes an extremely long time. Moreover, if a geological change occurs during actual construction, the previous optimal solution immediately becomes invalid, requiring a time-consuming global recalculation, making real-time response impossible.
[0005] Lack of long-term value assessment capability (short-sightedness): Heuristic algorithms typically evaluate based on the current or next stage's state, making it difficult to learn complex strategies of "delayed gratification." For example, to minimize the overall project duration, the current optimal strategy might be to temporarily sacrifice the rapid excavation of the pilot tunnel and allocate resources to prioritize opening a transverse passage on a critical path. Traditional algorithms struggle to capture such causal chains spanning long periods. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a dynamic decision-making method, device, system, and storage medium for synergistic optimization of long-term and short-term interests in construction organization.
[0007] To achieve the above objectives, the present invention provides the following solution: A dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization includes: Step S1: Construct a high-fidelity virtual simulation environment for tunnel construction; Step S2: Construct an intelligent decision-making model based on SoftActor-Critic that perceives the environmental state and makes optimal decisions; Step S3: Perform virtual training and real-time reasoning based on the intelligent decision-making model to achieve dynamic decision-making that optimizes the synergistic effects of long-term and short-term interests in construction organization.
[0008] As a preferred option, in step S1, a high-fidelity virtual simulation environment for tunnel construction is constructed based on the mileage coordinates of all active working faces along the tunnel, the current geological rock grade, and the global resource pool status, using nonlinear progress deduction based on material constraints, dynamic topology evolution determination, automatic breakthrough and resource release.
[0009] Preferably, step S3 includes: The virtual environment encapsulates the current mileage coordinates, resource reserves, and geological forecast information into a high-dimensional state vector. The data is input into the SAC intelligent decision-making model. Based on the state vector model, the SAC intelligent decision-making model outputs decision actions including face development, tunneling direction, and construction method selection through an Actor network. The virtual environment receives and executes this action, advancing the simulation time step. The virtual environment provides instant rewards based on the execution results. This leads to a positive and negative feedback mechanism in reinforcement learning for the intelligent agent. The interaction data quadruple at each step The data is stored in the experience replay pool at the bottom. During training, the agent randomly samples historical data from the replay pool to perform gradient descent training, thereby achieving continuous iteration and optimization of the policy.
[0010] This invention also provides a dynamic decision-making device for synergistic optimization of long-term and short-term benefits in construction organization, comprising: The first processing module is used to build a high-fidelity virtual simulation environment for tunnel construction. The second processing module is used to construct an intelligent decision-making model based on SoftActor-Critic that perceives the environmental state and makes optimal decisions. The third processing module is used to perform virtual training and real-time reasoning based on the intelligent decision-making model to achieve dynamic decision-making that optimizes the synergistic benefits of long-term and short-term construction organization.
[0011] As a preferred option, the first processing module constructs a high-fidelity virtual simulation environment for tunnel construction based on the mileage coordinates of all active working faces along the tunnel, the current geological rock grade, and the global resource pool status, using nonlinear progress deduction based on material constraints, dynamic topology evolution determination, automatic breakthrough and resource release.
[0012] Preferably, the third processing module includes: The state-aware unit is used to enable the virtual environment to encapsulate the current mileage coordinates, resource reserves, and geological forecast information into a high-dimensional state vector. The data is input into the SAC intelligent decision-making model. The decision execution unit enables the SAC intelligent decision model to output decision actions, including working face development, tunneling direction, and construction method selection, based on the state vector model and through the Actor network. The virtual environment receives and executes this action, advancing the simulation time step. The feedback loop unit is used to enable the virtual environment to provide immediate rewards based on the execution results. This leads to a positive and negative feedback mechanism in reinforcement learning for the intelligent agent. Experience replay unit, used to generate a quadruple of interaction data for each step. The data is stored in the experience replay pool at the bottom. During training, the agent randomly samples historical data from the replay pool to perform gradient descent training, thereby achieving continuous iteration and optimization of the policy.
[0013] The present invention also provides a dynamic decision-making system for synergistic optimization of long-term and short-term interests in construction organization, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization when executed by the processor.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention can construct a high-fidelity virtual simulation environment that includes physical constraints, resource coupling, and complex logic, and introduces an intelligent decision-making mechanism with long-term value assessment capabilities and millisecond-level inference speed. Attached Figure Description
[0015] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of a dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization, as described in an embodiment of the present invention. Figure 2 For simulation environment; Figure 3 It uses the SAC architecture; Figure 4 For virtual training and real-time inference processes. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] Example 1 like Figure 1 As shown, this invention provides a dynamic decision-making method for synergistic optimization of long-term and short-term benefits in construction organization, comprising: Step S1: Construct a high-fidelity virtual simulation environment for tunnel construction. This step aims to establish a digital construction site (corresponding to...) Figure 1 The core of the left-hand module lies not in the 3D display of its appearance, but in the precise simulation of its internal logic and physical constraints. Figure 2 This invention demonstrates the core computational logic and state evolution process of the tunnel construction virtual simulation environment within a single simulation time step; specifically, it includes the following steps: S101: Initialize Simulation TimeStep ) First, load the current time step. The system state snapshot contains the mileage coordinates of all active working surfaces, the current geological rock grade, the global resource pool status (material inventory, number of idle teams), and the topology connection matrix.
[0020] S102: Nonlinear Schedule Derivation Based on Material Constraints The resource-schedule tightly coupled calculation process aims to simulate the nonlinear impact of physical resource constraints on construction schedule: Geological Response and Demand Quantification: Traverse all active working face objects, and query the benchmark design efficiency based on the geological rock grade corresponding to the current coordinates of each working face. ) and material consumption per unit advance ( By summing up the data, the total theoretical material demand for the entire day is obtained: .
[0021] Supply and demand balance verification: Read the current inventory of the material warehouse object ( ), and determine theoretical requirements Is it greater than the current inventory? .
[0022] Nonlinear scaling factor calculation: If a shortage of materials is identified (yes), calculate the global nonlinear schedule scaling factor. ,at this time .
[0023] If resources are deemed sufficient (no), set the full-speed advancement coefficient. .
[0024] Physics engine update: Performs constrained physical displacement calculations based on scaling factors, updating the mileage coordinates of all working surfaces: This step ensures that the construction schedule is strictly limited by the availability of physical resources.
[0025] S103: Dynamic Topological Evolution Determination After the physical displacement is updated, the topology evolution monitoring process is entered in parallel to handle the logic of opening the main tunnel working face from the parallel pilot tunnel (parallel pilot tunnel): Spatial position tracking: Real-time acquisition of the coordinate vector of the working face object at the tunneling end of the horizontal guide vane. .
[0026] Geometric inclusion detection: Retrieve a preset set of inactive transverse channel nodes and determine whether the current guide advance interval geometrically covers a preset transverse channel node to be activated.
[0027] SAC agent interaction: If coverage is detected, construct the current state vector. The request is made to the SAC reinforcement learning agent to make a decision. The agent outputs a policy action. .
[0028] Action Decoding and Execution: Parsing Actions If the action is "Activate", then the topology fission mechanism is triggered. Instantiation: Instantiate a new working face object at the corresponding mileage position of the main tunnel.
[0029] Resource allocation: Deduct the corresponding construction team resources from the global resource pool and bind them to the new object.
[0030] Mount: Update the system topology map, mount the new working face to the main tunnel route list, and add it to subsequent time step iterations.
[0031] S104: Automatic Penetration and Resource Release Parallel to step S103, the automatic connection determination process is initiated to handle the complex logic when multiple working faces meet: Neighborhood detection algorithm: Traverse all adjacent working face objects on the same tunnel line and calculate the Euclidean distance for each pair of adjacent faces. .
[0032] Physical collision determination: Determines whether a collision exists. This refers to situations where physical spaces overlap or come into contact.
[0033] Tunneling Vector Analysis: If a collision occurs, the system calculates the dot product of the tunneling direction vectors of the two working faces. .
[0034] Scenario A: Two-way fusion: If (Opposite sign) indicates a connection between opposing work surfaces. The system executes the destruction logic, removes the two encountering work surface objects, generates a connected "completed road segment topology", and releases the two occupied construction teams back to the idle resource pool, marking their status as "idle".
[0035] Scenario B: One-way annexation: If (Same number) indicates that the process is in the same direction of catching up. The system retains the work surface attributes of the party with the progress advantage (the party with greater mileage or faster speed), cancels the work surface object of the party being caught up with, and releases the construction team resources of the party being caught up with.
[0036] S105: Multi-threaded state synchronization Wait for all logical branches of steps S103 (topology fission) and S104 (through fusion) to be completed, ensure that all object state changes (creation, destruction, merging) have been atomically committed in memory, and complete the system state synchronization of the current time step.
[0037] S106: Time-stepping iteration The system generates the global state vector for the next time step. The system calculates immediate rewards for reinforcement learning based on the construction progress, cost consumption, and violations at this time step. The simulation clock then advances to the next day. Repeat steps S101 to S106 above until the project is completed.
[0038] Step S2: Construct an intelligent decision-making model based on SoftActor-Critic (SAC) This step aims to build a brain capable of perceiving the aforementioned environmental states and making optimal decisions.
[0039] Markov Decision Process (MDP) Mapping: State space (State, ): Construct a high-dimensional feature vector to reflect the environment snapshot in real time. This includes: the normalized mileage of all active working faces; the predicted geological rock grade within a certain distance in front of each working face; the remaining quantity of various construction teams in the resource pool; the current inventory level of the material warehouse; and the opening status indicators of all potential cross passages (not reached / reached but pending decision / open).
[0040] Action space ): Defined as a multi-dimensional combined action vector. For each potential cross-channel node, the agent needs to output a combined instruction. These represent: whether to immediately activate the channel (0 / 1 discrete variables); selection of the main tunnel excavation direction (forward / reverse / bidirectional, discrete variables); and selection of the construction method (ordinary method / mechanized large machine method, discrete variables).
[0041] Reward function ): Design a multi-objective composite reward function to guide the agent.
[0042] The fixed negative reward (e.g., -1) for each step forces the agent to find the shortest path to complete the task as quickly as possible.
[0043] The negative incentive is based on the actual costs incurred on the day (construction costs + idle labor costs + equipment dispatch costs) to guide economic indicators.
[0044] The huge positive reward (completion bonus) given when the main tunnel is fully completed is the main source of long-term value.
[0045] : Penalties for actions taken by an agent that violate physical logic (such as forcibly starting work when there are no available teams) or safety regulations.
[0046] SAC network architecture and update mechanism (such as...) Figure 3 (as shown) Actor Policy Network ( ):like Figure 3 As shown, receiving status After feature extraction via a multi-layer fully connected network (MLP), the mean of the action distribution is output. and standard deviation To make the sampling process differentiable and support backpropagation, a reparameterization technique is employed, which involves first sampling noise from a standard normal distribution. Then through transformation Generate the final action.
[0047] Dual Critic Value Network ( ):like Figure 3 As shown, to address the common Q-value overestimation problem in deep reinforcement learning, two Critic networks with identical structures but independent parameters are established. They simultaneously receive state... and actions Estimate the long-term returns of the current strategy respectively. and When calculating the target value, take the smaller of the two. This provides a more conservative and stable value estimate.
[0048] Maximum entropy objective introduced: The core innovation of SAC lies in its objective function. The policy entropy is introduced in Temperature coefficient The importance of automatic entropy adjustment. This forces the agent to maintain a high degree of randomness in its actions during the early stages of training, and to explore a wide range of combinations of construction procedures (e.g., trying to open up the cross passage at the far end first), rather than getting stuck in a local optimum or empirical construction plan too early.
[0049] Step S3: Virtual Training and Real-Time Inference Workflow This step describes how the system transitions from the learning state to the application state (corresponding to...). Figure 3 (Two lanes).
[0050] Iterative training phase in a virtual environment (e.g.) Figure 4 (as shown) Interactive sampling: The agent performs "virtual construction" in a virtual simulation environment. Each day, the agent adjusts its actions based on the current state. Sampling actions through the Actor network It acts on the environment. The environment performs resource constraint deduction and topological evolution over a time step, transitioning to a new state. And provide feedback and rewards .
[0051] Experience storage: Store the interaction data quadruple for this step. Store it in the experience replay pool (ReplayBuffer).
[0052] Parameter update: After the experience pool accumulates to a certain size, a batch of data is randomly sampled. Gradient descent is used to minimize the Bellman error loss function. To update the parameters of the two Critic networks; by maximizing the objective function of the sum of the Q-value and the policy entropy. To update the parameters of the Actor network.
[0053] Iterative loop: The above process is repeated in parallel tens of thousands of times on a high-performance computing cluster until the agent's policy converges, which can stably obtain a high cumulative reward.
[0054] Real-time simulation and decision-making stage in actual engineering ( Figure 4 (as shown) Model Deployment: After training, the parameters of the Actor policy network are extracted and frozen, and then deployed to the decision support system at the engineering site.
[0055] Data input: Real-time collection of actual monitoring data from the construction site (such as the exact mileage of each working face today, the latest geological forecast results, and the current number of teams on duty).
[0056] Millisecond-level decision-making: The actual data is assembled into a state vector and input into the frozen Actor network. Since only one forward propagation calculation is required, the system can output the optimal construction instructions for the latest situation within milliseconds (e.g., the No. 3 cross passage must be started tomorrow, and a large fleet of excavators should be deployed to tunnel towards the greater mileage direction), for the project manager's decision-making reference.
[0057] This embodiment successfully transforms the complex tunnel construction organization problem into a reinforcement learning problem through the above steps. It utilizes a virtual environment to handle physical and logical constraints, and uses SAC agents to handle the trade-offs between long-term and short-term interests and uncertain planning, ultimately achieving dynamic intelligent optimization of the construction plan.
[0058] Example 2 This invention also provides a dynamic decision-making device for synergistic optimization of long-term and short-term benefits in construction organization, comprising: The first processing module is used to build a high-fidelity virtual simulation environment for tunnel construction. The second processing module is used to construct an intelligent decision-making model based on SoftActor-Critic that perceives the environmental state and makes optimal decisions. The third processing module is used to perform virtual training and real-time reasoning based on the intelligent decision-making model to achieve dynamic decision-making that optimizes the synergistic benefits of long-term and short-term construction organization.
[0059] As one embodiment of the present invention, the first processing module constructs a high-fidelity virtual simulation environment for tunnel construction based on the mileage coordinates of all active working faces along the tunnel, the current geological rock grade, and the global resource pool status, using nonlinear progress deduction based on material constraints, dynamic topology evolution determination, automatic breakthrough and resource release.
[0060] As one embodiment of the present invention, the third processing module includes: The state-aware unit is used to enable the virtual environment to encapsulate the current mileage coordinates, resource reserves, and geological forecast information into a high-dimensional state vector. The data is input into the SAC intelligent decision-making model. The decision execution unit enables the SAC intelligent decision model to output decision actions, including working face development, tunneling direction, and construction method selection, based on the state vector model and through the Actor network. The virtual environment receives and executes this action, advancing the simulation time step. The feedback loop unit is used to enable the virtual environment to provide immediate rewards based on the execution results. This leads to a positive and negative feedback mechanism in reinforcement learning for the intelligent agent. Experience replay unit, used to generate a quadruple of interaction data for each step. The data is stored in the experience replay pool at the bottom. During training, the agent randomly samples historical data from the replay pool to perform gradient descent training, thereby achieving continuous iteration and optimization of the policy.
[0061] Example 3 The present invention also provides a dynamic decision-making system for synergistic optimization of long-term and short-term interests in construction organization, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization when executed by the processor.
[0062] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization, characterized in that, include: Step S1: Construct a high-fidelity virtual simulation environment for tunnel construction; Step S2: Construct an intelligent decision-making model based on SoftActor-Critic that perceives the environmental state and makes optimal decisions; Step S3: Perform virtual training and real-time reasoning based on the intelligent decision-making model to achieve dynamic decision-making that optimizes the synergistic effects of long-term and short-term interests in construction organization.
2. The dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization as described in claim 1, characterized in that, In step S1, based on the mileage coordinates of all active working faces along the tunnel, the current geological rock grade, and the global resource pool status, a high-fidelity virtual simulation environment for tunnel construction is constructed using nonlinear progress deduction based on material constraints, dynamic topology evolution determination, automatic breakthrough, and resource release.
3. The dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization as described in claim 2, characterized in that, Step S3 includes: The virtual environment encapsulates the current mileage coordinates, resource reserves, and geological forecast information into a high-dimensional state vector. The data is input into the SAC intelligent decision-making model. Based on the state vector model, the SAC intelligent decision-making model outputs decision actions including face development, tunneling direction, and construction method selection through an Actor network. The virtual environment receives and executes this action, advancing the simulation time step. The virtual environment provides instant rewards based on the execution results. This leads to a positive and negative feedback mechanism in reinforcement learning for the intelligent agent. The interaction data quadruple at each step The data is stored in the experience replay pool at the bottom. During training, the agent randomly samples historical data from the replay pool to perform gradient descent training, thereby achieving continuous iteration and optimization of the policy.
4. A dynamic decision-making device for synergistic optimization of long-term and short-term interests in construction organization, characterized in that, include: The first processing module is used to build a high-fidelity virtual simulation environment for tunnel construction. The second processing module is used to construct an intelligent decision-making model based on SoftActor-Critic that perceives the environmental state and makes optimal decisions. The third processing module is used to perform virtual training and real-time reasoning based on the intelligent decision-making model to achieve dynamic decision-making that optimizes the synergistic benefits of long-term and short-term construction organization.
5. The dynamic decision-making device for synergistic optimization of long-term and short-term interests in construction organization as described in claim 4, characterized in that, The first processing module constructs a high-fidelity virtual simulation environment for tunnel construction based on the mileage coordinates of all active working faces along the tunnel, the current geological rock grade, and the global resource pool status. It uses nonlinear progress deduction based on material constraints, dynamic topology evolution determination, automatic breakthrough and resource release.
6. The dynamic decision-making device for synergistic optimization of long-term and short-term interests in construction organization as described in claim 5, characterized in that, The third processing module includes: The state-aware unit is used to enable the virtual environment to encapsulate the current mileage coordinates, resource reserves, and geological forecast information into a high-dimensional state vector. The data is input into the SAC intelligent decision-making model. The decision execution unit enables the SAC intelligent decision model to output decision actions, including working face development, tunneling direction, and construction method selection, based on the state vector model and through the Actor network. The virtual environment receives and executes this action, advancing the simulation time step. The feedback loop unit is used to enable the virtual environment to provide immediate rewards based on the execution results. This leads to a positive and negative feedback mechanism in reinforcement learning for the intelligent agent. Experience replay unit, used to generate a quadruple of interaction data for each step. The data is stored in the experience replay pool at the bottom. During training, the agent randomly samples historical data from the replay pool to perform gradient descent training, thereby achieving continuous iteration and optimization of the policy.
7. A dynamic decision-making system for synergistic optimization of long-term and short-term interests in construction organization, characterized in that, include: A memory and a processor, wherein the memory stores a computer program executed by the processor, the computer program, when executed by the processor, performs a dynamic decision-making method for synergistic optimization of long-term and short-term interests in construction organization as described in any one of claims 1-3.
Citation Information
Patent Citations
Shield tunneling attitude prediction method and system based on CNN-BiLSTM-STDAM combined model, and storage medium
CN120929759A
Slurry balance shield construction posture settlement collaborative optimization method and system based on reinforcement learning, electronic equipment and storage medium
CN121411141A