Vibroflotation pile construction progress intelligent simulation and optimization method based on layered reinforcement learning

By employing a multi-agent modeling and simulation method based on hierarchical reinforcement learning, the problem of controlling the construction progress of large-scale vibro-compacted piles was solved, thereby optimizing the construction progress and efficiently scheduling mechanical resources, and improving construction quality and efficiency.

CN121744931APending Publication Date: 2026-03-27TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies lack effective methods to optimize and control the construction progress of large-scale vibro-compacting piles. Especially under complex conditions, it is difficult to achieve reasonable scheduling of mechanical configuration and construction sequence within the flow unit, resulting in difficulties in construction progress control.

Method used

A hierarchical reinforcement learning-based approach is adopted to establish a multi-agent modeling and simulation model. The construction schedule optimization problem is decomposed into two levels: machine scheduling and intra-unit construction sequence, using the OC-MAPPO algorithm. Option-Critic and MAPPO algorithms are used to optimize the high-level and low-level tasks respectively, thereby realizing intelligent simulation and optimization of the construction scheme.

Benefits of technology

It effectively solved the problems of machinery scheduling and construction sequence in complex vibro-compaction pile construction, provided the optimal construction schedule plan, improved construction quality and efficiency, and reduced project costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744931A_ABST
    Figure CN121744931A_ABST
Patent Text Reader

Abstract

The invention discloses a hierarchical reinforcement learning-based vibroflotation pile construction progress intelligent simulation and optimization method, which comprises the following steps of: obtaining vibroflotation pile construction simulation modeling information, and generating a vibroflotation pile staging input file according to position and processing depth information in the vibroflotation pile construction simulation modeling information; the method comprises the following steps: determining vibroflotation pile construction simulation parameters, establishing a vibroflotation pile construction progress intelligent simulation model comprising a multi-agent Agent modeling and simulation method, establishing a vibroflotation pile construction progress simulation and optimization mathematical model in combination with hierarchical reinforcement learning, and solving the data model based on an OC-MAPPO hierarchical reinforcement learning algorithm; simulation parameters are input, the vibroflotation pile construction progress simulation and optimization model is operated, and an optimized vibroflotation pile construction progress scheme is obtained. According to the method, a complex construction progress optimization problem is decomposed into two levels, namely a pipeline unit mechanical scheduling level and an in-unit vibroflotation pile construction sequence, and a specific sub-problem is processed on each level, so that the complexity of an original problem is simplified, and the model solving efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of vibroflotation pile construction progress simulation, and particularly relates to a layered reinforcement learning-based intelligent simulation and optimization method for vibroflotation pile construction progress. BACKGROUND

[0002] The gravel vibroflotation pile is a ground treatment method for forming a composite foundation by combining the original ground soil with a pile formed by filling hard and coarse-grained materials such as sand, gravel, and cohesive soil into a hole formed in a soft foundation, so as to improve the bearing capacity of the foundation. The gravel vibroflotation pile has the characteristics of simple construction process, fast pile-forming speed, wide application range, good treatment effect, and low cost [1] . Through years of application and technical accumulation, the construction technology level, construction and equipment manufacturing capacity of the vibroflotation technology have been greatly improved, the treatment depth has exceeded 65 meters, and the technology has the technical reserves for deeper construction and the construction capacity for large-scale projects [2] . In order to improve the bearing capacity of the foundation under deep overburden conditions, the gravel vibroflotation pile has become an important means for foundation treatment under deep overburden conditions.

[0003] Although the treatment depth of the vibroflotation pile has been improved, the overall construction scale is limited, the types and number of construction machinery involved are small, and therefore the construction organization design is not complex, and the progress control is easy to implement. However, the construction of large-scale vibroflotation piles is a complex system involving vibroflotation pile arrangement, construction technology, construction machinery and equipment matching, and mutual cooperation. In order to meet the construction progress requirements, a construction scheme is usually adopted, in which flow operation units are configured for each unit on the basis of staging. Using scientific analysis methods and advanced technical means, various influencing factors in the construction of vibroflotation piles are systematically considered, and the constraints of various aspects are coordinated to reasonably organize and arrange the construction, which is beneficial to improving the construction quality, reducing the project cost, and ensuring the engineering progress. The current vibroflotation pile construction lacks a simulation and optimization method for vibroflotation pile construction under complex conditions, and cannot meet the progress control of large-scale vibroflotation pile construction. Therefore, it is necessary to study a layered reinforcement learning-based intelligent simulation and optimization method for vibroflotation pile construction progress, so as to realize the analysis and optimization of the vibroflotation pile construction scheme and ensure that the foundation treatment of water conservancy and hydropower projects meets the progress requirements.

[0004] Over the years, scholars at home and abroad have proposed a series of system simulation theories and methods for complex system simulation, and successfully applied them to a large number of engineering practices. Large-scale vibroflotation pile construction is a complex system, and the process of the system is determined by discrete events such as construction of each flow unit, which follows some complex human rules and can be regarded as a discrete event system, and the discrete event system simulation method is used for modeling. However, there are various entities such as mechanical equipment and project management in complex systems, each entity acts under certain rules and interacts with each other, and decisions need to be made when the system state changes. The agent-based modeling and simulation (ABMS) method is considered an effective way to study complex systems, and is a modeling and simulation methodology for complex systems. Multi-agent simulation can simulate the behavior, decision-making and interaction process of multiple entities in complex systems, and through the simulation of autonomous behavior and distributed decision-making of each entity, the interaction and dynamic evolution of each entity in the complex system are studied. For each entity in the vibroflotation pile construction system, multiple agents are established, and based on the multi-agent modeling and simulation method, the complex vibroflotation pile construction system simulation can be realized.

[0005] However, due to the involvement of mechanical configuration and vibroflotation pile construction sequence in the flow unit in the large-scale vibroflotation pile construction process, these are also the main aspects of construction progress control and optimization. Under the given workload, task division, resource allocation and multi-mechanical behavior coordination determine that the optimization problem of multi-agent collaborative decision-making in complex processes is difficult to solve directly, and hierarchical reinforcement learning (HRL) uses the idea of problem decomposition and divide-and-conquer, which is a potential effective way to solve large-scale reinforcement learning. HRL learns the strategy of each subtask based on task hierarchy, and combines the strategies of multiple subtasks to form an effective global strategy, which can effectively solve the problem of multi-agent collaborative decision-making optimization.

[0006] [Reference]

[0007] [1] Yu Hongzhi, Zhang Zhiwei, Wang Wenpeng. Vibroflotation engineering [M]. China Water Resources and Hydropower Press: 201903.302.

[0008] [2] Yin Changsheng, Yang Ruopeng, Zhu Wei, et al. Multi-agent hierarchical reinforcement learning review [J]. Journal of Intelligent Systems, 2020, 15(04): 646-655. SUMMARY

[0009] For the above prior art, the present application provides a hierarchical reinforcement learning-based vibroflotation pile construction progress intelligent simulation and optimization method, which establishes a complex vibroflotation pile construction progress model based on vibroflotation pile construction information and multi-agent modeling and simulation method, and realizes construction scheme optimization by using hierarchical reinforcement learning.

[0010] In order to solve the above technical problems, the present application provides a kind of based on hierarchical reinforcement learning's vibration pile construction progress intelligent simulation and optimization method, mainly includes the following steps:

[0011] Step 1) obtain vibration pile construction simulation modeling information, including vibration pile construction plan information obtained from construction organization design, and vibration pile treatment depth information obtained according to design depth requirement in combination with construction area geological model;And vibration pile staging input file is generated according to position and treatment depth information;

[0012] Step 2) determine vibration pile construction simulation parameters, including construction environment parameters, process parameters and resource allocation parameters;

[0013] Step 3) establish vibration pile construction progress intelligent simulation model, the vibration pile construction progress simulation model is a multi-agent Agent modeling and simulation method model, including the vibration pile construction ABM simulation system established and the multiple intelligent agents established according to different entities existing in process;

[0014] Step 4) based on the vibration pile construction progress simulation model established in step 3), vibration pile construction progress simulation and optimization mathematical model is established based on hierarchical reinforcement learning, and the data model is solved based on OC-MAPPO hierarchical reinforcement learning algorithm;

[0015] Step 5) input simulation parameters, run vibration pile construction progress simulation and optimization model, and obtain optimized vibration pile construction progress scheme.

[0016] Further, the vibration pile construction progress intelligent simulation and optimization method provided by the present application, wherein:

[0017] In step 1), the vibration pile construction plan information includes vibration pile construction staging plan, each staging construction progress constraint information and pile arrangement information;The vibration pile treatment depth information of each position includes pile length, stratum category involved and corresponding stratum thickness.

[0018] In step 2), the construction environment parameters include weather, daily effective working time;The process parameters include preparation activity, hole drilling, hole cleaning, construction efficiency and process duration of filling and vibration;The resource allocation parameters include mechanical equipment model, quantity, characteristic parameters and construction parameters involved, construction personnel information and gravel.

[0019] In step 3), the established ABM simulation system for vibro-replacement pile construction has the characteristics including entities, attributes, activities, system states, events and simulation clocks; a plurality of intelligent agents Agent established according to different entities existing in the process include a construction management intelligent agent Agent for coordinating, guiding and scheduling the construction process according to the vibro-replacement pile construction plan, a hole guiding intelligent agent Agent for accepting the instructions of the construction management intelligent agent Agent to guide the hole guiding machine to complete hole guiding and scheduling, and a vibro-replacement intelligent agent Agent for controlling the vibro-replacement machine to complete the vibro-replacement process of the planned vibro-replacement pile and scheduling; the plurality of intelligent agents Agent have the characteristics including autonomy, sociality and goal-oriented capability.

[0020] In step 4), based on the vibro-replacement pile construction progress simulation model established in step 3), a vibro-replacement pile construction progress simulation and optimization model is established by combining hierarchical reinforcement learning, and the optimization problem is divided into two levels of low level and high level; the low level is responsible for executing specific construction operations, and the optimization target of the low level is the construction scheme in the flow unit under the given mechanical configuration; the high level is responsible for the mechanical configuration and scheduling of the flow unit, and the optimization target of the high level is the mechanical configuration scheme under the condition of resource limitation; meanwhile, in step 4), the comprehensive objective function of the vibro-replacement pile construction progress simulation and optimization mathematical model is given.

[0021] Compared with the prior art, the present application has the following beneficial effects:

[0022] The present application establishes an ABM intelligent simulation model for the construction progress of vibro-replacement stone piles for the zoned flow construction of the vibro-replacement stone piles, and realizes intelligent simulation by setting intelligent agents Agent for a plurality of entities in the construction process; the complex construction progress optimization problem is divided into two levels of flow unit mechanical scheduling level and unit internal vibro-replacement pile construction sequence, and a hierarchical reinforcement learning algorithm of OC-MAPPO architecture is used to solve the complex condition vibro-replacement pile construction optimization scheme. According to the above theory and method, a hierarchical reinforcement learning-based vibro-replacement pile construction progress intelligent simulation and optimization method is proposed, which can effectively consider the flow unit mechanical scheduling and unit internal construction sequence, simulate and analyze the vibro-replacement pile construction process under the complex conditions of super-deep overburden layer, and obtain the most ideal vibro-replacement pile construction scheme. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is the hierarchical reinforcement learning-based vibro-replacement pile construction progress intelligent simulation and optimization solving process of the present application;

[0024] Figure 2 is the vibro-replacement pile construction process flowchart of the present application;

[0025] Figure 3 is the different-stage vibro-replacement pile construction simulation process time chart of the present application;

[0026] Figure 4 This invention provides a simulation progress diagram of phased and zoned construction of vibratory compaction piles. Detailed Implementation

[0027] This invention proposes an intelligent simulation and optimization method for vibro-compacted pile construction progress based on hierarchical reinforcement learning. The design concept addresses the shortcomings of existing research by comprehensively considering the constraints between vibro-compacted pile zoning and resource allocation. With construction progress and machinery utilization as optimization objectives, and based on multi-agent simulation theory, a hierarchical reinforcement learning method is used to establish an intelligent simulation and hierarchical optimization model for vibro-compacted pile construction progress, achieving both construction progress simulation and construction scheme optimization. This method simplifies the complexity of the original problem by decomposing the complex construction progress optimization problem into two levels: a machinery scheduling level and a vibro-compacted pile construction sequence within a unit. Specific sub-problems are addressed at each level, thereby improving the model's solution efficiency. The method of this invention mainly includes: first, acquiring information for the simulation modeling of vibro-compacted pile construction; second, determining the simulation parameters for vibro-compacted pile construction; then, based on multi-Agent modeling and simulation, establishing an ABM simulation model for the construction progress of multi-Agent vibro-compacted piles to achieve intelligent simulation of construction progress; further, for the task of zoned and sequential construction of vibro-compacted piles, using a hierarchical reinforcement learning method to establish a hierarchical optimization model for the simulation model of the construction progress of vibro-compacted piles; finally, inputting the simulation parameters, running the intelligent simulation and optimization model of the construction progress of vibro-compacted piles, and obtaining the most ideal construction organization scheme.

[0028] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the following embodiments are by no means intended to limit the present invention.

[0029] This invention proposes an intelligent simulation and optimization method for vibro-compacted pile construction progress based on hierarchical reinforcement learning. This method can effectively consider the mechanical scheduling of flow units and the construction sequence within units to obtain an ideal vibro-compacted pile construction scheme. The specific steps are as follows:

[0030] Step 1) Obtain simulation modeling information for vibro-compacted pile construction.

[0031] The simulation modeling information for vibro-compacted pile construction includes vibro-compacted pile construction plan information and processing depth information.

[0032] First, obtain the vibro-compacted pile construction plan information from the construction organization design, including the phased construction plan, the schedule constraints for each phase, and the pile layout information. Based on the vibro-compacted pile plan drawings and the phased construction plan table, obtain the plan layout scope, pile type, pile spacing, and relevant time constraints for each phase of vibro-compacted piles. Using the above information, perform the plan layout of each phase of vibro-compacted piles, obtain the location coordinates of each vibro-compacted pile, generate an information table of vibro-compacted piles to be constructed, and number them according to rows and columns.

[0033] Secondly, based on the geological model of the construction area, the pile treatment depth information at each location is obtained according to the design depth requirements, including pile length, type of stratum involved, and corresponding stratum thickness. According to the stratum location to be reached by the vibro-compacted piles as designed, the actual treatment depth, type of stratum involved, and thickness of each stratum are obtained from the geological model, corresponding to the location coordinates of each vibro-compacted pile.

[0034] Finally, input files for the phased construction of vibro-compacted piles are generated based on the location and treatment depth information. This involves summarizing the location coordinates of the vibro-compacted piles corresponding to the information obtained from the geological model into an information table for the vibro-compacted piles to be constructed, resulting in input files for each construction phase, including the vibro-compacted pile number, location coordinates, treatment depth, and the thickness of each stratum involved from top to bottom.

[0035] Step 2) Determine the simulation parameters for vibro-compacted pile construction.

[0036] The simulation parameters for vibro-compaction pile construction include three main categories: construction environment parameters, process parameters, and resource allocation parameters. The construction environment parameters include weather, daily effective working hours, and daily work shifts; the process parameters include the construction efficiency and duration of preparatory activities, pilot hole drilling, hole cleaning, and filling and vibro-compaction. The vibro-compaction pile construction process is detailed below. Figure 2 The resource configuration parameters include the model, quantity, characteristic parameters, safety distance and construction efficiency parameters of mechanical equipment, construction personnel information, and crushed stone materials, etc. The main mechanical equipment includes drilling machinery and vibratory compaction machinery. The simulation process time for different levels of vibratory compaction pile construction is shown in [link to simulation]. Figure 3 .

[0037] Step 3) Establish an intelligent simulation model for the construction progress of vibratory compaction piles.

[0038] The intelligent simulation model for the construction progress of vibro-compacted piles proposed in this invention is a multi-agent modeling and simulation model, which includes establishing an ABM simulation system for vibro-compacted pile construction and establishing multiple agents based on different entities existing in the process.

[0039] A simulation system for simulating the construction progress of vibro-compacted piles is constructed. This system includes features such as entities, attributes, activities, system states, events, and a simulation clock.

[0040] (1) Entity: refers to the object inside the system boundary. The constructed construction progress simulation system mainly includes construction management personnel who conduct construction organization design, drilling machinery and filling vibratory compaction machinery entities. Among them, the construction management personnel are responsible for formulating the vibratory compaction pile construction plan and coordinating, guiding and scheduling the construction process, while the drilling machinery and filling vibratory compaction machinery are responsible for receiving the instructions of the construction management personnel to arrive at the designated flow construction unit to cooperate in completing the construction and scheduling.

[0041] (2) Attributes: These refer to the characteristics of entities in the system. By designing intelligent agents for the entities in the system, the attributes of each entity are described, and the behavior of each entity is intelligently simulated. Corresponding to the entities in (1), this invention establishes a construction management agent, a borehole drilling agent, and a vibratory compaction agent.

[0042] (3) System state: refers to the description of all entities, attributes and activities in the system at a certain moment. At a certain moment, the system state is the set of all entity attributes in the system. The system status of vibratory compaction pile construction at any given time is determined by the vibratory compaction piles and mechanical equipment in each flow unit. The attributes and characteristics of a given moment.

[0043] (4) Activity: refers to a process that occupies a certain amount of time and / or resources and causes a change in the system state. Activities such as the construction process cycle of any unit vibratory pile and the coordination and scheduling of mechanical equipment cause changes in the system state.

[0044] (5) Event: An event is an instantaneous behavior that causes a change in the state of the system, which can be the start or end of an activity. This invention takes the start and end of any of the above-mentioned vibratory compaction pile construction activities as events that cause changes in the state of the system, thus promoting the development of construction simulation.

[0045] (6) Simulation Clock: This is a variable used to track the current simulation time, pushing the time from one value to the next. The simulation clock includes two types: the full-process simulation clock and the local simulation clock. At the start of the simulation, the full-process simulation clock is activated first, and the model is set to its initial state. When the full-process simulation clock reaches a specific flow construction unit, it retains its current state, while the local simulation clock is activated, entering the vibratory compaction pile construction cycle within the unit. The start time of the preparation activity is set to the zero point of the local simulation clock. Then, different simulation activity procedures are used as discrete events to advance the local simulation clock until all vibratory compaction pile construction in the current unit is completed, thus obtaining the current unit's construction period. Then, the system returns to the full-process simulation clock, and sends back the corresponding status, resource utilization, and other relevant information from the local simulation clock to the full-process simulation clock, saving it as the current simulation result. The full-process simulation clock then progresses, repeating the above steps until all unit construction is completed.

[0046] The agents established based on the different entities involved in the process include a construction management agent that formulates the vibro-compaction pile construction plan and coordinates, guides and schedules the construction process; a pilot hole agent that is responsible for receiving instructions from the construction management agent, arriving at the designated flow construction unit, instructing the pilot hole machine to complete the pilot hole and scheduling it; and a vibro-compaction agent that is responsible for controlling the vibro-compaction machine to complete the planned vibro-compaction pile and scheduling the vibro-compaction process. The characteristics of the established agents include autonomy, responsiveness, sociality and goal-oriented capabilities.

[0047] The autonomy refers to the ability of an agent to operate its core functions without human or other agent intervention, including autonomous perception (acquiring environmental information relevant to itself), autonomous decision-making (selecting actions based on internal rules, strategies, or learning mechanisms), autonomous action (performing movements, tasks, resource usage, etc.), and autonomous state management (updating its own state variables), thereby controlling its own behavior and internal state.

[0048] The reactivity refers to the agent's ability to respond actively and passively. The active response capability means that the agent can proactively carry out activities based on its own goals and beliefs. The passive response capability means that the agent can perceive the surrounding environment, respond to environmental changes, and take corresponding action plans.

[0049] The social nature refers to the ability of an agent to interact with other agents through a certain agent communication language, to communicate and cooperate, and to collaborate with each agent to achieve its own goals when each agent has different objectives. The goal-oriented capability refers to the agent's ability to handle complex and high-level tasks, including multi-step task decomposition, task planning, task scheduling, multi-objective optimization, etc., and to have the ability to divide and schedule tasks to meet system objectives.

[0050] The functional design and decision-making behavior of each Agent in this invention are as follows:

[0051] The main functions of the construction management agent are: to formulate vibro-compacting pile construction plans and to coordinate, guide, and schedule the construction process, including allocating machinery to vibro-compacting pile construction units and scheduling machinery between units. The construction management agent's decision-making behavior includes: dynamically allocating idle machinery to construction units based on changes in system status during construction until all vibro-compacting pile construction tasks are completed, and recording the time sequence of machinery allocation during the construction process.

[0052] The mechanical configuration scheme of the Agent unit actually determines the mechanical scheduling process and the order of construction of each unit.

[0053] The aforementioned pre-drilling agent is the construction entity for the pre-drilling process. Its main functions are: receiving instructions from the construction management agent and dispatching it to a designated unit, instructing the pre-drilling machinery to complete the pre-drilling process for vibro-compacted piles. The pre-drilling agent's decision-making behavior involves instructing the pre-drilling machinery to dynamically select pile numbers within the unit that have not yet undergone pre-drilling work, based on changes in the unit's status, until all vibro-compacted piles within the unit have completed the pre-drilling process. It records the machinery's construction path and time nodes during the construction process, and statistically analyzes the working time and workload. The pre-drilling agent's decision-making behavior effectively determines the construction sequence of the vibro-compacted pile pre-drilling process within the unit.

[0054] The vibro-compaction Agent is the construction entity of the vibro-compaction process. Its main functions are: receiving instructions from the construction management Agent, arriving at the designated unit, instructing the vibro-compaction machinery to complete the vibro-compaction process of the vibro-compaction piles, and scheduling operations. The vibro-compaction Agent's decision-making behavior: Within a unit, it instructs the vibro-compaction machinery to dynamically select piles that have completed the hole-cleaning process within the unit to complete the filling and vibro-compaction process until all vibro-compaction piles within the unit are vibro-compacted. It records the construction path and time nodes of the machinery during the construction process and statistically analyzes the working time and workload. The vibro-compaction Agent's decision-making behavior is essentially a collaborative decision-making behavior with the pilot-hole agent under the constraints of the construction process, thereby determining the completion time of each vibro-compaction pile within the unit.

[0055] Based on the acquired simulation modeling information of vibro-compacted pile construction, a construction progress simulation system and a multi-agent model are established. The intelligent simulation model of vibro-compacted pile construction progress realizes the simulation of the construction process from the formulation of vibro-compacted pile construction plan to the collaborative execution of the construction plan by multiple machines. The simulation clock of the running model simulates the construction process and can calculate the progress simulation results.

[0056] Step 4) Establish an intelligent simulation and optimization model for the construction progress of vibro-compacted piles. This invention proposes to use a hierarchical reinforcement learning algorithm to establish the simulation and optimization model for the construction progress of vibro-compacted piles. The solution process of this invention is as follows: Figure 1 The comprehensive objective function of the mathematical model for simulation and optimization of vibro-compacted pile construction progress is as follows:

[0057]

[0058] In the formula, For the comprehensive objective function, The construction critical path is the total construction period. The construction critical path is the path with the longest total duration consisting of multiple construction tasks. Any delay in any task on this path will delay the total construction period of the entire project. The unit is hours (h). Indicates the utilization rate of machinery; To maintain a safe distance; , For any two mechanical positions; , The weights are dynamic and change with the number of training steps. Since critical path duration, machine utilization, and collision-free safety distance are considered simultaneously, they are used... To adjust the weights, For the sake of the work option, As the weight for machinery utilization rate, For safety weights, .

[0059] Wherein, the objective function weights Using PID dynamic adjustment method:

[0060]

[0061] PID control term calculation:

[0062]

[0063] Error terms are calculated based on target completion rate:

[0064] In step 4) of this invention, based on the established intelligent simulation model of vibro-compacted pile construction progress, a reinforcement learning environment is provided to simulate multi-mechanical dynamic interactions and construction process constraints, avoid collisions and unreasonable construction, and provide system state space and action space:

[0065] Global state space of a reinforcement learning system:

[0066]

[0067] Mechanical state matrix : ;

[0068] The state of any machine includes the machine position coordinates. Mechanical category :0=Drill hole machinery ,1=Vibratory impact machinery Current working status : 0 = Idle, 1 = Working; Assign unit ID, where: -1 = Unassigned.

[0069] Element state matrix : ;

[0070] The state of any unit includes the construction stage it is in. : 0 = Not started, 1 = Under construction, 2 = Completed; Unit pilot hole work progress (normalized progress value): Unit vibration compaction work progress (normalized progress value): .

[0071] Vibro-compacted pile state matrix : ;

[0072] The arbitrary vibro-compactor pile state includes position coordinates. Current status : 0 = Not started, 1 = Drilling in progress, 2 = Drilling completed, awaiting vibratory compaction, 3 = Vibratory compaction in progress, 4 = Completed; Belongs to the construction flow unit ,

[0073] Time characteristics Simulate time steps, update the state space, make decisions and execute actions at each time step, and record data;

[0074] Action space that satisfies construction process constraints and filters out invalid actions:

[0075]

[0076] In the formula, For a set of valid actions; where, This indicates that vibratory compaction should not be activated until the pre-drilling work for the current pile is completed.

[0077] In step 4) of this invention, the complex problem of optimizing the construction progress of vibro-compacted piles is decomposed into the following two levels:

[0078] The high-level scheduling layer is responsible for the mechanical configuration and scheduling of the flow units, involving the construction management agent; including configuring machinery according to the flow units and scheduling between flow units; the task of the high-level layer is to dynamically allocate mechanical resources throughout the construction process to complete the construction needs of all units. The optimization objective of the high-level layer is the mechanical configuration scheme under resource constraints; the Option-Critic reinforcement learning algorithm is used.

[0079] The lower-level execution layer is responsible for executing specific construction operations, involving two mechanical agents; including receiving instructions from the construction management agent to dispatch to designated units, and executing the unit crushed stone pile construction process in a certain order; the task of the lower layer is to direct the mechanical actions to achieve specific construction tasks; the optimization objective of the lower layer is the construction plan within the flow unit under a given mechanical configuration; the MAPPO reinforcement learning algorithm is used.

[0080] This invention utilizes hierarchical reinforcement learning to solve the model, where the high-level scheduling layer obtains the state space of the construction management agent based on the global state:

[0081]

[0082] Mechanical state matrix : ;

[0083] The state of any machine includes the machine position coordinates. Mechanical category :0=Drill hole machinery ,1=Vibratory impact machinery Current working status : 0 = Idle, 1 = Working; Assign unit ID, where: -1 = Unassigned;

[0084] Element state matrix : ;

[0085] The state of any unit includes the construction stage it is in. : 0 = Not started, 1 = Under construction, 2 = Completed; Unit pilot hole work progress (normalized progress value): Unit vibration compaction work progress (normalized progress value): ;

[0086] This invention utilizes hierarchical reinforcement learning to solve the action space of the construction management agent in the high-level scheduling layer of the model:

[0087]

[0088] Among them, mechanical allocation strategy ; Distributed flow units Distribute the pilot hole machinery to the flow unit. Distribute vibratory compaction machinery to the production line unit. .

[0089] This invention utilizes hierarchical reinforcement learning to solve the comprehensive objective function of the mathematical model based on the simulation and optimization of the construction progress of vibratory compaction piles. In relation to the design of the higher-level scheduling layer, the reward function of the higher-level scheduling layer is as follows:

[0090]

[0091] In the formula, For high-level reward functions; As a key schedule incentive, we encourage the allocation of machinery to critical units to reduce the critical path duration; Incentives are given to improve machinery utilization and avoid idleness. To ensure safety and prevent collisions during cross-unit scheduling of machinery.

[0092] Among them, the reward function The PID dynamic adjustment method is adopted, with the same objective function. Maintain consistency.

[0093] In this invention, a hierarchical reinforcement learning model is used to solve a problem. The high-level scheduler employs the Option-Critic algorithm to implement machine allocation and scheduling tasks. The high-level scheduler obtains the state space of the construction management agent based on the global state and generates the action space. The process of the high-level scheduler using the Option-Critic algorithm is as follows:

[0094] A1. Clear the experience replay cache: Initialize the state space Randomly initialize the network parameters of the Option-Critic algorithm strategy. and Critic network parameters ;

[0095] B1. Looping trajectory generation (Episode), including:

[0096] B1-a) Option-Critic algorithm strategy: network input, current state Output the probability distribution of the Option policy: ;

[0097] B1-b) Strategy for selecting the current time step based on probability distribution sampling Execution, that is, allocating machinery to construction units in a flow process: ;

[0098] B1-c) The lower layer accepts the strategy of the higher layer to execute construction tasks and updates the unit state, and calculates the reward function of the higher layer. ;

[0099] B1-d) Critic neural network calculates the state value function That is, the long-term expected return under the current state: ;

[0100] B1-e) Calculate the Option value function That is, to execute the Option in the current state. Long-term expected returns: ;

[0101] B1-f) Temporal difference calculation of state value function TD objectives: ;

[0102] B1-g) Calculate the loss function of the Critic network: ;

[0103] B1-h) Minimize the gradient descent of the Critic network loss function and update the Critic network parameters. ;

[0104] B1-i) Calculate the network advantage function of the policy. : ,like This indicates that the current Option is better than the average strategy, and its probability should be increased;

[0105] B1-j) Advantage Function Policy gradient update policy network parameters ;

[0106] B1-k) proceeds to the next time step, and executes B1-b)~B1-g) cyclically according to the new state space.

[0107] B1-l) Termination condition: All units That is, to complete all unit construction tasks.

[0108] B1-m) records the complete simulation time step sequence: It is stored in the experience pool as a single trajectory (Episode).

[0109] C1. After completing every K trajectories, batch sample the most recently retained trajectory data from the experience pool, update the Critic network and policy network, and optimize the policy to the reward function. It tends to converge or reaches the preset maximum number of trajectories.

[0110] This invention utilizes hierarchical reinforcement learning to solve the model in which lower-level actuators, the orifice agent and the vibration agent, accept policies from higher-level agents. The system is dispatched to a designated flow unit, and the building status information of that unit is obtained based on the assigned unit to obtain the lower-level state space.

[0111]

[0112] unit Internal mechanical state matrix : ;

[0113] unit Internal vibration impact pile state matrix : ;

[0114] unit Intra-temporal characteristics : Simulation time step within the cell.

[0115] The low-level actuator vibration agent's action space includes:

[0116] The action space of the aperture agent: ;

[0117] Action space of the shock agent: ;

[0118] In the formula, This indicates that vibratory compaction should not be activated until the pre-drilling work for the current pile is completed.

[0119] In hierarchical reinforcement learning, the comprehensive objective function based on the simulation and optimization mathematical model of vibro-compaction pile construction progress is... Compared to low-level actuator design, the low-level actuator reward function is:

[0120]

[0121] In the formula, For low-level reward functions; As a progress incentive, the overall progress of each machine within the unit is encouraged to accelerate the overall construction progress within the unit; Utilization incentives are used to encourage increased machine efficiency and avoid idleness. For safety reasons, collisions should be avoided within the mechanical unit.

[0122] Among them, the reward function The PID dynamic adjustment method is adopted, with the same objective function. Maintain consistency.

[0123] The lower-level actuator uses the MAPPO algorithm to implement construction tasks within the flow unit, including:

[0124] A2. Clear the experience replay cache: Initialize the state space Randomly initialize the strategy parameters of the aperture agent and the vibration agent. , and centralized value network parameters ;

[0125] B2. Loop generation of trajectory (Episode), including:

[0126] B2-a) Input the current low-level state space into the MAPPO algorithm The network outputs the probability distribution of the actions of the hole-piercing machine and the vibratory impacting machine. : ;

[0127] B2-b) The pilot hole agent and vibratory compaction agent perform action sampling based on the action distribution p, that is, select the station number for construction at the current time step: ;

[0128] B2-c) Execute the action and calculate the lower-level reward function immediately. Store and transfer data:

[0129] B2-d) Centralized Critic network calculation of state value function : That is, the long-term expected return under the current state;

[0130] B2-e) Temporal difference calculation of state value function TD objectives: ;

[0131] B2-f) Calculate the critic network loss function: ;

[0132] B2-g) Minimize the gradient descent of the Critic network loss function and update the Critic network parameters. ;

[0133] B2-h) Each machine calculates its own motion dominance function: ;;Pruning mechanism gradient ascent update

[0134] B2-i) Each machine calculates its own importance sampling probability ratio: ;

[0135] B2-j) PPO-Clip operation: Limiting the probability ratio in... Between, when calculating Avoid using boundary values ​​within intervals to maintain exploration capabilities and training stability;

[0136] B2-k) Based on the advantage function after the clip operation The distributed policy gradient updates the policy parameters of each machine. , ;

[0137] B2-l) Proceed to the next time step, and execute B2-a) ~ B2-k) cyclically according to the new state space.

[0138] B2-m) Termination condition: All vibratory compaction piles within the unit. This means that the unit construction is completed;

[0139] The complete stored procedure transfer data (B2-n) is stored as a single episode in the experience pool.

[0140] C2. After completing a fixed number of trajectories, batch sample the most recently retained trajectory data from the experience pool, update the centralized Critic network, update the policy networks of the orifice agent and the vibration agent, and optimize to the reward function. It tends to converge or reaches the preset maximum number of trajectories.

[0141] In this invention, this hierarchical approach allows for a clearer distinction between the tasks of agents at different levels, enabling the use of algorithms at each level to solve the construction optimization problem more effectively. Each agent at each level focuses on its own task and, through collaboration and coordination, maximizes overall construction efficiency. To solve the optimization problem, this invention employs a hierarchical reinforcement learning algorithm with an OC-MAPPO hybrid architecture, primarily comprising high- and low-level state spaces, action spaces, and reward functions. The state space refers to the set of all possible states in the environment; the action space refers to the set of all possible actions the agent can take; and the reward function defines the immediate reward and the long-term global reward obtained by the agent after taking each action in each state.

[0142] Step 5) Solve the intelligent simulation and optimization model of the construction progress of vibro-compacted piles based on hierarchical reinforcement learning, and optimize the construction scheme of vibro-compacted piles.

[0143] First, initialize the model and input simulation parameters; second, run the simulation model for optimization using a hierarchical reinforcement learning algorithm and record the simulation optimization results; finally, adjust the algorithm parameters according to the algorithm execution results until the optimization objective is met, obtain the final construction scheme, and output the simulation results of the scheme.

[0144] In this example, using the phased and zoned construction design data of a vibro-compacting project, an intelligent simulation and optimization method for the construction progress of vibro-compacting piles based on hierarchical reinforcement learning is employed to obtain the optimized simulated phased and zoned construction progress of the vibro-compacting piles, as follows: Figure 4 As shown, Figure 4 The simulation prediction results of the optimized scheme are shown.

[0145] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many improvements and changes under the guidance of the present invention without departing from the spirit of the present invention, and these improvements and changes are all within the protection scope of the present invention.

Claims

1. A method for intelligent simulation and optimization of vibro-compacted pile construction progress based on hierarchical reinforcement learning, characterized in that, Includes the following steps: Step 1) Obtain simulation modeling information for vibro-compacted pile construction, including vibro-compacted pile construction plan information obtained from the construction organization design, and pile treatment depth information at each location obtained by combining the geological model of the construction area and according to the design depth requirements; and generate vibro-compacted pile phased input files based on the location and treatment depth information; Step 2) Determine the simulation parameters for vibro-compacted pile construction, including construction environment parameters, process parameters, and resource allocation parameters; Step 3) Establish an intelligent simulation model for the construction progress of vibro-compacted piles. The simulation model for the construction progress of vibro-compacted piles is a model of a multi-agent modeling and simulation method, which includes the established vibro-compacted pile construction ABM (Agent based modeling) simulation system and multiple intelligent agents established according to different entities existing in the process. Step 4) Based on the simulation model of the construction progress of the vibratory compaction pile established in Step 3), a mathematical model for simulation and optimization of the construction progress of the vibratory compaction pile is established by combining hierarchical reinforcement learning, and the data model is solved based on the OC-MAPPO hierarchical reinforcement learning algorithm. Step 5) Input simulation parameters, run the vibration compaction pile construction progress simulation and optimization model, and obtain the optimized vibration compaction pile construction progress scheme.

2. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 1), the vibro-compacting pile construction plan information includes the phased planning of vibro-compacting pile construction, the progress constraint information of each phase, and the pile layout information; the pile treatment depth information at each location includes the pile length, the type of stratum involved, and the corresponding stratum thickness.

3. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 2), the construction environment parameters include weather and daily effective working hours; the process parameters include the construction efficiency and process duration of preparation activities, pilot hole drilling, hole cleaning, and filling and vibratory compaction; the resource allocation parameters include the model, quantity, characteristic parameters and construction parameters of mechanical equipment, construction personnel information and crushed stone.

4. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 3), the ABM simulation system for vibro-compacted pile construction established has the following features: entities, attributes, activities, system states, events, and simulation clock. The entity refers to the object inside the system boundary, including all vibratory compaction piles and construction machinery and equipment that need to be constructed; The attribute refers to the characteristics of entities in the system, including the working efficiency and behavior rules of construction machinery; The activity refers to a process that occupies a certain amount of time and / or resources and causes a change in the system state; The system state refers to the description of all entities, attributes, and activities in the system at a certain moment; at a certain moment, the system state is the set of all entity attributes in the system. The event is an instantaneous action that causes a change in the state of the system; the event is divided into deterministic events or conditional events, the occurrence time of deterministic events is predetermined, and the conditional events begin after certain conditions are met; The simulation clock is a variable used to track the current simulation time, pushing the time from one value to the next.

5. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 3), multiple intelligent agents are established based on different entities in the process, including a construction management intelligent agent that coordinates, guides and schedules the construction process according to the vibro-compaction pile construction plan, a pilot hole intelligent agent that is responsible for receiving instructions from the construction management intelligent agent, arriving at the designated flow construction unit to instruct the pilot hole machine to complete the pilot hole and schedule it, and a vibro-compaction intelligent agent that is responsible for controlling the vibro-compaction machine to complete the planned vibro-compaction pile vibration process and scheduling the vibro-compaction.

6. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 5, characterized in that, The multiple intelligent agents mentioned above possess characteristics including autonomy, sociality, responsiveness, and goal-oriented capabilities; The autonomy refers to the ability of an agent to operate its core functions without human or other agent intervention, including autonomous perception, autonomous decision-making, autonomous action, and autonomous state management, thereby controlling its own behavior and internal state. The reactivity refers to the agent's ability to respond actively and passively. The active response capability means that the agent can proactively carry out activities based on its own goals and beliefs. The passive response capability means that the agent can perceive the surrounding environment, respond to changes in the environment, and take corresponding execution plans. The social nature refers to the ability of an agent to interact with other agents through a certain agent communication language, to communicate and cooperate, and to collaborate with each agent to achieve its own goals when each agent has different objectives. The goal-oriented capability refers to the ability of the agent to handle multi-step task decomposition, task planning, task scheduling, and multi-objective optimization, and to have the ability to divide and schedule tasks to meet system objectives.

7. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 6, characterized in that, The autonomous perception function is used to acquire environmental information related to itself; the autonomous decision-making function includes selecting actions based on internal rules, strategies, or learning mechanisms; the autonomous action function is used to perform movement, tasks, and resource usage; and the autonomous state management function is used to update its own state variables.

8. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 4), based on the simulation model of the construction progress of the vibratory compaction pile established in step 3), a simulation and optimization model of the construction progress of the vibratory compaction pile is established by combining hierarchical reinforcement learning, and the optimization problem is decomposed into two levels: low level and high level. The lower layer is responsible for executing specific construction operations, involving two mechanical agents; including receiving instructions from the construction management agent and scheduling to designated units, and executing the unit crushed stone pile construction process in a certain order; the task of the lower layer is to direct the mechanical actions to achieve specific construction tasks; the optimization objective of the lower layer is the construction plan within the flow unit under a given mechanical configuration; The upper layer is responsible for the mechanical configuration and scheduling of the flow unit, involving the construction management agent; including configuring machinery according to the flow unit and scheduling between flow units; the task of the upper layer is to dynamically allocate mechanical resources throughout the construction process to complete the construction needs of all units, and the optimization goal of the upper layer is the mechanical configuration scheme under resource constraints.

9. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 1, characterized in that, In step 4), the comprehensive objective function of the mathematical model for simulating and optimizing the construction progress of the vibro-compacted piles is as follows: ; In the formula, For the comprehensive objective function, The construction period for the critical path is the total construction period, expressed in hours. Indicates the utilization rate of machinery; To maintain a safe distance; , For any two mechanical positions; , The dynamic weights vary with the number of training steps, taking into account critical path duration, machine utilization, and safe distance for collision avoidance. To adjust the weights, among which, For the sake of the work option, As the weight for machinery utilization rate, For safety weights, ; The OC-MAPPO hierarchical reinforcement learning algorithm is used to solve the model, which includes a state space, an action space, and a reward function. The state space refers to the set of all possible states in the environment. The action space refers to the set of all possible actions that the agent can take. The reward function defines the immediate reward and the long-term reward from a global perspective after the agent takes each action in each state.

10. The intelligent simulation and optimization method for the construction progress of vibro-compacted piles according to claim 9, characterized in that, The critical path of construction refers to the path with the longest total duration consisting of multiple construction tasks. Any delay in any task on this path will delay the overall project duration.