A scheduling decision method and device for concrete arch dam pouring engineering

By employing a two-stage training method combining intelligent agent models and reinforcement learning, and integrating arch dam pouring progress simulation and scheduling decisions, the problem of insufficient dynamic adaptability and robustness in existing arch dam pouring scheduling methods has been solved. This has enabled automated and intelligent management of arch dam construction and optimized construction efficiency.

CN122113211APending Publication Date: 2026-05-29TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-01-19
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing arch dam casting scheduling methods are insufficient in terms of dynamic adaptability and robustness. They fail to effectively integrate construction process simulation with global progress benefits, making it difficult to provide feasible, optimizable, and adaptable intelligent scheduling solutions in complex dynamic scenarios.

Method used

An arch dam pouring progress simulation model based on intelligent agent model and discrete event simulation is adopted. Combined with a two-stage training method of reinforcement learning and behavior cloning, the cable crane pouring scheduling decision is generated through expert experience and reward reshaping, forming a closed-loop intelligent decision-making process from perception, simulation, decision-making to control.

Benefits of technology

It improves the robustness and scheduling efficiency of the model, realizes the automated and intelligent management of arch dam pouring construction, can learn autonomously and optimize long-term construction benefits, and adapt to real-time changes in the construction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113211A_ABST
    Figure CN122113211A_ABST
Patent Text Reader

Abstract

The application provides a scheduling decision method and device for concrete arch dam pouring engineering, and the method comprises the following steps: constructing an arch dam pouring progress simulation model through discrete event simulation and an agent-based model, pre-training the expert trajectory dataset by using behavior cloning, performing reinforcement learning on the behavior cloning pre-training model with reward remodeling to generate an intelligent scheduling decision model; perceiving the state variables required by the progress simulation model, initializing and running the progress simulation model; generating and executing the cable crane pouring scheduling decision result according to the state variables through the intelligent scheduling decision model. The two-stage training method fusing expert experience and reinforcement learning enables the model to autonomously learn and optimize long-term construction benefits, improves the robustness and scheduling efficiency of the model, and through offline training and online application of the model, a closed-loop intelligent decision process from perception, simulation, decision to control is formed, and reliable support is provided for the automation and intelligent management of arch dam pouring construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of arch dam casting simulation technology, and in particular to a scheduling decision-making method and apparatus for concrete arch dam casting projects. Background Technology

[0002] Concrete arch dams have become an important dam type in dam construction due to their excellent resistance to surcharges and their economic efficiency. However, arch dams are mostly built in high mountain and canyon areas, facing complex geological conditions such as high slopes and high ground stress, which places higher demands on construction technology. With the development of technologies such as the Internet of Things, digital twins, and artificial intelligence, arch dam construction is gradually moving from automation and digitalization to intelligentization. In this process, the main concrete pouring of the arch dam is a key construction link, subject to multi-dimensional constraints such as time, space, resources, and technology, and has significant uncertainties. Therefore, achieving scientific and intelligent pouring scheduling decisions is of great significance for ensuring construction progress and improving project efficiency.

[0003] Currently, several typical methods have been developed for the scheduling of arch dam pouring, but significant limitations remain. One type of method relies primarily on expert experience and the current construction status, using optimization techniques such as particle swarm optimization to weight and rank dam segment characteristics. However, this is essentially a static or semi-static ranking strategy, failing to fully consider the dynamic evolution of the dam section ranking as the project progresses, thus making it difficult to adapt to real-time changes during construction. Another type of method attempts to use intelligent algorithms such as deep Monte Carlo tree search to generate ranking schemes. While this considers the influence of decision sequences to some extent, it does not deeply integrate the temporal randomness and dynamic interaction of equipment during construction, potentially leading to insufficient robustness in actual implementation. Furthermore, some studies have analyzed the cable crane operation range and interference constraints through simulation models, but these focuses primarily on the feasibility of local operations or quality constraints, failing to systematically integrate construction process simulation with global optimization objectives such as schedule benefits, thus failing to provide effective decision support for overall project schedule optimization.

[0004] In summary, most existing arch dam pouring decision-making methods either lack dynamic adaptability, are insufficient in robustness under stochastic environments, or fail to organically combine construction process simulation with overall schedule benefits. Therefore, they struggle to provide feasible, optimizable, and adaptable intelligent scheduling solutions for real, complex, and dynamic construction scenarios. Consequently, there is an urgent need for a pouring scheduling method that deeply integrates real-time sensing information, dynamic simulation, and intelligent optimization decision-making to support closed-loop management of intelligent arch dam construction. Summary of the Invention

[0005] One objective of this invention is to provide a scheduling and decision-making method for concrete arch dam pouring projects. This method employs a two-stage training approach integrating expert experience and reinforcement learning, enabling the model to autonomously learn and optimize long-term construction benefits, improving model robustness and scheduling efficiency. Through offline training and online application of the model, a closed-loop intelligent decision-making process from perception, simulation, decision-making to control is formed, providing reliable support for the automated and intelligent management of arch dam pouring construction. Another objective of this invention is to provide a scheduling and decision-making device for concrete arch dam pouring projects. A further objective of this invention is to provide a computer-readable medium. A final objective of this invention is to provide a computer device.

[0006] To achieve the above objectives, this invention discloses a scheduling and decision-making method for concrete arch dam pouring projects, comprising: Based on agent-based model and discrete event simulation, a simulation model of the arch dam pouring progress is constructed according to the construction status data in the concrete arch dam pouring project. Collect dynamic observation state vectors and static physical property data, and initialize the arch dam pouring progress simulation model; By using a simulation model of the arch dam pouring progress and a pre-built intelligent scheduling decision model, scheduling is carried out based on dynamic observation state vectors and static physical attribute data to generate cable crane pouring scheduling decision results. The intelligent scheduling decision model is built based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

[0007] Preferably, the construction status data includes dam section characteristics, cable crane data, process parameters, and deck surface data; Based on agent-based models and discrete event simulations, a simulation model for the pouring progress of an arch dam is constructed using construction status data from the concrete arch dam pouring project. This model includes: Based on the intelligent agent model, micro-modeling is performed on cable crane data, and global advancement is carried out through discrete event simulation based on the cable crane single-cycle intelligent agent model and construction events to generate a simulation model framework. By injecting dam section features, cable crane data, process parameters, surface data, and construction constraints into the simulation model framework, a simulation model of the arch dam pouring progress is constructed.

[0008] Preferably, the construction constraints include time constraints, height difference constraints between adjacent dam sections, cantilever height difference constraints, maximum height difference constraints of the entire dam, height difference constraints of the first pouring, cable crane resource constraints, and concrete pouring temperature constraints.

[0009] Preferably, the method further includes: Collect expert trajectory datasets; Based on the expert trajectory dataset, supervised learning pre-training is performed on the policy network to construct a behavior clone pre-trained model; By integrating a near-end policy optimization algorithm that combines action masking with a pre-defined hybrid reward function, and based on collected simulation interaction data, the behavior clone pre-trained model is optimized to generate an intelligent scheduling decision model.

[0010] Preferably, the hybrid reward function includes sparse target reward and dense plasticity reward; Sparse target rewards are either rewards or penalties based on the actual project duration and the baseline project duration; Dense plasticity rewards are either rewards or penalties for each scheduling decision.

[0011] Preferably, the cable crane pouring scheduling decision results are generated by using a simulation model of the arch dam pouring progress and a pre-built intelligent scheduling decision model, based on dynamic observation state vectors and static physical attribute data, including: Feature extraction is performed on the dynamic observation state vector to obtain the state context feature vector; Map static physical attribute data to action embedding vectors; The state context feature vector and action embedding vector are fused to generate the cable crane pouring scheduling decision result.

[0012] This invention also discloses a scheduling and decision-making device for concrete arch dam pouring projects, comprising: The arch dam pouring progress simulation model building unit is used for agent-based model and discrete event simulation. It constructs an arch dam pouring progress simulation model based on the construction status data in the concrete arch dam pouring project. The state data acquisition unit is used to collect dynamic observation state vectors and static physical attribute data, and to initialize the arch dam pouring progress simulation model; The scheduling decision unit is used to schedule the cable crane pouring based on the dynamic observation state vector and static physical attribute data through the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model. The intelligent scheduling decision model is built based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

[0013] The present invention also discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0014] The present invention also discloses a computer device, including a memory and a processor, wherein the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions, wherein the processor executes the program to implement the method described above.

[0015] The present invention also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the method described above.

[0016] This invention is based on an intelligent agent model and discrete event simulation. It constructs a simulation model of the arch dam pouring progress based on construction status data during the concrete arch dam pouring project. It collects dynamic observation state vectors and static physical attribute data and initializes the simulation model. Through the arch dam pouring progress simulation model and a pre-built intelligent scheduling decision model, scheduling is performed based on the dynamic observation state vectors and static physical attribute data to generate cable crane pouring scheduling decision results. The intelligent scheduling decision model is constructed based on a two-stage intelligent agent training method that integrates expert experience and reinforcement learning. This method enables the model to learn autonomously and optimize long-term construction benefits, improving model robustness and scheduling efficiency. Through offline training and online application of the model, a closed-loop intelligent decision-making process from perception, simulation, decision-making to control is formed, providing reliable support for the automation and intelligent management of arch dam pouring construction. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart of a scheduling decision-making method for a concrete arch dam pouring project provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a scheduling decision-making method for a concrete arch dam pouring project, as provided in an embodiment of the present invention; Figure 3 A logic flowchart of a simulation model provided in an embodiment of the present invention; Figure 4 A graph of the reward function during training is provided as an embodiment of the present invention; Figure 5 A schematic diagram of the structure of a scheduling and decision-making device for a concrete arch dam pouring project provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To facilitate understanding of the technical solution provided in this application, the relevant content of the technical solution is described below. This invention simulates the arch dam pouring process and treats the cable crane as an intelligent agent, employing reinforcement learning to dynamically interact with the simulation. Through a scientific reward reshaping scheme and training in a stochastic environment, the dam segment sequencing and cable crane scheduling during the arch dam pouring process are achieved. This invention enables a cable crane scheduling method that balances real-time perception information, progress efficiency optimization, and robustness of the pouring scheme. The effectiveness of this method is verified in a simulation environment, thus yielding an intelligent pouring decision-making method suitable for intelligent arch dam construction.

[0021] The following example uses a scheduling and decision-making device for a concrete arch dam pouring project as the executing entity to illustrate the implementation process of the scheduling and decision-making method for a concrete arch dam pouring project provided in this embodiment of the invention. It is understood that the executing entity of the scheduling and decision-making method for a concrete arch dam pouring project provided in this embodiment of the invention includes, but is not limited to, a scheduling and decision-making device for a concrete arch dam pouring project.

[0022] Figure 1 A flowchart illustrating a scheduling decision-making method for concrete arch dam pouring projects provided in this embodiment of the invention is shown below. Figure 1 As shown, the method includes: Step 101: Based on the agent-based model and discrete event simulation, construct a simulation model of the arch dam pouring progress according to the construction status data in the concrete arch dam pouring project.

[0023] In this embodiment of the invention, the construction status data includes dam section characteristics, cable crane data, process parameters, and surface data, comprising two parts: static parameter settings and dynamic sensing data. The cable crane data describes the state changes of cable crane pouring using an agent-based model (ABM), which includes waiting, loading, heavy load, unloading, and empty return.

[0024] In this embodiment of the invention, the entire dam section pouring process (including joint grouting) is advanced using Discrete-Event Simulation (DES).

[0025] Step 102: Collect dynamic observation state vectors and static physical attribute data, and initialize the arch dam pouring progress simulation model.

[0026] In this embodiment of the invention, the initialization of the arch dam pouring progress simulation model is achieved by initializing some parameters of the model as initial boundary conditions, which are obtained from collected or stored fixed data.

[0027] In this embodiment of the invention, a candidate set of pourable dam sections is determined based on constraints such as intermittent period constraints, adjacent dam section elevation constraints, cantilever height constraints, maximum elevation difference constraints, and first pouring constraints. Illegal dam sections are filtered out by setting the probability of other dam sections being selected to a minimum value. Based on the arch dam pouring progress simulation model, dynamic observation state vectors and static physical attribute data are collected.

[0028] Step 103: Using the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model, scheduling is carried out based on the dynamic observation state vector and static physical attribute data to generate the cable crane pouring scheduling decision result. The intelligent scheduling decision model is built based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

[0029] In this embodiment of the invention, the intelligent scheduling decision model is constructed based on a two-stage agent training process involving behavior cloning and reinforcement learning with reward reshaping.

[0030] In this embodiment of the invention, a behavior cloning method is first used to pre-train the expert trajectory formed by expert experience. When the progress simulation model reaches the point where a pouring decision is needed, a reinforcement learning agent takes over. The agent's observation environment is derived from the simulation environment, and its decision output is to pour a certain dam section or wait. When the training reaches stability, the intelligent scheduling decision model is saved, and the robustness of the training results is verified.

[0031] In the intelligent scheduling decision-making model, small rewards or penalties are added to reshape the reward system by evaluating the state after pouring and judging whether the phased goals (such as through-holes) have been achieved. These rewards, along with the final completion reward, serve as the reward function for reinforcement learning. A function to record the expert's trajectory is designed, and a neural network with pre-trained network parameters is also designed.

[0032] In the technical solution provided by this invention, based on the model of intelligent agents and discrete event simulation, a simulation model of the arch dam pouring progress is constructed according to the construction status data in the concrete arch dam pouring project; dynamic observation state vectors and static physical attribute data are collected and the arch dam pouring progress simulation model is initialized; through the arch dam pouring progress simulation model and the pre-constructed intelligent scheduling decision model, scheduling is performed according to the dynamic observation state vectors and static physical attribute data to generate the cable crane pouring scheduling decision result. The intelligent scheduling decision model is constructed based on a two-stage intelligent agent training method that integrates behavior cloning and reinforcement learning with reward reshaping. It adopts a two-stage training method that integrates expert experience and reinforcement learning, enabling the model to learn autonomously and optimize long-term construction benefits, improve the model's robustness and scheduling efficiency. Through offline training and online application of the model, a closed-loop intelligent decision-making process from perception, simulation, decision-making to control is formed, providing reliable support for the automation and intelligent management of arch dam pouring construction.

[0033] Figure 2 A flowchart of a scheduling decision-making method for a concrete arch dam pouring project provided by an embodiment of the present invention is shown below. Figure 2 As shown, the method includes: Step 201: Based on the intelligent agent model, perform micro-modeling based on the cable crane data, and generate a simulation model framework by performing discrete event simulation, based on the cable crane single-cycle intelligent agent model and construction events.

[0034] In this embodiment of the invention, two simulation paradigms are combined to construct an interactive environment for reinforcement learning, namely, a simulation model framework. Specifically, it includes: Macro-process advancement: Discrete-Event Simulation (DES) is used to drive the simulation clock and event queue for the entire pouring construction (including joint grouting). DES is a technical paradigm for system modeling and analysis. Its core idea is that the system's state variables only change due to the occurrence of "events" at discontinuous, discrete points in time. This method drives the simulation process by maintaining a time-ordered event queue. The simulation clock jumps directly from the current event to the occurrence time of the next most recent event, thus efficiently simulating the dynamic processes of entities flowing through the system, contending for resources, and queuing.

[0035] Micro-resource modeling: An Agent-Based Model (ABM) is employed to finely describe the dynamic behavior and state transitions (such as waiting, loading, overloading, unloading, and empty return) of key resources (i.e., cable cranes). ABM is a bottom-up modeling approach that describes the system as a collection of numerous autonomous, interactive agents. Each agent follows an independent set of decision-making rules and takes actions based on its own state, its environment, and its interactions with other agents.

[0036] Environmental Functions: This hybrid model (DES+ABM) serves as the interactive environment for reinforcement learning agents, i.e., a simulation model framework, providing them with dynamic state observations and executing their decision-making instructions.

[0037] In this embodiment of the invention, the virtual simulation model does not cover the entire pouring construction process, and only controls the key aspects. The construction procedures involved in the virtual simulation environment are shown in Table 1: Table 1

[0038] The entire process of cable crane pouring is driven by discrete event simulation (DES), in which the cable crane, as an intelligent agent, has its loading, heavy load, unloading, empty return state changes and working time calculation logic implemented by an agent-based model (ABM). Figure 3 A logic flowchart of a simulation model provided in an embodiment of the present invention, such as... Figure 3 As shown, at the start of the simulation model, the state data of the dam section, cable crane, etc., are initialized until the stopping condition is met and the simulation ends. The process begins by searching for available cable cranes. If no available cable crane is found for pouring, there is a pause for a period of time. Available cable cranes search for pourable dam sections that meet various constraints, or pour a dam section simultaneously with other cable cranes according to rules. If no such dam section is found or the number of available cable cranes is insufficient to open a new section, there will also be a pause for a period of time. The selection of a new dam section to be poured is performed by a decision module that can receive information from the simulation model. Based on the received dam section state data and related information, the decision module selects the most suitable dam section from the set of pourable dam sections, or allows the cable crane to pause for an appropriate time to obtain better long-term benefits. If the cable crane and dam section are successfully matched, the cable crane agent will start continuously looping after receiving instructions from the decision module until the amount of concrete poured meets the pouring requirements of the target dam section. After pouring, the model updates its state. Before the stopping condition is met, the cable crane agent will return to a waiting state and perform pause and joint grouting checks, and re-pair with the dam section at an appropriate time. This process is continuously executed, and the appearance of the dam section continues to change. During the process, necessary parameters and variables can be recorded.

[0039] Step 202: Inject dam section features, cable crane data, process parameters, surface data, and construction constraints into the simulation model framework to construct a simulation model of the arch dam pouring progress.

[0040] In this embodiment of the invention, concrete pouring needs to consider the actual engineering conditions and meet the specifications. In the simulation model, these are quantified as construction constraints. These constraints include time constraints, elevation differences between adjacent dam sections, cantilever elevation differences, maximum elevation difference across the entire dam, initial pouring elevation differences, cable crane resource constraints, and concrete pouring temperature constraints. The specific constraint definitions are as follows: (1) Time constraint: To ensure that the concrete reaches the required strength, the minimum pouring interval must be met. When selecting dam segment i at time t, the interval between it and the previous pouring activity of that dam segment must satisfy:

[0041] Where t is time; To retrieve the most recent completion time of the pouring of dam segment i, obtained from the dynamic two-dimensional array L; This is the preset minimum pouring interval.

[0042] It is worth noting that the minimum pouring interval can be set according to actual needs, and this embodiment of the invention does not limit this.

[0043] In particular, for areas with complex structures such as openings and corridors, longer working windows are required due to the involvement of embedded parts and metal structure installation, thus strengthening the time constraints:

[0044] Where t is time; To retrieve the most recent completion time of the pouring of dam segment i, obtained from the dynamic two-dimensional array L; The design takes into account the intervals between construction processes such as metal structure installation.

[0045] (2) Elevation difference constraint between adjacent dam sections: In order to prevent construction interference and control the internal stress of the dam body, the pouring elevation H of any two adjacent dam sections i and j is limited. i and H j It must be maintained within the limits allowed by the regulations:

[0046] Among them, H i H is the pouring elevation of dam section i; j Here is the pouring elevation of dam segment j; i and j are two adjacent dam segments; and These are the minimum and maximum allowable elevation differences between adjacent dam sections as required by the construction specifications.

[0047] (3) Cantilever height constraint: Under the consideration of structural safety, the cantilever height of the dam section is limited:

[0048] Among them, H i The pouring elevation of dam section i; Let K be the capping elevation of the grouting area for the kth grouting cycle. The maximum cantilever height difference is required to ensure stress safety.

[0049] (4) Maximum elevation difference constraint of the entire dam: To control the uniform rise of the dam as a whole, at any time t, the set S of all dam sections that have started pouring. t (i.e., number of pours) The height difference between the highest and lowest points within the set of dam sections i must not exceed the upper limit:

[0050] Among them, H i H is the pouring elevation of dam section i; j S is the pouring elevation of dam section j; t This refers to all dam sections that have already begun pouring water; The maximum allowable elevation difference of the entire dam to ensure stress safety.

[0051] (5) Initial pouring height constraint: For pours that have never been poured before ( For the bank slope dam section, when it is first selected for pouring, its initial elevation difference must meet the following requirements to satisfy the structural requirements:

[0052] Among them, H i H is the pouring elevation of dam section i; j The pouring elevation of dam section j; Specific elevation difference limits considered in structural design.

[0053] (6) Cable car resource constraints: 1) Safety distance: Adjacent cable cranes operating simultaneously must maintain a minimum safety distance:

[0054] in, For cable machine c i The coordinates of the large vehicle; For cable machine c j The coordinates of the large vehicle; This is the weather impact coefficient, during windy weather. In normal weather ; This is the minimum safe distance reference value; For all adjacent cable cars that are currently in operation, the trolleys are grouped together.

[0055] 2) Number of Concrete Hoisting Machines Activated: To ensure the quality of poured concrete, the activation of any concrete hopper must meet the minimum required number of hoisting machines. The number of hoisting machines, n, must satisfy the following:

[0056] Where n is the number of cable cranes activated; A is the area of ​​the target storage area; h is the height of the poured concrete layer; v is the average pouring strength of the concrete; t set This refers to the initial setting time of the concrete.

[0057] 3) Operational Scope Coverage: To ensure that all cable cranes can cover the entire target warehouse surface when working together, the following two inequalities can be used as constraints:

[0058] Among them, C B The set of cable cars assigned to target warehouse B; and Let x and y be the minimum and maximum coordinates of the effective working range of any cable crane k, respectively. and These are the minimum and maximum coordinates of the target compartment surface B in the direction perpendicular to the cable car track, respectively.

[0059] (7) Concrete pouring temperature constraint: Based on the requirement to control temperature stress, the concrete pouring temperature T p It must be controlled within the prescribed range:

[0060] in, The concrete pouring temperature refers to the temperature of the old concrete when new concrete is placed on top after leveling, vibration, and the completion of the paving layer pouring. and These are the minimum and maximum temperatures required by the temperature control module, respectively.

[0061] In this embodiment of the invention, parameters are set for the model in four aspects: dam section characteristics, cable crane data, process parameters, and surface data, and then injected into the arch dam pouring progress simulation model. Some parameters do not change during the simulation and are considered static physical property data; others are input into the simulation model only as initialization conditions and change with the simulation state, constituting dynamic observation state vectors. These dynamic observation state vectors are obtained through the arch dam design or sensing module. The injected model parameters are shown in Table 2. Table 2

[0062] The simulation model based on the construction progress of arch dams constructed in this invention realizes high-fidelity dynamic simulation of key aspects such as cable crane state transition, dam section construction process and joint grouting, and can integrate multi-source sensing data and static design parameters in real time.

[0063] Step 203: Collect dynamic observation state vectors and static physical attribute data, and initialize the arch dam pouring progress simulation model.

[0064] In this embodiment of the invention, the dynamic observation state vector includes, but is not limited to, the sensing data in Table 2, namely: dam section elevation array, dam section highest grouting zone elevation array, dam section intermittent period array, cable crane real-time three-dimensional coordinates, cable crane load capacity, cable crane real-time speed vector, trolley position, solar radiation coefficient, surface microclimate temperature, concrete initial setting time, vibration capacity, and leveling capacity.

[0065] In this embodiment of the invention, the static physical attribute data includes, but is not limited to, the fixed parameters in Table 2, namely: dam segment number, number of dam segments, array of starting elevations of dam segments, array of ending elevations of dam segments, elevation-area relationship function of dam segments, elevation-width relationship function of dam segments, gallery number, gallery shape and location parameters, orifice number, orifice shape and location parameters, elevator shaft number, elevator shaft shape and location parameters, grouting zone number, array of starting elevations of grouting zones, array of ending elevations of grouting zones, number of cable cranes, cable crane number, cable crane trolley movement range, cable crane horizontal control range, cable crane vertical control range, cable crane canister capacity, cable crane trolley speed, full-load descent speed, full-load descent speed, and full-load descent speed. Lifting speed, trolley running speed, material supply platform elevation, wind speed and safety distance of cable cranes on the same level, wind speed and safety distance of cable cranes on different levels, surface heat transfer coefficient, thermal conductivity, specific heat, thickness of cast-in-place layer, height of cast-in-place layer in unconstrained area, height of cast-in-place layer in constrained area, minimum height difference between adjacent dam sections, maximum height difference between adjacent dam sections, maximum cantilever of non-orifice dam section, maximum cantilever of orifice dam section, minimum interval, minimum interval at orifice, minimum interval of gallery dam block, steel lining installation time, design specified value of concrete temperature of dam blocks on both sides of joint grouting area, minimum age of concrete at the top of joint grouting area, weight of joint grouting concrete, weight of consolidation grouting, interval time of consolidation grouting.

[0066] In this embodiment of the invention, step 203 specifically includes: Step 2031: Based on construction constraints, generate a candidate set of pourable dam sections at each decision point using the arch dam pouring progress simulation model.

[0067] In this embodiment of the invention, at each decision moment, the arch dam pouring progress simulation model checks whether each dam segment meets the construction constraints based on the construction constraints, and determines the set of dam segments that meet all construction constraints as the candidate set of pourable dam segments at the current decision moment.

[0068] Step 2032: Perform mask transformation on the candidate set of castable dam sections to obtain the action mask at the current decision moment.

[0069] In this embodiment of the invention, the candidate set of castable dam sections is converted into a binary action mask to obtain the action mask at the current decision moment. The action mask is essentially calculated based on the state and a preset function of the simulation model.

[0070] As an alternative approach, assume there are N dam sections in the system (numbered from 0 to N). 1) Adding a "wait" action (numbered N) brings the total action space to N+1. The system generates a binary vector of length N+1, which is the action mask.

[0071] For each dam segment number 'a': if the dam segment is in the candidate set of castable dam segments, the value at the corresponding position in the mask vector is set to 1 (indicating it is valid); otherwise, it is set to 0 (indicating it is non-compliant).

[0072] For the "wait" action: its mask value is always 1, and it is always considered an optional legal action.

[0073] In this embodiment of the invention, when calculating the strategy, the agent must use the action mask at the current decision moment to shield all "non-compliant" actions that do not meet the constraints, and make decisions only among the legal dam section options and the "wait" action.

[0074] The simulation model based on the construction progress of arch dams constructed in this invention not only provides a highly realistic and interactive decision-making environment for reinforcement learning agents, but also dynamically generates action masks through an embedded real-time constraint computing mechanism, thereby ensuring that each decision-making step strictly complies with construction specifications and safety requirements, laying a reliable foundation for subsequent intelligent optimization scheduling.

[0075] Step 204: Construct an intelligent scheduling decision model.

[0076] In this embodiment of the invention, step 204 specifically includes: Step 2041: Collect expert trajectory dataset.

[0077] In this embodiment of the invention, a rule-based heuristic strategy capable of simulating the decision-making logic of domain experts is constructed; the heuristic strategy is run in multiple simulation environments with different random settings, and its complete decision-making process is recorded, including the "dynamic observation state vector and static physical attribute data" at each decision point, the final "cable crane pouring scheduling decision result" and the "action mask" at that time, thereby forming an expert trajectory dataset.

[0078] Specifically, a weighted approach is used to normalize multiple features of the dam section, such as elevation, interval, and cantilever height, and then assigns fixed weights to calculate a comprehensive score. The dam section with the highest score is selected for pouring, forming a heuristic strategy. This heuristic strategy is then run in a high-fidelity arch dam pouring progress simulation model (DES+ABM). The entire dam pouring schedule is completed from start to finish in multiple different simulation scenarios (with different random seeds). The "dynamic observation state vector and static physical attribute data," "action mask," and "cable crane pouring schedule decision results" at each decision moment are recorded, thus forming an expert trajectory dataset.

[0079] In this embodiment of the invention, the expert trajectory dataset is used for behavioral cloning pre-training of the reinforcement learning policy network. The aim is to allow the neural network model to initially "imitate" the decision-making pattern of the heuristic policy, thereby obtaining a better initial policy, significantly reducing the time spent on blind exploration in subsequent reinforcement learning, and improving training efficiency and stability.

[0080] Step 2042: Based on the expert trajectory dataset, perform supervised learning pre-training on the policy network to construct a behavior clone pre-trained model.

[0081] In this embodiment of the invention, the collected expert trajectory dataset is used as a sample for supervised learning to pre-train the policy network. The goal of the training is to enable the network to output actions that best mimic the expert's choices when given the same state and action mask.

[0082] Furthermore, a policy network pre-trained through behavioral cloning is used as the initial decision-making model. When actual construction or simulation reaches a critical node requiring a pouring decision, the system initiates the decision-making process. At this time, the arch dam pouring progress simulation model is precisely initialized based on real-time perceived engineering conditions (such as the current elevation of each dam section, cable crane position and status, etc.), creating a virtual decision-making environment highly consistent with the current working conditions for the agent. The reinforcement learning agent then takes over this simulation environment and begins executing the decision-making task.

[0083] Step 2043: By integrating the near-end policy optimization (PPO) algorithm with action mask and the preset hybrid reward function, the behavior clone pre-trained model is optimized based on the collected simulation interaction data to generate an intelligent scheduling decision model.

[0084] In this embodiment of the invention, the simulation interaction data is collected and generated by the arch dam pouring progress simulation model, including the observation data and action space and constraints of the reinforcement learning agent. The observation data of the reinforcement learning agent is a multi-dimensional vector, provided by the progress simulation model at each decision moment. This vector contains key dynamic information affecting the decision, specifically including but not limited to: the current pouring elevation of all dam sections, the interval between the last pouring and the current pouring time for each dam section, the current global cantilever height, the current time of the simulation, and the global maximum elevation difference.

[0085] The action space is defined as a discrete set, where each action corresponds to selecting a specific dam segment to be poured, and additionally includes a special "waiting" action. At any decision point, the progress simulation model calculates all available dam segments to be poured based on the actual construction constraints and provides them to the agent in the form of a binary action mask vector to filter out illegal actions that cannot be selected.

[0086] In this embodiment of the invention, the PPO algorithm with fused action masks is combined with a behavior cloning pre-trained model. The core of the algorithm lies in integrating the action mask into the computation process, both during the pre-training phase and the policy update phase of reinforcement learning. The pseudocode is as follows: Input: Initial policy parameters θ Initial value function parameters .

[0087] Output: Optimized policy parameters θ.

[0088] Initialize θ←θ , ← .

[0089] For iterations 1, 2, ..., execute: B ← / / Initialize the trajectory buffer; / / --- Rolling Phase --- For t = 1 to T, execute: Observation state s t and action mask m t .

[0090] a t ~ π θold (·|s t , m / / Sample actions from the masked strategy; Execute a t Observation reward r t and the next state st+1 .

[0091] B ← B ∪ {(s t , a t , m t , r ,s t+1 )}.

[0092] End the loop; / / --- Update Phase --- Calculate the figure of merit estimates for all transitions in buffer B. and target value .

[0093] For epochs 1 to K, execute: For each mini-batch, by optimizing the objective function L(θ, To update θ and : ; End the loop.

[0094] End the loop.

[0095] In this embodiment of the invention, to guide the agent to learn efficient scheduling strategies, the reward function is designed as a hybrid structure, combining sparse final goal rewards with dense intermediate process shaping rewards. The hybrid reward function includes sparse goal rewards and dense plasticity rewards; the sparse goal rewards are rewards or penalties based on the actual project duration and the baseline project duration; the dense plasticity rewards are rewards or penalties for each scheduling decision.

[0096] Specifically, the sparse target reward is a final reward or penalty awarded based on a comparison between the final completion date and a preset baseline date at the end of the entire pouring task. As an option, a reward is given for actual completion dates earlier than the preset baseline date, and a penalty is given for actual completion dates later than the preset baseline date. The more the actual completion date is earlier than the preset baseline date, the larger the reward; the more the actual completion date is later than the preset baseline date, the larger the penalty.

[0097] Specifically, intensive plasticity rewards are small, immediate rewards or penalties given at each decision step in the simulation, based on the selected action. These rewards and penalties are designed to encourage good construction behavior patterns. For example, positive rewards are given for selecting dam sections that have never been poured or the lowest dam section in the poured surface; negative penalties are given for actions that result in excessive differences in the pouring surface or cause the cantilever height to approach the risk threshold; and phased target rewards are set for completing key construction nodes (such as through-holes).

[0098] In this embodiment of the invention, a behavior clone pre-trained model is used as the starting point. The model is then used to collect simulation interaction data online in a simulation environment. The model is iteratively updated by continuously optimizing the strategy through the PPO algorithm that integrates action masks and a hybrid reward function until the model converges, thereby constructing an intelligent scheduling decision model.

[0099] In this invention, during interaction with the simulation environment, the agent continuously explores and learns based on the PPO algorithm, constantly optimizing its decision-making strategy. The goal at this stage is to start by imitating expert behavior and, through reward signals from environmental feedback, autonomously discover a better scheduling pattern that surpasses expert experience. Once the policy network's performance converges and reaches a stable state, the optimal decision model is saved. To ensure the model's reliability and generalization ability, its robustness will be rigorously verified in multiple simulation scenarios with different random seeds, ensuring stable and efficient decision support under varying conditions.

[0100] Step 205: Extract features from the dynamic observation state vector to obtain the state context feature vector.

[0101] In this embodiment of the invention, in order to handle complex construction decision-making problems, the strategy network architecture adopted by the intelligent scheduling decision model integrates dynamic state information and static action attribute information.

[0102] Specifically, one branch of the policy network receives the dynamic observation state vector from the simulation environment, extracts the features of the dynamic observation state vector through a feature extractor, and obtains a state context feature vector that can characterize the current global operating condition.

[0103] Step 206: Map the static physical attribute data to action embedding vectors.

[0104] In this embodiment of the invention, another branch of the policy network processes static physical attribute data (e.g., the parity of a dam segment, whether it is a span dam segment, etc.). This branch maps these discrete, static attribute information into low-dimensional, dense action embedding vectors.

[0105] Step 207: Perform a fusion decision on the state context feature vector and action embedding vector to generate the cable crane pouring scheduling decision result.

[0106] Specifically, the state context feature vector is fused with the embedding vector of each action, and the fused information is scored by combining the action mask at the current decision moment. This score ultimately determines the probability of selecting each legal action in the current state, generating each candidate action and its corresponding legal probability. The candidate action with the highest legal probability is selected, and the cable crane pouring scheduling decision is generated based on the selected candidate action. The cable crane pouring scheduling decision includes the pouring sequence and corresponding pouring time for each dam section.

[0107] This invention starts with a policy network pre-trained using behavioral clones. Through continuous interaction with a high-fidelity progress simulation environment, it performs online optimization, ultimately forming a stable and efficient intelligent scheduling decision model. Its robustness is verified to guide actual construction. By running the intelligent scheduling decision model and conducting a complete simulation, a detailed cable crane pouring scheduling plan covering a future period can be generated. This plan not only clarifies the pouring sequence of each dam section but also includes precise timing selection. Finally, this series of optimized scheduling instructions can be directly output to the automated control system to guide on-site equipment to operate autonomously, or serve as high-level decision-making suggestions, providing a scientific basis for the command and decision-making of on-site management personnel.

[0108] The following specific example illustrates the scheduling and decision-making process for a concrete arch dam pouring project: Taking a large hydropower station in Southwest China as an example, the normal water level of this hydropower station is 825m, the flood control limit water level is 785m, the dead water level is 765m, and the flood control capacity is 7.5 billion cubic meters. 3 The reservoir capacity reached 10.436 billion cubic meters. 3 The reservoir has annual regulation capabilities; the power station has an installed capacity of 14,000 MW and an average annual power generation of 64.095 billion kWh, making it a giant hydropower project with a capacity of 10 million kW.

[0109] Step 1: Simulation begins at 00:00 on May 1st of the third year of arch dam construction, when 832 sections have been poured. Some model parameters are shown in Table 3 below: Table 3

[0110] The dam was equipped with 7 horizontal cable cranes, 3 for the high line and 4 for the low line, arranged in a double layer. The simulation ended when all sections were poured.

[0111] Step 2: Implement the masking function using the Maskable PPO algorithm from the open-source sb3-contrib library. Some reward function settings are shown in Table 4: Table 4

[0112] Table 5 shows some of the hyperparameters used by the Maskable PPO algorithm: Table 5

[0113] The pre-trained expert rules for this reinforcement learning method use a weighted approach. This involves normalizing features such as dam section elevation, interval, and cantilever height, then applying weights to calculate scores. The index of the dam section with the highest score under this weighted approach is used as the pouring target. Ten random seed environments with equal weights are used to record expert trajectories for behavior cloning.

[0114] Step 3: Start training and, after obtaining stable training results, test its effectiveness using different random seeds. Figure 4 A graph of the reward function during training is provided as an embodiment of the present invention, such as... Figure 4 As shown, the horizontal axis represents timesteps, ranging from 0 to 350,000 with an interval of 50,000; the vertical axis represents the reward function value, ranging from 124 to 146 with an interval of 2. Initially, there is a rapid decline followed by a rebound; then a rapid growth period occurs, with the reward quickly rising to a peak, before stabilizing and converging. This model is then saved as an intelligent scheduling decision model. A near-end policy optimization algorithm incorporating action masks and a pre-defined hybrid reward function are used for model optimization. This allows the model to quickly learn effective policies and achieve high performance, demonstrating excellent sample efficiency and convergence speed, thus verifying the algorithm's efficiency and robustness in this task.

[0115] Save it as an intelligent scheduling decision model, then execute the simulation. A portion of the scheduling actions are shown below: ...; Cable crane No. 2 began entering section 16 of the dam at 23:35:43 CST on Sep 18, 2019. Cable crane No. 6 began entering section 16 of the dam at 23:47:30 CST on Sep 18, 2019. Cable crane No. 5 began entering section 16 of the dam at 23:48:40 CST on Sep 18, 2019. Cable crane No. 3 began entering section 16 of the dam at 23:48:51 CST on Sep 18, 2019. Cable crane No. 2 began entering section 11 at 15:50:34 CST on September 19, 2019. Cable crane No. 6 began entering section 11 at 15:51:50 CST on September 19, 2019. Cable crane No. 5 began entering section 11 at 15:52:02 CST on September 19, 2019. Cable crane No. 3 began entering section 19 at 20:54:43 CST on September 19, 2019. Cable crane No. 7 began entering section 19 of dam at 20:55:48 CST on September 19, 2019. Cable crane No. 1 began entering section 19 of dam at 21:00:50 CST on September 19, 2019. Cable crane No. 4 began entering section 19 at 21:03:02 CST on September 19, 2019. Cable crane No. 2 began entering section 12 of the dam at 07:12:30 CST on September 20, 2019. Cable crane No. 6 began entering section 12 of the dam at 07:14:29 CST on Sep 20, 2019. Cable crane No. 5 began entering section 12 of the dam at 07:17:07 CST on September 20, 2019. ...

[0116] This invention enables intelligent pouring decisions during the intelligent construction of arch dams, providing control commands or optimization schemes for the pouring process, thereby reducing construction time and engineering costs and bringing considerable economic benefits.

[0117] It is worth noting that the acquisition, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. The user information in the embodiments of this application was obtained through legal and compliant means, and the acquisition, storage, use, and processing of user information have been authorized and agreed upon by the client.

[0118] It is worth noting that the information collected in this application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0119] It is worth noting that the technical solution provided in this application provides users with a corresponding operation entry point, allowing users to choose to agree to or reject the automated decision-making result; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0120] The technical solution of the scheduling decision-making method for concrete arch dam pouring projects provided in this invention embodiment is based on an intelligent agent model and discrete event simulation. A simulation model of the arch dam pouring progress is constructed according to the construction status data in the concrete arch dam pouring project. Dynamic observation state vectors and static physical attribute data are collected and initialized. Through the arch dam pouring progress simulation model and the pre-constructed intelligent scheduling decision-making model, scheduling is performed according to the dynamic observation state vectors and static physical attribute data, generating the cable crane pouring scheduling decision result. The intelligent scheduling decision-making model is constructed based on a two-stage intelligent agent training method that integrates expert experience and reinforcement learning. This method enables the model to learn autonomously and optimize long-term construction benefits, improving model robustness and scheduling efficiency. Through offline training and online application of the model, a closed-loop intelligent decision-making process from perception, simulation, decision-making to control is formed, providing reliable support for the automation and intelligent management of arch dam pouring construction.

[0121] Figure 5 This is a schematic diagram of a scheduling decision-making device for a concrete arch dam pouring project provided in an embodiment of the present invention. This device is used to execute the aforementioned scheduling decision-making method for a concrete arch dam pouring project. Figure 5 As shown, the device includes: an arch dam pouring progress simulation model construction unit 11, a status data acquisition unit 12, and a scheduling decision unit 13.

[0122] The arch dam pouring progress simulation model building unit 11 is used for agent-based model and discrete event simulation. Based on the construction status data in the concrete arch dam pouring project, it builds an arch dam pouring progress simulation model.

[0123] The state data acquisition unit 12 is used to acquire dynamic observation state vectors and static physical attribute data, and to initialize the arch dam pouring progress simulation model.

[0124] The scheduling decision unit 13 is used to schedule according to the dynamic observation state vector and static physical attribute data through the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model, and generate the cable crane pouring scheduling decision result. The intelligent scheduling decision model is built based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

[0125] In this embodiment of the invention, the construction status data includes dam segment characteristics, cable crane data, process parameters, and surface data. The arch dam pouring progress simulation model construction unit 11 is specifically used to perform micro-modeling based on the cable crane data using an intelligent agent model, and to generate a simulation model framework by performing discrete event simulation, global advancement based on the cable crane single-cycle intelligent agent model and construction events. The simulation model framework is then injected with dam segment characteristics, cable crane data, process parameters, surface data, and construction constraints to construct the arch dam pouring progress simulation model.

[0126] In this embodiment of the invention, the device further includes: a collection unit 14, a pre-training unit 15, and a model optimization unit 16.

[0127] Collection unit 14 is used to collect expert trajectory datasets.

[0128] The pre-training unit 15 is used to perform supervised learning pre-training on the policy network based on the expert trajectory dataset, and to build a behavior clone pre-trained model.

[0129] The model optimization unit 16 is used to optimize the behavior clone pre-trained model based on the collected simulation interaction data by using a near-end policy optimization algorithm that integrates action masking and a preset hybrid reward function to generate an intelligent scheduling decision model.

[0130] In this embodiment of the invention, the scheduling decision unit 13 is used to extract features from the dynamic observation state vector to obtain the state context feature vector; map the static physical attribute data into the action embedding vector; and perform a fusion decision on the state context feature vector and the action embedding vector to generate the cable crane pouring scheduling decision result.

[0131] In this embodiment of the invention, based on the model of intelligent agents and discrete event simulation, a simulation model of the arch dam pouring progress is constructed according to the construction status data in the concrete arch dam pouring project. Dynamic observation state vectors and static physical attribute data are collected and the arch dam pouring progress simulation model is initialized. Through the arch dam pouring progress simulation model and the pre-constructed intelligent scheduling decision model, scheduling is performed according to the dynamic observation state vectors and static physical attribute data to generate the cable crane pouring scheduling decision result. The intelligent scheduling decision model is constructed based on a two-stage intelligent agent training method that integrates behavior cloning and reinforcement learning with reward reshaping. It adopts a two-stage training method that integrates expert experience and reinforcement learning, enabling the model to learn autonomously and optimize long-term construction benefits, improve the model's robustness and scheduling efficiency. Through offline training and online application of the model, a closed-loop intelligent decision-making process from perception, simulation, decision-making to control is formed, providing reliable support for the automation and intelligent management of arch dam pouring construction.

[0132] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer device, specifically, a computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0133] This invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, they implement the steps of the above-described scheduling decision method for concrete arch dam pouring project. For a detailed description, please refer to the above-described embodiment of the scheduling decision method for concrete arch dam pouring project.

[0134] The following is for reference. Figure 6 It shows a schematic diagram of the structure of a computer device 600 suitable for implementing the embodiments of this application.

[0135] like Figure 6 As shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate tasks and processes based on programs stored in read-only memory (ROM) 602 or programs loaded from storage section 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the computer device 600. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0136] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal feedback (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed in storage section 608 as needed.

[0137] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611.

[0138] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0139] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0143] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0144] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0145] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0146] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0147] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0148] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0149] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A scheduling and decision-making method for concrete arch dam pouring projects, characterized in that, The method includes: Based on agent-based model and discrete event simulation, a simulation model of the arch dam pouring progress is constructed according to the construction status data in the concrete arch dam pouring project. Collect dynamic observation state vectors and static physical attribute data, and initialize the arch dam pouring progress simulation model; The cable crane pouring schedule decision results are generated by using the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model, based on the dynamic observation state vector and static physical attribute data. The intelligent scheduling decision model is constructed based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

2. The scheduling and decision-making method for concrete arch dam pouring projects according to claim 1, characterized in that, The construction status data includes dam section characteristics, cable crane data, process parameters, and deck surface data; The agent-based model and discrete event simulation, based on construction status data in the concrete arch dam pouring project, constructs a simulation model of the arch dam pouring progress, including: Based on the intelligent agent model, micro-modeling is performed on the cable crane data, and global advancement is carried out through discrete event simulation based on the cable crane single-cycle intelligent agent model and construction events to generate a simulation model framework. The dam section features, cable crane data, process parameters, surface data, and construction constraints are injected into the simulation model framework to construct the arch dam pouring progress simulation model.

3. The scheduling and decision-making method for concrete arch dam pouring projects according to claim 2, characterized in that, The construction constraints include time constraints, elevation differences between adjacent dam sections, cantilever elevation differences, maximum elevation differences across the entire dam, elevation differences during the first pour, cable crane resource constraints, and concrete pouring temperature constraints.

4. The scheduling and decision-making method for concrete arch dam pouring projects according to claim 1, characterized in that, The method further includes: Collect expert trajectory datasets; Based on the expert trajectory dataset, supervised learning pre-training is performed on the policy network to construct a behavior clone pre-trained model; By integrating a near-end policy optimization algorithm that incorporates action masking with a pre-defined hybrid reward function, and based on collected simulation interaction data, the behavior clone pre-trained model is optimized to generate the intelligent scheduling decision model.

5. The scheduling and decision-making method for concrete arch dam pouring projects according to claim 4, characterized in that, The hybrid reward function includes sparse target reward and dense plasticity reward; The sparse target reward is a reward or penalty based on the actual project duration and the baseline project duration; The dense plasticity reward is a reward or penalty for each scheduling decision.

6. The scheduling and decision-making method for concrete arch dam pouring projects according to claim 1, characterized in that, The process of scheduling based on the dynamic observation state vector and static physical attribute data, using the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model, generates cable crane pouring scheduling decision results, including: Feature extraction is performed on the dynamic observation state vector to obtain the state context feature vector; Map the static physical attribute data into action embedding vectors; The state context feature vector and action embedding vector are fused together to generate the cable crane pouring scheduling decision result.

7. A scheduling and decision-making device for concrete arch dam pouring projects, characterized in that, The device includes: The arch dam pouring progress simulation model building unit is used for agent-based model and discrete event simulation. It constructs an arch dam pouring progress simulation model based on the construction status data in the concrete arch dam pouring project. The state data acquisition unit is used to acquire dynamic observation state vectors and static physical attribute data, and to initialize the arch dam pouring progress simulation model. The scheduling decision unit is used to schedule the cable crane pouring based on the dynamic observation state vector and static physical attribute data through the arch dam pouring progress simulation model and the pre-built intelligent scheduling decision model. The intelligent scheduling decision model is constructed based on a two-stage agent training based on behavior cloning and reinforcement learning with reward reshaping.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the scheduling decision-making method for concrete arch dam pouring projects as described in any one of claims 1 to 6.

9. A computer device comprising a memory and a processor, the memory for storing information including program instructions, and the processor for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by the processor, they implement the scheduling decision-making method for concrete arch dam pouring projects as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the scheduling decision-making method for the concrete arch dam pouring project as described in any one of claims 1 to 6.