Earthwork Construction Dynamic Scheduling Method, Equipment and Storage Medium Based on Deep Reinforcement Learning

Through the dynamic scheduling method of earthwork construction based on deep reinforcement learning, the problems of local optimality and poor adaptability of earthwork construction scheduling in complex environments in the existing technology are solved, efficient integrated scheduling and dynamic environmental adaptability are achieved, and overall efficiency and scheduling quality are improved.

CN119624047BActive Publication Date: 2025-05-27ZHEJIANG HUADONG ENG DIGITAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510148941.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing earthwork construction scheduling methods are difficult to achieve efficient integrated scheduling in complex environments, and lack the flexibility to deal with sudden failures and dynamic environmental changes, resulting in the possibility of local optimum rather than global optimum in the scheduling scheme.

Method used

The dynamic scheduling method of earthwork construction based on deep reinforcement learning is adopted. By obtaining construction environment, process and equipment information, the integrated scheduling of construction equipment is expressed as Markov dynamic decision-making process, multi-objective reward function is designed, the agent is trained to generate an initial scheduling plan, and rescheduling is performed under hypothetical disturbance events.

Benefits of technology

It realizes efficient integrated scheduling in complex environments, improves the overall efficiency and adaptability of earthwork construction scheduling, can achieve a better balance between multiple goals, and improves the robustness and adaptability of scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624047B_ABST
    Figure CN119624047B_ABST
Patent Text Reader

Abstract

The present invention provides an earthwork construction dynamic scheduling method, device and storage medium based on deep reinforcement learning. The method includes: obtaining construction information, determining constraint conditions and the integrated scheduling objective of construction equipment, and setting an objective function; formulating the integrated scheduling problem of construction equipment as a Markov dynamic decision-making process to generate a first intelligent agent; designing a reward function based on the NSGA-II and DDQN algorithms to train the first intelligent agent; generating an initial scheduling plan and performing virtual scheduling based on this plan; supplementing constraint conditions based on assumed perturbation events, and generating a second intelligent agent through the Markov dynamic decision-making process; differentiating the operations in the initial scheduling plan, and designing encoding and decoding methods for the NSGA-II algorithm according to the differentiation results; designing a reward function based on the NSGA-II algorithm and the DDQN algorithm after encoding and decoding to train the second intelligent agent; deploying the intelligent agent after stable training in the actual construction environment to generate dynamic scheduling decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of earthwork construction scheduling, and in particular to a method, device and storage medium for dynamic scheduling of earthwork construction based on deep reinforcement learning. Background Art

[0002] Faced with a complex and ever-changing construction environment, how to rationally schedule multiple resources such as earthwork and construction machinery to maximize efficiency and reduce energy consumption has become a key issue that needs to be addressed urgently.

[0003] There are many similarities between earthwork construction machinery scheduling and flexible workshop scheduling, both involving complex tasks and resource arrangements. Multi-resource integrated scheduling is even more complex and difficult, requiring comprehensive consideration of multiple objectives such as construction efficiency, energy consumption control, and resource utilization. Current research often uses linear programming, heuristic rules, and meta-heuristic algorithms to solve. With the increase in construction site workload and the expansion of construction scale, many existing intelligent scheduling algorithms often focus on single-objective optimization, such as maximizing construction efficiency or minimizing energy consumption, and it is difficult to take into account multiple objectives at the same time, resulting in local optimal rather than global optimal scheduling solutions in practical applications.

[0004] In addition, in actual construction environments, we often face various emergencies such as sudden failures of construction equipment. The current scheduling method lacks the ability to flexibly adjust the scheduling plan when dealing with these complex and uncertain situations, and needs to further improve its robustness and adaptability.

[0005] Therefore, new methods are needed to improve the problems of poor adaptability and low efficiency of existing earthwork construction scheduling methods. Summary of the invention

[0006] The present invention focuses on solving the problems of local optimality and poor adaptability to dynamic environments in earthwork construction scheduling in the prior art. The purpose is to construct a dynamic scheduling method for earthwork construction based on deep reinforcement learning to achieve efficient integrated scheduling in complex environments, and to deeply consider the impact of disturbance events on the scheduling system, thereby effectively improving the overall efficiency and adaptability of earthwork construction scheduling.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention discloses a dynamic scheduling method for earthwork construction based on deep reinforcement learning, comprising the following steps:

[0009] Obtain information about the construction environment, construction process, and construction equipment, determine constraints and integrated scheduling objectives for construction equipment based on the information obtained, and set an objective function;

[0010] The integrated scheduling problem of construction equipment is formulated as a Markov dynamic decision process, and a first intelligent agent is generated for solving the integrated scheduling problem of construction equipment;

[0011] Designing a first multi-objective reward function based on the NSGA-Ⅱ algorithm and the DDQN algorithm, and training the first agent according to the first multi-objective reward function;

[0012] Generate an initial scheduling plan through the trained first agent, and perform virtual scheduling based on the initial scheduling plan;

[0013] Supplementing the constraint conditions based on the assumed disturbance event, generating a second intelligent agent for solving the integrated scheduling problem of construction equipment under disturbance conditions through a Markov dynamic decision process;

[0014] Differentiate the jobs in the initial scheduling plan, design the encoding and decoding methods of the NSGA-Ⅱ algorithm based on the differentiation results, retain the unaffected parts of the initial scheduling plan, and reschedule the affected parts;

[0015] Design a second multi-objective reward function based on the NSGA-Ⅱ algorithm and the DDQN algorithm after the design of the codec, and train the second agent according to the second multi-objective reward function;

[0016] The first and second agents, after stable training, are deployed in the actual construction environment to generate dynamic scheduling decisions for construction equipment under real fault events.

[0017] Furthermore, the construction environment information includes construction zones and layer information to be constructed, the construction process information includes operation phase information of each construction zone, and the construction equipment information includes information of construction equipment that can be used for scheduling operations.

[0018] Furthermore, the constraints include:

[0019] During different operation phases, construction equipment needs to go to the construction zone, and the location of the construction zone is known;

[0020] The construction equipment corresponding to each operation stage can be used to construct any construction zone, but the construction time is different and has been determined;

[0021] Once a construction device starts working in a construction zone, the work will not be interrupted;

[0022] The same construction zone can only carry out the subsequent operation phase after the current operation phase is completed;

[0023] After the construction equipment finishes the current construction work, it will immediately start from the current position to the construction area where the next task is located to carry out the work.

[0024] Furthermore, the integrated scheduling objectives of the construction equipment include:

[0025] Minimize the operation time of construction equipment, the corresponding objective function is ; Minimize the total idle time of construction equipment, the corresponding objective function is ; Minimize the total moving distance, the corresponding objective function is ;in, k The equipment is numbered in the order of stage 1 to 3 according to the number of construction equipment available in each stage. For construction equipment k The construction completion time, m is the total number of construction equipment; and All are array types. For construction equipment k Construction time records, For construction equipment k The moving time record.

[0026] Furthermore, the integrated scheduling problem of construction equipment is expressed as a Markov dynamic decision process, and the first intelligent agent generated for solving the integrated scheduling problem of construction equipment includes:

[0027] The state space is designed based on the two objects of construction partition and construction equipment, including several construction partition state information and construction equipment state information. Each state information involves the current state s and the next state ;

[0028] An action space is designed for the intelligent agent, wherein the action space includes a plurality of composite scheduling rules formed by arranging and combining a plurality of scheduling rules for construction zoning and stratification and construction equipment allocation.

[0029] Furthermore, the multi-objective reward function designed based on the NSGA-Ⅱ algorithm and the DDQN algorithm includes:

[0030] Set multiple levels of reward function values;

[0031] The minimum distance solution of the NSGA-II algorithm and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three objective function values ​​corresponding to each step of scheduling;

[0032] Compare the objective function values ​​of the two algorithms, select the smaller value and compare it with the objective function value in the current state as the input of the reward function, and the reward function outputs the reward function value according to the comparison result;

[0033] Add the three reward function values ​​​​to get the final reward value.

[0034] Furthermore, the constraints supplemented based on the assumed disturbance events include:

[0035] Construction equipment failures occur randomly and can occur on any piece of construction equipment at any point in time;

[0036] Once a construction equipment fails, the operation on the construction equipment will be suspended, while the operation of other construction equipment will not be affected;

[0037] When the operation on a construction layer is interrupted due to a construction equipment failure, the original construction equipment will continue to operate on the construction layer after the failed construction equipment is repaired;

[0038] The repair time of faulty construction equipment can be estimated in advance.

[0039] Furthermore, the jobs in the initial scheduling scheme are differentiated into:

[0040] The jobs in the initial scheduling plan are classified into three categories: G1, G2, and G3, where:

[0041] G1 includes the jobs that have been completed before the failure event occurs, and the jobs that have been started on normal construction equipment. These jobs will maintain their original scheduling status during rescheduling;

[0042] G2 includes the operations that were being carried out on the faulty construction equipment when the fault event occurred. They will continue after the fault is repaired, so the status will not be adjusted during rescheduling;

[0043] G3 includes the operations that have not entered the next construction phase at the time of the failure event. These operations need to rearrange the batch sequence and construction equipment allocation during rescheduling;

[0044] The mobile tasks of construction equipment are divided into three groups: A1, A2, and A3.

[0045] A1 includes tasks that have been completed before the failure event occurs or tasks that are being moved by construction equipment that is operating normally, and their status remains unchanged during rescheduling;

[0046] A2 includes tasks where the faulty equipment is being moved at the time of the fault event. For such tasks, the time required for other normal equipment to reach the next target point at the time of the fault is compared with the repair time of the faulty equipment. If the repair time is shorter, the task will continue to be completed by the faulty equipment; if the arrival time of other equipment is shorter, the equipment that arrives the fastest will be selected to take over the movement. At this time, the construction equipment allocation needs to be rescheduled;

[0047] A3 includes tasks that have not started to move to the next target at the time of the failure event. Their hierarchical sequence and equipment allocation need to be adjusted during rescheduling.

[0048] Furthermore, the encoding and decoding methods of the NSGA-Ⅱ algorithm are designed according to the differentiation results, including:

[0049] During the encoding process, the start time of all operation sequences in the initial scheduling plan is fully traversed, and the construction phase, construction partition number and layer number corresponding to the operation sequences classified as G1 and G2 are recorded to construct a fixed operation sequence, and the construction equipment number used in this operation sequence is recorded, so that the construction equipment sequence remains unchanged during the rescheduling process;

[0050] According to the impact of the fault on the construction equipment, the earliest available time of the construction equipment and the earliest available time for the construction equipment to move are initialized, and the operations and construction equipment affected by the fault are rescheduled according to the differentiation results to obtain the adjusted operation sequence and construction equipment sequence, and a decoding operation is performed.

[0051] Furthermore, based on the NSGA-Ⅱ algorithm and DDQN algorithm after the design of the codec, the second multi-objective reward function is designed, including:

[0052] Set multiple levels of reward function values;

[0053] The minimum distance solution of the NSGA-II algorithm after the design of the codec and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three objective function values ​​corresponding to each step of scheduling;

[0054] Compare the objective function values ​​of the two algorithms, select the smaller value and compare it with the objective function value in the current state as the input of the reward function, and the reward function outputs the reward function value according to the comparison result;

[0055] Add the three reward function values ​​​​to get the final reward value.

[0056] Further, training the second agent according to the second multi-objective reward function includes:

[0057] Multiple examples are randomly generated, each of which includes a different number of construction zones and layers, as well as the number of construction equipment at different construction stages;

[0058] Initialize the parameters of the Q network in the DDQN algorithm, initialize the values ​​corresponding to all states and actions, and clear the experience replay set;

[0059] Extract state features based on the current construction environment and input the features into the neural network to calculate the value of all possible actions in the current state;

[0060] Use the evaluation network to select the action with the greatest value, and determine the value of the action in the target network based on the selected action to update the Q network;

[0061] Select the action with the highest value from the target network again, execute the selected action, observe the reward and next state fed back by the environment, and update the reward and state;

[0062] Repeat the above steps for each construction partition layer until all construction partition layers are completed, completing the training of the agent.

[0063] Repeat the above training process until the number of training reaches the set number of iterations to complete the training of the agent.

[0064] The second aspect of the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores computer instructions, and the processor executes the earthwork construction dynamic scheduling method as described in the first aspect above by executing the computer instructions.

[0065] The third aspect of the present invention further discloses a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the earthwork construction dynamic scheduling method as described in the first aspect above is executed.

[0066] The beneficial effects of the present invention are as follows:

[0067] 1) Excellent comprehensive distribution performance: When dealing with multi-objective problems, conventional algorithms in the prior art can only produce individual solutions that perform well in a single objective, while the method of the present invention exhibits more excellent characteristics in terms of comprehensive distribution performance, and can achieve a better balance between multiple objectives, avoiding the one-sided pursuit of optimization of a single objective while ignoring overall performance.

[0068] 2) Reward signal optimization guidance: The method of the present invention innovatively introduces a comparison mechanism with the local optimal solution in the reward function. In this way, more valuable reward signals can be accurately obtained, thereby providing a clear and effective guidance for the algorithm to optimize towards a more ideal target, accelerating the optimization process and improving the optimization quality.

[0069] 3) Significant anti-disturbance effect advantage: Compared with the existing technology, the method of the present invention has a strong ability to deal with random disturbance events. In complex and changeable practical application scenarios, it can ensure the stability and reliability of the algorithm and maintain good optimization effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 The figure is a schematic flow chart of a method for dynamic scheduling of earthwork construction according to an embodiment of the present invention.

[0071] Figure 2 It is a schematic diagram of the earthwork construction filling area process shown in an embodiment of the present invention.

[0072] Figure 3 It is a schematic diagram of the initial scheduling process shown in an embodiment of the present invention.

[0073] Figure 4 The figure is a schematic diagram of a dynamic scheduling process shown in an embodiment of the present invention.

[0074] Figure 5 Schematic diagram of dynamic coding of NSGA-Ⅱ chromosome shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0076] See also Figure 1 The embodiment of the present invention discloses a dynamic scheduling method for earthwork construction based on deep reinforcement learning, comprising the following steps:

[0077] S1. Obtain the construction environment, construction process and construction equipment information, determine the constraints and integrated scheduling objectives of the construction equipment based on the acquired information, and set the objective function.

[0078] As a preferred implementation scheme, in this embodiment, the construction environment information includes construction zones and layer information to be constructed, the construction process information includes operation phase information of each construction zone, and the construction equipment information includes information of construction equipment that can be used to schedule operations.

[0079] In one illustrated example, a construction site to be constructed is defined as n There are 3 construction zones with different engineering quantities that need to be constructed, and each construction zone is constructed in layers. The construction process of the construction zone is divided into three stages, and the stage numbers are named j , the types of operations in each stage are different, so the corresponding types of construction equipment are also different. Figure 2 As shown in the figure: First, in the paving stage, the dump truck transports the earth to the construction area and dumps it, then in the paving stage, the bulldozer evenly spreads the earth in the construction area, and finally in the rolling stage, the roller compacts the earth. For the convenience of subsequent explanation, it is assumed that each stage j There are Construction equipment can be used for scheduling operations, that is, in Phase 1 (material laying phase), Dump truck, stage 2 (paving stage) has bulldozers, stage 3 (paving stage) Road roller.

[0080] As a preferred implementation scheme, in this embodiment, the constraints summarized for the construction environment include:

[0081] ① Construction equipment needs to go to the construction zone in different operation stages, and the location of the construction zone is known;

[0082] ② The construction equipment corresponding to each operation stage can be used to construct any construction zone, but the construction time is different and has been determined;

[0083] ③ Once a construction device starts working in a construction zone, it will not be interrupted;

[0084] ④ The construction of the subsequent operation phase can only be carried out after the current operation phase is completed in the same construction zone;

[0085] ⑤ After the construction equipment completes the current construction work, it will immediately start from the current position to the construction area where the next task is located to carry out the work.

[0086] It should be noted that the above constraints are only some examples, and other constraints can be set according to different construction environments. Therefore, the above constraints do not constitute a limitation on the implementation scheme of the present invention.

[0087] In the scheduling problem involved in the present invention, by selecting construction equipment for different construction stages of each construction zone, scheduling optimization is performed according to the construction sequence and equipment allocation. In this process, the goals of construction efficiency, energy consumption control and resource utilization may conflict with each other: shortening the operation time of construction equipment can directly improve construction efficiency, but blindly pursuing this goal may lead to frequent mobilization of construction equipment, thereby increasing the total moving distance and energy consumption, affecting the overall benefit. Reducing the total idle time of construction equipment means maximizing the utilization rate of construction equipment, which can avoid resource waste, but this may require more sophisticated scheduling arrangements, and the balance with other goals will be broken if it is not careful. For example, if the idle time is compressed too much to make the construction equipment too compact, it may cause unreasonable planning of the mobile path of the construction equipment, resulting in a significant increase in the total moving distance. Based on this, the present invention aims to comprehensively consider the inherent connection and mutual constraint relationship between the various goals through an innovative scheduling method, accurately weigh the priority and degree of realization of each goal under different construction scenarios and requirements, and build a scientific and reasonable multi-objective optimization scheduling strategy to achieve the balance of construction equipment in the three dimensions of operation time, idle time and moving distance, thereby improving the overall economy, efficiency and sustainability of earthwork construction.

[0088] As a preferred implementation scheme, in this embodiment, the integrated scheduling objectives of the construction equipment set include:

[0089] ① Minimize the operation time of construction equipment. The corresponding objective function is: ; ② Minimize the total idle time of construction equipment. The corresponding objective function is ; ③ Minimize the total moving distance, the corresponding objective function is ;

[0090] in, k Number the construction equipment. The equipment can be numbered in the order of stage 1 to 3 according to the number of construction equipment available in each construction stage defined above; For construction equipment k The construction completion time, m is the total number of construction equipment, For construction equipment k Construction time records, For construction equipment k Mobile time records, and All are array types.

[0091] S2. The integrated scheduling problem of construction equipment is expressed as a Markov dynamic decision process (MDP), thereby generating a first intelligent agent for solving the integrated scheduling problem of construction equipment.

[0092] The MDP problem needs to be constructed based on the environment and the agent. The elements that the environment needs to have include state, action, strategy and reward. The agent that solves the MDP problem interacts with the environment in the following way: the agent perceives the initial state of the environment, takes action according to the strategy, the environment is affected by the action and enters a new state, and a reward (positive or negative) is fed back to the agent. Then the agent adopts a new strategy based on the new state to continue to interact with the environment, and this cycle repeats.

[0093] As a preferred implementation scheme, in this embodiment, the construction environment involves two objects, construction partitioning and layering, and construction equipment, for the scheduling problem. In order to facilitate the intelligent agent to perceive the state of the construction site, a state space is designed for these two objects. S , contains all possible states in the scheduling process, including 4 construction partition state information and 4 construction equipment state information, each of which involves the current state s and the next state .

[0094] In an example shown, the four construction zone status information are: average construction zone layered completion rate , average construction completion rate , standard deviation of average construction completion rate , Current maximum construction time ; The 4 construction equipment status information are: total idle time of construction equipment , Average idle rate of construction equipment , Standard deviation of the average idle rate of construction equipment , Total moving distance of the equipment The specific formulas are:

[0095] in, Zoning for construction i The stratified completion rate can be obtained by Calculate and obtain, For partition i The number of layers currently completed, Zoning for construction i The number of layers, n is the number of partitions; Zoning for construction i The completion rate of the job can be obtained by Calculate and obtain, Zoning for construction i Layering p The number of completed stages, a is the total number of stages; For construction stage j Construction equipment k The idle rate, For stage j machine k The completion time, For stage j machine k The start time, For operation In the machine k Processing time on.

[0096] After obtaining the state, the intelligent agent must select the appropriate construction layer and construction equipment according to the corresponding scheduling rules. For this purpose, an action space A is designed for the intelligent agent, which includes multiple composite scheduling rules formed by arranging and combining several scheduling rules for construction zoning and stratification and construction equipment allocation.

[0097] In an example, for the integrated scheduling problem of equipment construction and movement, 3, 2, and 3 scheduling rules are designed for construction zoning and stratification, equipment allocation, and equipment movement, respectively. The three scheduling rules are arranged and combined to form 3×2×3=18 composite scheduling rules, which become the final action space.

[0098] Specifically, the scheduling rules for construction zoning and stratification include selecting the zoning and stratification with the longest estimated construction time, selecting the zoning and stratification with the longest estimated remaining construction time, and FIFO (First Input First Output); the scheduling rules for equipment allocation include selecting the earliest available construction equipment and selecting the construction equipment with the longest idle time; the scheduling rules for equipment movement include selecting the earliest available equipment, selecting the equipment that reaches the next stage layer the fastest, and selecting the equipment with the smallest total moving distance.

[0099] S3. Design a first multi-objective reward function based on the NSGA-Ⅱ algorithm and the DDQN algorithm, and train the first agent according to the first multi-objective reward function.

[0100] In the training process of an intelligent agent, the design of the reward function is a key step in guiding the agent because it has a great influence on the convergence of the agent. The reward function is designed to evaluate actions and optimize strategies.

[0101] As a preferred implementation scheme, the reward function design method in this embodiment is:

[0102] First, set the multi-level reward function values;

[0103] Then, the minimum distance solution of the NSGA-II algorithm and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three scheduling objective function values ​​of each scheduling step;

[0104] Then, the objective function values ​​of the two algorithms are compared, and the smaller value is selected to be compared with the objective function value in the current state as the input of the first multi-objective reward function, and the first multi-objective reward function outputs the reward function value according to the comparison result;

[0105] Finally, the three reward function values ​​are added together as the final reward signal.

[0106] This strategy ensures that the reward signal at each step is optimal, reducing the instability of the DDQN algorithm, which helps the algorithm converge faster and guides the agent to find a better solution.

[0107] In one example, the example of setting multiple levels of reward function values ​​is: +3 represents significant improvement, +1 represents general improvement, -1 represents general deterioration, and -3 represents significant deterioration. By setting multiple levels of reward values, the learning process of the agent can be guided more accurately to avoid it falling into a local optimum or conducting ineffective exploration.

[0108] In one illustrated example, the first multi-objective reward function is set to include:

[0109] The first reward function, the input of which is the smaller of the maximum construction time under the minimum distance solution state obtained based on the NSGA-II algorithm and the maximum construction time under the minimum distance solution state in the Pareto solution set generated in each iteration, and the comparison result with the maximum construction time in the current state, and the output is the first reward function value assigned based on the comparison result; for example, if the maximum construction time in the current state is less than 90% of the smaller one, the reward value is +3; if it is greater than 90% of the smaller one but less than the smaller one, the reward value is +1; if they are the same, the reward value is 0; if it is greater than the smaller one but less than 1.1 times of the smaller one, the reward value is -1; if it is greater than 1.1 times of the smaller one, the reward value is -3.

[0110] The second reward function, the input of which is the smaller of the total idle time of construction equipment in the minimum distance solution state obtained based on the NSGA-II algorithm and the total idle time of construction equipment in the minimum distance solution state in the Pareto solution set generated in each iteration, and the total idle time of construction equipment in the current state; the output is the second reward function value assigned based on the comparison result; for example, if the total idle time of construction equipment in the current state is less than 90% of the smaller one, the reward value is +3; if it is greater than 90% of the smaller one but less than the smaller one, the reward value is +1; if they are the same, the reward value is 0; if it is greater than the smaller one but less than 1.1 times of the smaller one, the reward value is -1; if it is greater than 1.1 times of the smaller one, the reward value is -3.

[0111] The third reward function, the input of which is the smaller of the total driving distance of the device in the minimum distance solution state obtained by the NSGA-II algorithm and the total driving distance of the device in the minimum distance solution state in the Pareto solution set generated in each iteration, and the total driving distance of the device in the current state, and the output is the third reward function value assigned based on the comparison result; for example, if the total driving distance of the device in the current state is less than 90% of the smaller one, the reward value is +3; if it is greater than 90% of the smaller one but less than the smaller one, the reward value is +1; if they are the same, the reward value is 0; if it is greater than the smaller one but less than 1.1 times of the smaller one, the reward value is -1; if it is greater than 1.1 times of the smaller one, the reward value is -3.

[0112] Finally, the three reward function values ​​are added together as the final reward signal.

[0113] S4. Generate an initial scheduling plan through the trained first agent, and perform virtual scheduling based on the initial scheduling plan. Specifically, it can simulate the construction equipment going to the layered construction according to the initial scheduling plan, update the construction site status, and feedback the scheduling reward to the agent.

[0114] See also Figure 3 In an example shown, the agent first obtains the status of the construction site, and then selects the appropriate target layer and construction equipment for scheduling according to the initial scheduling plan. The selected construction equipment goes to the assigned target layer according to the scheduling result. If it is exactly at the target position, it does not need to move. After arriving at the location, it determines whether the equipment of the previous process is under construction in the layer. If so, it waits, otherwise it starts construction. After the construction is completed, the state of the construction site environment is updated, and the reward of this scheduling result is calculated according to the reward function, and then fed back to the agent.

[0115] S5. Based on the assumed disturbance event supplementary constraints, a second intelligent agent is generated through a Markov dynamic decision process to solve the integrated scheduling problem of construction equipment under disturbance conditions.

[0116] The above process shown in this embodiment is the process of generating an initial scheduling plan without considering unexpected situations such as equipment failure. However, in the actual production environment, random disturbance events such as sudden failure of construction equipment are inevitable. The occurrence of these events may make the originally formulated scheduling plan lose its effectiveness. Therefore, it is necessary to dynamically generate a rescheduling plan after the disturbance factor occurs. The dynamic scheduling process is as follows: Figure 4 shown.

[0117] As a preferred implementation scheme, in this embodiment, the constraints supplemented based on the assumed disturbance event include:

[0118] ⑥ Construction equipment failures occur randomly and can occur on any construction equipment at any time;

[0119] ⑦ Once a construction equipment fails, the operation on the construction equipment will be suspended, while the operation of other construction equipment will not be affected;

[0120] ⑧ When the operation on a construction layer is interrupted due to a construction equipment failure, the original construction equipment will continue to operate on the construction layer after the faulty construction equipment is repaired;

[0121] ⑨ The repair time of faulty construction equipment can be estimated in advance.

[0122] Similarly, the above-mentioned additional constraints are only some examples, and other constraints can be set according to different construction environments. Therefore, the above-mentioned constraints do not constitute a limitation on the embodiments of the present invention.

[0123] Based on the above supplemented constraints, a second agent for solving the integrated scheduling problem of construction equipment under disturbance conditions can be generated through the Markov dynamic decision process. The specific generation method can refer to the generation of the first agent above, which will not be repeated here.

[0124] S6. Differentiate the jobs in the initial scheduling plan, and design the encoding and decoding methods of the NSGA-Ⅱ algorithm based on the differentiation results to retain the unaffected parts of the initial scheduling plan and only reschedule the affected parts.

[0125] In the technical concept of the present invention, when generating a rescheduling plan for processing a fault event, unaffected jobs such as tasks that have been scheduled before the fault occurs will not be changed in the rescheduling plan, so the jobs in the original initial scheduling plan need to be distinguished.

[0126] As a preferred implementation scheme, in this embodiment, distinguishing jobs in the initial scheduling scheme includes:

[0127] The jobs in the initial scheduling plan are classified into three categories: G1, G2, and G3, where:

[0128] G1 includes the jobs that have been completed before the failure event occurs, and the jobs that have been started on normal construction equipment. These jobs will maintain their original scheduling status during rescheduling;

[0129] G2 includes the operations that were being carried out on the faulty construction equipment when the fault event occurred. They will continue after the fault is repaired, so the status will not be adjusted during rescheduling;

[0130] G3 includes the operations that have not entered the next construction phase at the time of the failure event. These operations need to rearrange the batch sequence and construction equipment allocation during rescheduling;

[0131] Similarly, the moving tasks of construction equipment are also grouped and divided into three groups: A1, A2, and A3.

[0132] A1 includes tasks that have been completed before the failure event occurs or tasks that are being moved by construction equipment that is operating normally, and their status remains unchanged during rescheduling;

[0133] A2 includes tasks that are being moved on the faulty device at the time of the fault event. For such tasks, the time required for other normal devices to reach the fault point at the time of the fault is compared with the repair time of the faulty device. If the repair time is shorter, the task will continue to be completed by the faulty device; if the arrival time of other devices is shorter, the fastest arriving device will be selected to take over the movement. At this time, the allocation of mobile devices needs to be rescheduled;

[0134] A3 includes tasks that have not started to move to the next stage at the time of the failure event. Their hierarchical sequence and equipment allocation need to be adjusted during rescheduling.

[0135] In the technical concept of the present invention, in response to the occurrence of a fault event, the scheduling scheme adopts a pre-reactive scheduling method, in which the initial scheduling scheme has been determined in advance. In the rescheduling scheme, part of the initial scheduling scheme remains unchanged, and only the affected part is changed. To achieve this goal, the present invention designs the encoding and decoding method of the NSGA-Ⅱ algorithm based on the above-mentioned distinction results to retain the unaffected part of the initial scheduling scheme and only perform rescheduling on the affected part.

[0136] As a preferred implementation scheme, in this embodiment, the encoding and decoding method of the NSGA-Ⅱ algorithm is designed according to the distinction result, including:

[0137] Coding of NSGA-Ⅱ algorithm:

[0138] During the encoding process, the start construction time of all job sequences in the initial scheduling plan is fully traversed, and the construction phases, construction partition numbers and layer numbers corresponding to the job sequences classified as G1 and G2 are recorded to construct a fixed job sequence and record the construction equipment number used in this job sequence so that the construction equipment sequence remains unchanged during the rescheduling process.

[0139] In an example, the chromosome encoding of the designed NSGA-II algorithm is shown as follows: Figure 5 As shown: First, randomly generate chromosome sequences of construction operations and construction equipment as a rescheduling scheme. The chromosome arrangement order of the construction operation sequence in the rescheduling scheme represents the execution order of the jobs. Next, the start time of all construction job sequences in the initial scheduling plan is traversed and compared with the rescheduling time point. Through this step, the job sequence that is not affected in the initial scheduling plan is obtained as .

[0140] In order to ensure the effectiveness of the scheduling plan, this embodiment adopts a specific replacement rule to integrate the unaffected operation sequences in the initial scheduling plan into the rescheduling plan. Since the construction process involved in this embodiment includes 3 stages, each partition layer must appear 3 times in the construction operation sequence chromosome. When the unaffected construction operation sequence in the initial scheduling plan is replaced with the rescheduling plan, the extra construction operation sequence in the rescheduling plan needs to be replaced accordingly. For example, replace the second gene in the chromosome of the rescheduling plan operation sequence. Replaced with the initial schedule When the rescheduling scheme is executed, the first gene after the second gene in the chromosome Replace with ,Right now Figure 5 By analogy, all the unaffected job sequences in the initial scheduling scheme are replaced with the job sequences of the rescheduling scheme according to this rule, and the final legal and effective job sequence chromosome encoding is obtained.

[0141] The construction equipment sequence chromosomes are encoded according to the stage, and crossover mutation can only be performed within the stage. Figure 5 In the example, there are three unaffected job sequences in the initial scheduling plan. ,two and a , and when The first occurrence corresponds to the construction on the construction equipment in the first stage, and the second occurrence corresponds to the construction on the construction equipment in the second stage. Therefore, the construction equipment corresponding to the unchanged operation sequence in the initial scheduling plan is the white part. The white part in the initial scheduling plan construction equipment chromosome is combined with the dotted box part in the rescheduling plan construction equipment chromosome to obtain the final construction equipment sequence chromosome encoding.

[0142] Decoding of NSGA-Ⅱ algorithm:

[0143] According to the impact of the fault on the construction equipment, the earliest available time of the construction equipment and the earliest available time for the construction equipment to move are initialized, and the operations and construction equipment affected by the fault are rescheduled according to the differentiation results to obtain the adjusted operation sequence and construction equipment sequence, and a decoding operation is performed.

[0144] In response to the occurrence of failure events, according to the above-mentioned scheduling operation classification, the construction operations and equipment movement tasks are divided into three categories, G1, G2, and G3, and three groups, A1, A2, and A3, at the time of rescheduling. It involves the construction allocation and movement allocation of batches of equipment, so it is necessary to determine the earliest available time of all construction equipment.

[0145] In one illustrative example, the decoding process is described as follows:

[0146] 1) Initialization of relevant information

[0147] The earliest available time for initializing construction equipment. The earliest available time for faulty construction equipment will depend on the time it takes to repair the fault and the time required to complete the remaining construction tasks of the current job. Construction equipment can only be put back into use after the fault is repaired and the remaining part of the current job is properly handled; the earliest available time for construction equipment that is in construction operation is the moment when the current job is completed; the earliest available time for idle construction equipment is the time of rescheduling;

[0148] Initialize the earliest available time for construction equipment movement. The earliest available time for the movement of faulty equipment depends on whether the distribution task in A2 is replaced by the movement of other normally working equipment. Compare the time when the fault occurs, the time required for the remaining normally working equipment to reach the location of the faulty equipment and the repair time of the faulty equipment. If the repair time is less than the time required for the remaining normally working equipment to reach the location of the faulty equipment, the distribution task will continue to move forward after the faulty equipment is restored, and its equipment allocation cannot be changed during rescheduling. The earliest available time for the movement of the faulty equipment is the time after the fault is repaired and the current forward task is completed; if the repair time is greater than the time required for the remaining normally working equipment to reach the location of the faulty equipment, the same type of equipment with the shortest time to reach the location of the faulty equipment is arranged to take over this forward task, that is, the equipment allocation needs to be rescheduled, and the earliest available time for the movement of the faulty equipment is the time when the fault is repaired; the earliest available time for the movement of equipment in motion is the time when the movement is completed; the earliest available time for the movement of idle equipment is the rescheduling time.

[0149] 2) Execute rescheduling

[0150] For jobs in G1, their status does not change during rescheduling; for jobs in G2, their construction equipment allocation does not change during rescheduling; for jobs in G3, the batch sequence and construction equipment allocation need to be rescheduled.

[0151] For the forward task in A1, its status does not change during rescheduling; for the forward task in A2, if the forward task will continue to move forward after the faulty device is restored, its equipment allocation does not change during rescheduling; if the device with the shortest time to reach the location of the faulty device is scheduled to take over this forward task, the equipment movement allocation needs to be rescheduled; for the forward task in A3, the hierarchical sequence and equipment movement allocation need to be rescheduled.

[0152] 3) Complete the decoding process and obtain the corresponding target value.

[0153] For the initial scheduling plan, the construction equipment failure time is 1771.9, the failure maintenance time is 500, and the construction equipment is available at time 2271.9. The failure times during the equipment movement are 2147.9, 1211.9 and 2114.2, and the failure maintenance time is 500, so the equipment is available at times 2647.9, 1711.9 and 2614.2. According to the above decoding process, the Gantt chart of the rescheduling plan is generated.

[0154] S7. Design a second multi-objective reward function based on the NSGA-Ⅱ algorithm and DDQN algorithm after the design of the encoding and decoding, and train the second agent according to the second multi-objective reward function.

[0155] As a preferred implementation scheme, in this embodiment, the second multi-objective reward function is designed based on the NSGA-Ⅱ algorithm and DDQN algorithm after the design of the codec, including:

[0156] Set multiple levels of reward function values;

[0157] The minimum distance solution of the NSGA-II algorithm after the design of the codec and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three objective function values ​​corresponding to each step of scheduling;

[0158] Compare the objective function values ​​of the two algorithms, select the smaller value to compare with the objective function value in the current state, and use it as the input of the second multi-objective reward function. The second multi-objective reward function outputs a reward function value according to the comparison result.

[0159] Add the three reward function values ​​together as the final reward.

[0160] The design of the second multi-objective reward function and the specific calculation process of the reward function value can refer to the setting and calculation process of the first multi-objective reward function for the first agent, the only difference is that the NSGA-II algorithm after the design codec is used here instead of the original NSGA-II algorithm. Therefore, no further explanation will be given.

[0161] In the rescheduling scenario, compared with the simple reward function based on result comparison, the reward function based on ratio comparison adopted by the present invention can provide more accurate and sensitive feedback, helping the agent to better learn and adapt to the dynamic environment. This design enables the agent to more effectively capture subtle changes in the environment state or task requirements and adjust the agent's behavior accordingly.

[0162] As a preferred implementation, in this embodiment, using the second multi-objective reward function to train the second agent includes:

[0163] Multiple examples are randomly generated, each of which includes a different number of construction zones and layers, as well as the number of construction equipment at different construction stages;

[0164] Initialize the parameters of the Q network in the DDQN algorithm, initialize the values ​​corresponding to all states and actions, and clear the experience replay set;

[0165] Extract state features based on the current construction environment and input the features into the neural network to calculate the value of all possible actions in the current state;

[0166] Use the evaluation network to select the action with the greatest value, and determine the value of the action in the target network based on the selected action to update the Q network;

[0167] Select the action with the highest value from the target network again, execute the selected action, observe the reward and next state fed back by the environment, and update the reward and state;

[0168] Repeat the above steps for each construction partition layer until all construction partition layers are completed, completing the training of the agent.

[0169] Repeat the above training process until the number of training reaches the set number of iterations to complete the training of the agent.

[0170] In an illustrative example, the training process is as follows:

[0171] Three examples were randomly generated, and the number of partitions in each example was 6, 9, and 12 respectively. The number of layers in each partition in each example was 4-10. The number of construction equipment in the first construction stage in each example was 4, 8, and 12 respectively. The number of construction equipment in the second construction stage in each example was 5, 10, and 15 respectively. The number of construction equipment in the third construction stage in each example was 3, 6, and 9 respectively.

[0172] After each training session, a set of experience values ​​is obtained, and each set of experience values ​​obtained after each training session is put into the experience replay collection. Each set of experience values ​​contains the reward value of each scheduling. During the training process, when the experience replay collection is full, if two values ​​in a set of experience values ​​obtained in this training session are greater than the corresponding two values ​​in a previous set of experience values, the set of experience values ​​obtained in this training session will replace the previous set of experience values.

[0173] After the training is completed, the empirical values ​​of the three objectives under different parameter combinations are compared, and the parameter combination corresponding to the best set of empirical values ​​is used as the optimal parameter combination of the agent. The agent with the best parameter combination is the trained agent. The optimal parameter combination is: the population size is 125, the crossover rate CR is 0.3, the mutation rate MR is 0.03, the learning rate a is 0.001, the discount factor y is 0.001, and the greed rate s is 0.1.

[0174] S8. Deploy the first agent and the second agent after stable training in an actual construction environment to generate dynamic scheduling decisions for construction equipment under real fault events.

[0175] In an illustrated example, the trained first agent and second agent are deployed to the construction site. Based on the Python simulation platform, the construction and movement of the construction site equipment are integrated and scheduled according to the actual construction situation of the construction site, and the specific integrated scheduling is represented by a Gantt chart. The Gantt chart is a project management tool that displays the schedule and progress of each task in the project in the form of a bar chart. Each task is represented by a corresponding bar on the Gantt chart. The length of the bar represents the duration of the task and the position of the bar represents the start and end time of the task. Through the Gantt chart, managers can clearly understand the schedule of each task in the project and the relationship between them, so as to understand the overall construction operation of the construction site.

[0176] Another embodiment of the present invention further provides a computer device, including a memory and a processor, wherein the memory stores computer instructions, and the processor executes the earthwork construction dynamic scheduling method disclosed in the above-mentioned embodiment by executing the computer instructions.

[0177] Another embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the earthwork construction dynamic scheduling method disclosed in the above embodiment is executed.

[0178] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or a server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the described method.

[0179] It should be noted that some embodiments of the present invention are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0180] The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the present invention.

Claims

1. A dynamic scheduling method for earthwork construction based on deep reinforcement learning, characterized in that: The steps include: Obtain information about the construction environment, construction process, and construction equipment, determine constraints and integrated scheduling objectives for construction equipment based on the information obtained, and set an objective function; The integrated scheduling problem of construction equipment is formulated as a Markov dynamic decision process, and a first intelligent agent is generated for solving the integrated scheduling problem of construction equipment; The first multi-objective reward function is designed based on the NSGA-Ⅱ algorithm and the DDQN algorithm, including: Set multiple levels of reward function values; The minimum distance solution of the NSGA-II algorithm and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three objective function values ​​corresponding to each step of scheduling; Compare the objective function values ​​of the two algorithms, select the smaller value and compare it with the objective function value in the current state as the input of the reward function, and the reward function outputs the reward function value according to the comparison result; Add the three reward function values ​​as the final reward value; Training a first agent according to the first multi-objective reward function; Generate an initial scheduling plan through the trained first agent, and perform virtual scheduling based on the initial scheduling plan; Supplementing the constraint conditions based on the assumed disturbance event, generating a second intelligent agent for solving the integrated scheduling problem of construction equipment under disturbance conditions through a Markov dynamic decision process; Differentiate the jobs in the initial scheduling plan, design the encoding and decoding methods of the NSGA-Ⅱ algorithm based on the differentiation results, retain the unaffected parts of the initial scheduling plan, and reschedule the affected parts; Design a second multi-objective reward function based on the NSGA-Ⅱ algorithm and the DDQN algorithm after the design of the codec, and train the second agent according to the second multi-objective reward function; The first and second agents, after stable training, are deployed in the actual construction environment to generate dynamic scheduling decisions for construction equipment under real fault events.

2. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 1, characterized in that: The construction environment information includes construction zones and layer information to be constructed, the construction process information includes operation phase information of each construction zone, and the construction equipment information includes information of construction equipment that can be used for scheduling operations.

3. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 2 is characterized in that: The constraints include: During different operation phases, construction equipment needs to go to the construction zone, and the location of the construction zone is known; The construction equipment corresponding to each operation stage can be used to construct any construction zone, but the construction time is different and has been determined; Once a construction device starts working in a construction zone, the work will not be interrupted; The same construction zone can only carry out the subsequent operation phase after the current operation phase is completed; After the construction equipment finishes the current construction work, it will immediately start from the current position to the construction area where the next task is located to carry out the work.

4. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 2, characterized in that: The integrated scheduling objectives of the construction equipment include: Minimize the operation time of construction equipment, the corresponding objective function is ; Minimize the total idle time of construction equipment, the corresponding objective function is ; Minimize the total moving distance, the corresponding objective function is ; in, k Number the equipment in the order of stage 1 to 3 according to the number of construction equipment available at each construction stage; For construction equipment k The construction completion time, m is the total number of construction equipment; and All are array types. For construction equipment k Construction time records, For construction equipment k The moving time record.

5. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 4, characterized in that: The integrated scheduling problem of construction equipment is expressed as a Markov dynamic decision process, and the first intelligent agent generated to solve the integrated scheduling problem of construction equipment includes: The state space is designed based on the two objects of construction partition and construction equipment, including several construction partition state information and construction equipment state information. Each state information involves the current state s and the next state ; An action space is designed for the intelligent agent, wherein the action space includes a plurality of composite scheduling rules formed by arranging and combining a plurality of scheduling rules for construction zoning and stratification and construction equipment allocation.

6. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 1, characterized in that: The additional constraints based on the assumed disturbance events include: Construction equipment failures occur randomly and can occur on any piece of construction equipment at any point in time; Once a construction equipment fails, the operation on the construction equipment will be suspended, while the operation of other construction equipment will not be affected; When the operation on a construction layer is interrupted due to a construction equipment failure, the original construction equipment will continue to operate on the construction layer after the failed construction equipment is repaired; The repair time of faulty construction equipment can be estimated in advance.

7. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 1, characterized in that: Differentiating jobs in the initial scheduling plan includes: The jobs in the initial scheduling plan are classified into three categories: G1, G2, and G3, where: G1 includes the jobs that have been completed before the failure event occurs, and the jobs that have been started on normal construction equipment. These jobs will maintain their original scheduling status during rescheduling; G2 includes the operations that were being carried out on the faulty construction equipment when the fault event occurred. They will continue after the fault is repaired, so the status will not be adjusted during rescheduling; G3 includes the operations that have not entered the next construction phase at the time of the failure event. These operations need to rearrange the batch sequence and construction equipment allocation during rescheduling; The mobile tasks of construction equipment are divided into three groups: A1, A2, and A3. A1 includes tasks that have been completed before the fault event occurs or tasks that are being moved by construction equipment that is operating normally, and their status remains unchanged during rescheduling; A2 includes tasks where the faulty equipment is being moved at the time of the fault event. For such tasks, the time required for other normal equipment to reach the next target point at the time of the fault is compared with the repair time of the faulty equipment. If the repair time is shorter, the task will continue to be completed by the faulty equipment; if the arrival time of other equipment is shorter, the equipment that arrives the fastest will be selected to take over the movement. At this time, the construction equipment allocation needs to be rescheduled; A3 includes tasks that have not started moving to the next target at the time of the failure event. Their hierarchical sequence and equipment allocation need to be adjusted during rescheduling.

8. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 7, characterized in that: The encoding and decoding methods of the NSGA-Ⅱ algorithm are designed according to the differentiation results, including: During the encoding process, the start time of all operation sequences in the initial scheduling plan is fully traversed, and the construction phase, construction partition number and layer number corresponding to the operation sequences classified as G1 and G2 are recorded to construct a fixed operation sequence, and the construction equipment number used in this operation sequence is recorded, so that the construction equipment sequence remains unchanged during the rescheduling process; According to the impact of the fault on the construction equipment, the earliest available time of the construction equipment and the earliest available time for the construction equipment to move are initialized, and the operations and construction equipment affected by the fault are rescheduled according to the differentiation results to obtain the adjusted operation sequence and construction equipment sequence, and a decoding operation is performed.

9. The earthwork construction dynamic scheduling method based on deep reinforcement learning as claimed in claim 8, characterized in that: The second multi-objective reward function designed based on the NSGA-Ⅱ algorithm and DDQN algorithm after the design of the codec includes: Set multiple levels of reward function values; The minimum distance solution of the NSGA-II algorithm after the design of the codec and the minimum distance solution of the DDQN algorithm are backtracked to obtain the three objective function values ​​corresponding to each step of scheduling; Compare the objective function values ​​of the two algorithms, select the smaller value and compare it with the objective function value in the current state as the input of the reward function, and the reward function outputs the reward function value according to the comparison result; Add the three reward function values ​​​​to get the final reward value.

10. The earthwork construction dynamic scheduling method based on deep reinforcement learning according to claim 9, characterized in that: Training the second agent according to the second multi-objective reward function includes: Multiple examples are randomly generated, each of which includes a different number of construction zones and layers, as well as the number of construction equipment at different construction stages; Initialize the parameters of the Q network in the DDQN algorithm, initialize the values ​​corresponding to all states and actions, and clear the experience replay set; Extract state features based on the current construction environment and input the features into the neural network to calculate the value of all possible actions in the current state; Use the evaluation network to select the action with the greatest value, and determine the value of the action in the target network based on the selected action to update the Q network; Select the action with the highest value from the target network again, execute the selected action, observe the reward and next state fed back by the environment, and update the reward and state; Repeat the above steps for each construction partition layer until all construction partition layers are completed, completing the training of the agent. Repeat the above training process until the number of training reaches the set number of iterations to complete the training of the agent.

11. A computer device, characterized in that: It comprises a memory and a processor, wherein the memory stores computer instructions, and the processor executes the earthwork construction dynamic scheduling method as described in any one of claims 1 to 10 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions, and when the computer instructions are executed by a computer, the earthwork construction dynamic scheduling method as described in any one of claims 1 to 10 is executed.

Citation Information

Patent Citations

  • Zero-carbon building optimization design method based on deep reinforcement learning

    CN114692265A

  • Equipment manufacturing workshop intelligent scheduling method and system based on deep reinforcement learning

    CN116542445A