Resource-balanced construction scheduling method and apparatus based on deep reinforcement learning, and device

The neural network model constructed through deep reinforcement learning solves the problem of resource balance optimization in construction schedule, realizes balanced allocation and efficiency improvement of resources during construction, and shortens the construction period.

WO2025157071A1PCT designated stage Publication Date: 2025-07-31QINGYUNARCH (BEIJING) INNOVATION TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2025/072853
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-16
Publication Date
2025-07-31

Smart Images

  • Figure CN2025072853_31072025_PF_FP_ABST
    Figure CN2025072853_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of building construction scheduling, and provides a resource-balanced construction scheduling method and apparatus based on deep reinforcement learning, and a device. The method comprises: acquiring project information and resource demand information corresponding to at least one sample construction project; by taking resource balance and scheduling efficiency as optimization goals, respectively constructing a single-step reward function and a project total reward function; constructing a deep neural network model, performing reinforcement learning on the deep neural network model on the basis of construction status data corresponding to a current construction time step, the project information and the resource demand information, to output a decision corresponding to a next construction time step, and updating model parameters on the basis of the single-step reward function; when construction scheduling is completed, updating the model parameters on the basis of the project total reward function; and traversing all the sample construction projects, and repeatedly executing the step of updating the model parameters to obtain a trained construction scheduling model. The present application can implement construction scheduling with resource balance as the goal.
Need to check novelty before this filing date? Find Prior Art

Description

Resource-balanced construction scheduling method, device and equipment based on deep reinforcement learning

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202410111166.4, filed on January 26, 2024, entitled “Resource-balanced construction scheduling method, device and equipment based on deep reinforcement learning”, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present application relates to the technical field of construction scheduling, and in particular to a resource-balanced construction scheduling method, device, and equipment based on deep reinforcement learning. Background Art

[0004] Construction scheduling is a crucial aspect of construction engineering. Automated construction scheduling technology can comprehensively consider multiple factors to provide engineers with scientific reference plans, significantly freeing up labor and narrowing the experience gap between newcomers and experienced engineering experts. Automated construction scheduling typically begins by defining a Resource Constrained Project Scheduling Problem (RCPSP). Solving the RCPSP yields a construction project schedule that considers both process priority constraints and resource finiteness.

[0005] In existing technologies, reinforcement learning algorithms can be used to solve RCPSP, with the optimization goal of minimizing project duration or cost. Construction scheduling for engineering projects is achieved by making decisions about worker hours and material allocation. However, for construction optimization problems that consider resource balancing constraints, existing reinforcement learning algorithms approach the optimal overall decision by achieving the optimal decision at each step. However, the resource balancing goal considers resource stability throughout the entire construction project. The difference between the optimal decision at each step and the optimal overall decision is significant, and existing reinforcement learning algorithms are unable to solve construction scheduling optimization with resource balancing as the optimization goal. Summary of the Invention

[0006] The present application provides a resource-balanced construction scheduling method, device and equipment based on deep reinforcement learning, which is used to solve the defect of the existing technology that it is unable to solve the construction scheduling optimization with resource balancing as the optimization goal.

[0007] This application provides a resource-balanced construction scheduling method based on deep reinforcement learning, including:

[0008] Acquire project information and resource requirement information corresponding to at least one sample construction project; the resource requirement information is used to represent a mapping relationship between processes, resource types, and resource requirements in the sample construction project;

[0009] Taking resource balance and scheduling efficiency as optimization goals, a single-step reward function and a project total reward function are respectively constructed; based on the project information, the resource demand information, the single-step reward function and the project total reward function, a deep neural network model is constructed, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model, a third sub-model and a fourth sub-model constructed based on a recurrent neural network, and a main sub-model constructed based on a deep neural network;

[0010] Based on the deep neural network model, construction status data corresponding to the current construction time step is obtained. Based on the construction status data corresponding to the current construction time step, the project information, and the resource demand information, reinforcement learning is performed on the deep neural network model. The main sub-model outputs a decision corresponding to the next construction time step, and the model parameters of the deep neural network model are updated based on the single-step reward function. The construction status data is used to characterize the process completion progress and resource ownership corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model.

[0011] When the next construction time step is less than or equal to the construction duration threshold, repeatedly executing the single-step decision-making step, and after the construction schedule is completed, updating the model parameters of the current iteration round based on the project total reward function;

[0012] Traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters to obtain a trained construction scheduling model, and schedule the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to the target construction project at each construction time step.

[0013] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, the resource ownership in the construction status data includes material resource ownership, worker resource ownership, and reusable equipment resource ownership;

[0014] The method of acquiring construction status data corresponding to a current construction time step based on the deep neural network model, performing reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource demand information, wherein the main sub-model outputs a decision corresponding to the next construction time step, including:

[0015] When the current iteration round is less than or equal to the preset iteration round, obtaining the process completion progress and resource ownership corresponding to the current construction time step based on the deep neural network model;

[0016] When the process completion progress is less than the preset resource requirement, the process completion progress is input into the first sub-model and a process vector is output;

[0017] Inputting the material resource ownership into the second sub-model, the worker resource ownership into the third sub-model, and the reusable equipment resource ownership into the fourth sub-model, and outputting a resource vector;

[0018] The process vector and the resource vector are input into the main sub-model, and the decision of the next construction time step is output.

[0019] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, the updating of the model parameters of the deep neural network model based on the single-step reward function includes:

[0020] If the decision satisfies a feasible strategy condition, the decision is executed, and based on the single-step reward function, a scheduling reward, a progress reward, and a cost reward corresponding to the execution of the decision are determined; the feasible strategy condition includes a storage space constraint condition and a construction space constraint condition;

[0021] Determining a single-step reward value corresponding to the decision based on the scheduling reward, the progress reward, and the cost reward;

[0022] Based on the single-step reward value, model parameters of the deep neural network model are updated.

[0023] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, the method also includes:

[0024] When the decision does not meet the feasible strategy conditions, the construction schedule of the current iteration round is determined to be ended, and based on the total project reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current construction time step are rolled back to the model parameters corresponding to the previous construction time step.

[0025] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, the method also includes:

[0026] When the completion progress of the process is equal to the preset resource demand, the construction schedule of the current iteration round is determined to be completed, and based on the total project reward function, the total project reward value of the current iteration round is determined, and based on the total project reward value, the model parameters of the current iteration round are updated.

[0027] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, determining the total project reward value of the current iteration round based on the total project reward function includes:

[0028] Based on the total project reward function, respectively determining the material resource fluctuation rate, worker resource fluctuation rate, and cost waste rate corresponding to the current iteration round;

[0029] Based on the material resource fluctuation rate, the worker resource fluctuation rate and the cost waste rate, a total reward value of the project corresponding to the current iteration round is determined.

[0030] According to the resource-balanced construction scheduling method based on deep reinforcement learning provided by this application, the method also includes:

[0031] When the next construction time step is greater than the construction period threshold, the construction schedule is determined to be completed, and based on the project total reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current iteration round are rolled back to the model parameters corresponding to the previous iteration round.

[0032] This application also provides a resource-balanced construction scheduling device based on deep reinforcement learning, including:

[0033] An acquisition module, configured to acquire project information and resource requirement information corresponding to at least one sample construction project; the resource requirement information is used to characterize the mapping relationship between the process, resource type, and resource requirement in the sample construction project;

[0034] A construction module is used to construct a single-step reward function and a project total reward function respectively with resource balance and scheduling efficiency as optimization objectives; based on the project information, the resource demand information, the single-step reward function and the project total reward function, a deep neural network model is constructed, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model, a third sub-model and a fourth sub-model constructed based on a recurrent neural network, and a main sub-model constructed based on a deep neural network;

[0035] A first updating module is configured to obtain, based on the deep neural network model, construction status data corresponding to the current construction time step, perform reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource demand information, wherein the main sub-model outputs a decision corresponding to the next construction time step, and updates the model parameters of the deep neural network model based on the single-step reward function; the construction status data is used to represent the process completion progress and resource ownership corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model;

[0036] A second updating module is configured to repeatedly execute the single-step decision-making step when the next construction time step is less than or equal to the construction duration threshold, and update the model parameters of the current iteration round based on the project total reward function after the construction schedule is completed;

[0037] The scheduling module is used to traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters, obtain a trained construction scheduling model, and schedule the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to the target construction project at each construction time step.

[0038] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for resource-balanced construction scheduling based on deep reinforcement learning as described above is implemented.

[0039] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements any of the above-described resource-balanced construction scheduling methods based on deep reinforcement learning.

[0040] The resource-balanced construction scheduling method, apparatus, and equipment based on deep reinforcement learning provided herein obtain project information corresponding to each sample construction project, as well as resource demand information consisting of mappings between process steps, resource types, and resource demand quantities. With resource balance and scheduling efficiency as optimization goals, a single-step reward function and a project total reward function are constructed. Based on the resource demand information, the single-step reward function, and the project total reward function, a deep neural network model is constructed. First through fourth sub-models process different types of resource training data in the construction status data, respectively, and the main sub-model outputs the decision corresponding to the next construction time step. During training, the deep neural network model performs reinforcement learning through interaction with the environment, enabling the construction scheduling model to efficiently adapt to the complex data corresponding to the environment and the target construction project. When updating the model parameters of the deep neural network model, the single-step reward function is used to ensure that the decision corresponding to each construction time step is optimal, while the project total reward function is used to ensure that the decisions corresponding to all construction time steps, as a whole, meet the requirements of resource balance and scheduling efficiency. When the trained construction scheduling model is used to schedule the target construction project, resource balance is achieved among different resource types during the construction scheduling process, improving the rationality of the construction scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] FIG1 is a flow chart of a resource-balanced construction scheduling method based on deep reinforcement learning provided in an embodiment of the present application;

[0043] FIG2 is a schematic diagram of the training process of the deep neural network model provided in an embodiment of the present application;

[0044] FIG3 is a schematic diagram of the structure of a deep neural network model provided in an embodiment of the present application;

[0045] FIG4 is a schematic diagram showing how the total reward value of a project varies with iteration rounds, as provided in an embodiment of the present application;

[0046] FIG5 is a schematic diagram showing changes in the procurement volume of material resources with construction time steps according to an embodiment of the present application;

[0047] FIG6 is a schematic diagram showing how material warehouse utilization varies with project time, according to an embodiment of the present application;

[0048] FIG7 is a schematic diagram of the structure of a resource-balanced construction scheduling device based on deep reinforcement learning provided in an embodiment of the present application;

[0049] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0051] In view of the problem that the existing technology cannot solve the construction scheduling optimization with resource balancing as the optimization goal, the embodiment of the present application provides a resource-balanced construction scheduling method based on deep reinforcement learning. FIG1 is a flow chart of the resource-balanced construction scheduling method based on deep reinforcement learning provided by the embodiment of the present application. As shown in FIG1, the method includes:

[0052] Step 110: Obtain project information and resource requirement information corresponding to at least one sample construction project; the resource requirement information is used to represent the mapping relationship between the process, resource type and resource requirement in the sample construction project.

[0053] Optionally, the project information corresponding to the sample construction project may include construction information, the total number of processes, the construction period threshold, the material loss coefficient and the available space, etc. Taking the sample construction project as the construction of a large sports stadium as an example, the construction information of the large sports stadium may include: the length from north to south of the large sports stadium is 360m, the width from east to west is 270m, the large sports stadium is oval in shape as a whole, with a construction area of ​​about 128,000 square meters, the main stadium is 6 stories high, and the partial construction is 8 stories high. The actual amount of concrete used in the large sports stadium is about 346,000 tons, and the amount of steel structure used is about 16,000 tons. The total number of processes of the large sports stadium is 118. The construction period threshold is 611 natural days, that is, the maximum construction period specified for the construction of the large sports stadium is 611 natural days. In the material loss coefficient, since concrete cannot be used overnight, the loss coefficient of concrete δ con =1.00, loss coefficient δ of materials other than concrete k = 0.02. Available space spa includes: Material warehouse space spa1 is 5000m 2 、Workers' dormitory area spa2 is 2400m 2 and available construction space spa3 is 2800m 2 Among them, the area of ​​a single room in the workers' dormitory is 30m 2, a total of 80. In addition, the project information corresponding to the sample construction project may also include: the space occupancy coefficient of the resource, the unit cost of the dispatch vehicle, the capacity of the dispatch vehicle, and the cost of different types of resources, etc. This embodiment of the application does not limit this.

[0054] Table 1

[0055] Optionally, the resource requirement information may include a resource requirement table and a construction rate table corresponding to each process. The resource requirement table includes a mapping relationship between the process, the first resource type, and the resource demand. The first resource type includes material resources and reusable equipment resources. The resource requirement table can be understood as an N×(M1+M3) matrix, where N is the total number of processes in the sample construction project, M1 represents the type of material resources, and M3 represents the type of reusable equipment resources. Taking the sample construction project of building a large sports stadium as an example, there are 21 types of resources involved, including: 8 types of material resources, 8 types of worker resources, and 5 types of reusable equipment resources. The resource requirement table corresponding to each process is shown in Table 1.

[0056] Furthermore, the construction rate table includes a mapping between worker resource types and construction rates. This construction rate table can be understood as a vector of length M2, where M2 represents the worker resource type. For example, with eight worker resource types, the construction rate table is shown in Table 2.

[0057] It should be noted that M=M1+M2+M3, where M represents the total number of resource types. Taking 21 resource types, including 8 material resources, 8 worker resources, and 5 reusable equipment resources, as an example, M=21, M1=8, M2=8, and M3=5.

[0058] Table 2

[0059] Step 120: Taking resource balance and scheduling efficiency as optimization goals, construct a single-step reward function and a project total reward function respectively; based on the project information, the resource demand information, the single-step reward function and the project total reward function, construct a deep neural network model, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model, a third sub-model and a fourth sub-model constructed based on a recurrent neural network, and a main sub-model constructed based on a deep neural network.

[0060] Specifically, after obtaining the project information and resource demand information corresponding to each sample construction project, a single-step reward function and a project total reward function are constructed, respectively, with resource balance and scheduling efficiency as optimization goals. These two reward functions can determine the reward value obtained by the deep neural network model through interaction with the resource scheduling environment after outputting the decision, that is, the reward value obtained after executing the decision. The single-step reward function can determine the single-step reward value obtained after executing the decision at construction time step t, where:

[0061] 1. The single-step reward function includes three parts: scheduling reward, progress reward, and cost reward, where:

[0062] 1) The scheduling reward can be understood as the cost of each new scheduling event, which will generate a scheduling reward. This scheduling reward is a negative number. The scheduling event includes material scheduling, worker resource scheduling, and reusable equipment resource scheduling, where:

[0063] Material scheduling is accomplished through dispatch vehicles, and the use of each dispatch vehicle incurs additional costs. Using formula (1), the material scheduling reward can be calculated based on the capacity of the dispatch vehicle, the unit cost of the dispatch vehicle, and the material resource procurement quantity for the kth resource at the tth construction time step. Formula (1) is:

[0064] Among them, cost_delta_res1 k,t represents the material scheduling reward for the kth resource at the tth construction time step, α is the unit cost of the dispatch vehicle, m is the capacity of the dispatch vehicle, delta_res1 k,t represents the material resource procurement quantity for the kth resource at the tth construction time step, that is, the procurement quantity required at the end of the tth construction time step. The material resources purchased with this material resource procurement quantity will be used at the t+1th construction time step. k It represents the ratio coefficient between the quantity of the k-th resource and the space it occupies. Since it is the scheduling of material resources, 1≤k≤M1, and k is an integer. It means rounding up, that is, the number of dispatch vehicles depends on the space required for the purchase volume of material resources.

[0065] Each scheduling of worker resources and reusable equipment resources will increase the cost. Using formula (2), the scheduling reward of worker resources and reusable equipment resources can be calculated based on the scheduling cost coefficient of the k-th resource and the scheduling amount of worker resources and reusable equipment resources of the k-th resource at the t-th construction time step. Formula (2) is: cost_delta_res2 k,t =β k *|delta_res2 k,t |

[0066] Among them, cost_delta_res2 k,t It represents the scheduling reward for the kth resource in the tth construction time step. Since it is the scheduling of worker resources and reusable equipment resources, M1+1≤k≤M, and k is an integer, β k Indicates the scheduling cost coefficient of the k-th resource, delta_res2 k,t Represents the scheduling quantity of worker resources and reusable equipment resources for the kth resource at the tth construction time step.

[0067] After determining the above-mentioned material scheduling rewards, as well as the scheduling rewards for worker resources and reusable equipment resources, the scheduling reward corresponding to the t-th construction time step can be calculated using formula (3), which is:

[0068] Among them, r_delta_res t represents the scheduling reward corresponding to the t-th construction time step, c k Indicates the procurement cost of the kth resource, res_require k,i represents the resource demand of the i-th process for the k-th resource, It represents the cost corresponding to the resource requirements of all processes, which is used to eliminate the impact of the cost dimension and facilitate the training of deep neural network models.

[0069] 2) The progress reward can be understood as when the construction period of the target construction project exceeds the construction period threshold, it means that the construction of the target construction project has failed. Therefore, a large and negative progress reward is required. During the actual training of the deep neural network model, when the deep neural network model is initially trained, the probability of successful execution of the target construction project is small, resulting in the progress reward often being the same negative value during the previous training process of the deep neural network model, resulting in the progress reward being relatively sparse, making the deep neural network model difficult to train. Therefore, a progress reward for the sparse part and a progress reward for the dense part are set in the progress reward, and the progress reward is determined based on the sum of the progress reward for the sparse part and the progress reward for the dense part. The progress reward for the sparse part is shown in formula (4), which is:

[0070] Among them, r_pro spar,t represents the progress reward of the sparse part, T max Indicates the construction duration threshold.

[0071] The progress reward of this intensive part is shown in formula (5), which is:

[0072] Among them, r_pro den,t represents the progress reward of the dense part corresponding to the t-th construction time step, prfk,i.t It represents the completion progress of the kth resource in the i-th process at the t-th construction time step, that is, the consumption of the kth resource in the i-th process at the t-th construction time step.

[0073] After determining the progress reward of the sparse part and the progress reward of the dense part, the formula (6) can be used to calculate the sum of the progress reward of the sparse part and the progress reward of the dense part to determine the progress reward corresponding to the t-th construction time step. Formula (6) is: r_pro t =r_pro spar,t +r_pro den,t

[0074] Among them, r_pro t Represents the progress reward corresponding to the t-th construction time step.

[0075] 3) The cost incentive is used to prevent over-purchasing. The cost incentive may include the material procurement cost, as well as the cost of workers and reusable equipment, where:

[0076] The material purchase cost can be calculated using formula (7): res_cost1 k,t =c k *delta_res1 k,t

[0077] Among them, since the purchased resources are material resources, 1≤k≤M1, and k is an integer, res_cost1 k,t Represents the material procurement cost for the kth resource at the tth construction time step.

[0078] The cost of the worker and the reusable equipment can be calculated using formula (8): res_cost2 k,t =d k *res_exsit k,t

[0079] Among them, since it is the employment of workers and the leasing of reusable equipment resources, therefore, M1+1≤k≤M, and k is an integer, res_cost2 k,t represents the cost of the kth resource at the tth construction time step, d k Represents the cost of hiring or renting the k-th resource, res_exsit k,t It represents the resource ownership of the kth resource at the end of the tth construction time step. The resource ownership is the resource stock at the t+1th construction time step, that is, the resource corresponding to the resource ownership is used in the t+1th construction time step.

[0080] After determining the material procurement cost, the labor and reusable equipment costs, use formula (9) to calculate the sum of the material procurement cost and the labor and reusable equipment costs to determine the procurement cost corresponding to the t-th construction time step. Formula (9) is:

[0081] Among them, res_cost_total t Indicates the purchase cost.

[0082] After determining the procurement cost, the cost reward for the t-th construction time step can be calculated using formula (10), which is:

[0083] Among them, r_res_cost t represents the cost reward for the t-th construction time step, prf t Indicates the completion progress corresponding to the t-th construction time step, that is, the resource consumption of the t-th construction time step, through The impact of the inherent resource requirements of the process can be eliminated.

[0084] 4) After determining the scheduling reward, progress reward, and cost reward, the single-step reward function corresponding to the t-th construction time step can be determined using formula (11) based on the weighted sum of the scheduling reward, progress reward, and cost reward. Formula (11) is: t =ω1*r_delta_res t +ω2*r_pro t +ω3*r_res_cost t

[0085] Among them, r t represents the single-step reward function corresponding to the t-th construction time step, ω1 represents the weight corresponding to the scheduling reward of the t-th construction time step, ω2 represents the weight corresponding to the progress reward of the t-th construction time step, and ω3 represents the weight corresponding to the cost reward of the t-th construction time step. ω1, ω2, and ω3 can be given by the user.

[0086] 2. The total reward function of this project includes three parts: material resource volatility, worker resource volatility, and cost waste rate.

[0087] 1) Using formula (12), calculate the material resource fluctuation rate, formula (12) is:

[0088] Among them: rate_change1 represents the material resource fluctuation rate after the completion of the target construction project, 1≤k≤M1, and k is an integer, T represents the total construction period, delta_res1 k,t-1 represents the material resource procurement quantity for the kth resource at the t-1th construction time step, Indicates the resource demand of the kth resource in all processes, through The dimensional influence can be eliminated, and material resources can be balanced based on the scheduling quantity.

[0089] 2) Using formula (13), calculate the worker resource volatility, formula (13) is:

[0090] Among them, rate_change2 represents the worker resource fluctuation rate after the completion of the target construction project, and res_exist k,t-1 represents the resource ownership of the kth resource at the end of the t-1th construction time step, spa2 represents the area of ​​workers' dormitories, M1+1≤k≤M1+M2, and k is an integer. In formula (13), the dimensionless effect of the workers' dormitories is eliminated by the workers' resources, and the workers' resources are balanced by the resource stock.

[0091] 3) Using formula (14), calculate the cost waste rate, formula (14) is:

[0092] Among them, rate_waste represents the cost waste rate, which is used to prevent excessive purchases.

[0093] 4) After determining the material resource volatility, worker resource volatility, and cost waste rate, use formula (15) to determine the total project reward value, which is:

[0094] When the total construction period exceeds the construction duration threshold, the target construction project fails, and the total project reward is set to 0. Alternatively, when the decision output by the deep neural network model fails to meet the feasible strategy conditions within the construction constraints, the total project reward is set to 0. That is, if the decision does not simultaneously meet the storage space constraints and the construction space constraints within the feasible strategy conditions, construction cannot proceed. Therefore, the total project reward in this case is set to 0. Other cases can be understood as situations where the decision meets the feasible strategy conditions within the construction constraints, that is, the decision satisfies both the storage space constraints and the construction space constraints within the feasible strategy conditions. ω4 represents the weight corresponding to the material resource volatility, ω5 represents the weight corresponding to the worker resource volatility, and ω6 represents the weight corresponding to the cost waste rate. Q1, Q2, and Q3 are preset values ​​that can be set by the user and adjusted based on the values ​​of the material resource volatility, worker resource volatility, and cost waste rate to facilitate deep neural network model training.

[0095] At the same time, a deep neural network model can be constructed, which includes a first sub-model, a second sub-model, a third sub-model, a fourth sub-model and a main sub-model. The first sub-model can be a convolutional neural network (CNN), and the three models from the second to the fourth sub-models are recurrent neural networks (RNN). The first to fourth sub-models are in a parallel relationship, and the main sub-model can be a deep neural network (DNN). The output data of the first to fourth sub-models are the input data of the main sub-model.

[0096] Furthermore, after determining the single-step reward function and the total project reward function, the loss function corresponding to the deep neural network model can be constructed based on the above single-step reward function or the total project reward function, that is, the negative of the above single-step reward function or the negative of the total project reward value is determined as the loss function corresponding to the deep neural network model. When the negative of the single-step reward function is determined as the loss function corresponding to the deep neural network model, the model parameters corresponding to the current construction time step can be updated according to the loss function. When the negative of the total project reward value is determined as the loss function corresponding to the deep neural network model, the model parameters of the current iteration round in the deep neural network model can be updated according to the loss function. After determining the loss function, using formula (16), the weight corresponding to the previous iteration round can be updated using the gradient descent method according to the selected optimizer. Formula (16) is:

[0097] Among them, w new represents the new weight corresponding to the last iteration round after the update, w represents the weight corresponding to the last iteration round before the update, lr represents the learning rate, Represents the gradient of the loss function, which can be the negative of the single-step reward function or the negative of the total reward value of the project.

[0098] After building the deep neural network model, the model parameters of the deep neural network model can be initialized based on project information and resource requirement information. Specifically, after building the deep neural network model, the project information and resource requirement information are obtained, and in response to relevant user operations, the preset parameters related to the deep neural network model entered by the user are obtained. The preset parameters include the weight corresponding to the scheduling reward ω1, the weight corresponding to the progress reward ω2, the weight corresponding to the cost reward ω3, the weight corresponding to the material resource volatility ω4, the weight corresponding to the worker resource volatility ω5, the weight corresponding to the cost waste rate ω6, three preset values ​​Q1, Q2, and Q3, the learning rate lr, the optimizer Optimizer, and the preset number of iterations epochs. The current iteration epoch is set to 0, and the model parameters corresponding to the deep neural network model are randomly initialized.

[0099] Step 130: Based on the deep neural network model, obtain the construction status data corresponding to the current construction time step, perform reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information and the resource demand information, the main sub-model outputs the decision corresponding to the next construction time step, and updates the model parameters of the deep neural network model based on the single-step reward function; the construction status data is used to characterize the process completion progress and resource ownership corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model.

[0100] Specifically, the deep neural network model includes first, second, and fourth sub-models and a main sub-model. The first sub-model processes the progress of a process's completion. This progress can be time series data, a tensor of N×M1×S dimensions, obtained by taking S construction time steps from the cumulative progress of the process after a construction time step. Of the three resources, material resources are consumed and require dispatch vehicles to dispatch them, so they have spatial attributes. Worker resources are not consumed, but different types of workers have different construction efficiencies and require dormitories for accommodation. Therefore, worker resources have both efficiency and spatial attributes. Reusable equipment resources are not consumed but require constructable space for storage. Therefore, they have spatial attributes. Due to the different characteristics of the three resources, the second sub-model processes the material resource ownership corresponding to material resources, the third sub-model processes the worker resource ownership corresponding to worker resources, and the fourth sub-model processes the reusable equipment resource ownership corresponding to reusable equipment resources. The material resource ownership is a tensor of M1×S, the worker resource ownership is a tensor of M2×S, and the worker resource ownership is a tensor of M3×S. Different models can extract data features of different categories, solving the problem of insufficient data boundaries that can be processed in existing technologies and achieving better solution results. After the first to fourth sub-models have determined the corresponding output data, all output data can be input into the main sub-model, which outputs the decision corresponding to the next construction time step. After calculating the single-step reward value based on the single-step reward function, the model parameters of the deep neural network model for the current construction time step are updated based on this single-step reward value.

[0101] Furthermore, the resource ownership in the construction status data includes material resource ownership, worker resource ownership, and reusable equipment resource ownership;

[0102] The method of acquiring construction status data corresponding to a current construction time step based on the deep neural network model, performing reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource demand information, wherein the main sub-model outputs a decision corresponding to the next construction time step, including:

[0103] When the current iteration round is less than or equal to the preset iteration round, obtaining the process completion progress and resource ownership corresponding to the current construction time step based on the deep neural network model;

[0104] When the process completion progress is less than the preset resource requirement, the process completion progress is input into the first sub-model and a process vector is output;

[0105] Inputting the material resource ownership into the second sub-model, the worker resource ownership into the third sub-model, and the reusable equipment resource ownership into the fourth sub-model, and outputting a resource vector;

[0106] The process vector and the resource vector are input into the main sub-model, and the decision of the next construction time step is output.

[0107] Specifically, FIG2 is a schematic diagram of the training process of the deep neural network model provided by the embodiment of the present application. As shown in FIG2, when training the deep neural network model, it is first determined whether the current iteration round epoch is less than or equal to the preset iteration round epochs. If the current iteration round epoch is greater than the preset iteration round epochs, it indicates that the iteration stop condition is met and the training of the deep neural network model is terminated. If the current iteration round epoch is less than or equal to the preset iteration round epochs, it indicates that the training of the deep neural network model is not terminated, and the construction status data of the current construction time step can be obtained. The construction status data state0 may include the process completion progress pr t and resource ownership res_exist t , and pr t =0,res_exist t =0, t=0.

[0108] Table 3: Architecture of the first sub-model

[0109] Afterwards, determine whether the process completion progress of the current construction time step is less than the preset resource demand, so as to determine whether the sample construction project has been completed. If the process completion progress of the current construction time step is less than the preset resource demand, it indicates that there is a surplus of resource demand, and it can be known that the processes under the current construction time step have not been completed. At this time, the acquired construction status data can be input into the deep neural network model to obtain the decision of the next construction time step, that is, the allocation amount of M kinds of resources is allocated to each process in the next construction time step. Figure 3 is a structural diagram of the deep neural network model provided in an embodiment of the present application. As shown in Figure 3, the process of determining the decision corresponding to the next construction time step includes the following steps:

[0110] 1) The first sub-model includes zero-padding layers, convolutional layers, pooling layers, and fully connected layers. After inputting the process completion progress into the first sub-model, the zero-padding layer pads the process completion progress to a tensor of [256, 256, 64]. Subsequently, through convolution, pooling, and fully connected operations, the final output is a process vector of length 512. The specific architecture of the first sub-model and the data size of each layer output are shown in Table 3.

[0111] In the fully connected layer of the first sub-model, the hyperbolic tangent function tanh is used as the activation function. The output value from the input layer to the first hidden layer can be calculated by formula (17): H1=tanh(W1s+b1)

[0112] Where H1 represents the output value from the input layer to the first hidden layer, s represents the input value of the first hidden layer, tanh(*) represents the hyperbolic tangent function, and Right now, W1 represents the weight corresponding to the first hidden layer, and b1 represents the bias value.

[0113] 2) Each model in the second to fourth sub-models includes 4 RNN hidden layers, 3 fully connected layers and 1 splicing layer. The first RNN hidden layer uses the rectified linear unit ReLU as the activation function and outputs a vector of length 64. The output value of the first RNN hidden layer can be calculated by formula (18), which is: h1=ReLU(W xh1 x t +W hh1 h1+bh1)

[0114] Among them, h1 represents the output value of the first RNN hidden layer, W xh1 Represents the weight matrix from the input layer to the hidden state, W hh1 represents the weight matrix from hidden state to hidden state, bh1 represents the bias vector of hidden state, W xh1 x t +W hh1 h1+bh1 represents the input vector of the first layer of neurons, which is used to determine the hidden state of the current construction time step. ReLU(*) represents the rectified linear unit, ReLU(θ)=max(0,θ), that is, ReLU(W xh1 x t +W hh1 h1+bh1)=max(0,W xh1 x t +W hh1 h1+bh1) is used to update the hidden state calculated in the current construction time step to the hidden state corresponding to the next construction time step, that is, to transfer information from the current construction time step to the next construction time step.

[0115] The second to fourth RNN hidden layers all use the rectified linear unit (ReLU) as the activation function, and each hidden layer in the second to fourth RNN hidden layers outputs a vector of length 128. Subsequently, the fully connected layer of each model in the second to fourth sub-models uses the rectified linear unit (ReLU) as the activation function and outputs a vector of length 128. The three 128-length vectors are then concatenated through a concatenation layer to obtain a vector of length 384. Subsequently, two fully connected layers are used, both using the rectified linear unit (ReLU) as the activation function, and each fully connected layer outputs a resource vector of length 256.

[0116] The 512-length process vector output by the first submodel and the 256-length resource vectors output by the second through fourth submodels are then fed into the main submodel. Through the concatenation layer within the main submodel, a 768-length vector is concatenated. Three fully connected layers are then used, all using rectified linear units (ReLUs) as activation functions: the first fully connected layer outputs a 768-length vector, the second a 512-length vector, and the third a 512-length vector. The output layer of the main submodel uses rectified linear units (ReLUs) as activation functions, outputting an N×M-length vector. This N×M-length vector is then reshaped and converted into an N×M matrix, which represents the decision.

[0117] Furthermore, updating the model parameters of the deep neural network model based on the single-step reward function includes:

[0118] If the decision satisfies a feasible strategy condition, the decision is executed, and based on the single-step reward function, a scheduling reward, a progress reward, and a cost reward corresponding to the execution of the decision are determined; the feasible strategy condition includes a storage space constraint condition and a construction space constraint condition;

[0119] Determining a single-step reward value corresponding to the decision based on the scheduling reward, the progress reward, and the cost reward;

[0120] Based on the single-step reward value, model parameters of the deep neural network model are updated.

[0121] Specifically, after determining the decision, when it is determined that the process completion progress of the current construction time step is less than the preset resource demand, and when it is determined that the decision output by the deep neural network model meets the feasible strategy conditions, the decision can be executed in the resource scheduling environment, that is, the decision is used as the input data in the resource scheduling environment, and after executing the decision using formula (19), the process observation data corresponding to the next construction time step is obtained, that is, the construction state data corresponding to the next construction time step. Formula (19) is: State t+1 =Transition(State t , Action t )

[0122] Among them, Transition(*) indicates the change information of the construction status data corresponding to the current construction time step after the decision is executed in the resource scheduling environment, and Action t Indicates the decision corresponding to the current construction time step, State t Represents the construction status data of the current construction time step, State t+1 It represents the process observation data corresponding to the next construction time step or the construction status data corresponding to the next construction time step.

[0123] When executing this decision, the project manager must, at the end of the current construction time step, determine the allocation of M resources to each process in the next construction time step. Specifically, if the resource allocations for this decision exceed the construction unit's existing resources at the next construction time step, the construction unit must immediately purchase, hire, or lease the remaining resources to ensure smooth construction in the next construction time step. Conversely, if the resource allocations for this decision are less than the construction unit's existing resources, the construction unit must immediately return or terminate the workers. Furthermore, workers and reusable equipment are reusable; already deployed workers and reusable equipment are released at the end of the current construction time step and are available for use in the next construction time step. Furthermore, while deploying more workers can speed up process execution, dormitory space is limited at the end of the current construction time step, and workers are paid based on workdays. Therefore, the project manager can terminate workers at any time. Similar to workers, reusable equipment resources can terminate leases for reusable equipment at any time. On the other hand, for consumable material resources, the material resources that have been invested in the construction cannot be used in the subsequent construction process; the material resources that have not been used up on the day will have a certain degree of loss, and the loss rate remains unchanged throughout the project. The unconsumed material resources can continue to be used in the next construction time step. When the warehouse has sufficient storage, due to the balance requirements of resources and the cost of scheduling, the project manager has the motivation to purchase material resources that will not be used on the day in advance. Based on the above resource utilization assumptions, formula (20) can be used to calculate the resource purchase quantity that needs to be purchased at the beginning of the next construction time step. Formula (20) is:

[0124] Among them, res k,i,t Indicates the usage of the k-th resource by the i-th process at the t-th construction time step or the next construction time step, res_exist k,t-1 It represents the resource ownership corresponding to the k-th resource at the t-1th construction time step or the end of the current construction time step, 1≤k≤M, and k is an integer.

[0125] Afterwards, to determine the corresponding resource ownership at the end of the next construction time step, it is also necessary to determine the resource consumption at the end of the next construction time step, that is, the process completion progress at the end of the current construction time step, which can also be understood as the process completion amount at the beginning of the next construction time step. There is a process that is affected by the supply of three resources during the completion process. Specifically, using formula (21), calculate the maximum construction progress allowed for the i-th process due to the k-th resource. The k-th resource belongs to the material resource. Formula (21) is: prf_mat k,i,t =res k,i,t

[0126] Among them, prf_mat k,i,t represents the maximum construction progress allowed for the $i$-th process at the current construction time step or the $t$-th construction time step due to the $k$-th material resource, res k,i,t represents the allocation amount of the $k$-th resource allocated to the $i$-th process at the current construction time step or the $t$-th construction time step in the decision-making, where $1\leq k\leq M1$ and $k$ is an integer.

[0127] Using Equation (22), calculate the maximum construction progress allowed for the $i$-th process due to the $k$-th resource. This $k$-th resource belongs to the worker resource. Equation (22) is:

[0128] Among them, represents the maximum construction progress allowed for the $i$-th process at the current construction time step or the $t$-th construction time step due to the $(k - M1)$-th worker resource, v k represents the construction efficiency of the $k$-th worker.

[0129] Using Equation (23), calculate the maximum construction progress allowed for the $i$-th process due to the $k$-th resource. This $k$-th resource belongs to the reusable equipment resource. Equation (23) is:

[0130] Among them, prf_equ k,i,t represents the maximum construction progress allowed for the $i$-th process at the current construction time step or the $t$-th construction time step due to the $k$-th reusable equipment resource, res k,i,t represents the allocation amount of the $k$-th reusable equipment resource allocated to the $i$-th process at the current construction time step or the $t$-th construction time step in the decision-making, where $1 < k\leq M1$ and $k$ is an integer.

[0131] In addition, if the preceding process of a certain process is not satisfied, normal construction cannot be carried out even if the resource availability is sufficient. Therefore, using Equation (24), calculate the supply progress of the preceding process of the $i$-th process. Equation (24) is:

[0132] Among them, $j$ represents the $j$-th preceding process, $i$ represents the $i$-th succeeding process. After the $j$-th preceding process is completed before the current construction time step or the $t$-th construction time step, the $i$-th succeeding process can be executed in the next construction time step or the $(t + 1)$-th construction time step. prf_prior k,i,t represents the supply progress of the preceding process of the $k$-th resource in the $i$-th succeeding process at the current construction time step or the $t$-th construction time step.

[0133] After that, Equation (25) can be used to calculate the consumption amount res_consume of the $k$-th material resource by the $i$-th process in the next construction time step according to the minimum value among the above four progressions k,i,t , Equation (25) is: res_consumek,i,t =min(prf_mat k,i,t ,prf_wor k,i,t ,prf_equ k,i,t ,prf_prior k,i,t )

[0134] Where 1≤k≤M1, and k is an integer, the consumption of the k-th material resource by the i-th process in the next construction time step is res_consume k,i,t , the process completion progress prf of the kth resource in the i-th process in the next construction time step can be determined k,i,t , that is, prf k,i,t =res_consume k,i,t .

[0135] After determining the process completion progress, according to the process completion progress prf of the kth resource of the i-th process in the next construction time step k,i,t The cumulative completed progress pr of the kth resource in the i-th process at the current construction time step k,i,t-1 The sum of the completed progress pr of the kth resource in the i-th process at the end of the next construction time step is determined k,i,t , that is, pr k,i,t =pr k,i,t-1 +prf k,i,t .

[0136] Afterwards, the resource ownership corresponding to the material resources at the end of the next construction time step can be calculated using formula (26), which is:

[0137] Among them, res_exist k,t Indicates the resource ownership at the end of the tth construction time step or the next construction time step, res_exist k,t-1 Indicates the resource ownership of the kth resource at the end of the t-1th construction time step or the current construction time step, delta_res k,t res_consume represents the resource purchase quantity of material resources that need to be purchased at the beginning of the tth construction time step or the next construction time step. k,i,t represents the resource consumption at the end of the tth construction time step or the next construction time step, δ k It represents the loss coefficient of the k-th resource. Since only material resources will be consumed, 1≤k≤M1, and k is an integer.

[0138] Using formula (27), calculate the resource ownership corresponding to the labor resources at the end of the next construction time step. Formula (27) is: res_exist k,t =res_exist k,t-1+delta_res k,t

[0139] where res_exist k,t represents the resource ownership corresponding to the labor resources at the end of the next construction time step or the t-th construction time step, res_exist k,t-1 represents the resource ownership of the k-th resource at the end of the current construction time step or the (t - 1)-th construction time step in the construction status data, and delta_res k,t is the resource employment volume of the worker resources to be hired at the start of the t-th construction time step or the next construction time step, M1 < k ≤ M1 + M2, and k is an integer.

[0140] The resource ownership corresponding to the labor resources at the end of the next construction time step can be calculated using Equation (28). Equation (28) is: res_exist k,t = res_exist k,t-1 +delta_res k,t

[0141] where res_exist k,t represents the resource ownership corresponding to the reusable equipment resources at the end of the next construction time step or the t-th construction time step, res_exist k,t-1 represents the resource ownership of the k-th resource at the end of the current construction time step or the (t - 1)-th construction time step in the construction status data, and delta_res k,t is the resource employment volume of the reusable equipment resources to be leased at the start of the t-th construction time step or the next construction time step, M 1+ M2 < k ≤ M, and k is an integer.

[0142] It should be noted that the above three resource ownerships are the process observation data corresponding to the next construction time step.

[0143] After that, Equation (3) can be used to calculate the scheduling reward corresponding to the execution of this decision, Equation (6) can be used to calculate the progress reward corresponding to the execution of this decision, and Equation (10) can be used to calculate the cost reward corresponding to the execution of this decision. After determining the scheduling reward, progress reward, and cost reward, Equation (11) can be used to calculate the single-step reward value corresponding to this decision. Based on this single-step reward value, the corresponding loss function value can be determined, and the model parameters of the deep neural network model can be updated according to this loss function value. Through the single-step reward function, the training of the deep neural network model can be accelerated to ensure that the decision corresponding to each construction time step is the optimal decision.

[0144] Optionally, the feasible strategy conditions include storage space constraint conditions and construction space constraint conditions. The judgment formula corresponding to the storage space constraint conditions can include: material warehouse space conditions and the spatial conditions of workers' dormitories If the decision satisfies both the material warehouse space condition and the worker dormitory space condition, it indicates that the decision satisfies the storage space constraint condition. The judgment formula corresponding to the construction space constraint condition can be: If both the storage space constraint and the construction space constraint are satisfied, the decision is a feasible strategy. If at least one constraint is not satisfied, the decision is an infeasible strategy.

[0145] Furthermore, as shown in FIG2 , the method further includes:

[0146] When the decision does not meet the feasible strategy conditions, the construction schedule of the current iteration round is determined to be ended, and based on the total project reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current construction time step are rolled back to the model parameters corresponding to the previous construction time step.

[0147] Specifically, when it is determined that the process completion progress of the current construction time step is less than the preset resource demand, and it is determined that the decision output by the deep neural network model does not meet the feasible strategy conditions, the T>T in formula (15) can be used. max Or if the decision is an infeasible strategy, the total reward value of the project is determined to be 0, and it is indicated that the decision planning of the next construction time step is wrong. Therefore, after determining the total reward value of the project, the model parameters corresponding to the current construction time step are rolled back to the model parameters corresponding to the previous construction time step, and the construction scheduling of the current iteration round is repeated.

[0148] Furthermore, as shown in FIG2 , the method further includes:

[0149] When the completion progress of the process is equal to the preset resource requirement, the construction schedule of the current iteration round is determined to be completed, the total project reward value of the current iteration round is determined based on the total project reward function, and the model parameters of the current iteration round are updated based on the total project reward value.

[0150] Specifically, if the process completion progress of the current time step is equal to the preset resource demand, it means that the resource consumption after the end of the current time step has reached the preset resource demand, indicating that the construction project to be tested has been completed in the current iteration cycle. The other cases in formula (15) can be used to calculate the total reward value of the project, and the model parameters of the current iteration round are updated according to the total reward value of the project.

[0151] Furthermore, determining the total reward value of the project in the current iteration round based on the total reward function of the project includes:

[0152] Based on the total project reward function, respectively determining the material resource fluctuation rate, worker resource fluctuation rate, and cost waste rate corresponding to the current iteration round;

[0153] Based on the material resource fluctuation rate, the worker resource fluctuation rate and the cost waste rate, a total reward value of the project corresponding to the current iteration round is determined.

[0154] Specifically, after the construction schedule of the current iteration round is determined, the material resource fluctuation rate can be calculated using formula (12), the worker resource fluctuation rate can be calculated using formula (13), and the cost waste rate can be calculated using formula (14). After determining the material resource fluctuation rate, worker resource fluctuation rate, and cost waste rate, the total project reward value can be calculated using formula (15).

[0155] Step 140: When the next construction time step is less than or equal to the construction period threshold, the single-step decision step is repeated. After the construction schedule is completed, the model parameters of the current iteration round are updated based on the total project reward function.

[0156] Specifically, if the next construction time step is less than or equal to the construction period threshold, it indicates that the construction time step does not exceed the fixed maximum construction period. The above process observation data can be determined as the construction status data corresponding to the next construction time step. The single-step decision step is repeated until the current iteration round is greater than the preset iteration round. The scheduling of the sample construction project is determined to be completed, and the material resource fluctuation rate is calculated using formula (12), the worker resource fluctuation rate is calculated using formula (13), and the cost waste rate is calculated using formula (14). After determining the material resource fluctuation rate, the worker resource fluctuation rate and the cost waste rate, the total project reward value can be calculated using formula (15). The model parameters of the current iteration round are updated according to the total project reward value.

[0157] Furthermore, as shown in FIG2 , the method further includes:

[0158] When the next construction time step is greater than the construction period threshold, the construction schedule is determined to be completed, and based on the project total reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current iteration round are rolled back to the model parameters corresponding to the previous iteration round.

[0159] Specifically, if the next construction time step is greater than the construction period threshold, it indicates that the sample construction project has failed. Therefore, using T>T in formula (15), maxOr if the decision is an infeasible strategy, the total reward value of the project is determined to be 0, and it indicates that all decision planning errors in the current iteration round. Therefore, after determining the total reward value of the project, the model parameters corresponding to the current iteration round are rolled back to the model parameters corresponding to the previous iteration round, and the construction scheduling of the current iteration round is repeated.

[0160] Step 150: traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters, obtain a trained construction scheduling model, and schedule the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to each construction time step of the target construction project.

[0161] Specifically, after the deep neural network model is trained based on any sample construction project, other sample construction projects can be traversed and the deep neural network model can be trained based on other sample construction projects. After the termination condition is met, a trained construction scheduling model is obtained.

[0162] Next, the project information and resource requirement information corresponding to the target construction project are obtained. Based on the construction scheduling model, the target construction project is scheduled with resource balance and scheduling efficiency as the optimization goals, resulting in a target scheduling strategy for the target construction project at each construction time step. It should be noted that the target construction project can be any one of the sample construction projects.

[0163] It should be noted that each of the above-mentioned iteration rounds represents the length of time it takes to complete the construction scheduling of the sample construction project. By iteratively training the deep neural network model for a preset number of iteration rounds, the model parameters in the deep neural network model are continuously adjusted so that the construction scheduling model after training can adapt to the sample construction project. If a new construction project is replaced, the deep neural network model needs to be retrained. Each construction time step can be daily, which is not limited in the embodiments of the present application.

[0164] Optionally, the server hardware platform used for training can include an Intel i9 13900k CPU, an NVIDIA GeForce RTX 4090 GPU, 128GB of DDR4 3600MHz memory, and Ubuntu 22.04LTS. The training software environment can include Pytorch 1.12 and Python 3.10. FP16 mixed-precision acceleration is used during training, and the total training time is approximately 5 days and 14 hours.

[0165] For example, the target construction project is to build a large sports stadium, and the reward weights ω1=0.7, ω2=0.25, ω3=0.05, ω4=0.35, ω5=0.65, ω6=0.3, Q1=1.1, Q2=1.1, Q3=1.8, epochs=55000 rounds, without involving batch size, and using Adam optimizer for training. The initial learning rate is 0.001, and the learning rate decay method is cosine annealing. For example, S=10 is set. Taking the 201st day to the 210th day as an example, the input data process completion progress of the first sub-model in the construction scheduling model is the process progress from the 201st day to the 210th day, and is expressed in terms of the demand for material resources. The process completion progress is a 118×8×10-dimensional tensor, and from the 201st day to the 210th day, the process completion progress of each process is monotonically increasing. The input data of the second to fourth sub-models are the resource holdings of the three resources from day 201 to day 210. The output data of the main sub-model is the decision at the end of day 210, which is a 118×21-dimensional vector.

[0166] Optionally, Figure 4 is a schematic diagram of the total project reward value as the number of iterations provided in an embodiment of the present application. After the deep neural network model training is completed, it can be seen from Figure 4 that the total project reward value converges at approximately 0.843550. After approximately 27,700 rounds of training, the total project reward value begins to gradually exceed 0.5; after approximately 35,600 rounds of training, the total project reward value begins to gradually exceed 0.8; after approximately 40,000 rounds of training, the total project reward value tends to converge. In the early stages of training, decisions often output infeasible strategies, so the total project reward value is often 0; in the later stages of training, the frequency of outputting infeasible strategies decreases. In the 40,001-55,000 rounds of training, the decision success rate of its output is 99.9734%, which can be considered that almost no infeasible strategies are output.

[0167] Optionally, Figure 5 is a schematic diagram of the change in the material resource procurement volume with the construction time step provided in the embodiment of the present application. As shown in Figure 5, after the training is completed, in the final decision obtained, the actual construction period of the project is 588 days. Due to the different dimensions of different resources, the resource procurement volumes in Figure 5 are presented in the form of relative values. In each construction stage, the corresponding required material procurement is relatively uniform. When switching between different processes, there is a small jump in the daily resource procurement volume; when switching between different construction stages, the daily resource procurement volume changes rapidly to a new level. The decision will choose to purchase in advance within a period of time before the actual construction begins to achieve a balance in resource procurement.

[0168] Alternatively, Figure 6 is a diagram illustrating how warehouse utilization varies over time, as provided in an embodiment of the present application. As shown in Figure 6 , during part of the mid-term period of the target construction project, warehouse utilization was close to 100%. Throughout the project, warehouse utilization exceeded 85% for 19.2% of the time. Concrete was only purchased and used on the same day, so there was no inventory.

[0169] When comparing all target scheduling strategies provided in the examples of this application with actual construction, the total project reward value corresponding to all target scheduling strategies provided in the examples of this application is 0.843550, while the total project reward value of the actual project is 0.662935, which is a 27.82% improvement over the actual project. The reward is divided into a resource balance component and a project cost component, with the resource balance component being the main indicator. In the resource balance component, the material balance indicator is 24.26% higher than the actual project, and the worker balance indicator is 35.66% higher than the actual project. This can effectively achieve resource balance, prevent project rushing, and thus improve project quality. In addition, in the project cost component, the material cost of this application is 1.24% higher than the actual project, and the rental cost of workers and reusable equipment is 11.82% lower than the actual project, ultimately resulting in a 0.79% reduction in the total cost of the target construction project compared to the actual project. The increase in material costs mainly comes from the small amount of waste caused by the decision-making of advance procurement, while the decision-making effectively reduces the rental costs of both workers and reusable equipment by rationally arranging the workers and reusable equipment. In addition, the total construction period provided by the embodiment of this application is 2.49% shorter than the actual construction period of the actual project and 3.76% shorter than the maximum construction period specified. The construction period is not the optimization target of this application, but by considering the construction efficiency of workers and rationally arranging personnel, it also contributes to the reduction of construction period.

[0170] The resource-balanced construction scheduling method based on deep reinforcement learning provided in this application obtains project information corresponding to each sample construction project, as well as resource demand information consisting of mappings between process steps, resource types, and resource demand quantities. With resource balance and scheduling efficiency as optimization goals, a single-step reward function and a project total reward function are constructed. Based on the resource demand information, the single-step reward function, and the project total reward function, a deep neural network model is constructed. First through fourth sub-models process different types of resource training data in the construction status data, respectively, and the main sub-model outputs the decision corresponding to the next construction time step. During training, the deep neural network model performs reinforcement learning through interaction with the environment, enabling the construction scheduling model to efficiently adapt to the complex data corresponding to the environment and the target construction project. When updating the model parameters of the deep neural network model, the single-step reward function is used to ensure that the decision corresponding to each construction time step is optimal, while the project total reward function is used to ensure that the decisions corresponding to all construction time steps, as a whole, meet the requirements of resource balance and scheduling efficiency. When the trained construction scheduling model is used to schedule the target construction project, resource balance is achieved among different resource types during the construction scheduling process, thereby improving the rationality of the construction scheduling. In addition, the process adopted in the embodiment of the present application is complete, comprehensive, and the target scheduling strategy for all construction time steps provided has high reference value. It fully considers factors such as the adjustability of resource allocation, the variability of construction period with resource allocation, and the huge differences in resource requirements of different processes. It has high practical value in actual applications.

[0171] The resource-balanced construction scheduling device based on deep reinforcement learning provided in this application is described below. The resource-balanced construction scheduling device based on deep reinforcement learning described below and the resource-balanced construction scheduling method based on deep reinforcement learning described above can be referenced to each other.

[0172] The present application also provides a resource-balanced construction scheduling device based on deep reinforcement learning. FIG7 is a schematic structural diagram of the resource-balanced construction scheduling device based on deep reinforcement learning provided by the present application. As shown in FIG7 , the resource-balanced construction scheduling device 700 based on deep reinforcement learning includes: an acquisition module 710, a construction module 720, a first update module 730, a second update module 740, and a scheduling module 750, wherein:

[0173] An acquisition module 710 is configured to acquire project information and resource requirement information corresponding to at least one sample construction project; the resource requirement information is used to represent a mapping relationship between processes, resource types, and resource requirements in the sample construction project;

[0174] A construction module 720 is configured to construct a single-step reward function and a project total reward function, respectively, with resource balance and scheduling efficiency as optimization objectives; and construct a deep neural network model based on the project information, the resource requirement information, the single-step reward function, and the project total reward function, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model, a third sub-model, and a fourth sub-model constructed based on a recurrent neural network, and a main sub-model constructed based on a deep neural network;

[0175] A first updating module 730 is configured to obtain, based on the deep neural network model, construction status data corresponding to the current construction time step, perform reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource demand information, so that the main sub-model outputs a decision corresponding to the next construction time step, and updates the model parameters of the deep neural network model based on the single-step reward function; the construction status data is used to represent the process completion progress and resource ownership corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model;

[0176] A second updating module 740 is configured to repeatedly execute the single-step decision-making step if the next construction time step is less than or equal to the construction duration threshold, and update the model parameters of the current iteration round based on the project total reward function after the construction schedule is completed;

[0177] The scheduling module 750 is used to traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters, obtain a trained construction scheduling model, and schedule the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to the target construction project at each construction time step.

[0178] The resource-balanced construction scheduling device based on deep reinforcement learning provided in the embodiments of the present application obtains project information corresponding to each sample construction project, as well as resource demand information consisting of mappings between process steps, resource types, and resource demand quantities. With resource balance and scheduling efficiency as optimization goals, it constructs a single-step reward function and a project total reward function. Based on the resource demand information, the single-step reward function, and the project total reward function, a deep neural network model is constructed. First through fourth sub-models process different types of resource training data in the construction status data, respectively, and the main sub-model outputs a decision corresponding to the next construction time step. During training, the deep neural network model performs reinforcement learning through interaction with the environment, enabling the construction scheduling model to efficiently adapt to the complex data corresponding to the environment and the target construction project. When updating the model parameters of the deep neural network model, the single-step reward function is used to ensure that the decision corresponding to each construction time step is optimal, while the project total reward function is used to ensure that the decisions corresponding to all construction time steps, as a whole, meet optimal resource balance and scheduling efficiency. When the trained construction scheduling model is used to schedule the target construction project, resource balance is achieved among different resource types during the construction scheduling process, thereby improving the rationality of the construction scheduling. In addition, the process adopted in the embodiment of the present application is complete, comprehensive, and the target scheduling strategy for all construction time steps provided has high reference value. It fully considers factors such as the adjustability of resource allocation, the variability of construction period with resource allocation, and the huge differences in resource requirements of different processes. It has high practical value in actual applications.

[0179] Optionally, the resource ownership in the construction status data includes material resource ownership, worker resource ownership, and reusable equipment resource ownership.

[0180] Optionally, the first updating module 730 is specifically configured to:

[0181] When the current iteration round is less than or equal to the preset iteration round, obtaining the process completion progress and resource ownership corresponding to the current construction time step based on the deep neural network model;

[0182] When the process completion progress is less than the preset resource requirement, the process completion progress is input into the first sub-model and a process vector is output;

[0183] Inputting the material resource ownership into the second sub-model, the worker resource ownership into the third sub-model, and the reusable equipment resource ownership into the fourth sub-model, and outputting a resource vector;

[0184] The process vector and the resource vector are input into the main sub-model, and the decision of the next construction time step is output.

[0185] Optionally, the first updating module 730 is specifically configured to:

[0186] If the decision satisfies a feasible strategy condition, the decision is executed, and based on the single-step reward function, a scheduling reward, a progress reward, and a cost reward corresponding to the execution of the decision are determined; the feasible strategy condition includes a storage space constraint condition and a construction space constraint condition;

[0187] Determining a single-step reward value corresponding to the decision based on the scheduling reward, the progress reward, and the cost reward;

[0188] Based on the single-step reward value, model parameters of the deep neural network model are updated.

[0189] Optionally, the second updating module 740 is further configured to:

[0190] When the decision does not meet the feasible strategy conditions, the construction schedule of the current iteration round is determined to be ended, and based on the total project reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current construction time step are rolled back to the model parameters corresponding to the previous construction time step.

[0191] Optionally, the second updating module 740 is further configured to:

[0192] When the completion progress of the process is equal to the preset resource demand, the construction schedule of the current iteration round is determined to be completed, and based on the total project reward function, the total project reward value of the current iteration round is determined, and based on the total project reward value, the model parameters of the current iteration round are updated.

[0193] Optionally, the second updating module 740 is further configured to:

[0194] Based on the total project reward function, respectively determining the material resource fluctuation rate, worker resource fluctuation rate, and cost waste rate corresponding to the current iteration round;

[0195] Based on the material resource fluctuation rate, the worker resource fluctuation rate and the cost waste rate, a total reward value of the project corresponding to the current iteration round is determined.

[0196] Optionally, the second updating module 740 is further configured to:

[0197] When the next construction time step is greater than the construction period threshold, the construction schedule is determined to be completed, and based on the project total reward function, the total project reward value of the current iteration round is determined to be 0, and the model parameters corresponding to the current iteration round are rolled back to the model parameters corresponding to the previous iteration round.

[0198] FIG8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in FIG8 , the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the resource-balanced construction scheduling method based on deep reinforcement learning.

[0199] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0200] On the other hand, the present application also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the resource-balanced construction scheduling method based on deep reinforcement learning provided by the above methods.

[0201] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the resource-balanced construction scheduling method based on deep reinforcement learning provided by the above-mentioned methods.

[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A resource-balanced construction scheduling method based on deep reinforcement learning, comprising: Obtaining project information and resource requirement information corresponding to at least one sample construction project; The resource requirement information is used to characterize the mapping relationship among processes, resource types, and resource requirements in the sample construction project; Taking resource balance and scheduling efficiency as optimization objectives, respectively constructing a single-step reward function and a total project reward function; based on the project information, the resource requirement information, the single-step reward function, and the total project reward function, constructing a deep neural network model, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model constructed based on a recurrent neural network, a third sub-model, and a fourth sub-model, and a main sub-model constructed based on a deep neural network; Based on the deep neural network model, obtaining construction status data corresponding to the current construction time step, and performing reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource requirement information, the main sub-model outputs a decision corresponding to the next construction time step, and updates the model parameters of the deep neural network model based on the single-step reward function; the construction status data is used to characterize the process completion progress and resource ownership corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model; When the next construction time step is less than or equal to the construction duration threshold, repeatedly execute the single-step decision step, and after the construction scheduling is completed, update the model parameters of the current iteration round based on the total project reward function; Traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters, obtain a trained construction scheduling model, and perform construction scheduling on the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to each construction time step of the target construction project.

2. The resource-balanced construction scheduling method based on deep reinforcement learning according to claim 1, wherein, The resource ownership in the construction status data includes material resource ownership, worker resource ownership, and reusable equipment resource ownership; Based on the deep neural network model, obtaining construction status data corresponding to the current construction time step, and performing reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource requirement information, the main sub-model outputs a decision corresponding to the next construction time step, including: When the current iteration round is less than or equal to the preset iteration round, based on the deep neural network model, obtaining the process completion progress and resource ownership corresponding to the current construction time step; When the process completion progress is less than the preset resource requirement, inputting the process completion progress into the first sub-model to output a process vector; Respectively inputting the material resource ownership into the second sub-model, inputting the worker resource ownership into the third sub-model, and inputting the reusable equipment resource ownership into the fourth sub-model to output a resource vector; Input the process vector and the resource vector into the main sub-model to output the decision for the next construction time step.

3. The resource-balanced construction scheduling method based on deep reinforcement learning according to claim 2, wherein, Updating the model parameters of the deep neural network model based on the single-step reward function includes: When the decision meets the feasible policy conditions, execute the decision, and based on the single-step reward function, respectively determine the scheduling reward, progress reward, and cost reward corresponding to the execution of the decision; the feasible policy conditions include storage space constraint conditions and construction space constraint conditions; Based on the scheduling reward, the progress reward, and the cost reward, determine the single-step reward value corresponding to the decision; Based on the single-step reward value, update the model parameters of the deep neural network model.

4. The resource-balanced construction scheduling method based on deep reinforcement learning according to claim 3, wherein, The method further includes: When the decision does not meet the feasible policy conditions, determine that the construction schedule for the current iteration round ends, based on the total project reward function, determine that the total project reward value for the current iteration round is 0, and roll back the model parameters corresponding to the current construction time step to the model parameters corresponding to the previous construction time step.

5. The resource-balanced construction scheduling method based on deep reinforcement learning according to claim 2, wherein, The method further includes: When the completion progress of the process is equal to the preset resource demand, determine that the construction schedule for the current iteration round ends, based on the total project reward function, determine the total project reward value for the current iteration round, and based on the total project reward value, update the model parameters for the current iteration round.

6. The resource-balanced construction scheduling method based on deep reinforcement learning according to claim 5, wherein, Determining the total project reward value for the current iteration round based on the total project reward function includes: Based on the total project reward function, respectively determine the material resource volatility, worker resource volatility, and cost waste rate corresponding to the current iteration round; Based on the material resource volatility, the worker resource volatility, and the cost waste rate, determine the total project reward value corresponding to the current iteration round.

7. The resource-balanced construction scheduling method based on deep reinforcement learning according to any one of claims 1-6, wherein, The method further includes: When the next construction time step is greater than the construction period threshold, determine that the construction schedule ends, and based on the total project reward function, determine that the total project reward value for the current iteration round is 0, and roll back the model parameters corresponding to the current iteration round to the model parameters corresponding to the previous iteration round.

8. A resource-balanced construction scheduling device based on deep reinforcement learning, comprising: An acquisition module for acquiring project information and resource demand information corresponding to at least one sample construction project; The resource demand information is used to represent the mapping relationship among processes, resource types, and resource demand quantities in the sample construction project; A construction module for respectively constructing a single-step reward function and a total project reward function with resource balance and scheduling efficiency as optimization objectives; based on the project information, the resource demand information, the single-step reward function, and the total project reward function, construct a deep neural network model, wherein the deep neural network model includes a first sub-model constructed based on a convolutional neural network, a second sub-model constructed based on a recurrent neural network, a third sub-model, and a fourth sub-model, and a main sub-model constructed based on a deep neural network; The first update module is configured to obtain construction status data corresponding to the current construction time step based on the deep neural network model, perform reinforcement learning on the deep neural network model based on the construction status data corresponding to the current construction time step, the project information, and the resource requirement information, output a decision corresponding to the next construction time step by the main sub-model, and update the model parameters of the deep neural network model based on the single-step reward function; the construction status data is used to characterize the progress of the completed processes and the resource possession corresponding to the current construction time step in the sample construction project, and different types of resource training data in the construction status data are respectively input into the first sub-model to the fourth sub-model; The second update module is configured to repeatedly execute the single-step decision step when the next construction time step is less than or equal to the construction duration threshold, and update the model parameters of the current iteration round based on the total project reward function after the construction scheduling ends; The scheduling module is configured to traverse each of the sample construction projects, repeatedly execute the step of updating the model parameters, obtain a trained construction scheduling model, and perform construction scheduling on the target construction project based on the construction scheduling model, and output the target scheduling strategy corresponding to each construction time step of the target construction project.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the resource-balanced construction scheduling method based on deep reinforcement learning according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by the processor, it implements the resource-balanced construction scheduling method based on deep reinforcement learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent production scheduling dynamic scheduling method based on deep reinforcement learning

    CN114154821A

  • Workshop scheduling method adaptive to machine state based on deep reinforcement learning

    CN114219274A

  • Prefabricated part production scheduling optimization method and system based on reinforcement learning

    CN115204497A

  • Scheduling optimization method combining production robustness and resource balance of fabricated components

    CN116663861A

  • Resource balance construction scheduling method, device and equipment based on deep reinforcement learning

    CN117634859A

Cited By

  • Project plan dynamic adjustment system and method based on reinforcement learning

    CN121032063A

  • Stadium building construction resource scheduling method and system based on large model

    CN121235419A

  • Cold rolling processing line production schedule optimization method based on knowledge graph

    CN122311821A

  • Remote wind and light new energy and flexibility resource sequential collaborative development time sequence decision-making method

    CN122453215A