Plan generation method, device and electronic device based on improved deep sub-Q network

By improving the deep subQ network method, the annual plan generation problem of the construction planning is broken down into single-objective optimization sub-problems, and the deep subQ network is trained through parameter migration strategies, the complexity of multi-objective sequential decision-making problems in the existing technology is solved, and accurate solution and decision support for dynamic multi-objective optimization are achieved.

CN118485239BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410548535.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-05-16
Estimated Expiration
2044-05-06

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the problem of multi-objective sequential decision-making in the generation of annual construction plans, especially when dealing with complex situations of dynamics, uncertainties and multi-objective optimization.

Method used

The improved deep subQ network (DQN) method is adopted to decompose the construction plan into multiple single-objective optimization subproblems, and the deep subQ network is trained through parameter migration strategies to solve the Pareto frontier that meets the constraints, thereby generating an annual plan.

Benefits of technology

It achieves accurate solution to dynamic multi-objective optimization problems, provides decision-making support, and improves the planning efficiency in the implementation of construction plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118485239B_ABST
    Figure CN118485239B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a plan generation method, device and electronic device based on an improved deep sub-Q network, the method comprising: obtaining a construction plan including the construction of multiple projects that need to be completed within a construction period; decomposing the construction plan into multiple objective optimization problems that need to be optimized through a decomposition strategy, and obtaining multiple corresponding sub-objective functions and constraints; applying a deep sub-Q network based on multiple objective functions to solve the Pareto front that meets the constraints; generating an annual plan for the construction plan based on the Pareto front. Through the above method, the embodiment of the present invention can accurately solve dynamic multi-objective optimization problems and provide decision support for plan compilation during the execution of construction planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of network technology, and in particular to a plan generation method, device and electronic device based on an improved deep sub-Q network. Background Art

[0002] The construction plan is a phased and systematic task deployment based on the development strategy concept, and is a medium- and long-term design for construction and development. The implementation of the construction plan mainly relies on construction projects. Through the preparation of annual plans, strengthening project management, organizing joint promotion, implementing process control, and conducting execution evaluation, the deployment of the construction plan is transformed into actual construction results. It can be seen that the preparation of the annual plan is the first step in the implementation of the construction plan and an important basis for the implementation of the construction plan. Specifically, in the process of planning implementation, according to the superior decision-making instructions and the actual situation and tasks, from the comprehensive perspective of project construction results and resource costs, based on the implementation of the annual plan in the previous year, the construction tasks, progress nodes, target indicators and funding resources of the planned construction projects in each year are rolled out. It is essentially a multi-objective sequential decision-making problem that needs to be solved by iterative optimization methods.

[0003] The general planning scheme generation problem mainly adopts the traditional project portfolio selection (PPS) method to obtain the scheme sequence by solving the planning model under the constraints. The preparation stage is based on the final value cost of each construction project and the final total value and total cost obtained by their aggregation. It focuses on static indicators and mainly solves the problem of "what to build". In the planning execution stage, its particularity is mainly reflected in the characteristics of "building and using" of construction projects, that is, some construction projects can pass the test after completing the annual plan and be put into use within a certain scale and scope, which helps to improve capabilities and contribute value. The preparation of the annual plan cannot only focus on the value and cost after the completion of the construction, and simply formulate it based on the status and evaluation of the project after the completion of the entire planning. Instead, it should start from the timeline of the entire planning cycle, focus on the dynamic changes of various indicators of the construction project as they are in the state, capture the dynamic characteristics of the cumulative cost and value changes in each stage caused by different project start times and changes in construction status, and mainly solve the problem of "how to build".

[0004] When abstracting the annual construction plan preparation task in the actual business process into the annual construction plan generation problem that can be described by mathematical models, the following three characteristics of the problem need to be considered: First, the solution to the problem has a sequential characteristic. Compared with the Project Portfolio Selection (PPS) solution, the annual plan not only considers the issue of whether to select or not select a project, but also the issue of start time and start order; second, the dynamic nature of the evaluation indicators. The different states of the project will directly affect the value cost and cumulative value cost of different stages. The objective function expression constructed based on this also has the dynamic characteristics of multi-stage recursion; third, the diversity of goals. The preparation of the annual plan is a multi-objective optimization problem under the comprehensive requirements of construction effectiveness and resource constraints. At present, the research on complex project planning problems mainly uses dynamic programming, heuristic algorithms, and optimization algorithms to solve them, but lacks consideration of the sequential characteristics of the problem solution. Reinforcement learning optimizes strategies through continuous trial and feedback based on learning experience, so that the system can maximize long-term benefits. It can handle dynamic and uncertain problems and has great reference significance and application potential in the preparation of annual construction plans. However, there is currently no research on the simultaneous processing of multiple goals. There is no relevant research on the generation of annual construction plans that has the characteristics of solution sequence, dynamic evaluation indicators, and diverse goals. Summary of the invention

[0005] In view of the above problems, the embodiments of the present invention provide a plan generation method, device and electronic device based on an improved deep sub-Q network, which overcome the above problems or at least partially solve the above problems.

[0006] According to one aspect of an embodiment of the present invention, a plan generation method based on an improved deep sub-Q network is provided, the method comprising: obtaining a construction plan including the construction of multiple projects that need to be completed within a construction period; decomposing the construction plan into multiple objective optimization problems that need to be optimized through a decomposition strategy, and obtaining corresponding multiple sub-objective functions and constraints; applying a deep sub-Q network based on multiple objective functions to solve a Pareto front that meets the constraints; and generating an annual plan for the construction plan based on the Pareto front.

[0007] Optionally, the decomposition strategy is used to decompose the construction plan into multiple objectives that need to be optimized, and corresponding multiple sub-objective functions and constraints are obtained, including: dividing the status of each project into three states: unbuilt, under construction, and built; decomposing the construction plan into a cumulative cost optimization problem and a cumulative value optimization problem through a decomposition strategy, and obtaining a first sub-objective function representing the minimization of the cumulative cost and a second sub-objective function representing the maximization of the cumulative value respectively; determining constraints based on the construction plan, the constraints including: the various costs generated by each project each year cannot exceed the funding ceiling of that year, each project can only be in one state each year, and there are mutually exclusive constraints between projects with a preceding relationship.

[0008] Optionally, decomposing the construction plan into a first sub-objective function representing minimization of cumulative cost and a second sub-objective function representing maximization of cumulative value through a decomposition strategy includes: obtaining the first objective sub-function according to the status of each project in each year during the execution of the construction plan and the cost associated with the status:

[0009]

[0010] Among them, S represents the state set of the project, I represents the project set, T represents the construction period, f1 represents the first objective sub-function, Indicates whether project i is in state s in year j, c ij represents the cost generated by project i in the jth year of the entire construction cycle; the second objective subfunction is obtained according to the status of each project in each year of the construction cycle during the execution of the construction plan and the value related to the status:

[0011]

[0012] Among them, v ij It represents the value generated by project i in the jth year of the entire construction cycle.

[0013] Optionally, before applying the deep sub-Q network based on the multiple objective functions to solve the Pareto front that meets the constraints, the method includes: defining an action according to the state change of any project during the execution of the construction plan, and forming an action set with all action sets of all projects; encoding the state of each project in each year during the construction cycle of the construction plan to generate a state matrix, and any item x in the state matrix ij Represents the status of project i in year j.

[0014] Optionally, the application of a deep sub-Q network based on multiple objective functions to solve a Pareto front that meets the constraints includes: decomposing the construction planning problem into N single-objective sub-problems, and determining the weight vector and the total objective function of each single-objective sub-problem; training the N deep sub-Q networks based on a parameter migration strategy to obtain network parameters of the deep sub-Q networks corresponding to the N sub-problems; solving the Markov decision process corresponding to each sub-problem using the trained deep sub-Q network based on multiple objective functions to obtain a Pareto front that meets the constraints.

[0015] Optionally, the method of training N deep sub-Q networks based on the parameter migration strategy to obtain network parameters of the deep sub-Q networks corresponding to the N sub-problems includes: starting from the first sub-problem, training the first deep sub-Q network from the initial parameters, and obtaining the trained network parameters of the first deep sub-Q network; training the next deep sub-Q network using the trained network parameters of the previous deep sub-Q network as the initial parameters of the next deep sub-Q network, and obtaining the trained network parameters of the next deep sub-Q network, until the training of all N deep sub-Q networks is completed.

[0016] Optionally, the deep sub-Q network trained based on multiple objective functions solves the Markov decision process corresponding to each subproblem to obtain a Pareto front that meets the constraints, including: for any subproblem, randomly selecting and executing an action from the action set to obtain a next state matrix; calculating a reward based on the current state matrix and the next state matrix based on multiple objective functions; traversing all actions in the action set, and if the current state matrix is ​​a terminal state, taking the current state matrix as a solution in the Pareto front; traversing all subproblems to obtain a Pareto front that meets the constraints.

[0017] Based on the same inventive concept, a plan generation device based on an improved deep sub-Q network is provided, including: a plan acquisition unit, used to obtain a construction plan including the construction of multiple projects that need to be completed within a construction period; a target constraint unit, used to decompose the construction plan into multiple target optimization problems that need to be optimized through a decomposition strategy, and obtain the corresponding multiple sub-objective functions and constraints; a solving unit, used to apply a deep sub-Q network based on multiple objective functions to solve the Pareto front that meets the constraints; an annual planning unit, used to generate an annual plan for the construction plan based on the Pareto front.

[0018] Based on the same inventive concept, an embodiment of the present invention further proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method when executing the program.

[0019] Based on the same inventive concept, an embodiment of the present invention further proposes a computer storage medium, in which at least one executable instruction is stored, and the executable instruction enables a processor to execute the aforementioned method.

[0020] The embodiment of the present invention obtains a construction plan including the construction of multiple projects that need to be completed within a construction period; decomposes the construction plan into multiple objective optimization problems that need to be optimized through a decomposition strategy, and obtains corresponding multiple sub-objective functions and constraints; applies a deep sub-Q network based on the multiple objective functions to solve the Pareto front that meets the constraints; generates an annual plan for the construction plan according to the Pareto front, which can accurately solve dynamic multi-objective optimization problems and provide decision support for plan preparation during the execution of the construction plan.

[0021] The above description is only an overview of the technical solution of the embodiment of the present invention. In order to more clearly understand the technical means of the embodiment of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0023] Figure 1 A schematic diagram of a flow chart of a plan generation method based on an improved deep sub-Q network provided in an embodiment of the present invention is shown;

[0024] Figure 2 A schematic diagram of a construction planning generation process according to an embodiment of the present invention is shown;

[0025] Figure 3 A state matrix and an action schematic diagram in an embodiment of the present invention are shown;

[0026] Figure 4 A schematic diagram of a decomposition strategy and sub-problems in an embodiment of the present invention is shown;

[0027] Figure 5 A schematic diagram of parameter migration according to an embodiment of the present invention is shown;

[0028] Figure 6 An example diagram of weight vector distribution corresponding to different decomposition methods in an embodiment of the present invention is shown;

[0029] Figure 7An example diagram of a training effect curve of a DSubQN model with different decomposition scales in an embodiment of the present invention is shown;

[0030] Figure 8 An example diagram showing the spatial distribution of non-dominated solution sets of different algorithms on different data sets in the embodiments of the present invention is shown;

[0031] Fig. 9 A schematic diagram of the structure of a plan generation device based on an improved deep sub-Q network provided by an embodiment of the present invention is shown;

[0032] Fig.10 A schematic diagram of an electronic device in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0033] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.

[0034] Figure 1 FIG. 2 is a flow chart showing a plan generation method based on an improved deep sub-Q network provided by an embodiment of the present invention. Figure 1 As shown, the plan generation method based on the improved deep sub-Q network is applied to the server, and the plan generation method based on the improved deep sub-Q network includes:

[0035] Step S11: Obtain a construction plan including multiple projects that need to be completed within the construction period.

[0036] In the embodiment of the present invention, there are various types of projects in the construction plan, among which projects that undertake construction of sites, facilities, systems, etc. have the characteristics of dynamic iteration and construction and use. Assuming that in the embodiment of the present invention, the construction implementation of the project has three states of "unbuilt-under construction-completed", then the annual plan generation problem of the construction plan can be described as follows: When executing the construction plan, dynamically consider the value and cost brought by the different construction conditions of the project each year, and reasonably arrange the progress of the project each year while satisfying various constraints, so as to achieve the relatively optimal value cost and cumulative value cost of the annual plan at each time node. Figure 2 As shown, the solution to the annual plan generation problem of construction planning can be composed horizontally of a separate construction plan for each project, or vertically of the corresponding plans for each year in the plan. Therefore, the solution generation process has the characteristics of both multi-objective combinatorial optimization and the dynamic characteristics of sequential decision-making.

[0037] In step S11, a construction plan to be completed within the construction period is obtained, and the construction plan includes multiple projects to be constructed. The construction period is in years, and the corresponding plan to be generated is an annual plan. Of course, in other embodiments of the present invention, the construction period is in months, and the corresponding plan to be generated is a monthly plan, which is not specifically limited here. The construction plan can be a military construction plan or a city construction plan, etc. The embodiment of the present invention takes the annual plan of the military construction plan as an example for explanation.

[0038] Step S12: Decomposing the construction plan into multiple objective optimization problems to be optimized through a decomposition strategy, and obtaining multiple corresponding sub-objective functions and constraints.

[0039] The variables and parameters involved in the embodiments of the present invention are mathematically expressed and described as shown in Table 1. The value ranges of the decision variables in Table 1 are as follows:

[0040]

[0041]

[0042] Table 1 Symbols and usage instructions

[0043]

[0044]

[0045] In the embodiment of the present invention, the status of each project is optionally divided into three states: unbuilt, under construction, and built. Then, the construction plan is decomposed into a cumulative cost optimization problem and a cumulative value optimization problem through a decomposition strategy, and a first sub-objective function representing minimization of cumulative cost and a second sub-objective function representing maximization of cumulative value are obtained respectively. The first objective sub-function is obtained according to the status of each project in each year during the execution of the construction plan and the cost associated with the status:

[0046]

[0047] Among them, S represents the state set of the project, I represents the project set, T represents the construction period, f1 represents the first objective sub-function, Indicates whether project i is in state s in year j, c ij represents the cost of project i in the jth year of the entire construction cycle. The first objective subfunction represents minimizing the cumulative cost during the execution of the construction plan. The first objective subfunction can be equivalent to the following recursive form:

[0048]

[0049] Among them, |T| represents the execution period of the construction plan.

[0050] The second objective sub-function is obtained according to the status of each project in each year during the construction period and the value related to the status during the execution of the construction plan:

[0051]

[0052] Among them, v ij It represents the value generated by project i in the jth year of the entire construction cycle. The second objective subfunction represents the cumulative value of maximizing the execution of the military construction plan. In order to make the optimization direction of the first objective subfunction consistent with that of the second objective subfunction, the cumulative value is negative. The second objective subfunction can be equivalent to the following recursive form:

[0053]

[0054] The embodiment of the present invention also determines constraints based on the construction plan, and the constraints include: cost constraints, rationality constraints and relationship constraints. The cost constraints limit the various costs generated by each project each year to not exceed the annual funding limit:

[0055]

[0056] The rationality constraint restricts each project to only be in one state each year:

[0057]

[0058] The rationality constraint also requires that once any project in the annual construction plan is started, it must be completed, regardless of the situation where it is abandoned halfway or partially built:

[0059]

[0060]

[0061] The rationality constraint also restricts the order of states, that is, when a project is completed, it will always be in the built state:

[0062]

[0063] Relationship constraints describe the mutually exclusive relationship constraints between projects with a close relationship in the construction plan:

[0064]

[0065] Relationship constraints also describe constraints between mutually exclusive items in the construction plan:

[0066]

[0067] Step S13: applying a deep sub-Q network based on the multiple objective functions to solve the Pareto front that meets the constraints.

[0068] In the embodiment of the present invention, the dynamic annual construction plan generation problem is considered, and the connectivity and rationality between the various stages are required to be higher, and the state of each stage is related to the state of the previous stage. In addition, the objective function of the problem can be written in a recursive form, which is a typical sequential decision problem. In this problem, the decision is made based on the current state, rather than the entire historical state sequence. The current state contains all decision-related information (project progress, current cumulative cost, current cumulative value), and the past state may be irrelevant to the current decision or can be completely summarized by the current state. In the mathematical model abstracted in this way, the current state does only depend on the state of the previous stage, and can be regarded as multiple parallel Markov decision processes. In view of this feature, it can be improved on the basis of the deep Q network (Deep Q-network, DQN), and a decomposition strategy can be adopted to obtain the sub-Q network (SubQnetwork) corresponding to multiple sub-problems. An improved DSubQN multi-objective optimization algorithm based on the SubQ network is proposed to solve the annual plan generation problem of construction planning.

[0069] Before step S13, optionally, the state change of any project during the execution of the construction plan is defined as an action, and the set of all actions of all projects forms an action set; the state of each project in each year of the construction cycle of the construction plan is encoded to generate a state matrix, and any item x in the state matrix ij Indicates the status of project i in year j. In the embodiment of the present invention, based on the characteristics of multi-objective dynamic optimization problems, the deep sub-Q network is combined to design states, actions and rewards. Figure 3 As shown, the state is represented as a matrix X of |I|×|T|, where 0 in the matrix represents that project i is not started in the jth year of the entire construction cycle, 1 represents that project i is under construction in the jth year of the entire construction cycle, and 2 represents that project i is completed in the jth year of the entire construction cycle. The state matrices at different stages can correspond to the construction plans at different stages, and the action set contains |I| actions, which represent changing the i-th project from unstarted to under construction. The reward R is expressed as the reward for state X at stage k. k Make a i After that, the objective function value changes in the optimization direction.

[0070] The basis for decision-making in each step of reinforcement learning and deep reinforcement learning methods is the reward of each step and the Q value fitted by the deep neural network, but both are scalars, so reinforcement learning and deep reinforcement learning are generally used for single-objective optimization problems. In the embodiment of the present invention, there are two objective functions for the problem of generating the annual plan of the military construction plan, and the dimensions are different, the optimization directions are opposite, and there is a conflict between the two objectives, which cannot be maximized or minimized at the same time. If the normalized weighted processing is a single reward value, this conflict relationship may be ignored. At the same time, since the decision maker's preferences may be uncertain or fuzzy in practical problems, it will bring difficulties to the weights in weighted processing, and a single reward value corresponds to a single solution, while the particularity of the problem of compiling the annual plan of the military construction plan requires retaining multiple high-quality solutions for selection. Therefore, considering the conflict between objectives, the uncertainty of preferences and the requirements of diverse solutions, it is not appropriate for the embodiment of the present invention to directly express the two objective functions as rewards. Therefore, it is necessary to make targeted improvements to the deep sub-Q network method to make it suitable for multi-objective optimization problems.

[0071] The embodiment of the present invention adopts a decomposition strategy to deal with multiple objectives in the problem. The core idea is to transform the multi-objective optimization problem into a series of single-objective optimization sub-problems, and then use the information of a certain number of adjacent problems to optimize these sub-problems simultaneously. After decomposing the multi-objective optimization main problem into a series of single-objective sub-problems, it is necessary to establish a sub-Q network for each sub-problem to speed up the solution of the main problem. The Q value represents the expectation of the total reward after the intelligent agent selects a certain action until the final state, and the weighted aggregation of the objective function can more intuitively reflect the reward obtained at each step in the Markov decision process.

[0072] In step S13, optionally, the construction planning problem is first decomposed into N single-objective sub-problems, and the weight vector and the total objective function of each single-objective sub-problem are determined. Specifically, given N uniformly distributed weight vectors λ1,λ2,...,λ N , where λ N =(λ n1 ,λ n2 ,…,λ nM ) T , N represents that the multi-objective dynamic optimization main problem is decomposed into N single-objective dynamic project cluster optimization sub-problems, and M represents the number of objective functions in the original multi-objective dynamic optimization main problem. Specifically, in the embodiment of the present invention, M=2. The nth sub-problem can be expressed as:

[0073]

[0074] like Figure 4As shown, the arrow vectors represent N uniformly distributed weight vectors λ1,λ2,...,λ that conform to the decision maker's preferences. N , the direction of the arrow represents the direction of sub-problem optimization, so an arrow vector corresponds to a decomposed sub-problem. The dotted line perpendicular to the weight vector is the contour line of the vector. n The target point with the smallest mapping length is the optimal solution to the subproblem, then the solution set composed of the optimal solutions of N subproblems is It can be approximated as the Pareto frontier, thus solving multi-objective dynamic optimization problems.

[0075] In an embodiment of the present invention, when solving the Pareto frontier, N deep sub-Q networks are first trained based on the parameter migration strategy to obtain the network parameters of the deep sub-Q networks corresponding to the N sub-problems. Since the multi-objective dynamic optimization main problem of generating the annual plan for the construction plan is decomposed into N single-objective dynamic optimization sub-problems, when the DQN model is used to solve the Markov decision process corresponding to each sub-problem, there will be a corresponding deep sub-Q network for approximating the Q value. For the N deep sub-Q networks corresponding to the N sub-problems, if each network is trained from the initial state to obtain the optimal parameters, the time required is too long, reducing the feasibility of its practical application. Figure 4 It can be seen that for two adjacent sub-problems, the weight vectors are similar and the optimal solutions are similar, that is, the expectation of the sum of the final state rewards of adjacent sub-problems is also similar. Based on this feature, the Q-value fitting process of each sub-problem in the reinforcement learning solution can be solved by referring to the relevant information of the successful fitting of the Q-value of its adjacent sub-problem. This idea is the idea of ​​transfer learning.

[0076] The embodiment of the present invention adopts the most basic parameter migration strategy, that is, the trained Q network model parameters are migrated from the current sub-problem to the next sub-problem in order according to the adjacent relationship of the sub-problems. Optionally, starting from the first sub-problem, the first deep sub-Q network is trained from the initial parameters to obtain the trained network parameters of the first deep sub-Q network; the trained network parameters of the previous deep sub-Q network are used as the initial parameters of the next deep sub-Q network to train the next deep sub-Q network, and the trained network parameters of the next deep sub-Q network are obtained until the training of all N deep sub-Q networks is completed. Figure 5 As shown, θ is used to represent the parameters of the deep sub-Q network that has not been trained, corresponding to all weights and biases, θ * It indicates that the network parameters of the deep sub-Q network have reached the optimal level after training, and θ nTo represent the network parameters of the deep sub-Q network for the nth sub-problem. The figure still uses long solid arrows to represent N uniformly distributed weight vectors λ1,λ2,...,λ that meet the decision maker's preferences. N , each vector corresponds to a decomposed sub-problem, and the dotted arrows represent the iterative process of the parameters in the whole process. The whole iterative training process can be described as follows: starting from the first sub-problem, the sub-Q (SubQ) network starts training with the initial parameters θ1. After the training, its network parameters have been trained and updated to near the optimal. Then, the initial parameters θ2 of the second sub-problem SubQ network model can be set as the optimal network parameters of the first sub-problem. Start training and update parameters based on the previous subproblem, and so on, until the SubQ network parameters corresponding to all N subproblems are obtained. That is, the SubQ network parameters of the neighboring subproblem are transferred to the next adjacent subproblem, and the relevant information of the successful fitting of the Q value of the adjacent subproblem is used to reduce the training time of the model.

[0077] After completing the training of all deep sub-Q networks, the Markov decision process corresponding to each sub-problem is solved based on the application of the trained deep sub-Q network based on multiple objective functions to obtain a Pareto front that meets the constraints. Optionally, for any sub-problem, randomly select and execute an action from the action set to obtain the next state matrix; calculate the reward based on the current state matrix and the next state matrix based on multiple objective functions; traverse all actions in the action set, and if the current state matrix is ​​a terminal state, use the current state matrix as a solution in the Pareto front; traverse all sub-problems to obtain a Pareto front that meets the constraints. The following is the specific process of the algorithm:

[0078]

[0079]

[0080] The specific description of the algorithm is as follows:

[0081] Step 1: Decompose the main problem into N sub-problems and obtain the weight vector λ1,λ2,...,λ corresponding to each sub-problem N And the objective function g ws (X|λ n ). Let the first subproblem be the current subproblem n, and the first year be the current stage j.

[0082] Step 2: Initialize SubQ Network Q n If n = 1, the SubQ network parameters are randomly initialized to θ n, if n>1, then according to the parameter migration strategy, the target SubQ network of the previous subproblem is The optimal parameters Set to the SubQ network Q of this subproblem n The initial parameter θ n . Initialize the target SubQ network Parameters Initialize experience pool

[0083] Step 3: At the current stage in the current subproblem, record the current state matrix as According to the ∈-greedy strategy, a random action is chosen with probability ∈ and an action is chosen with probability 1-∈ The selected action is denoted as α k .

[0084] Step 4: Execute action α k , observe the environment, adjust the corresponding project to the under-construction state, and get the next step status Get instant rewards

[0085] Step 5: Stored in experience pool middle.

[0086] Step 6: From the Experience Pool Randomly sample (X′, α′, R′, X″) in, where X′ represents the current state matrix, α′ represents the currently executed action, R′ represents the reward obtained by executing the current action, and X″ represents the state matrix after executing the current action. Set:

[0087]

[0088] (y′-Q n (X′,α′)) 2 The loss function is used to train the SubQ network Q n And update the network parameters θ n .

[0089] Repeat steps 3 to 6 for a total of K times. If the current stage is the last year of the construction cycle, that is, is the terminal state, then add the current state matrix to the Pareto frontier x * , the current target SubQ network Parameters Recorded as The optimal parameters Record the next subproblem as the current subproblem and go to step 2. Otherwise, record the next year in the construction cycle as the current stage and go to step 3.

[0090] Step S14: Generate an annual plan for the construction plan based on the Pareto front.

[0091] The Pareto front includes N solutions, and the corresponding construction plan can have N high-quality solutions. One of the solutions can be selected according to the actual situation to determine the status of each project in each year during the construction cycle, the actions to be performed, and then the annual plan can be generated, and the military construction can be carried out according to the generated annual plan.

[0092] The following examples illustrate the experimental results of the plan generation method based on the improved deep sub-Q network in an embodiment of the present invention. By fitting the data set statistics of all major investment planning projects approved between 1996 and 2012, five example data sets of military construction projects with sizes of 100, 200, 300, 400 and 500 were constructed, corresponding to data sets 1, 2, 3, 4 and 5 respectively. The embodiment of the present invention uses data set 1 containing 100 projects to demonstrate the training effect. The larger data sets 2-5 with a scale of 200-500 are used to compare and analyze the various performances of the deep sub-Q network (DSubQN) algorithm and other algorithms. According to the designed deep sub-Q network algorithm, the problem is decomposed into three different scales, namely 20 sub-problems, 50 sub-problems and 100 sub-problems, and the weight vectors corresponding to the three decomposition methods are respectively Then there is

[0093]

[0094] The weight vector distribution corresponding to each decomposition method is as follows Figure 6 As shown, the corresponding algorithms are recorded as DSubQN-20, DSubQN-50 and DsubQN-100. The basic parameters of the SubQ network model used in the algorithm are unified, and the specific settings are shown in Table 2.

[0095] Table 2DsubQN model parameter settings

[0096]

[0097] For the SubQ networks in the model, all DSubQ networks are trained using the gradient descent-based Adam optimizer, with the learning rate set to 0.01 and the epoch set to 10000. In particular, the weights of the first subproblem are randomly initialized, and for the training of the deep SubQ networks corresponding to the second and subsequent subproblems, the initial weights are initialized using a neighborhood-based parameter migration strategy. Representative algorithms in various fields are selected for comparison with the proposed deep SubQ optimization algorithm, including the second-generation non-dominated sorting genetic algorithm (NSGA-II) and the multi-objective evolutionary algorithm based on decomposition (MOEA / D). NSGA-II has been successfully applied in the optimization and decision-making of defense project capability combinations. Compared with NSGA-I and NSGA-III, the NSGA-II algorithm uses a congestion method to protect the diversity of solutions. This method does not require artificial setting of additional parameters, while NSGA-I and NSGA-III need to set the shared radius and the number of target segments respectively. These two parameters have an important impact on the diversity protection mechanism of the solution. Different values ​​will have a greater impact on the solution. Therefore, the more stable NSGA-II is selected as the comparison algorithm. MOEA / D is a classic heuristic method that uses a decomposition strategy to solve multi-objective optimization problems. It can be used as a suitable comparison for the embodiment of the present invention to improve the DQN algorithm based on the decomposition strategy to solve multi-objective optimization problems. The parameter settings of the two reference models are shown in Table 3.

[0098] Table 3 Parameter settings

[0099]

[0100]

[0101] The embodiment of the present invention mainly analyzes the effects and time of model training of DSubQN-20, DSubQN-50 and DSubQN-100 on a dataset 1 of size 100, and conducts comparative analysis to verify the learning ability of the model training stage proposed in this study.

[0102] Since DSubQN-20, DSubQN-50 and DSubQN-100 decompose the problem into 20, 50 and 100 sub-problems respectively, the corresponding SubQ networks have 20, 50 and 100 respectively. The SubQ networks corresponding to each problem must be trained in this way. The number is too large and the training effect cannot be displayed intuitively at the same time. Therefore, the first, last and middle sub-problems decomposed by each DSubQN are selected, and the corresponding weight vectors are (1,0) T 、(0.5,0.5) T and (0,1) T , that is, the training data of SubQ-20 on the 1st, 11th and 20th sub-problems, the training data of SubQ-50 on the 1st, 26th and 50th sub-problems, and the training data of SubQ-100 on the 1st, 51st and 100th sub-problems are used as representatives. Figure 7 The visualization in Figure 2 compares the model convergence performance of the DSubQN algorithm with three decomposition scales during the training process. The convergence performance of the model is measured by the deviation indicator e.

[0103]

[0104] The deviation index e defines the quality of a solution. f1(x) and f2(x) represent the objective function values ​​corresponding to the decision variable set x. * ) and f2(x * ) represents the decision variable set x corresponding to the optimal solution among all feasible solutions obtained according to the decomposed sub-problems * The corresponding objective function value. Obviously, the value range of e is [0,100]. When e=0, it means that the decision variable set x is the best solution among all known feasible solutions, and the quality of the solution decreases as the value of e increases. It should be noted that e=0 does not mean that the solution must be the optimal solution in the entire solution space, because the number of decomposed sub-problems is limited, and it is very difficult to find the optimal solution in the entire solution space. However, when comparing the relative advantages and disadvantages of several algorithms, the deviation index e can meet the needs. Another benefit of introducing the deviation index e is that this processing method makes the dominant direction of the solution on the horizontal and vertical axes on the coordinate axis the minimum direction, that is, the direction of the origin, which makes it easier to observe the relationship between the advantages and disadvantages of the solutions, which will be used in the next section.

[0105] like Figure 7 As shown, Figure 7 (a) shows the training effect curves of DSubQN-20, DSubQN-50 and DSubQN-100 on the first sub-problem; Figure 7(b) shows the training effect curves of DSubQN-20 on the 11th sub-problem, DSubQN-50 on the 26th sub-problem, and DSubQN-100 on the 51st sub-problem; Figure 7 (c) shows the training effect curves of DSubQN-20 on the 20th sub-problem, DSubQN-50 on the 50th sub-problem, and DSubQN-100 on the 100th sub-problem. Figure 7 (a) It can be seen that in the first sub-problem, the convergence performance of the model is almost the same. Only due to the different randomly selected initial solutions, the convergence curve fluctuates in a small range. This is because the weight vectors of the first sub-problem decomposed by the three models are equal to the weight vector of sub-problem 1. Therefore, the weights of the objective function are the same, the initial parameters of the SubQ network are the same, and the convergence curves of the training results are also the same. Figure 7 In (b), the three models gradually show differences. Compared with DSubQ-20 and DSubQ-50, the convergence speed of the DSubQN-100 model begins to dominate. Thanks to the high-quality random initial solution (the lower starting point in the curve), the convergence speed of DSubQ-50 is slightly faster than that of DSubQ-20 in the early stage. This is because the weights of subproblem 2 correspond to the 11th, 26th, and 51st subproblems decomposed by the three models, respectively. The number of parameter migrations is different, so the convergence performance of the model is also different. This difference is in Figure 7 (c) is more obvious, although compared to Figure 7 The convergence speed of the three models in (a) has been accelerated, but the convergence speed of DSubQN-100 has been improved more. It can be seen that the model converges faster for the sub-problems at the end of the sequence, because the migration strategy gives the initial parameters and the gap between adjacent problems is not large, so the convergence speed is fast, which also shows that the migration strategy improves the model training speed.

[0106] Since all three models can eventually obtain stable convergence ability and acceptable convergence speed through training, the decomposition scale is not the larger the better. For each sub-problem, the number of iterations is set to 10,000. For the three models, 2×10 iterations are required for each training. 5 , 5×10 5 and 10 6 times. Therefore, the larger the scale of the decomposed subproblem, the longer the training time is required. However, judging from the final stable deviation index range of the three figures, the final training effect does not increase in direct proportion to the increase in the scale of the decomposed subproblem. The scale of the decomposed subproblem of DSubQN-100 is doubled compared to DSubQN-50, but the final convergence effect is only slightly better than DSubQN-50.

[0107] After completing the training of the SubQ network models in the three DSubQN algorithms on the dataset 1 of scale 100, the trained models are used to solve the annual construction plan generation problems corresponding to the datasets 2, 3, 4 and 5 of scales 200, 300, 400 and 500 respectively. Figure 8 The Pareto frontiers obtained by MOEA / D, NSGA-II and three DSubQN algorithms are intuitively displayed using different scattered points in the two-dimensional coordinate axis. Figure 8 (a)- Figure 8 (d) shows the experimental results of datasets 2, 3, 4, and 5. In order to more conveniently observe the dominance relationship, the algorithm performance and the spatial position of the marked solution are measured according to the cost deviation e1 and value deviation e2 in the previous section.

[0108] like Figure 8 As shown in (a), although dataset 2 is the smallest of the four datasets, the non-dominated solution set obtained by the proposed DSubQN algorithm has obviously dominated the non-dominated solution sets obtained by the MOEA / D algorithm and the NSGA-II algorithm. At the same time, on dataset 2, it can be found that increasing the decomposition scale of the subproblem also makes the DSubQN algorithm show better convergence.

[0109] like Figure 8 As shown in (b), (c) and (d), on other large-scale examples, the MOEA / D algorithm and the NSGA-II algorithm are still difficult to converge. For the three DSubQN algorithms, as the problem size increases, blindly increasing the decomposition scale of the subproblem begins to produce negative benefits on the improvement of algorithm convergence. Specifically, in the three examples, the non-dominated solution set of the DSubQN-50 algorithm dominates the non-dominated solution set of the DSubQN-20 algorithm, which means that increasing the decomposition scale of the subproblem from 20 to 50 has a positive benefit, but in Figure 8 In (b), the curves corresponding to the non-dominated solution set obtained by the DSubQN-50 algorithm and the curves corresponding to the non-dominated solution set obtained by the DSubQN-100 algorithm are intertwined, making it difficult to judge the convergence simply from the picture. Figure 8 In (c), a significant portion of the solutions in the set of non-dominated solutions obtained by the DSubQN-50 algorithm already dominate the non-dominated solutions obtained by the DSubQN-50 algorithm. Figure 8In (d), the convergence of the DSubQN-50 algorithm is almost completely better than that of the DSubQN-100 algorithm. At this time, increasing the decomposition scale of the subproblem from 50 to 100 has a negative benefit in terms of convergence. However, it should be noted that in the examples corresponding to the four data sets, the diversity shown by the DSubQN-100 algorithm is the best, which is also related to the characteristics of the DSubQN algorithm. Generally speaking, the larger the decomposition scale of the subproblem, the better the diversity of the algorithm.

[0110] In addition, the hypervolume (HV) metric and running time are used to quantitatively evaluate the convergence and diversity of the algorithm as Figure 8 As a supplement, the data pair consisting of the largest cost deviation and the largest value deviation is selected as the reference point. In this way, the larger the hypervolume index formed by the non-dominated solution set obtained by each algorithm and the reference point, the better. The specific results and solution time are shown in Table 4.

[0111] Table 4 Experimental results of each algorithm on different datasets

[0112]

[0113] It should be noted that the time corresponding to the three DSubQN algorithms listed in Table 4 includes the training time of the algorithm, that is, the time in the table is the sum of the training time and the solution time. However, as a reusable model, DSubQN can repeatedly solve multiple data set problems of the same type with only one training, so the actual solution time will be less than the data in the table. The HV index listed in Table 4 confirms the analysis results of the algorithm performance in the previous article. DSubQN-100 only obtains the optimal HV index value on data set 2, and the HV index values ​​obtained on data sets 3, 4 and 5 are all smaller than DSubQN-50. And because the decomposition scale of DSubQN-100 is too large, its total time consumption has always been the largest. DSubQN-20 and DSubQN-50 can effectively control the running time of the algorithm by reasonably controlling the scale of the decomposed sub-problems, and both can obtain better results than traditional multi-objective optimization methods.

[0114] Therefore, in the actual application of the annual plan of military construction planning, it is necessary to set a suitable decomposition scale of the problem. It is not blindly decomposing the larger the better. It is necessary to find a suitable critical point and comprehensively consider diversity, convergence and running time. Appropriate control of the scale of decomposed sub-problems may achieve better results in some examples. This requires decision makers to grasp the balance of all parties involved and comprehensively consider and determine the plan. In addition, it is necessary to pay attention to the generalization and consistency of the problem, data and algorithm design. Although the algorithm is a reusable model, if the same type of problem or data cannot be guaranteed, the model needs to be retrained each time it is solved, which increases the time cost. This requires decision makers to maintain the continuity and consistency of their work.

[0115] The plan generation method based on the improved deep sub-Q network of the embodiment of the present invention can solve the multi-objective dynamic optimization problem with a recursive objective function. The multi-objective dynamic optimization main problem is decomposed into multiple single-objective dynamic optimization sub-problems through a decomposition strategy, so that the DQN model is used to solve the Markov decision process corresponding to each sub-problem. The objective function is expressed in a recursive form, and the state is encoded into a matrix representation of the construction plan through the annual construction plan optimization model of the project cluster based on dynamic programming, so that the model can directly output alternative construction plans. The embodiment of the present invention also adopts the idea of ​​transfer learning in the training process of the model to shorten the training time of the model. The experimental results of the example study part show the good performance of the plan generation method based on the improved deep sub-Q network of the embodiment of the present invention compared with the traditional multi-objective optimization algorithm and dynamic programming method.

[0116] In order to make the scale of the action space controllable and the algorithm's solution time more stable, the embodiments of the present invention can also refer to the pre-processing clustering method, hierarchically cluster the action space, and recommend actions based on the tree-structured policy network to achieve generalization processing for extremely large-scale action spaces.

[0117] In summary, the plan generation method based on the improved deep sub-Q network in the embodiment of the present invention obtains a construction plan that needs to be completed within the construction period, including the construction of multiple projects; decomposes the construction plan into multiple objective optimization problems that need to be optimized through a decomposition strategy, and obtains corresponding multiple sub-objective functions and constraints; applies a deep sub-Q network based on multiple objective functions to solve the Pareto front that meets the constraints; generates an annual plan for the construction plan based on the Pareto front, which can accurately solve dynamic multi-objective optimization problems and provide decision support for plan preparation during the execution of the construction plan.

[0118] The above specific embodiments of the present invention are described. In some cases, the actions or steps recorded in the embodiments of the present invention can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0119] Based on the same concept, an embodiment of the present invention also provides a plan generation device based on an improved deep sub-Q network. Applied to a server. Fig. 9 As shown, the plan generation device based on the improved deep sub-Q network includes: a plan acquisition unit, a target constraint unit, a solution unit and an annual planning unit.

[0120] A planning acquisition unit is used to acquire the construction plan including the construction of multiple projects that need to be completed within the construction period;

[0121] The target constraint unit is used to decompose the construction plan into multiple target optimization problems to be optimized through a decomposition strategy, and obtain multiple corresponding sub-target functions and constraint conditions;

[0122] A solving unit, configured to solve a Pareto front that meets the constraint conditions by applying a deep sub-Q network based on the multiple objective functions;

[0123] The annual planning unit is used to generate an annual plan for the construction plan based on the Pareto front.

[0124] For the convenience of description, the above device is described as various modules according to their functions. Of course, when implementing the embodiment of the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0125] The device of the above embodiment is applied to the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0126] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of the above embodiments is implemented.

[0127] An embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the method described in any of the above embodiments.

[0128] Fig.10A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1001, a memory 1002, an input / output interface 1003, a communication interface 1004, and a bus 1005. The processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are connected to each other through the bus 1005 in the device.

[0129] The processor 1001 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided by the method embodiment of the present invention.

[0130] The memory 1002 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1002 can store an operating system and other application programs. When the technical solution provided by the method embodiment of the present invention is implemented by software or firmware, the relevant program code is stored in the memory 1002 and called and executed by the processor 1001.

[0131] The input / output interface 1003 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0132] The communication interface 1004 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0133] The bus 1005 includes a path that transmits information between various components of the device (eg, the processor 1001 , the memory 1002 , the input / output interface 1003 , and the communication interface 1004 ).

[0134] It should be noted that, although the above device only shows the processor 1001, the memory 1002, the input / output interface 1003, the communication interface 1004 and the bus 1005, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present invention, and does not necessarily include all the components shown in the figure.

[0135] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0136] This application is intended to cover all such substitutions, modifications and variations that fall within the broad scope of all embodiments. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of this disclosure.

Claims

1. A plan generation method based on an improved deep sub-Q network, characterized in that: The method comprises: Obtain the construction plan including multiple projects that need to be completed within the construction period; Decomposing the construction plan into multiple target optimization problems to be optimized through a decomposition strategy, and obtaining corresponding multiple sub-objective functions and constraints, including: dividing the status of each project into three states: unbuilt, under construction, and built; decomposing the construction plan into a cumulative cost optimization problem and a cumulative value optimization problem through a decomposition strategy, and obtaining a first sub-objective function representing minimization of cumulative cost and a second sub-objective function representing maximization of cumulative value respectively; determining constraints according to the construction plan, the constraints including: various costs generated by each project each year cannot exceed the funding ceiling of the year, each project can only be in one state each year, and there are mutually exclusive constraints between projects with a close predecessor relationship; The decomposition strategy is used to decompose the construction plan into a cumulative cost optimization problem and a cumulative value optimization problem, and a first sub-objective function representing minimization of the cumulative cost and a second sub-objective function representing maximization of the cumulative value are obtained respectively, including: The first sub-objective function is obtained according to the status of each project in each year during the execution of the construction plan and the cost associated with the status: Among them, S represents the state set of the project, I represents the project set, T represents the construction period, f1 represents the first sub-objective function, Indicates whether project i is in state s in year j, c ij represents the cost incurred by project i in the jth year of the entire construction period; The second sub-objective function is obtained according to the status of each project in each year during the construction period during the execution of the construction plan and the value related to the status: Among them, v ij represents the value generated by project i in the jth year of the entire construction cycle; Based on the multiple objective functions, a deep sub-Q network is applied to solve a Pareto frontier that meets the constraints, including: decomposing the construction planning problem into N single-objective sub-problems, and determining the weight vector and the total objective function of each single-objective sub-problem; training the N deep sub-Q networks based on a parameter migration strategy to obtain network parameters of the deep sub-Q networks corresponding to the N sub-problems; based on the multiple objective functions, the deep sub-Q network trained is applied to solve the Markov decision process corresponding to each sub-problem to obtain a Pareto frontier that meets the constraints; An annual plan for the construction planning is generated based on the Pareto front.

2. The method according to claim 1, characterized in that Before applying the deep sub-Q network based on the multiple objective functions to solve the Pareto front that meets the constraint conditions, the method includes: A state change of any project during the execution of the construction plan is defined as an action, and the set of all actions of all projects forms an action set; According to the status of each project in each year of the construction cycle of the construction plan, a status matrix is ​​generated. Any item x in the status matrix ij Represents the status of project i in year j.

3. The method according to claim 2, characterized in that The method of training N deep sub-Q networks based on the parameter migration strategy to obtain network parameters of the deep sub-Q networks corresponding to the N sub-problems includes: Starting from the first sub-problem, train the first deep sub-Q network from the initial parameters, and obtain the network parameters after the first deep sub-Q network is trained; The trained network parameters of the previous deep sub-Q network are used as the initial parameters of the next deep sub-Q network to train the next deep sub-Q network, and the trained network parameters of the next deep sub-Q network are obtained until the training of all N deep sub-Q networks is completed.

4. The method according to claim 2, characterized in that: The deep sub-Q network trained based on the application of the multiple objective functions solves the Markov decision process corresponding to each sub-problem to obtain a Pareto front that meets the constraint conditions, including: For any sub-problem, randomly select and execute an action from the action set to obtain the next state matrix; A reward calculated based on the multiple objective functions according to the current state matrix and the next state matrix; Traverse all actions in the action set, and if the current state matrix is ​​a terminal state, take the current state matrix as a solution in the Pareto frontier; Traverse all subproblems and obtain the Pareto front that meets the constraints.

5. A plan generation device based on an improved deep sub-Q network, characterized in that: The device comprises: A planning acquisition unit is used to acquire the construction plan including the construction of multiple projects that need to be completed within the construction period; The target constraint unit is used to decompose the construction plan into multiple target optimization problems to be optimized through a decomposition strategy, and obtain corresponding multiple sub-objective functions and constraint conditions, including: dividing the status of each project into three states: unbuilt, under construction, and built; decomposing the construction plan into a cumulative cost optimization problem and a cumulative value optimization problem through a decomposition strategy, and respectively obtaining a first sub-objective function representing minimization of cumulative cost and a second sub-objective function representing maximization of cumulative value; determining constraint conditions according to the construction plan, the constraint conditions including: various costs generated by each project each year cannot exceed the funding ceiling of the year, each project can only be in one state each year, and there are mutually exclusive constraints between projects with a close predecessor relationship; The decomposition strategy is used to decompose the construction plan into a cumulative cost optimization problem and a cumulative value optimization problem, and a first sub-objective function representing minimization of the cumulative cost and a second sub-objective function representing maximization of the cumulative value are obtained respectively, including: The first sub-objective function is obtained according to the status of each project in each year during the execution of the construction plan and the cost associated with the status: Among them, S represents the state set of the project, I represents the project set, T represents the construction period, f1 represents the first sub-objective function, Indicates whether project i is in state s in year j, c ij represents the cost incurred by project i in the jth year of the entire construction period; The second sub-objective function is obtained according to the status of each project in each year during the construction period during the execution of the construction plan and the value related to the status: Among them, v ij represents the value generated by project i in the jth year of the entire construction cycle; A solving unit, used for applying a deep sub-Q network to solve a Pareto frontier that meets the constraints based on multiple objective functions, including: decomposing the construction planning problem into N single-objective sub-problems, and determining the weight vector and the total objective function of each single-objective sub-problem; training the N deep sub-Q networks based on a parameter migration strategy to obtain network parameters of the deep sub-Q networks corresponding to the N sub-problems; solving the Markov decision process corresponding to each sub-problem based on the deep sub-Q network trained by multiple objective functions to obtain a Pareto frontier that meets the constraints; The annual planning unit is used to generate an annual plan for the construction plan based on the Pareto front.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

7. A computer storage medium, characterized in that: The storage medium stores at least one executable instruction, and the executable instruction enables the processor to execute the method as described in any one of claims 1-4.