Multi-stage decision model construction method based on dynamic programming

By constructing a resource allocation state transition network using dynamic programming algorithms, identifying key nodes, and optimizing the decision-making model, the problem that traditional multi-stage decision-making models cannot cope with environmental changes is solved, and efficient resource allocation and decision matching are achieved.

CN120875402AInactive Publication Date: 2025-10-31BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511012162.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multi-stage decision-making models fail to respond promptly to changes in the external environment, resulting in poor decision-making flexibility, inefficient resource allocation, and an inability to achieve optimal matching.

Method used

By dividing the multi-stage decision problem into sub-problems using dynamic programming, a resource allocation state transition network is constructed to identify key nodes, detect path redundancy or missing paths, and optimize the decision model by adjusting strategies in conjunction with the dynamic environment.

Benefits of technology

It improves the flexibility and accuracy of the decision-making process, ensures that resource allocation matches the environment, reduces resource waste and decision lag, and enhances the overall system adaptability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875402A_ABST
    Figure CN120875402A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of resource optimization decision, in particular to a multi-stage decision model construction method based on dynamic programming, which comprises the following steps: modeling based on resource constraint and an objective function, dividing a multi-stage decision problem into sub-problems, extracting logic association between states, identifying key nodes, and detecting path redundancy or missing. And evaluating resource allocation mode requirements, optimizing connection consistency between states, adjusting a decision model structure, identifying new mode attributes, and obtaining an updated multi-stage decision model. According to the method, the problems of key node identification and redundant path detection are effectively solved by accurately dividing the multi-stage decision problem and combining dynamic planning and analysis of the state transition path, the flexibility and accuracy of the decision process are improved, it is ensured that the decision is efficiently matched with external fluctuation all the time, decision lagging and low efficiency are avoided, resource allocation is continuously optimized, and the efficiency is improved. Coordination and consistency of all links are ensured, and the problems of resource waste and lag caused by the fact that real-time adjustment cannot be achieved in a traditional method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of resource optimization decision-making technology, and in particular to a method for constructing a multi-stage decision-making model based on dynamic programming. Background Technology

[0002] The field of resource optimization decision-making technology mainly involves maximizing overall benefits through rational allocation and scheduling under limited resource constraints. This technology is widely applied in scenarios such as production planning, logistics scheduling, energy allocation, supply chain management, and transportation. Core aspects include resource constraint modeling, setting the optimization objective function, designing solution strategies, and making decisions based on dynamic environments. The overall technical system is based on mathematical modeling, optimization algorithms, constraint theory, and decision support, emphasizing global and local collaborative decision-making mechanisms under multi-factor, multi-variable, and multi-stage conditions. Traditional multi-stage decision model construction methods, when facing resource allocation problems with phased progression characteristics, involve dividing the problem into stages, establishing decision models for each stage, and connecting stages using preset rules. This is achieved using linear programming models, dynamic programming based on state transitions, or condition tree structures based on empirical rules to construct and update the multi-stage structure. The process includes constructing single-stage optimization functions for static resource allocation constraints, designing stage transition conditions based on static historical data, and constructing decision path sequences using static parameters to complete the overall model framework.

[0003] Traditional multi-stage decision-making models in existing technologies use preset rules and static data to connect stages, failing to consider dynamic changes in the external environment. This results in poor decision-making flexibility and an inability to respond promptly to new demands and unforeseen events during resource allocation. Especially when facing complex multi-stage resource scheduling problems, traditional methods struggle to effectively identify potential path redundancy or omissions, leading to inefficient resource allocation. Due to fixed decision paths and static resource allocation constraints, they cannot adequately respond to real-time changes, exhibiting significant decision-making lag, which prevents optimal resource allocation and hinders the maximization of overall benefits. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for constructing a multi-stage decision model based on dynamic programming. The technical solution is as follows: On the one hand, a method for constructing a multi-stage decision-making model based on dynamic programming is provided, which includes: S1: Based on resource constraints and objective function modeling, the multi-stage decision problem is divided into several sub-problems. The state transition path of each sub-problem is analyzed through dynamic programming algorithm, and a resource allocation state transition network is established. S2: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, construct a state transition graph, analyze the paths between states in the graph, identify key nodes of resource allocation, detect potential path redundancy or missing values ​​in the graph, and obtain state transition correction values. S3: Call the state transition correction value, combine it with the dynamic environment adjustment strategy, analyze the impact of environmental changes on resource allocation, evaluate the need for new resource allocation mode, compare the matching degree between the existing decision model and environmental changes, screen the decision parts that need to be optimized, and obtain the decision optimization priority. S4: Invoke the decision optimization priority, analyze the logical derivation relationship between the states to be optimized, optimize the logical consistency of the connection between the states, adjust the decision model structure, and output the dynamic decision update sequence.

[0005] Optionally, the resource allocation state transition network includes state transition distribution, state association strength, and resource constraint dependency; the state transition correction value includes path deviation, state stability offset, and graph structure adjustment magnitude; the decision optimization priority includes resource adaptation coefficient, environment matching level, and optimization urgency; and the dynamic decision update sequence includes state adjustment order, logical change type, and constraint set.

[0006] Optionally, the steps of the resource allocation state transition network are as follows: S101: Based on resource constraints and objective function modeling, the multi-stage decision problem is decomposed by decomposition technology. The resource allocation problem in each stage is transformed into a state transition problem. The features of the states and their dependencies are extracted to obtain state transition distribution data. S102: Based on the state transition distribution data, analyze the logical relationship between any two states, use the state transition probability to measure the proximity between states, filter state pairs with a correlation higher than a preset threshold, form a state association network, and obtain state association strength distribution data. S103: Call the state association strength distribution data to analyze the resource constraint dependency between state pairs, combine with the state transition distribution to evaluate the implicit association between states, and obtain the resource allocation state transition network.

[0007] Optionally, the objective function is a mathematical expression used to measure and optimize the effectiveness of problem-solving in an optimization problem. The preset threshold refers to a manually set value used to distinguish and filter state pairs that have a sufficiently strong logical relationship in the analysis.

[0008] Optionally, the step of setting the state transition correction value specifically includes: S201: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, treat each resource allocation state as a node, form edges between nodes through logical links, calculate the shortest path between nodes, identify potential key nodes for state transition, and obtain state transition distribution data. S202: Based on the state transition distribution data, analyze the path deviation between nodes, detect redundancy or missing state transition paths in the graph, judge the rationality of state transitions in combination with environmental information, calculate the path deviation degree and state stability offset, and obtain the state transition correction value.

[0009] Optionally, the steps for prioritizing decision optimization are as follows: S301: Based on the state transition correction value, analyze the impact of dynamic environmental changes on resource allocation, extract key dimensions of environmental changes, calculate environmental change weights, and obtain environmental change distribution data; S302: Based on the environmental change distribution data, assess the demand for new resource allocation models, compare the matching degree between existing decision-making models and environmental changes, screen the decision parts that need to be optimized, calculate the optimization urgency value, and obtain the decision optimization priority.

[0010] Optionally, the environmental change weight is used in dynamic environmental change analysis to quantify the degree of influence of differentiated dimensions or factors on resource allocation patterns. The optimization urgency value is a quantitative indicator that measures the priority or urgency of resource allocation decisions that need to be optimized.

[0011] Optionally, the steps of the dynamic decision update sequence are as follows: S401: Call the decision optimization priority, analyze the logical deduction relationship between the states to be optimized, construct a logical deduction chain, optimize the logical consistency of the connection between the states, and obtain a logical deduction network; S402: Based on the logical derivation network, adjust the decision model structure, optimize the logical relationship between states, identify the logical consistency during decision updates, and obtain the dynamic decision update order; S403: Invoke the dynamic decision update sequence, extract its logical features through the context information of the state to be optimized, calculate the logical consistency score of the decision optimization task, sort the optimization tasks according to the score results, obtain the task execution list, and output the dynamic decision update sequence.

[0012] Optionally, the method also includes step S5: S5: Based on the dynamic decision update sequence, identify new patterns and their attributes in resource optimization decision-making, screen patterns that meet environmental requirements, match logical deduction rules, perform decision model optimization and adjustment, and obtain the updated multi-stage decision model. The updated multi-stage decision-making model includes new pattern entries, attribute mapping rules, and decision-level expansion paths.

[0013] Optionally, the steps of the updated multi-stage decision model are as follows: S501: Call the dynamic decision update sequence, extract new patterns and their attributes, analyze the logical relationship between the patterns and existing decision models, filter patterns that meet environmental requirements, and obtain a list of candidate patterns. S502: Based on the candidate pattern list, match the logical deduction rules, analyze the logical compatibility between patterns, filter the patterns that conform to environmental changes, determine the pattern expansion order, and obtain the logical deduction adaptation sequence. S503: Invoke the logical deduction adaptation sequence, optimize the decision model according to the environmental adaptability adjustment plan, and combine the logical deduction results to perform decision model optimization and adjustment to obtain the updated multi-stage decision model.

[0014] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: By precisely dividing the decision problem into multi-stage segments and combining dynamic programming algorithms to analyze the state transition paths of each sub-problem, the problems of identifying key nodes and detecting redundant paths in resource allocation can be effectively solved, optimizing the flexibility and accuracy of the decision-making process. By incorporating dynamic environmental adjustment strategies, the needs and changes in resource allocation can be evaluated in real time, enabling timely adaptation to fluctuations in the external environment and ensuring that the decision-making process and resource allocation maintain a consistently high degree of matching, avoiding decision lag and inefficiency caused by ignoring environmental changes. Simultaneously, by dynamically updating the decision path and model structure, resource allocation can be continuously optimized, improving the adaptability and effectiveness of the overall system, ensuring coordination and consistency among all aspects of the decision-making process, and reducing resource waste and decision lag caused by the inability to adjust in real time using traditional segmentation methods. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a detailed flowchart of S1 of the present invention; Figure 3 This is a detailed flowchart of the S2 process of the present invention; Figure 4 This is a detailed flowchart of the S3 process of the present invention; Figure 5 This is a detailed flowchart of the S4 process of the present invention; Figure 6 This is a detailed flowchart of S5 of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0017] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0018] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0019] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0021] Please see Figure 1 This invention provides a method for constructing a multi-stage decision model based on dynamic programming. The processing flow of this method may include the following steps: S1: Based on resource constraints and objective function modeling, the multi-stage decision problem is divided into several sub-problems. The state transition path of each sub-problem is analyzed through dynamic programming algorithm, and a resource allocation state transition network is established. S2: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, construct a state transition graph, analyze the paths between states in the graph, identify key nodes of resource allocation, detect potential path redundancy or missing values ​​in the graph, and obtain state transition correction values. S3: Call the state transition correction value, combine it with the dynamic environment adjustment strategy, analyze the impact of environmental changes on resource allocation, evaluate the need for new resource allocation mode, compare the matching degree between the existing decision model and environmental changes, screen the decision parts that need to be optimized, and obtain the decision optimization priority. S4: Call the decision optimization priority, analyze the logical derivation relationship between the states to be optimized, optimize the logical consistency of the connection between the states, adjust the decision model structure, and output the dynamic decision update sequence. S5: Based on the dynamic decision update sequence, identify new patterns and their attributes in resource optimization decision-making, screen patterns that meet environmental requirements, match logical deduction rules, perform decision model optimization and adjustment, and obtain the updated multi-stage decision model; The resource allocation state transition network includes state transition distribution, state association strength, and resource constraint dependency. The state transition correction value includes path deviation, state stability offset, and graph structure adjustment magnitude. The decision optimization priority includes resource adaptation coefficient, environment matching level, and optimization urgency. The dynamic decision update sequence includes state adjustment order, logical change type, and constraint set. The updated multi-stage decision model includes new pattern entries, attribute mapping rules, and decision level expansion path.

[0022] Specifically, such as Figure 2 As shown, the specific steps of the resource allocation state transition network are as follows: S101: Based on resource constraints and objective function modeling, the multi-stage decision problem is decomposed by decomposition technology. The resource allocation problem in each stage is transformed into a state transition problem. The features of the states and their dependencies are extracted to obtain state transition distribution data. The objective function is a mathematical expression used in optimization problems to measure and optimize the effectiveness of problem-solving. Based on resource constraints and objective function modeling, the multi-stage decision-making problem is decomposed into multiple independent state transition problems. Resource constraints include a total power generation capacity of 5000 MW, a maximum transmission capacity of 1000 MW, a fuel supply cost of 200 yuan per MWh, and an environmental emission limit of 10 kg of CO2 emissions per MWh. The objective function is set as minimizing the total operating cost, which includes fuel costs, operation and maintenance costs, and environmental penalties. By analyzing electricity demand, generation output, and transmission capacity at different time stages (e.g., peak, off-peak, and low-peak periods), the resource allocation problem at each time stage is transformed into a state transition problem. Features of the current state (e.g., current power system load, available generating units) are extracted. For example, load status is divided into high load (4000-5000MW), medium load (2000-4000MW), and low load (1000-2000MW). Generator unit operating status is divided into operating, standby, and maintenance. Transmission line status is divided into normal and congested. Through historical data analysis, the dependencies between states are determined. For example, under high load, the probability of generator unit operating status changing from standby to operating is 0.8, and the probability of transmission line changing from normal to congested is 0.3. By quantifying the characteristics and dependencies, state transition distribution data is obtained. For example, within a specific time period, the probability of changing from "low load - some units standby" to "medium load - all units operating" is 0.6, while generating 500 tons of carbon dioxide emissions and 80,000 yuan of fuel costs.

[0023] S102: Based on the state transition distribution data, analyze the logical relationship between any two states, use the state transition probability to measure the proximity between states, filter state pairs with a correlation higher than a preset threshold, form a state association network, and obtain state association strength distribution data. Preset thresholds refer to manually set values ​​used to distinguish and filter state pairs that have a sufficiently strong logical relationship in the analysis; Based on state transition distribution data, for example, analyzing the correlation from "State A: Load 3000MW, Generator A in operation" to "State B: Load 3500MW, Generator B started," state transition probabilities are used to measure the proximity between states. If the transition probability from State A to State B is 0.75, while the transition probability from State A to State C is 0.25, then the proximity between State A and State B is significantly higher than that between State A and State C. State pairs with a correlation higher than a preset threshold are selected. The preset threshold is set at 0.6, and this threshold is based on historical power system operation data and expert opinions. Based on experience, by analyzing the operational data of the past year, it was found that when the state transition probability is below 0.6, the correlation between states is weak and has little impact on decision-making, while when it is above 0.6, it is considered to have a sufficiently strong correlation. For example, if the transition probability from state X to state Y is 0.8, which is above the threshold of 0.6, then (state X, state Y) is considered a correlated state pair. If the transition probability from state Z to state W is 0.3, which is below the threshold of 0.6, then it is not considered a correlated state pair. Through this screening process, a state correlation network is formed, and finally, the state correlation strength distribution data is obtained.

[0024] S103: Call the state association strength distribution data, analyze the resource constraint dependency between state pairs, combine with the state transition distribution, evaluate the implicit association between states, and obtain the resource allocation state transition network. The system utilizes state association strength distribution data. For example, for a state pair (insufficient power generation, low fuel inventory), it analyzes the dependence of insufficient power generation on fuel inventory. Combining this with state transition distribution, it assesses implicit associations between states. For instance, although there is no direct state transition between "generator failure" and "high grid load," an implicit association can be found by analyzing the common association with the intermediate state "insufficient reserve capacity." Regarding resource constraint dependence, it is quantitatively evaluated by setting specific indicators. For example, when a state transition leads to a power shortage exceeding 100MW, its resource constraint dependence is defined as 0.9; when the power shortage is between 50MW and 100MW, its resource constraint dependence is defined as 0.7; and when the power shortage is below 50MW, its resource constraint dependence is defined as 0.5. The assessment of implicit associations is done by calculating indicators such as the number of common neighbors and path length. For example, if two states are connected by fewer than three intermediate states and share at least two common downstream states, a significant implicit association is considered to exist. Through analysis, the resource allocation state transition network is finally obtained.

[0025] Specifically, such as Figure 3 As shown, the specific steps for correcting the state transition value are as follows: S201: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, treat each resource allocation state as a node, form edges between nodes through logical links, calculate the shortest path between nodes, identify potential key nodes for state transition, and obtain state transition distribution data. The resource allocation state transition network is invoked. For example, the logical association from "increasing power load" to "increasing generator output" is used. An edge set is generated based on the state transition frequency. For instance, in the past month, the state transition from "increasing power load" to "increasing generator output" occurred 50 times, while the state transition from "decreasing power load" to "decreasing generator output" occurred 40 times. Each resource allocation state is treated as a node; for example, "increasing power load" is one node, and "increasing generator output" is another. Nodes are logically linked to form edges, and the weight of each edge is set to the state transition frequency. The higher the frequency of occurrence, the greater the edge weight. The shortest path between nodes is calculated, and potential key nodes for state transitions are identified. For example, the shortest path from the initial state to the target state is calculated using the Dijkstra algorithm, and nodes connecting multiple key sub-processes in the path are identified. For example, in a power dispatching scenario, "inaccurate load forecasting" leads to "insufficient reserve capacity," which in turn triggers "regional power rationing." In this path, "insufficient reserve capacity" is a key node because it connects the forecasting and execution stages. Through this process, state transition distribution data is obtained, which includes nodes, edges and their weights, as well as the identified key nodes.

[0026] S202: Based on state transition distribution data, analyze the path deviation between nodes, detect redundancy or missing state transition paths in the graph, combine environmental information to judge the rationality of state transition, calculate path deviation degree and state stability offset, and obtain state transition correction value. The state transition correction value is calculated using the following formula: ; in, Represents the state transition correction value. Representing the Path deviation, Representing the Expected stability of the path, Representing the Movement deviation value on the path, Representing the Redundancy on the path Represents the maximum path deviation. Represents the minimum path deviation. Represents a constant adjustment factor; Based on state transition distribution data, the first step is to analyze path deviations between nodes. Path deviation refers to the difference between the actual state transition and the expected state transition, which can be obtained by collecting historical state transition data and comparing it with the expected model. For example, in a power system, if a device transitions from a normal operating state to a fault state, the actual transition time may differ from the expected time; the path deviation is this time difference. Next, the redundancy and missing paths in the state transition graph are analyzed. Redundant paths refer to paths in the network that have multiple feasible paths, but these paths lead to wasted resources or increased computational complexity. Missing paths refer to paths that should exist but were not identified due to data acquisition deficiencies. To address this, algorithms are used to detect redundancy and missing paths. Redundancy or missing paths are identified by comparing existing paths with ideal paths in the network. In power load dispatching scenarios, some backup generator startup paths are redundant, while missing paths represent omissions in emergency response paths. The rationality of state transitions is assessed by incorporating environmental information. Environmental factors such as temperature, humidity, and load changes affect operating states. When judging path rationality, external factors must be considered. For example, during periods of high summer load, paths that reduce industrial power consumption should be avoided; instead, paths that add backup generators should be prioritized. Finally, path deviation and state stability offset are calculated. Path deviation measures the difference between the actual and expected path, while state stability offset reflects the impact of state transitions on stability. A combined measurement of both reveals the effectiveness of the current state transition. Through a series of analyses and calculations, the final state transition correction value is obtained, reflecting the rationality and stability of the path, thus providing a basis for subsequent optimization decisions.

[0027] According to the formula The specific meanings and calculation logic of each parameter are as follows: This represents the state transition correction value, used to measure the stability of the path; Representing the Path deviation measures the difference between the actual path execution and the expected outcome. Representing the The expected stability of a path is based on the expected value of historical data or models; The value representing the path's movement deviation reflects the changes in the path due to external factors; It is the redundancy of the path, which measures whether the path contains redundant parts; and These represent the maximum and minimum path deviations, respectively, and are used to measure the range of deviations within a path set. It is an adjustment constant used to optimize the influence of different parameters. The operators include absolute value (ensuring that the path deviation direction is independent), division (measuring the ratio of path deviation to stability), multiplication (comprehensively considering the relationship between path stability, redundancy and movement deviation), square root (balancing the weights of movement deviation and redundancy), and summation (comprehensively evaluating all paths). Assume the path parameter is: the deviation of path 1. Expected stability , shift deviation Redundancy ; Deviation of Path 2 Expected stability , shift deviation Redundancy ; Deviation of Path 3 Expected stability , shift deviation Redundancy ; The calculation result is: ; The state transition correction value is calculated to be... This indicates that there is a certain deviation in path stability, which requires further optimization. By introducing a comprehensive consideration of path deviation, stability, redundancy, and movement deviation, this formula can more accurately assess the rationality of the path and reflect the rationality of the path selection through the correction value, which helps to optimize resource allocation and path selection. The calculation and quantification process of the parameter values ​​is clearer through the actual data table, ensuring the accuracy and reliability of the correction value calculation.

[0028] Table 1: Path Parameter Data Table

[0029] The table data provides a clear view of the deviation, stability, movement deviation, and redundancy of different paths. The data directly affects the calculation of correction values. Table 1 lists the path parameter data, specifically showing parameters such as path deviation, stability, movement deviation, and redundancy, providing a reliable basis for calculating correction values ​​in practical applications.

[0030] Specifically, such as Figure 4 As shown, the specific steps for prioritizing decision optimization are as follows: S301: Based on the state transition correction value, analyze the impact of dynamic environmental changes on resource allocation, extract the key dimensions of environmental changes, calculate the weight of environmental changes, and obtain environmental change distribution data; Environmental change weighting is used in dynamic environmental change analysis to quantify the degree of influence of differentiated dimensions or factors on resource allocation patterns. Based on state transition correction values, such as current environmental states including electricity market price fluctuations, extreme weather events (e.g., a surge in electricity demand due to sustained high temperatures), and policy and regulatory adjustments (e.g., stricter carbon emission restrictions), key dimensions of environmental change are extracted. For example, for electricity market price fluctuations, the key dimensions are "price fluctuation magnitude" (high, medium, low) and "duration" (short, medium, long); for extreme weather events, the key dimensions are "temperature increase magnitude" (e.g., 5 degrees Celsius above average temperature is considered a significant increase) and "duration"; for policy and regulatory adjustments, the key dimensions are "degree of restriction" and "speed of implementation." Environmental change weights are then calculated. For example, when the electricity market... When the price fluctuation amplitude is "high", the weight is set to 0.4; when the duration is "long", the weight is set to 0.3; and when both occur simultaneously, the environmental change weight is set to 0.7 (0.4 + 0.3). The weights are set based on historical data analysis and expert experience. By analyzing the operating data of the power system under different environmental changes over the past five years, the impact of various environmental factors on the model's stability and cost is assessed. For example, historical data shows that the impact of extreme high temperatures on the power system load is twice that of price fluctuations, so it is given a higher weight. By quantifying the key dimensions, the distribution data of environmental changes is obtained, which includes various environmental changes and their corresponding weights.

[0031] S302: Based on environmental change distribution data, assess the need for new resource allocation models, compare the matching degree between existing decision-making models and environmental changes, screen the decision parts that need to be optimized, calculate the optimization urgency value, and obtain the decision optimization priority. The optimization urgency value is a quantitative indicator that measures the priority or urgency of resource allocation decisions that need to be optimized. Based on environmental change distribution data, for example, in the calculation of environmental change weights, if environmental changes such as "extreme high temperature" and "fuel price increase" are detected, their environmental change weights reach 0.7 and 0.6 respectively. In this case, the assessment needs to introduce a new resource allocation model, such as a "load-side response priority model" or a "multi-fuel source dispatch model," and compare the matching degree between the existing decision-making model and environmental changes. For example, the ability of the current power dispatch model to cope with extreme high temperature scenarios (such as the model's dispatch efficiency and number of power outages in historical high temperature events) is compared with the model's performance under normal conditions to screen out the decision-making parts that need to be optimized. For example, if the accuracy of the current model's load forecasting module decreases by 15% during high temperature weather, then it is determined that load forecasting needs to be optimized. In the decision-making section, the urgency value for optimization is calculated. For example, the urgency value of optimization for a decision-making component can be quantified by its impact on overall performance and the severity of current environmental changes. If the optimization of the load forecasting module can reduce power curtailment by 5%, and the current power curtailment risk is high, then the urgency value is set to 0.9. The urgency value is a quantitative indicator that measures the priority or urgency of resource allocation decisions that need to be optimized. For example, when the urgency value is higher than 0.8, it indicates that the optimization of this decision-making component has a high priority and needs to be carried out immediately. When the urgency value is between 0.5 and 0.8, it indicates a medium priority. When the urgency value is lower than 0.5, it indicates a low priority. Through this process, the decision optimization priority is obtained.

[0032] Specifically, such as Figure 5 As shown, the specific steps of the dynamic decision update sequence are as follows: S401: Call the decision optimization priority, analyze the logical deduction relationship between the states to be optimized, construct the logical deduction chain, optimize the logical consistency of the connection between the states, and obtain the logical deduction network; Prioritizing decision optimization, for example, if "insufficient reserve capacity" is a state to be optimized, its derivation from "generator set failure" or "load forecast deviation" is analyzed, and a logical deduction chain is constructed. For example, from the logical relationship of "inaccurate load forecast" → "actual load exceeds expectations" → "insufficient reserve capacity" → "activation of emergency plan", the logical consistency of the connection between states is optimized. For example, by checking for breaks or contradictions in the logical chain, such as "sufficient fuel supply" leading to "generator set shutdown", inconsistencies are corrected. For example, if it is found that the logical connection between "generator set failure" and "insufficient reserve capacity" is inconsistent due to the lack of consideration for the start-up time of the standby unit, the parameter of "standby unit start-up time" is introduced to correct it. Through these steps, a logical deduction network is finally obtained. This network clearly shows the causal relationship of the state to be optimized, providing a logical basis for subsequent decision adjustments.

[0033] S402: Based on the logical derivation network, adjust the decision model structure, optimize the logical relationship between states, identify the logical consistency during decision updates, and obtain the dynamic decision update order; Based on the logical derivation network, for example, if the logical derivation network reveals a lag in information transmission between the "prediction module" and the "scheduling module", the model structure is adjusted, and a "real-time data interface" is introduced to reduce the lag. Logical consistency during decision updates is identified. For example, when adjusting the "generation plan generation module", its output is checked to ensure compatibility with the input logic of the "transmission constraint module", ensuring that the newly introduced decision logic does not conflict with existing logic. The consistency between the newly introduced "considering new energy output fluctuations" decision logic and the "grid stability control" logic is verified. If they are inconsistent, further model adjustments are needed to ensure compatibility. Through this process, the dynamic decision update sequence is obtained, which clarifies the update order of each part of the decision model. For example, after updating the load forecast model, the generator scheduling model must be updated first, followed by the transmission line management model, to ensure the logical consistency of the entire method.

[0034] S403: Call the dynamic decision update sequence, extract its logical features through the context information of the state to be optimized, calculate the logical consistency score of the decision optimization task, sort the optimization tasks according to the score results, obtain the task execution list, and output the dynamic decision update sequence. The dynamic decision update sequence is invoked. For example, for the state to be optimized, "grid congestion," its contextual information is extracted, including current line load, power demand in adjacent areas, and generator output distribution. Logical features such as "bottleneck capacity of congested lines" and "overcapacity of upstream power generation" are identified from this information. A logical consistency score is then calculated for the decision optimization task. This score is determined by evaluating the degree of impact of the optimization task on key logical chains in the logical derivation network. For example, an optimization task that can resolve key breakpoints in the logical derivation chain receives a higher score. For instance, a task that addresses "inaccurate load forecasting leading to scheduling errors" receives a higher logical consistency score. The priority score is set to 0.9, while the score for tasks that solve minor problems, such as adjusting the report generation format, is set to 0.2. The optimization tasks are prioritized according to the score results. For example, tasks with a score higher than 0.8 are listed as high priority, those between 0.5 and 0.8 are medium priority, and those below 0.5 are low priority. This results in a task execution list, for example, Task 1: Load forecasting model update (priority: high), Task 2: Reserve capacity optimization (priority: medium), Task 3: Report format adjustment (priority: low). The output is a dynamic decision update sequence, which is a detailed list containing all optimization tasks and their execution priorities.

[0035] Specifically, such as Figure 6 As shown, the steps of the updated multi-stage decision model are as follows: S501: Call the dynamic decision update sequence, extract new patterns and their attributes, analyze the logical relationship between the patterns and existing decision models, filter patterns that meet environmental requirements, and obtain a list of candidate patterns. The process involves invoking dynamic decision update sequences, such as identifying "machine learning-based real-time load forecasting models" with attributes like "5% improvement in forecast accuracy" and "10% increase in computational resource consumption." It also analyzes the logical relationship between these models and existing decision models, examining how the "machine learning-based real-time load forecasting model" integrates with existing "traditional forecasting models" and "power dispatching models," and whether adjustments to the interfaces of existing models are necessary. Models that meet environmental requirements are then selected. For example, in the current environment of volatile electricity market prices, models that effectively address price fluctuations, such as those demonstrating excellent real-time price response, are chosen. For instance, a "real-time load response model" can guide users to adjust their electricity consumption behavior based on electricity price levels, aligning with the current environment's requirement to improve power system flexibility. Through rigorous evaluation, the new models are ensured to be highly compatible with current environmental trends, ultimately resulting in a list of candidate models.

[0036] S502: Based on the candidate pattern list, match the logical deduction rules, analyze the logical compatibility between patterns, filter the patterns that conform to environmental changes, determine the pattern expansion order, and obtain the logical deduction adaptation sequence. Based on the candidate mode list, for example, for the candidate mode "energy storage system optimized scheduling", its logical derivation rules are matched with those of "peak-valley electricity price arbitrage" and "grid peak shaving", and the logical compatibility between modes is analyzed. For example, if "energy storage system optimized scheduling mode" and "renewable energy grid connection optimization mode" coexist, it is necessary to analyze their compatibility in terms of dispatch command issuance and power allocation to ensure that they work together rather than interfere with each other. For example, if there is a conflict between energy storage system scheduling and renewable energy grid connection, which leads to grid instability, it is necessary to adjust their compatibility to ensure that the energy storage system charges first when the grid load is low and discharges first when the load is high, in order to match the volatility of renewable energy. Modes that conform to environmental changes are selected. For example, in an environment where wind power generation is unstable, modes that can effectively smooth wind power fluctuations are selected first. The mode expansion order is determined. For example, "short-term load forecasting mode" is deployed first, followed by "real-time unit combination optimization mode" to ensure that the deployment of each mode can provide the necessary data and foundation for subsequent modes, and finally, a logical derivation adaptation sequence is obtained.

[0037] S503: Call the logical deduction adaptation sequence, optimize the decision model according to the environmental adaptability adjustment plan, and combine the logical deduction results to perform the decision model optimization adjustment to obtain the updated multi-stage decision model; The system invokes logical derivation adaptation sequences and optimizes the decision-making model based on environmental adaptability adjustment plans. For example, if the adaptation sequence indicates the need to introduce a "flexible load control module," then the module's parameter settings are adjusted based on environmental adaptability factors such as the current grid's actual flexible load access capacity and user response willingness. Combining the logical derivation results, the decision-making model is optimized and adjusted. For instance, if the logical derivation results show that improving "load forecasting accuracy" is key to improving the effectiveness of "unit optimized scheduling," then resources are concentrated on training and deploying the load forecasting model, rather than blindly adjusting parameters. Through these steps, an updated multi-stage decision-making model is finally obtained. This model can effectively cope with dynamic environmental changes, ensuring the efficiency of resource allocation and the stability of the system. For example, after introducing a new market mechanism, the updated multi-stage decision-making model can more accurately predict market responses and adjust power generation plans and transmission arrangements accordingly to ensure power supply and demand balance and minimize operating costs.

[0038] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for constructing a multi-stage decision-making model based on dynamic programming, characterized in that, Includes the following steps: S1: Based on resource constraints and objective function modeling, the multi-stage decision problem is divided into several sub-problems. The state transition path of each sub-problem is analyzed through dynamic programming algorithm, and a resource allocation state transition network is established. S2: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, construct a state transition graph, analyze the paths between states in the graph, identify key nodes of resource allocation, detect potential path redundancy or missing values ​​in the graph, and obtain state transition correction values. S3: Call the state transition correction value, combine it with the dynamic environment adjustment strategy, analyze the impact of environmental changes on resource allocation, evaluate the need for new resource allocation mode, compare the matching degree between the existing decision model and environmental changes, screen the decision parts that need to be optimized, and obtain the decision optimization priority. S4: Invoke the decision optimization priority, analyze the logical derivation relationship between the states to be optimized, optimize the logical consistency of the connection between the states, adjust the decision model structure, and output the dynamic decision update sequence.

2. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 1, characterized in that, The resource allocation state transition network includes state transition distribution, state association strength, and resource constraint dependency. The state transition correction value includes path deviation, state stability offset, and graph structure adjustment magnitude. The decision optimization priority includes resource adaptation coefficient, environment matching level, and optimization urgency. The dynamic decision update sequence includes state adjustment order, logical change type, and constraint set.

3. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 1, characterized in that, The specific steps of the resource allocation state transition network are as follows: S101: Based on resource constraints and objective function modeling, the multi-stage decision problem is decomposed by decomposition technology. The resource allocation problem in each stage is transformed into a state transition problem. The features of the states and their dependencies are extracted to obtain state transition distribution data. S102: Based on the state transition distribution data, analyze the logical relationship between any two states, use the state transition probability to measure the proximity between states, filter state pairs with a correlation higher than a preset threshold, form a state association network, and obtain state association strength distribution data. S103: Call the state association strength distribution data to analyze the resource constraint dependency between state pairs, combine with the state transition distribution to evaluate the implicit association between states, and obtain the resource allocation state transition network.

4. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 3, characterized in that, The objective function is a mathematical expression used in optimization problems to measure and optimize the effectiveness of problem-solving. The preset threshold refers to a manually set value used to distinguish and filter state pairs that have a sufficiently strong logical relationship in the analysis.

5. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 3, characterized in that, The specific steps for setting the state transition correction value are as follows: S201: Call the resource allocation state transition network, extract the logical relationships between states, generate an edge set according to the state transition frequency, treat each resource allocation state as a node, form edges between nodes through logical links, calculate the shortest path between nodes, identify potential key nodes for state transition, and obtain state transition distribution data. S202: Based on the state transition distribution data, analyze the path deviation between nodes, detect redundancy or missing state transition paths in the graph, judge the rationality of state transitions in combination with environmental information, calculate the path deviation degree and state stability offset, and obtain the state transition correction value.

6. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 5, characterized in that, The specific steps for prioritizing decision optimization are as follows: S301: Based on the state transition correction value, analyze the impact of dynamic environmental changes on resource allocation, extract key dimensions of environmental changes, calculate environmental change weights, and obtain environmental change distribution data; S302: Based on the environmental change distribution data, assess the demand for new resource allocation models, compare the matching degree between existing decision-making models and environmental changes, screen the decision parts that need to be optimized, calculate the optimization urgency value, and obtain the decision optimization priority.

7. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 6, characterized in that, The environmental change weight is used in dynamic environmental change analysis to quantify the degree of influence of differentiated dimensions or factors on resource allocation patterns. The optimization urgency value is a quantitative indicator that measures the priority or urgency of resource allocation decisions that need to be optimized.

8. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 6, characterized in that, The specific steps of the dynamic decision update sequence are as follows: S401: Call the decision optimization priority, analyze the logical deduction relationship between the states to be optimized, construct a logical deduction chain, optimize the logical consistency of the connection between the states, and obtain a logical deduction network; S402: Based on the logical derivation network, adjust the decision model structure, optimize the logical relationship between states, identify the logical consistency during decision updates, and obtain the dynamic decision update order; S403: Invoke the dynamic decision update sequence, extract its logical features through the context information of the state to be optimized, calculate the logical consistency score of the decision optimization task, sort the optimization tasks according to the score results, obtain the task execution list, and output the dynamic decision update sequence.

9. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 1, characterized in that, The method also includes step S5: S5: Based on the dynamic decision update sequence, identify new patterns and their attributes in resource optimization decision-making, screen patterns that meet environmental requirements, match logical deduction rules, perform decision model optimization and adjustment, and obtain the updated multi-stage decision model. The updated multi-stage decision-making model includes new pattern entries, attribute mapping rules, and decision-level expansion paths.

10. The method for constructing a multi-stage decision-making model based on dynamic programming according to claim 9, characterized in that, The steps of the updated multi-stage decision-making model are as follows: S501: Call the dynamic decision update sequence, extract new patterns and their attributes, analyze the logical relationship between the patterns and existing decision models, filter patterns that meet environmental requirements, and obtain a list of candidate patterns. S502: Based on the candidate pattern list, match the logical deduction rules, analyze the logical compatibility between patterns, filter the patterns that conform to environmental changes, determine the pattern expansion order, and obtain the logical deduction adaptation sequence. S503: Invoke the logical deduction adaptation sequence, optimize the decision model according to the environmental adaptability adjustment plan, and combine the logical deduction results to perform decision model optimization and adjustment to obtain the updated multi-stage decision model.

Citation Information

Cited By

  • Container configuration reconstruction quantitative analysis method fusing file semantics and cross-stage dependence

    CN121523693A