Water storage scheduling rule extraction method based on reinforcement learning
By using a reservoir scheduling optimization model based on reinforcement learning, the problem of reservoir scheduling being unable to adapt to different water inflow conditions was solved, and the comprehensive benefits of reservoirs in dry and wet years were improved, providing a better scheduling decision-making scheme.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA POWER CONSRTUCTION GRP GUIYANG SURVEY & DESIGN INST CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-12
AI Technical Summary
The existing reservoir water storage and scheduling system is unable to adapt to different water inflow conditions, resulting in low power generation efficiency in dry years and high flood control risks in wet years, making it difficult to achieve a balanced scheduling of all parties' needs under the condition of limited water storage capacity.
A multi-objective optimization scheduling model is established using a reinforcement learning-based approach. Combining the objective functions of maximizing total power generation and minimizing flood control risk, an agent is trained using a reinforcement learning algorithm to extract water storage scheduling rules. Considering various constraints of the reservoir, the scheduling decision is optimized.
By maximizing total power generation and minimizing flood control risks during the reservoir's operation period under different inflow conditions, the overall benefits of the reservoir are improved, and the reduced benefits during high-water-level operation caused by single-objective scheduling are avoided.
Smart Images

Figure CN122022233A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of reservoir scheduling technology and relates to a method for extracting water storage scheduling rules based on reinforcement learning. Background Technology
[0002] With the continuous construction and commissioning of new reservoirs, the proportion of water storage to runoff during the storage period has increased, resulting in significant water reduction in the middle and lower reaches of the Yangtze River. This leads to insufficient capacity for public-interest water management such as drought relief and ecological restoration in the basin. Under limited water storage capacity, how to scientifically explore various engineering problems by considering the inflow to each reservoir, combined with factors such as flood control, power generation, water storage, and navigation.
[0003] Current real-time operation and scheduling of reservoirs are based on conventional scheduling according to the designed water storage scheduling line, which is difficult to adapt to different inflow volumes. This results in low power generation efficiency in dry years and high flood control risk in wet years. Therefore, this method employs a reinforcement learning-based water storage scheduling rule extraction approach to obtain optimized scheduling rules that balance flood control risk and reservoir power generation efficiency under different inflow conditions during the water storage period. Summary of the Invention
[0004] The purpose of this invention is to overcome the problem that existing reservoir water storage scheduling methods are difficult to adapt to different inflow conditions. This invention provides a water storage scheduling rule extraction method based on reinforcement learning algorithm, which takes maximizing the total power generation and minimizing the flood control risk rate during the scheduling period as the objectives, and extracts optimized water storage scheduling rules for reservoir operation during the water storage period, providing a reference for reservoir operation during the water storage period. This invention comprehensively considers the overall benefits of the water storage period and the high water level operation period, effectively helping reservoir managers to make scheduling decisions that meet the needs of all parties under different inflow conditions.
[0005] The present invention achieves the above objectives using the following technical solution: A method for extracting water storage optimization scheduling rules based on reinforcement learning includes the following steps: (1) Establish an optimized scheduling model for advance water storage with the objectives of maximizing total power generation and minimizing flood risk during the scheduling period; (2) Using the inflow rate over many years as input, the agent interacts with the water storage scheduling environment and is trained based on reinforcement learning algorithm to extract water storage scheduling rules; (3) Perform simulated scheduling based on the obtained scheduling rules.
[0006] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, the objective function of the advance water storage optimization scheduling model in step (1) is: The reservoir generates the most electricity during the scheduling period: In the formula: E is the total power generation during the reservoir scheduling period, in hundreds of millions of kWh; E(t) is the multi-year average power generation at time t, in hundreds of millions of kWh; the reservoir scheduling period refers to the reservoir impoundment period and the subsequent high water level operation period; T is the total step size of the scheduling period, and Y is the total number of years; The time step is in days; K is the power output coefficient; Q(y,t) is the power generation flow of the reservoir at time t in year y, in meters. 3 / s; H(y,t) is the net head of water generated by the reservoir at time t in year y, in meters.
[0007] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, the objective function of the advance water storage optimization scheduling model in step (1) is: Reservoirs pose the least risk for flood control: In the formula: FCR(t) represents the maximum flood control risk over many years at time t during the reservoir's impoundment period; FCR(y,t) represents the flood control risk of the reservoir at time t in year y; V(y,t) represents the reservoir capacity at time t in year y; m 3 VS(t) represents the reservoir capacity corresponding to the phased flood control limit water level at time t in year y, m. 3 VU represents the normal water level of the reservoir, in meters. 3 VL represents the reservoir capacity corresponding to the flood control limit water level, in meters. 3 ; This is the penalty coefficient.
[0008] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, in step (2), the reinforcement learning framework for training the agent includes the agent and the environment. The agent interacts with the environment at discrete time steps based on the Markov decision process to acquire knowledge, and performs agent training, i.e., value valuation update, with the goal of maximizing value valuation, and finally obtains the optimal action strategy. The reinforcement learning algorithm, as the agent, interacts with the environment containing cascade reservoir scheduling knowledge to acquire knowledge samples. In the formula: s t The initial state of the agent at time t is determined by the scheduling time t and the water level Z. t Composition, water level Z t The discrete values within the reservoir water level constraint range; a t Let r be the action of the agent at time t; t For the action reward during time period t, r t The value is the power generation benefit E(t) generated by the scheduling.
[0009] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, the value obtained by an agent taking an action in a certain state during the Markov decision-making process is called the action value, denoted as the Q-value. The reinforcement learning algorithm updates the Q-value to approximate its optimal value in order to find the optimal strategy for solving multi-stage problems. The optimal Q-values at the beginning and end of the time period satisfy the Bellman equation: In the formula: q*(s t ,a t ) represents state s t Take action a t The value obtained later; γ is the discount rate, used to control the impact of future earnings on the present; P st,st+1 This indicates that the cascade reservoirs are in state s. t Transition to the next state s t+1 The Markov state transition probability is used to describe the randomness of the flow in a state, where S is the set of states.
[0010] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, the ε-greedy policy is used in reinforcement learning to determine the decision action at each stage. The action selection probability expression in the ε-greedy policy is as follows: In the formula: Represents the state s at time t t The probability of randomly selecting an action; Represents the state s at time t t The probability of choosing the action that yields the highest evaluation value; ε is the greed rate.
[0011] In the aforementioned reinforcement learning-based method for extracting water storage optimization scheduling rules, in step (3), Using the test period data from the inflow data, and through reinforcement learning-based water storage optimization scheduling rules obtained from training, considering water balance constraints, reservoir capacity curve constraints, discharge facility discharge capacity constraints, upper and lower limits and water level fluctuation constraints, output constraints, and boundary condition constraints, the dual objective of the water storage scheduling model is transformed into a single objective for simulating reservoir scheduling. The comprehensive benefit index is expressed as: In the formula: R is the comprehensive benefit index; a is the weight.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The extracted reservoir water storage optimization scheduling rules are applicable to various water inflow conditions during the water storage period. They also take into account the total power generation and flood risk rate during the scheduling period, and can provide a reference for reservoir water storage scheduling decisions.
[0013] (2) The total benefits of the water storage period and the subsequent high water level operation period are taken into account, avoiding the reduction of the overall benefits of the reservoir during the high water level operation period due to overemphasis on the benefits of the water storage period. Attached Figure Description
[0014] Figure 1 This is a flowchart of the present invention; Figure 2 This is a diagram of the simulated scheduling process in 1965. Figure 3 This is a comparison chart of overall benefits. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0016] Example. A method for extracting water storage optimization scheduling rules based on reinforcement learning, as shown in the attached figure. Figure 1 As shown, the present invention includes the following process: (1) Establish an optimized scheduling model for advance water storage with the objectives of maximizing total power generation and minimizing flood risk during the scheduling period; (2) Using the inflow rate over many years as input, the agent interacts with the water storage scheduling environment and is trained based on reinforcement learning algorithm to extract water storage scheduling rules; (3) Perform simulated scheduling based on the obtained scheduling rules.
[0017] The specific techniques and steps used in the above method are as follows: Step (1) Establish an optimized scheduling model for advance water storage with the objectives of maximizing total power generation and minimizing flood risk during the scheduling period. The objective of this method for water storage scheduling is to advance the start-up time of reservoirs while satisfying complex constraints and comprehensively considering the ecological, flood control, and navigation requirements. Based on the water storage scheduling diagram, it seeks an optimal trajectory that balances the power generation and flood control of cascade reservoirs. To this end, the following multi-objective water storage scheduling model is established. The objective function of the water storage scheduling model is: (1) The total power generation is greatest during the reservoir's operation period: In the formula: E and E(t) are the total power generation during the reservoir scheduling period (hundred million kWh) and the multi-year average power generation at time t (hundred million kWh), respectively. In this method, the reservoir scheduling period refers to the reservoir impoundment period and the subsequent high water level operation period; T is the total step size of the scheduling period, and Y is the total number of years. The time step is 1 day in this method; K is the power output coefficient; Q(y,t) and H(y,t) are the power generation flow (m³) of the reservoir at time t in year y, respectively. 3 / s) and power generation head (m).
[0018] (2) Reservoirs pose the least risk for flood control: In the formula: FCR(t) and FCR(y,t) represent the multi-year maximum flood control risk at time t during the reservoir's impoundment period and the flood control risk at time t in year y, respectively; V(y,t) represents the reservoir capacity (m³) at time t in year y. 3 ); VS(t) represents the reservoir capacity (in meters) corresponding to the phased flood control limit water level at time t in year y. 3 VU and VL represent the reservoir capacity (in meters) corresponding to the normal water level and the flood control limit water level, respectively. 3 ); This is the penalty coefficient.
[0019] (2) Using the inflow rate over many years as input, the agent interacts with the water storage scheduling environment and is trained based on reinforcement learning algorithm to extract water storage scheduling rules. The main components of a reinforcement learning framework are an agent and an environment. The agent, based on a Markov decision process, interacts with the environment at discrete time steps to acquire knowledge. The agent is trained (i.e., its value is updated) with the goal of maximizing its value assessment, ultimately obtaining the optimal action policy. The reinforcement learning algorithm, acting as the agent, interacts with the environment, which contains knowledge about cascade reservoir scheduling, to acquire knowledge samples. In the formula: s t The initial state of the agent at time t is determined by the scheduling time t and the water level Z. t Composition, of which water level Z t The discrete values within the reservoir water level constraint range; a t Let r be the action of the agent at time t; t The reward for actions during time period t is the power generation benefit E(t) generated by the scheduling.
[0020] In the decision-making process, the value gained by an agent taking an action in a given state is called the action value, also known as the Q-value. Reinforcement learning algorithms update the Q-value, approximating its optimal value to find the optimal policy for multi-stage problems. The optimal Q-values at the beginning and end of the time interval satisfy the Bellman equation: In the formula: q*(s t ,a t ) represents state s tTake action a t The value obtained later; γ is the discount rate, used to control the impact of future earnings on the present; P st,st+1 This indicates that the cascade reservoirs are in state s. t Transition to the next state s t+1 The Markov state transition probability is used to describe the randomness of the flow in a state, where S is the set of states.
[0021] In reinforcement learning, the decision-making action at each stage is determined by exploring exploitation strategies. In these strategies, the agent uses current information to select the best action to obtain high-value immediate gains, but is prone to getting trapped in local optima; the agent also explores other actions, but the results are uncertain. To fully utilize existing experience in reservoir management, an ε-greedy exploitation strategy is adopted, which does not rely on any specific environmental information. The action selection probability expression in the ε-greedy strategy is as follows.
[0022] In the formula: Represents the state s at time t t The probability of randomly selecting an action; Represents the state s at time t t The probability of choosing the action with the highest evaluation value; ε is the greed rate, which is the probability of choosing an action in an exploratory manner.
[0023] (3) Based on the obtained scheduling rules, simulate scheduling and analyze whether the power generation and flood risk rate of the obtained operation trajectory are reasonable.
[0024] Using the test period data from the inflow data, and through the reinforcement learning-based water storage optimization scheduling rules obtained from the training, considering water balance constraints, reservoir capacity curve constraints, discharge facility discharge capacity constraints, upper and lower limits and water level fluctuation constraints, output constraints, and boundary condition constraints, the dual objective of the water storage scheduling model is transformed into a single objective for simulated water storage scheduling: In the formula: R is the comprehensive benefit index; a is the weight.
[0025] To analyze the rationality of the power generation and flood risk rate obtained from the obtained operating trajectory, simulation scheduling was conducted using the same test period data and water storage scheduling model, based on the original designed water storage scheduling line. The comprehensive scheduling benefits of the two schemes were compared. The specific simulation and comparison results are as follows: When using the above method for simulation, taking the Wudongde Reservoir, the leading reservoir in the downstream cascade reservoirs of the Jinsha River basin, as an example, the study period is from 1950 to 2020. The three years with the highest water inflow (2006, 2011, and 1992), the three years with the lowest water inflow (1965, 1966, and 1954), and the three years with moderate water inflow (2012, 2018, and 1986) are selected as the verification period (9 years). The remaining 62 years are used as the calibration period. Within the given scheduling period, based on the actual scheduling needs of the Wudongde Reservoir, and considering the requirements for comprehensive utilization and other constraints, the scheduling model uses the maximum total power generation and the minimum flood control risk rate as its objective functions. The optimal scheduling rules for water storage are extracted through reinforcement learning algorithm training.
[0026] Based on the obtained optimized water storage scheduling rules, the reservoir water storage scheduling process was simulated and operated during the verification period. The results were then evaluated and compared with the original design scheme. (See...) Figure 3 Taking 1965 as an example, the simulation scheduling process diagram is shown below. Figure 2 ; Analysis of comprehensive evaluation indicators shows that the operational trajectory simulated by the scheduling rules in this embodiment is superior to the simulation based on the original design water storage scheduling line. Compared to the original design scheme, the comprehensive benefit indicators of this scheduling rule are improved by an average of 2.82% in dry years, 4.98% in normal years, and 6.79% in wet years, with an average improvement of 4.86% during the verification period. This provides a certain reference for reservoir water storage scheduling and operation.
[0027] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the methods and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for extracting water storage optimization scheduling rules based on reinforcement learning, characterized in that: Includes the following steps: (1) Establish an optimized scheduling model for advance water storage with the objectives of maximizing total power generation and minimizing flood risk during the scheduling period; (2) Using the inflow rate over many years as input, the agent interacts with the water storage scheduling environment and is trained based on reinforcement learning algorithm to extract water storage scheduling rules; (3) Perform simulated scheduling based on the obtained scheduling rules.
2. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 1, characterized in that: In step (1), the objective function of the advance water storage optimization scheduling model is: The reservoir generates the most electricity during the scheduling period: ; ; In the formula: E is the total power generation during the reservoir scheduling period, in hundreds of millions of kWh; E(t) is the multi-year average power generation at time t, in hundreds of millions of kWh; the reservoir scheduling period refers to the reservoir impoundment period and the subsequent high water level operation period; T is the total step size of the scheduling period, and Y is the total number of years; The time step is in days; K is the power output coefficient; Q(y,t) is the power generation flow of the reservoir at time t in year y, in meters. 3 / s; H(y,t) is the net head of water generated by the reservoir at time t in year y, in meters.
3. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 2, characterized in that: In step (1), the objective function of the advance water storage optimization scheduling model is: Reservoirs pose the least risk for flood control: ; ; In the formula: FCR(t) represents the maximum flood control risk over many years at time t during the reservoir's impoundment period; FCR(y,t) represents the flood control risk of the reservoir at time t in year y; V(y,t) represents the reservoir capacity at time t in year y; m 3 VS(t) represents the reservoir capacity corresponding to the phased flood control limit water level at time t in year y, m. 3 VU represents the normal water level of the reservoir, in meters. 3 VL represents the reservoir capacity corresponding to the flood control limit water level, in meters. 3 ; This is the penalty coefficient.
4. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 3, characterized in that: In step (2), the reinforcement learning framework for training the agent includes the agent and the environment. The agent interacts with the environment at discrete time steps based on the Markov decision process to acquire knowledge, and the agent is trained with the goal of maximizing the value estimate, i.e., updating the value estimate, and finally obtaining the optimal action policy. The reinforcement learning algorithm, as the agent, interacts with the environment containing knowledge of cascade reservoir scheduling to acquire knowledge samples. ; In the formula: s t The initial state of the agent at time t is determined by the scheduling time t and the water level Z. t Composition, water level Z t The discrete values within the reservoir water level constraint range; a t Let r be the action of the agent at time t; t For the action reward during time period t, r t The value is the power generation benefit E(t) generated by the scheduling.
5. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 4, characterized in that: In the Markov decision-making process, the value obtained by an agent in a certain state by taking a certain action is called the action value, denoted as Q value. Reinforcement learning algorithms update Q-values to approximate the optimal value in order to find the optimal strategy for solving multi-stage problems; the optimal Q-values at the beginning and end of each time period satisfy the Bellman equation: ; In the formula: q*(s t ,a t ) represents state s t Take action a t The value obtained later; γ is the discount rate, used to control the impact of future earnings on the present; P st,st+1 This indicates that the cascade reservoirs are in state s. t Transition to the next state s t+1 The Markov state transition probability is used to describe the randomness of the flow in a state, where S is the set of states.
6. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 5, characterized in that: In reinforcement learning, an ε-greedy policy is used to determine the decision action at each stage. The probability expression for action selection in the ε-greedy policy is as follows: ; In the formula: Represents the state s at time t t The probability of randomly selecting an action; Represents the state s at time t t The probability of choosing the action that yields the highest evaluation value; ε is the greed rate.
7. The method for extracting water storage optimization scheduling rules based on reinforcement learning according to claim 6, characterized in that: In step (3), Using the test period data from the inflow data, and through reinforcement learning-based water storage optimization scheduling rules obtained from training, considering water balance constraints, reservoir capacity curve constraints, discharge facility discharge capacity constraints, upper and lower limits and water level fluctuation constraints, output constraints, and boundary condition constraints, the dual objective of the water storage scheduling model is transformed into a single objective for simulating reservoir scheduling. The comprehensive benefit index is expressed as: ; In the formula: R is the comprehensive benefit index; a is the weight.