Long-term scheduling decision-making method, system, equipment and storage medium for cascade hydropower stations
By applying the DDPG deep reinforcement learning algorithm in cascade hydropower stations, it is transformed into the Markov algorithm decision-making process, and the problems of modeling difficulties and local optimal solutions in the existing hydropower station scheduling methods are solved, more efficient and refined scheduling decisions are achieved, and power generation benefits are improved.
Patent Information
- Application Number
- CN202411199246.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The existing hydropower station scheduling methods are difficult to effectively consider all influencing factors, have poor ability to adapt to environmental changes, and traditional mathematical planning and machine learning methods have problems such as low computational efficiency and easy to fall into local optimality.
Deep deterministic strategy gradient (DDPG) deep reinforcement learning algorithm is adopted to transform the long-term optimization and scheduling problems of cascade hydropower stations into Markov algorithm decision-making process. By obtaining basic data and operational conditions, a long-term optimization and scheduling model is built and the results of long-term scheduling decisions are output.
It improves the timeliness of cascade hydropower station scheduling and the availability of results, enhances the refinement of the scheduling strategy, improves the power generation efficiency, and avoids local optimal solutions.
Smart Images

Figure CN118691128B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydropower station scheduling, and in particular, to a long-term scheduling decision-making method, system, device and storage medium for cascade hydropower stations. Background Art
[0002] Hydropower station scheduling is an important guarantee for making full use of water resources and improving power generation efficiency. In recent years, with the cascade rolling development and centralized commissioning of hydropower bases in the southwestern region of China, the scale of hydropower scheduling has expanded rapidly, and the grid's refined requirements for hydropower have increased day by day. The timeliness and result availability of hydropower scheduling face great challenges.
[0003] The optimal scheduling of cascade hydropower stations is a typical discrete, multi-dimensional, non-linear, large-scale, multi-time and multi-space decision-making optimization problem, and its modeling and solution are very difficult. When establishing a mathematical model with existing methods, it is difficult to consider all influencing factors, and the adaptability to changing environments is poor. In terms of optimization and solution, it mainly includes traditional mathematical programming methods and traditional machine learning methods. Mathematical programming methods have problems such as time-consuming calculation, high memory occupancy, and low calculation efficiency; while traditional machine learning algorithms have problems such as being easily limited to local optima and poor reproducibility of results.
[0004] Currently, with the rapid development of big data and artificial intelligence algorithms, reinforcement learning and even deep reinforcement learning algorithms have been preliminarily applied in the optimal scheduling of cascade hydropower stations. However, at present, most use the deep reinforcement learning method of Deep Q-Network (DQN) to construct an optimal scheduling model for cascade hydropower stations. Since the method outputs discrete action values, it affects the refinement degree of the scheduling strategy, and thus affects the economy of the operation of cascade hydropower stations. Under this background, in order to further refine the intelligent scheduling decision-making of cascade hydropower stations and improve the economy of the operation of cascade hydropower stations, it is necessary to study new deep reinforcement learning methods.
[0005] The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0006] The present invention provides a long-term scheduling decision-making method, system, device and storage medium for cascade hydropower stations, thus effectively solving the problems in the background art.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is: a long-term scheduling decision-making method and system for cascade hydropower stations, including:
[0008] Obtain the basic data and operation conditions of the cascade hydropower stations, and construct a long-term optimal scheduling model for the cascade hydropower stations;
[0009] Transform the long-term optimal scheduling problem of cascade hydropower stations into a Markov algorithm decision-making process;
[0010] Use the deep deterministic policy gradient (DDPG) deep reinforcement learning algorithm to solve the Markov decision-making process, and obtain the long-term scheduling decision-making scheme for each hydropower station in the cascade hydropower stations;
[0011] Based on the actual cascade hydropower stations, output the long-term scheduling decision-making results.
[0012] Furthermore, the acquisition of the basic data and operation conditions of the cascade hydropower stations includes: the highest and lowest water levels, the maximum and minimum reservoir capacities, the maximum power generation flow rate, and the maximum output of each hydropower station; use the reservoir capacity and water level data at each moment to fit the reservoir capacity-water level curve of each hydropower station.
[0013] Furthermore, in the long-term optimal scheduling model of the cascade hydropower stations, the highest total power generation within the scheduling period of the cascade hydropower stations is used as the objective function, and the water balance, reservoir capacity, outflow, power generation flow rate, unit output, water level limit, reservoir capacity-water level relationship, and initial water level are used as constraints.
[0014] Furthermore, the objective function of the long-term optimal scheduling model of the cascade hydropower stations includes:
[0015] ;
[0016] In the formula, is the total power generation within the scheduling period of the cascade hydropower stations, with the unit of kWh; is the number of hydropower stations in the cascade basin; is the number of time periods within a scheduling period, which is 365 days in a year here; Δt is the number of hours in a time scale, with the unit of h; i is the hydropower station number; is the time period number; P i,t is the hydropower station i at t the average power generation power during the time period, with the unit of kW; k i is the hydropower station i 's output coefficient, which takes the value of 8.5 here; is the hydropower station i at t the average power generation diversion flow rate during the time period, with the unit of m³ / s; H i,t is the hydropower station i at t the average net head for power generation during the time period, with the unit of m;
[0017] Furthermore, the constraints of the long-term optimal scheduling model for the cascade hydropower stations include:
[0018] Water balance constraint:
[0019] ;
[0020] In the formula, and are the initial and final reservoir storages of the hydropower station at time period, with the unit of m³; at respectively and are the average inflow and average outflow discharges of the hydropower station at time period, with the unit of m³ / s; at respectively
[0021] The water balance constraint between two hydropower stations is:
[0022] ;
[0023] In the formula, is the outflow discharge of the upper-level hydropower station at time period, with the unit of m³ / s; respectively is the natural runoff between the hydropower station and the upper-level hydropower station at time period ;
[0024] Reservoir storage constraint:
[0025] ;
[0026] In the formula, and are the minimum and maximum values of the reservoir storage of the hydropower station at time period, with the unit of m³; at respectively
[0027] Outflow discharge constraint:
[0028] ;
[0029] In the formula, and are the minimum and maximum values of the outflow discharge of the hydropower station at time period, with the unit of m³ / s; at respectively
[0030] Power generation discharge constraint:
[0031] ;
[0032] In the formula, and are the minimum and maximum values of the power generation diversion flow of the hydropower station during period, with the unit of m³ / s;
[0033] Unit output constraint:
[0034] ;
[0035] In the formula, is the maximum value of the allowable output of the hydropower station during period, with the unit of kW;
[0036] Water level limit constraint:
[0037] ;
[0038] In the formula, and are the minimum and maximum values of the net head for power generation of the hydropower station during period, with the unit of m;
[0039] Reservoir storage - water level relationship constraint:
[0040] ;
[0041] The storage - water level curve reflects the functional correspondence between the reservoir storage and the water head of the hydropower station during period;
[0042] Initial water level constraint:
[0043] ;
[0044] In the formula, is the initial water head of the hydropower station with the unit of m.
[0045] Furthermore, transforming the long - term optimal scheduling problem of cascade hydropower stations into a Markov algorithm decision - making process includes the following steps:
[0046] According to the characteristics of the long - term scheduling problem of cascade hydropower stations, define the state, action, state transition function, and reward function in reinforcement learning, and construct a learning process, where,
[0047] State : used to describe the state information of cascade hydropower stations at the current moment, defined as:
[0048] ;
[0049] In the formula, is t the state quantity at time t , which includes time i , as well as the incoming flow, reservoir capacity and water level height of the hydropower station at this time;
[0050] Action : Action is the power generation diversion flow of each hydropower station, expressed as:
[0051] ;
[0052] State transition function: When the incoming water from the upstream at this time is determined, the reservoir capacity and water level at the next time can be determined by controlling the power generation flow at this time. Therefore, the operating environment of cascade hydropower stations is also determined, and the state transition probability of the Markov process is 1, that is:
[0053] ;
[0054] In the formula, when the agent executes action s in state a , a definite environmental state at the next time can be obtained, and this state transition process satisfies various set constraints;
[0055] Reward function : In the long-term optimal scheduling of cascade hydropower stations, the reward value is used to reflect the objective function value of a certain time period, and further the sum of all reward values within the scheduling period is used to represent the objective function within the scheduling period of cascade hydropower stations, expressed as:
[0056] ;
[0057] In the formula, during the scheduling process of cascade hydropower stations, the reward value of each time period is the sum of the power generation of each hydropower station in this time period. When the environmental state of the cascade hydropower station and the action value are determined, the corresponding environmental reward value can be generated for the learning of optimal scheduling.
[0058] Furthermore, solving the Markov decision process by using the deep deterministic policy gradient DDPG deep reinforcement learning algorithm includes the following steps:
[0059] Step 1: Randomly initialize the network parameters of the prediction policy network and the prediction value network and ;
[0060] Step 2: Use the prediction network parameters and to initialize the weight parameters of the target prediction network and the target value network and ;
[0061] Step 3: Initialize the capacity of the experience pool D ;
[0062] Step 4: Initialize the state at the initial moment of the cascade hydropower stations in the environment s 0 ;
[0063] Step 5: Input the state at the initial moment s 0 into the prediction policy network and use the exploration noise N to generate the corresponding action ; where is the prediction policy network with network parameters and is used to output the corresponding policy action;
[0064] Step 6: Input the generated execution action a 0 into the cascade hydropower station environment to obtain the reward r 0 and the state at the next moment s 1 data, and store ( s 0 ,a 0 ,r 0 , s 1 ) as a set of data into the experience pool;
[0065] Step 7: Use the state quantity at the next moment to loop Steps 5 and 6 to generate a series of data groups and save them in the experience pool;
[0066] Step 8: When the number of data groups reaches a certain amount, randomly sample n multidimensional data groups ( s t ,a t ,r t , s t+1 ) as training data and substitute them into the target network to calculate the target value: ; where is the reward value generated by the scheduling action at the corresponding moment, representing the power generation of the cascade hydropower station at a moment in the optimal scheduling of the cascade hydropower station. is the discount factor, used to weaken the influence of future rewards on the total current state rewards;
[0067] Step 9: Update the parameters of the predictive value network by minimizing the loss function : ; where MSE represents the mean square error, used to calculate and measure the error size between the target value and the predicted value;
[0068] Step 10: Update the parameters of the predictive policy network by maximizing the objective function to update the parameters of the predictive policy network : ;
[0069] Step 11: Soft update the parameters of the target policy network and the target value network: , where τ is a hyperparameter much smaller than 1;
[0070] Step 12: Use the updated network parameters to continue to loop through Steps 5 and 6 to update the data groups in the experience pool;
[0071] Step 13: Use the updated experience pool data to continue to take Steps 8 to 11 to update the network parameters;
[0072] Step 14: Loop through Steps 12 and 13 until the network parameters meet the requirements and can achieve the final scheduling decision requirements.
[0073] The present invention also includes a long-term scheduling decision system for a cascade hydropower station, using the method as described above, including:
[0074] A modeling module, used to obtain the basic data and operating conditions of the cascade hydropower station, and construct a long-term optimal scheduling model for the cascade hydropower station;
[0075] A decision-making module, used to transform the long-term optimal scheduling problem of the cascade hydropower station into a Markov algorithm decision-making process;
[0076] A solving module, used to solve the Markov decision-making process by using the deep deterministic policy gradient DDPG deep reinforcement learning algorithm to obtain the long-term scheduling decision schemes of each power station in the cascade hydropower station;
[0077] An application module, used to output the long-term scheduling decision results based on the actual cascade hydropower station.
[0078] The present invention further includes a computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method is implemented.
[0079] The present invention further includes a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0080] The beneficial effects of the present invention are as follows: Based on a data-driven method for scheduling decisions, the present invention can effectively solve the problems of difficult modeling, easy to fall into local optimal solutions, and poor policy flexibility existing in the existing mathematical model-based methods; at the same time, compared with the existing Deep Q-Learning deep reinforcement learning algorithm, its continuous action policy can make the scheduling policy more refined and effectively improve the power generation efficiency of cascade hydropower stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0082] Figure 1 is a flowchart of the method of the present invention;
[0083] Figure 2 is a schematic structural diagram of the system of the present invention;
[0084] Figure 3 is a flowchart of the long-term scheduling intelligent decision-making method for cascade hydropower stations based on the deep deterministic policy gradient algorithm;
[0085] Figure 4 is the storage capacity - water level curve of Jinping-I and Ertan hydropower stations;
[0086] Figure 5 is the DDPG algorithm process;
[0087] Figure 6 is the water level of Jinping-I after intelligent scheduling;
[0088] Figure 7 is the generated flow rate of Jinping-I after intelligent scheduling;
[0089] Figure 8 is the generated power of Jinping-I after intelligent scheduling;
[0090] Figure 9 is the water level of Ertan after intelligent scheduling;
[0091] Figure 10 is the Ertan power generation flow after intelligent scheduling;
[0092] Figure 11 is the Ertan power generation power after intelligent scheduling;
[0093] Figure 12 is the structural schematic diagram of the computer device. Specific implementation manners
[0094] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0095] Embodiment 1
[0096] As Figure 1 shown: A long-term scheduling decision method for cascade hydropower stations includes:
[0097] Obtain the basic data and operation conditions of the cascade hydropower stations, and construct a long-term optimal scheduling model for the cascade hydropower stations;
[0098] Convert the long-term optimal scheduling problem of the cascade hydropower stations into a Markov algorithm decision process;
[0099] Use the DDPG (Deep Deterministic Policy Gradient) deep reinforcement learning algorithm to solve the Markov decision process, and obtain the long-term scheduling decision schemes of each power station in the cascade hydropower stations;
[0100] Based on the actual cascade hydropower stations, output the long-term scheduling decision results.
[0101] By making scheduling decisions through a data-driven method, the problems of difficult modeling, easy to fall into local optimal solutions, and poor policy flexibility existing in the existing mathematical model-based methods can be effectively solved; at the same time, compared with the existing Deep Q-Learning deep reinforcement learning algorithm, its continuous action policy can make the scheduling policy more refined, and can effectively improve the power generation efficiency of cascade hydropower stations.
[0102] The deep reinforcement learning algorithm based on Deep Deterministic Policy Gradient (DDPG) is a reinforcement learning algorithm that performs well in continuous action spaces. It can model and optimize high-dimensional and complex action spaces. At the same time, due to the continuity of actions, it can effectively improve the long-term operation economy of cascade hydropower stations. The DDPG algorithm uses deep neural networks to approximate the value function and policy function, can handle large-scale state and action spaces, and has strong expressive ability. At the same time, it adopts the deterministic policy gradient method, which can converge to a better policy during training and find a solution close to the optimal policy.
[0103] In this embodiment, the basic data and operation conditions of the cascade hydropower station are obtained, including: the highest and lowest water levels, the maximum and minimum reservoir capacities, the maximum power generation flow rate, and the maximum output of each hydropower station; the reservoir capacity-water level curve of each hydropower station is fitted using the reservoir capacity and water level data at each moment.
[0104] In the long-term optimal scheduling model of the cascade hydropower station, the objective function is to maximize the total power generation within the scheduling period of the cascade hydropower station, and the constraints include water balance, reservoir capacity, discharge flow, power generation flow, unit output, water level limit, reservoir capacity-water level relationship, and initial water level.
[0105] The objective function of the long-term optimal scheduling model of the cascade hydropower station includes:
[0106] ;
[0107] In the formula, is the total power generation within the scheduling period of the cascade hydropower station, with the unit of kWh; is the number of hydropower stations in the cascade basin; is the number of time periods within a scheduling period, which is 365 days in a year here; Δt is the number of hours in a time scale, with the unit of h; i is the hydropower station number; is the time period number; P i,t is the average power generation of hydropower station i at t time period, with the unit of kW; k i is the output coefficient of hydropower station i , which takes a value of 8.5 here; is the average power generation diversion flow of hydropower station i at t time period, with the unit of m³ / s; H i,t is the average net head for power generation of hydropower station i at t time period, with the unit of m;
[0108] Among them, the constraints of the long-term optimal scheduling model of cascade hydropower stations include:
[0109] Water balance constraint:
[0110] ;
[0111] In the formula, and are the initial and final reservoir storages of the hydropower station at the time period, with the unit of m³; and are the average inflow and average outflow discharges of the hydropower station at the time period, with the unit of m³ / s; at the time period, respectively;
[0112] The water balance constraint between two hydropower stations is:
[0113] ;
[0114] In the formula, is the outflow discharge of the upper-level hydropower station at the time period, with the unit of m³ / s; is the natural runoff between the hydropower station and the upper-level hydropower station at the time period;
[0115] Reservoir storage constraint:
[0116] ;
[0117] In the formula, and are the minimum and maximum values of the reservoir storage of the hydropower station at the time period, with the unit of m³; respectively;
[0118] Outflow discharge constraint:
[0119] ;
[0120] In the formula, and are the minimum and maximum values of the outflow discharge of the hydropower station at the time period, with the unit of m³ / s; respectively;
[0121] Power generation discharge constraint:
[0122] ;
[0123] In the formula, and are respectively the minimum and maximum values of the flow rate diverted for power generation by the hydropower station during period, with the unit of m³ / s;
[0124] Unit output constraint:
[0125] ;
[0126] In the formula, is the maximum value of the allowable output of the hydropower station during period, with the unit of kW;
[0127] Water level limit constraint:
[0128] ;
[0129] In the formula, and are respectively the lowest and highest values of the net head for power generation of the hydropower station during period, with the unit of m;
[0130] Reservoir storage - water level relationship constraint:
[0131] ;
[0132] The storage - water level curve reflects the functional correspondence between the reservoir storage and the water head of the hydropower station during period;
[0133] Initial water level constraint:
[0134] ;
[0135] In the formula, is the initial water head of the hydropower station with the unit of m.
[0136] In this embodiment, the long - term optimal scheduling problem of cascade hydropower stations is transformed into a Markov algorithm decision process, including the following steps:
[0137] According to the characteristics of the long - term scheduling problem of cascade hydropower stations, define the state, action, state - transition function, and reward function in reinforcement learning, and construct a learning process, where,
[0138] State : Used to describe the state information of the cascade hydropower station at the current moment, defined as:
[0139] ;
[0140] In the formula, is t the state quantity at time t , which includes time i , as well as the inflow, storage capacity, and water level height of the hydropower station at this time;
[0141] Action : Action is the power generation diversion flow of each hydropower station, expressed as:
[0142] ;
[0143] State transition function: When the upstream water inflow at this time is determined, the storage capacity and water level at the next time can be determined by controlling the power generation flow at this time. Therefore, the operating environment of cascade hydropower stations is also determined, and the state transition probability of the Markov process is 1, that is:
[0144] ;
[0145] In the formula, when the agent executes action s in state a , a definite environmental state at the next time can be obtained, and this state transition process satisfies various set constraints;
[0146] Reward function : In the long-term optimal scheduling of cascade hydropower stations, the reward value is used to reflect the objective function value of a certain time period, and further the sum of all reward values within the scheduling period is used to represent the objective function within the scheduling period of cascade hydropower stations, expressed as:
[0147] ;
[0148] In the formula, during the scheduling process of cascade hydropower stations, the reward value of each time period is the sum of the power generation of each hydropower station in this time period. When the environmental state of the cascade hydropower station and the action value are determined, the corresponding environmental reward value can be generated for the learning of optimal scheduling.
[0149] Among them, the Markov decision process is solved using the deep deterministic policy gradient DDPG deep reinforcement learning algorithm, including the following steps:
[0150] Step 1: Randomly initialize the network parameters of the prediction policy network and the prediction value network and ;
[0151] Step 2: Using the prediction network parameters and to initialize the weight parameters of the target prediction network and the target value network and ;
[0152] Step 3: Initialize the capacity of the experience pool D ;
[0153] Step 4: Initialize the state at the initial moment of the cascade hydropower stations in the environment s 0 ;
[0154] Step 5: Input the state at the initial moment s 0 into the prediction policy network and generate the corresponding action using the exploration noise N ; where is the prediction policy network with network parameters for outputting the corresponding policy action;
[0155] Step 6: Input the generated execution action a 0 into the cascade hydropower station environment to obtain the reward r 0 and the state at the next moment s 1 data, and store ( s ,a 0 ,r 0 , s 0 1 ) as a set of data into the experience pool;
[0156] Step 7: Using the state quantity at the next moment, loop through Step 5 and Step 6 to generate a series of data sets and save them in the experience pool;
[0157] Step 8: When the number of data sets reaches a certain amount, randomly sample n multidimensional data sets ( s t ,a t ,r t , s t+1 ) as training data and substitute them into the target network to calculate the target value: ; where is the reward value generated by the corresponding moment scheduling action, representing the power generation amount of the cascade hydropower stations at a moment in the optimal scheduling of cascade hydropower stations, is a discount factor used to weaken the impact of future rewards on the total current state rewards;
[0158] Step 9: Update the parameters of the predictive value network by minimizing the loss function : ; where MSE represents the mean square error, which is used to calculate and measure the error magnitude between the target value and the predicted value;
[0159] Step 10: Update the parameters of the predictive policy network by maximizing the objective function to update the parameters of the predictive policy network : ;
[0160] Step 11: Soft update the parameters of the target policy network and the target value network: , where τ is a hyperparameter much smaller than 1;
[0161] Step 12: Continue to loop Steps 5 and 6 using the updated network parameters to update the data groups in the experience pool;
[0162] Step 13: Continue to take Steps 8 to 11 to update the network parameters using the updated experience pool data;
[0163] Step 14: Loop Steps 12 and 13 until the network parameters meet the requirements and can achieve the final scheduling decision requirements.
[0164] As Figure 2 shown, this embodiment also includes a long-term scheduling decision system for cascade hydropower stations. Using the method as described above, it includes:
[0165] A modeling module for obtaining the basic data and operation conditions of the cascade hydropower station and constructing a long-term optimal scheduling model for the cascade hydropower station;
[0166] A decision-making module for transforming the long-term optimal scheduling problem of the cascade hydropower station into a Markov algorithm decision-making process;
[0167] A solving module for solving the Markov decision-making process using the deep deterministic policy gradient DDPG deep reinforcement learning algorithm to obtain the long-term scheduling decision schemes for each power station in the cascade hydropower station;
[0168] An application module for outputting the long-term scheduling decision results based on the actual cascade hydropower station.
[0169] Embodiment 2
[0170] As Figure 3 shown, the long-term scheduling intelligent decision-making method and system for cascade hydropower stations based on the deep deterministic policy gradient algorithm include the following steps:
[0171] S1. Investigate the basic data and operation conditions of cascade hydropower stations, and establish a long-term optimal operation model for cascade hydropower stations;
[0172] The research object is the cascade hydropower stations in the lower reaches of the Yalong River Basin, mainly including 5 hydropower stations: Jinping-I, Jinping-II, Guandi, Ertan, and Tongzilin. Among them, Jinping-I Hydropower Station has annual operation, Ertan Reservoir has seasonal regulation, and Jinping-II Reservoir, Guandi Reservoir, and Tongzilin Reservoir all have daily regulation. Therefore, during the long-term operation process, long-term optimal operation is carried out for Jinping-I Hydropower Station and Ertan Hydropower Station. The optimal operation time is the whole year from June 1, 2020 to May 31, 2021, and the time scale is set as one day.
[0173] The basic data of Jinping-I and Ertan Hydropower Stations include the highest and lowest water levels, the maximum and minimum reservoir capacities, the maximum power generation flow rate, and the maximum output, as shown in Table 1. The reservoir capacity-water level curves of the two hydropower stations are shown in Figure 4 .
[0174] Table 1 Basic Data of Jinping-I and Ertan Hydropower Stations
[0175] Power Station Jinping-I Hydropower Station Ertan Hydropower Station Highest Water Level (m) 1880 1200 Lowest Water Level (m) 1800 1155 Maximum Storage Capacity (100 million m³) 77.6 57.9 Minimum Storage Capacity (100 million m³) 28.5 24.2 Maximum Generation Discharge (m³ / s) 2024 2226 Maximum Output (MW) 3600 3300
[0176] Establish a long-term optimal operation model for cascade hydropower stations, and the objective function is:
[0177] (1)
[0178] In the formula, is the total power generation during the operation period of the cascade hydropower stations, kWh; is the number of hydropower stations in the cascade basin; is the number of time periods in an operation period, which is 365 days in a year here; Δt is the number of hours in a time scale, h; i is the hydropower station number; is the time period number; P i,t is the average power generation i of hydropower station t during the k i is the output coefficient of hydropower station i , which takes a value of 8.5 here; is the average water intake flow rate for power generation i of hydropower station t during the H i,t is the average net head for power generation i of hydropower station t during the
[0179] The constraints include:
[0180] (1) Water balance constraint:
[0181] (2)
[0182] In the formula, and are the initial and final reservoir capacities of the hydropower station at time period in , in m³; and are the average inflow and average outflow discharges of the hydropower station at time period in , in m³ / s.
[0183] The water balance constraint between the two hydropower stations is:
[0184] (3)
[0185] In the formula, is the outflow discharge of the upper-level hydropower station at time period , in m³ / s.
[0186] (2) Reservoir capacity constraint
[0187] (4)
[0188] In the formula, and are the minimum and maximum reservoir capacities of the hydropower station at time period in , in m³.
[0189] (3) Outflow discharge constraint
[0190] (5)
[0191] In the formula, and are the minimum and maximum outflow discharges of the hydropower station at time period in , in m³ / s.
[0192] (4) Power generation flow constraint
[0193] (6)
[0194] In the formula, and are the minimum and maximum power generation diversion flows of the hydropower station at time period in , in m³ / s.
[0195] (5) Unit output constraint
[0196] (7)
[0197] In the formula is the maximum allowable output of the hydropower station during period, in kW.
[0198] (6) Water level limit constraint
[0199] (8)
[0200] In the formula and are respectively the minimum and maximum values of the net head for power generation of the hydropower station during period, in m.
[0201] (7) Reservoir storage - water level relationship constraint
[0202] (9)
[0203] The storage - water level curve reflects the functional correspondence between the reservoir storage and the water head of the hydropower station during period.
[0204] (8) Initial water level constraint
[0205] (9)
[0206] In the formula is the initial water head of the hydropower station in m.
[0207] S2, transforms the long - term optimal scheduling problem of cascade hydropower stations into a Markov decision process;
[0208] According to the characteristics of the long - term scheduling problem of cascade hydropower stations, the Markov decision process defines the state, action, state transition function, and reward function in reinforcement learning to construct a learning process. Among them,
[0209] State : Used to describe the state information of cascade hydropower stations at the current moment, providing reference guidance for the scheduling decision of the intelligent agent through the state information of cascade hydropower stations, and is defined as:
[0210] (10)
[0211] In the formula is tThe state variables at a moment, which mainly include the moment t , and the inflow, storage capacity, and water level of the hydropower station at this moment i .
[0212] Action : The long-term scheduling of cascade hydropower stations realizes the long-term optimal scheduling decision of cascade hydropower stations by controlling the power generation flow of each hydropower station. Therefore, the action is the power generation diversion flow of each hydropower station, expressed as:
[0213] (11)
[0214] State transition function: When the upstream inflow is determined at this moment, by controlling the power generation flow at this moment, the storage capacity and water level at the next moment can be determined. Therefore, the operating environment of the cascade hydropower station is also determined, and the state transition probability of the Markov process is 1, that is:
[0215] (12)
[0216] In the formula, when the agent executes the action s in the state a , a definite environmental state at the next moment can be obtained, and this state transition process satisfies various set constraints.
[0217] Reward function : In the long-term optimal scheduling of cascade hydropower stations, the reward value is used to reflect the objective function value in a certain time period, and further the sum of all reward values within the scheduling period is used to represent the objective function of the cascade hydropower station within the scheduling period; (13)
[0218] In the formula, during the scheduling process of the cascade hydropower station, the reward value for each time period is the sum of the power generation of each hydropower station in this time period. When the environmental state of the cascade hydropower station and the action value are determined, the corresponding environmental reward value can be generated for the learning of optimal scheduling.
[0219] S3. Use the DDPG deep reinforcement learning algorithm to solve the Markov decision process to obtain the long-term scheduling decision scheme for each power station in the cascade hydropower station;
[0220] The specific steps are shown in Figure 5 , specifically:
[0221] Step 1: Randomly initialize the network parameters of the prediction policy network and the prediction value network and ;
[0222] Step 2: Initialize the target prediction network using the prediction network parameters and the weight parameters of the target value network and ;
[0223] Step 3: Initialize the capacity of the experience pool D ;
[0224] Step 4: Initialize the state at the initial moment of the cascade hydropower stations in the environment s 0 ;
[0225] Step 5: Input the state at the initial moment s 0 into the prediction policy network and generate the corresponding action using the exploration noise N ; ;
[0226] Step 6: Input the generated execution action a 0 into the cascade hydropower station environment to obtain the reward r 0 and the state at the next moment s 1 data, and store ( s 0 ,a 0 ,r 0 , s 1 ) as a set of data into the experience pool;
[0227] Step 7: Using the state quantity at the next moment, loop through Step 5 and Step 6 to generate a series of data groups and save them in the experience pool;
[0228] Step 8: When the number of data groups reaches a certain amount, randomly sample n multidimensional data groups ( s t ,a t ,r t , s t+1 ) as training data and substitute them into the target network to calculate the target value: ;
[0229] Step 9: Update the prediction value network parameters by minimizing the loss function : ;
[0230] Step 10: Use the maximized objective function to update the parameters of the prediction policy network : ;
[0231] Step 11: Soft-update the parameters of the target policy network and the target value network: , where τ is a hyperparameter much smaller than 1;
[0232] Step 12: Use the updated network parameters to continue looping through Steps 5 and 6 to update the data groups in the experience pool;
[0233] Step 13: Use the updated experience pool data to continue taking Steps 8 to 11 to update the network parameters;
[0234] Step 14: Loop through Steps 12 and 13 until the network parameters meet the requirements and can achieve the final scheduling decision requirements.
[0235] S4. Use DDPG to make long-term scheduling intelligent decisions for Yalong River Jinping-I Hydropower Station and Ertan Hydropower Station for the whole year from June 1, 2020 to May 31, 2021. The time scale of the scheduling is set to one day. The hyperparameter settings of DDPG are shown in Table 2.
[0236] Table 2 Hyperparameter values of the DDPG algorithm
[0237] Sequence Hyperparameter Value Specific Description 1 episode num 500 Total Number of Time Loops 2 learning rate <![CDATA[3×10 -9 > Learning Rate 3 optimizer Adam Selected Optimizer 4 max_buffer 40000 Size of Experience Replay Buffer 5 batch size 64 Number of Samples for Batch Sampling 6 target update rate 0.005 Target Network Update Rate 7 discount factor 0.996 Discount Coefficient 8 exploration noise (0,0.1) Exploration Noise
[0238] The scheduling decision results of Yalong River Jinping-I Hydropower Station and Ertan Hydropower Station for the whole year from June 1, 2020 to May 31, 2021 are shown in Figure 6 and Figure 11 .
[0239] From Figures 6 to 8It can be seen that at the initial stage of the dispatch, the generated power flow of Jinping-I Hydropower Station under DDPG dispatch is much greater than the actual generated power flow, and it is in the state of operating at the maximum generated power flow. At the same time, the rising speed of the reservoir water level corresponding to it is slower than that of the actual reservoir water level. In this way, the reservoir capacity can be utilized more fully for power generation before the flood season comes; at the middle stage of the dispatch, the generated power flow under DDPG dispatch is less than the actual generated power flow, and its dispatch reservoir water level remains at a relatively high level compared with the actual water level; at the end stage of the dispatch, the generated power flow of Jinping-I Hydropower Station under DDPG dispatch begins to increase and is greater than the actual generated power flow, and the corresponding water level also begins to decline from the high water level, so as to make full use of the reservoir capacity for power generation. From the perspective of the generated power, at the initial stage of the dispatch, the generated power of Jinping-I Hydropower Station under DDPG dispatch is greater than the actual generated power, and then the generated power begins to decline, which is related to the reduction of the generated power flow at the end of the flood season. At the middle stage of the dispatch, the generated power under DDPG dispatch is less than the actual generated power, so as to maximize the utilization of the water level advantage for power generation; at the end stage of the dispatch, the generated power under DDPG dispatch increases rapidly and far exceeds the actual generated power.
[0240] It can be seen from Figures 9 to 11 that at the initial stage of the dispatch, the generated power flow of Ertan Hydropower Station under DDPG dispatch is greater than the actual generated power flow. Therefore, the rising speed of its dispatch water level is slower than that of the actual reservoir water level; at the middle stage of the dispatch, the generated power flow of Ertan Hydropower Station under DDPG dispatch begins to decrease compared with the actual generated power flow, so as to keep the water level at a higher position; at the end stage of the dispatch, the generated power flow of Ertan Hydropower Station under DDPG dispatch begins to increase rapidly and is greater than the actual generated power flow, so as to make full use of the reservoir water volume for power generation. From the perspective of the generated power, at the initial stage of the dispatch, the generated power curve of Ertan is greater than the actual generated power at the initial stage of the dispatch; while at the middle stage of the dispatch, both the generated power under DDPG dispatch and the actual generated power begin to decline, and the generated power under DDPG dispatch is lower; at the end stage of the dispatch, DDPG dispatch increases the generated power and continuously maintains it at a relatively high level of operation, increasing the final generated power value.
[0241] Furthermore, the DQN method is further adopted for the long-term dispatch intelligent decision-making of the Yalong River cascade hydropower stations. The generated power of the cascade hydropower stations under each method is shown in Table 3.
[0242] Table 3 Generated power of cascade hydropower stations under actual dispatch, DQN and DDPG algorithms (unit: 100 million kWh)
[0243] Power Station Actual Scheduling DQN Scheduling DDPG Scheduling Jinping-I 188.85 190.05 195.23 Ertan 158.20 192.08 193.62 Total of Cascade 347.05 382.13 388.85
[0244] As can be seen from Table 3, in the intelligent scheduling decision based on the DDPG deep reinforcement learning algorithm, the total annual power generation of cascade hydropower stations is 3,888.5 kWh respectively, which is 4.18 billion kWh higher than the annual power generation of the historical actual scheduling decision and 672 million kWh higher than the annual power generation of the DQN scheduling decision. Compared with the actual scheduling decision, the method of the present invention can significantly improve the power generation benefit of cascade hydropower stations; at the same time, compared with the existing DQN deep reinforcement learning scheduling decision, since the present invention adopts continuous action values, the scheduling strategy is more refined and has better power generation benefits.
[0245] Please refer to Figure 12 the structural schematic diagram of the computer device provided by the embodiment of the present application shown in. A computer device 400 provided by an embodiment of the present application includes: a processor 410 and a memory 420. The memory 420 stores a computer program executable by the processor 410. When the computer program is executed by the processor 410, the above method is executed.
[0246] An embodiment of the present application also provides a storage medium 430. A computer program is stored on the storage medium 430. When the computer program is run by the processor 410, the above method is executed.
[0247] Among them, the storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviation: SRAM), electrically erasable programmable read-only memory (abbreviation: EEPROM), erasable programmable read-only memory (abbreviation: EPROM), programmable read-only memory (abbreviation: PROM), read-only memory (abbreviation: ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0248] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The meaning of "a plurality" is two or more unless otherwise specifically defined.
[0249] In the present invention, unless otherwise clearly defined and limited, the terms "installed", "connected", "connected to", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0250] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not have to be directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0251] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0252] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0253] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0254] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0255] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A long-term dispatch decision method for cascade hydropower stations, characterized in that: The steps include: Obtain the basic data and operation status of cascade hydropower stations and build a long-term optimization scheduling model for cascade hydropower stations; The long-term optimal dispatching problem of cascade hydropower stations is transformed into a Markov algorithm decision-making process; The Markov decision process is solved using the deep deterministic policy gradient (DDPG) deep reinforcement learning algorithm to obtain the long-term scheduling decision plan for each power station in the cascade hydropower station; Based on actual cascade hydropower stations, output long-term scheduling decision results; The long-term optimization scheduling problem of cascade hydropower stations is converted into a Markov algorithm decision process, including the following steps: According to the characteristics of the long-term scheduling problem of cascade hydropower stations, the state, action, state transfer function and reward function in reinforcement learning are defined to construct a learning process, in which: Status S t : It is used to describe the status information of the cascade hydropower station at the current moment and is defined as: In the formula, S t is the state quantity at time t, which includes time t and the inflow flow of hydropower station i at that time Storage capacity V i,t And the water level H i,t ; Action a t : Action a t is the power generation reference flow of each hydropower station, expressed as: In the formula, is the average power generation reference flow of hydropower station i in period t; State transfer function: When the upstream water flow is determined at this moment, the reservoir capacity and water level at the next moment can be determined by controlling the power generation flow at this moment. Therefore, the operating environment of the cascade hydropower station is also determined, and the state transfer probability of the Markov process is 1, that is: p(s′|s,a)=1; In the formula, after the agent performs action a in state s, it can obtain a certain environment state s′ at the next moment, and the state transition process satisfies various set constraints; Reward function r t (s t , a t ): In the long-term optimal dispatching of cascade hydropower stations, the reward value is used to reflect the objective function value of a certain period of time, and the sum of all reward values in the dispatching period is further used to represent the objective function in the dispatching period of the cascade hydropower station, which is expressed as: Where N is the number of hydropower stations in the cascade basin, P i,t is the average power generation of hydropower station i in period t, in kW; k i is the output coefficient of hydropower station i; during the dispatching process of cascade hydropower stations, the reward value of each time period is the sum of the power generation of each hydropower station in this time period. When the environmental state H of the cascade hydropower station is determined i,t And the action value After that, the corresponding environmental reward value can be generated for learning of optimal scheduling; The method of solving the Markov decision process using a deep deterministic policy gradient (DDPG) deep reinforcement learning algorithm comprises the following steps: Step 1: Randomly initialize the prediction strategy network μ θ (s) and the predicted value network Q w (s,a) network parameters w and θ; Step 2: Use the prediction network parameters w and θ to initialize the weight parameters w′ and θ′ of the target prediction network and the target value network; Step 3: Initialize the experience pool capacity D; Step 4: Initialize the initial state s0 of the cascade hydropower station in the environment; Step 5: Input the initial state s0 into the prediction strategy network a = μ θ (s), and use the exploration noise N to generate the corresponding action a0 = μ θ (s0)+N; where μ θ (s) is a prediction strategy network with a network parameter of θ, which is used to output corresponding strategy actions; Step 6: Input the generated execution action a0 into the cascade hydropower station environment, obtain the reward r0 and the next state s1 data, and store (s0, a0, r0, s1) as a set of data into the experience pool; Step 7: Using the state quantity at the next moment, loop through steps 5 and 6 to generate a series of data groups and save them in the experience pool; Step 8: When the number of data sets reaches the set number, randomly sample n multidimensional data sets (s t ,a t ,r t ,s t+1 ) as training data and substitute it into the target network to calculate the target value: y = r + γQ w′ (s′,μ θ′ (s′)); where r is the reward value generated by the scheduling action at the corresponding moment, which represents the power generation of the cascade hydropower station at a moment in the optimal scheduling of cascade hydropower stations, and γ is the discount coefficient, which is used to weaken the impact of future rewards on the sum of current state rewards; Step 9: Update the prediction value network parameter w by minimizing the loss function: L(w) = MSE[Q w (s,a),r+γQ w′ (s′,μ θ′ (s′))]; where MSE stands for mean square error, which is used to calculate the error between the target value and the predicted value; Step 10: Maximize the objective function Q w (s, a) to update the prediction strategy network parameters θ: Loss = -Q w (s, a); Step 11: Soft update target policy network and target value network parameters: Where τ is a hyperparameter much smaller than 1; Step 12: Use the updated network parameters to continue looping steps 5 and 6 to update the data set in the experience pool; Step 13: Continue to take steps 8 to 11 to update the network parameters using the updated experience pool data; Step 14: loop step 12 and step 13 until the network parameters meet the requirements and the final scheduling decision requirements can be achieved; The basic data and operation status of the cascade hydropower stations are obtained, including the highest and lowest water levels, the maximum and minimum reservoir capacities, the maximum power generation flow and the maximum output of each hydropower station; the reservoir capacity and water level data at each moment are used to fit the reservoir capacity-water level curve of each hydropower station; In the long-term optimization scheduling model of the cascade hydropower station, the maximum total power generation within the scheduling period of the cascade hydropower station is taken as the objective function, and the water balance, reservoir capacity, outflow flow, power generation flow, unit output, water level limit, reservoir capacity and water level relationship and initial water level are taken as constraints; The objective function of the long-term optimal dispatching model of the cascade hydropower stations includes: Where E is the total power generation of the cascade hydropower station in the dispatching cycle, in kWh; N is the number of hydropower stations in the cascade basin; T is the number of time periods in a dispatching cycle, here there are 365 days in a year; Δt is the number of hours in a time scale, in h; i is the hydropower station number; t is the time period number; P i,t is the average power generation of hydropower station i in period t, in kW; k i is the output coefficient of hydropower station i, which is 8.5 here; is the average power generation flow of hydropower station i in period t, in m 3 / s;H i,t is the average net water head of power generation at hydropower station i in period t, in m; The constraints of the long-term optimal dispatch model of the cascade hydropower stations include: Water balance constraints: Where V i,t and V i,t+1 are the initial and final storage capacities of hydropower station i in period t, in m 3 ; and are the average inflow and outflow of hydropower station i in period t, respectively, in m 3 / s; The water balance constraint between the two hydropower stations is: In the formula, is the outflow of the upper-level hydropower station in period t, in m 3 / s;q i,t is the natural runoff between hydropower station i and the upper hydropower station i-1 in period t; Reservoir capacity constraints: In the formula, and are the minimum and maximum reservoir capacities of hydropower station i in period t, in m 3 ; Outbound traffic constraints: In the formula, and are the minimum and maximum outflows of hydropower station i in period t, in m 3 / s; Power generation flow constraints: In the formula, and are the minimum and maximum values of the power generation reference flow of hydropower station i in period t, in m 3 / s; Unit output constraints: In the formula, is the maximum output allowed by hydropower station i in period t, in kW; Water level limit constraints: In the formula, and are the minimum and maximum values of the net water head of power generation of hydropower station i in period t, in m; Reservoir capacity and water level relationship constraints: H i,t =f(V i,t ); The reservoir capacity and water level curve reflects the functional relationship between the reservoir capacity and the hydraulic head of hydropower station i in period t; Initial water level constraint: In the formula, is the initial water head of hydropower station i, in m.
2. A long-term dispatching decision system for cascade hydropower stations, characterized in that: Using the method as claimed in claim 1, comprising: Modeling module, used to obtain basic data and operation status of cascade hydropower stations and build a long-term optimization scheduling model for cascade hydropower stations; A decision-making module, which is used to transform the long-term optimal scheduling problem of cascade hydropower stations into a Markov algorithm decision-making process; A solution module is used to solve the Markov decision process using a deep deterministic policy gradient (DDPG) deep reinforcement learning algorithm to obtain a long-term scheduling decision plan for each power station in the cascade hydropower station; The application module is used to output long-term scheduling decision results based on actual cascade hydropower stations.
3. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method as claimed in claim 1 is implemented.
4. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method as claimed in claim 1 is implemented.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
Electricity and natural gas comprehensive energy system online scheduling method based on deep strategy optimization
CN112290535A