Multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm
By combining deep reinforcement learning and deterministic policy gradient algorithm, the target weights of reservoir scheduling are dynamically adjusted, which solves the problem of insufficient flexibility of scheduling strategies caused by fixed weight design, realizes adaptive optimization of reservoir scheduling in complex environments, and improves the adaptability and robustness of scheduling strategies.
Patent Information
- Application Number
- CN202510940749.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing deep reinforcement learning methods use fixed weights to design incentive functions in reservoir operation, resulting in insufficient flexibility in the operation strategy and making it difficult to achieve real-time dynamic coordinated optimization of the operation objectives of each reservoir according to the environmental status.
A multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm is adopted. By dynamically adjusting the target weights and combining flood control benefits, power generation benefits and water resource benefits, a multi-objective optimization scheduling model is established. The deep deterministic policy gradient algorithm is used for training to achieve multi-objective optimization intelligent scheduling of reservoirs.
The adaptive optimization scheduling capability of the reservoir scheduling strategy under non-stationary hydrological conditions is realized, the scheduling priority can be dynamically switched, and the adaptability and robustness of the scheduling strategy are improved. Especially in the event of sudden floods or extreme droughts, different scheduling objectives can be more effectively taken into account to ensure the safety and efficiency of scheduling.
Smart Images

Figure CN120430478B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of reservoir optimization intelligent scheduling, and specifically relates to a reservoir multi-objective optimization intelligent scheduling method based on deep reinforcement learning and deterministic policy gradient algorithm. Background Art
[0002] Traditional reservoir operation methods mainly include rule-based or experience-based operation and model-based optimization operation. In rule-based or experience-based operation methods, operation decisions are heavily dependent on historical data and pre-set rule curves, and experience-based operation methods often rely on the subjective experience of the operator to make operation decisions. Therefore, these methods have weak adaptability to complex or extreme hydrological conditions and cannot respond quickly and effectively to situations such as sudden floods or droughts. In addition, in multi-objective reservoir systems, the data processing and computational complexity of traditional operation methods are huge, making it difficult to simultaneously consider multiple objectives such as water supply, flood control, power generation, and ecology, and it is difficult to accurately characterize the nonlinear relationships between relevant factors.
[0003] In recent years, artificial intelligence (AI) technologies, particularly deep reinforcement learning (DL), have been introduced into the field of reservoir operation, enabling autonomous learning and optimization of operation strategies through interaction with the environment. Previous studies have shown that DL methods can transform reservoir operation problems into Markov decision processes. Deep neural networks are used to autonomously learn the mapping between reservoir capacity dynamics, water level demand, grid load, and other information, and reservoir water release decisions, thereby improving decision-making accuracy and efficiency in reservoir system operation. These DL-based research efforts often employ a linearly weighted approach to combine multiple operation objectives into a single incentive function, balancing the priorities of each objective using pre-set fixed weights.
[0004] However, most current scheduling research based on deep reinforcement learning (DRL) uses fixed weight coefficients when designing the activation function, resulting in a lack of flexibility in the learned strategies. In actual reservoir operation, conditions such as water inflow, water level, and water demand change over time. However, fixed-weight strategies maintain constant weights after training and lack the ability to dynamically adjust the importance of optimization objectives based on real-time water conditions. Fixed-weight activation function designs cannot meet the requirements of dynamic target adjustment, making it difficult for existing strategies to quickly adapt to changes in water inflow or water storage, and thus difficult to achieve real-time coordinated optimization of various scheduling objectives.
[0005] In summary, it is difficult for existing technologies to take into account both flexibility and multi-objective optimization requirements, and there is an urgent need for a scheduling method that can dynamically adjust the objective weights. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm, aiming to solve the problem that the deep reinforcement learning method in the existing technology generally adopts fixed weights to design the incentive function, resulting in insufficient flexibility of the learned scheduling strategy and difficulty in achieving real-time dynamic coordinated optimization of the scheduling objectives of each reservoir according to the environmental status.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is: a multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm, comprising:
[0008] Step 1: Collect the original hydrological data and reservoir characteristic curve of the target reservoir, and preprocess the original hydrological data;
[0009] Step 2: According to the main functions and scheduling requirements of the target reservoir, the optimization objectives are determined from the three aspects of flood control benefits, power generation benefits and water resource benefits, and a multi-objective optimization scheduling model including objective functions and constraints is established;
[0010] Step 3: Map the multi-objective optimization scheduling model into a Markov decision process, thereby establishing a corresponding reinforcement learning MDP environment;
[0011] Step 4: Select a deep deterministic policy gradient algorithm to conduct interactive training with the environment and build a multi-objective optimization intelligent scheduling model for reservoirs;
[0012] Step 5: Deploy the trained reservoir multi-objective optimization intelligent scheduling model in the scheduling system of the target reservoir to perform multi-objective optimization intelligent scheduling on the reservoir.
[0013] In a preferred solution, in step one, preprocessing the original hydrological data includes preprocessing missing values and outliers. For missing values in the original hydrological data, the missing values are filled by interpolation using the nearest neighbor method. For outliers in the original hydrological data, the outliers are corrected by the nearest neighbor method.
[0014] In the preferred solution, in the step 2, in the multi-objective optimization scheduling model, the flood control benefit adopts maximizing the annual flood interception volume as the objective function, the power generation benefit adopts maximizing the annual power generation as the objective function, and the water resource benefit adopts minimizing the abandoned water volume as the objective function;
[0015] The constraints include outflow restrictions, water level restrictions, water balance principle, power station output restrictions, water level and storage capacity relationship, and the relationship between outflow and downstream water level.
[0016] In a preferred solution, in step three, the multi-objective optimization scheduling model is mapped into a Markov decision process, including constructing a complete MDP environment, defining the state space, action space, incentive function and state transition law function.
[0017] In a preferred solution, the state space is expressed as:
[0018] ;
[0019] in: Indicates the target reservoir at the current moment t water storage capacity; represents the inflow of the target reservoir at the current moment; Indicates that the outflow data of the target reservoir at the previous moment is used; Indicates the target reservoir in the future at the current moment N Inflow flow forecast value within a period;
[0020] ;
[0021] In the above formula: Respectively 、 2. 、 ,..., Inflow flow forecast value within the time period.
[0022] In a preferred solution, the calculation steps of the activation function are:
[0023] 1) Set incentive functions for flood control benefits, power generation benefits, water resource benefits, water level and flow respectively 、 、 、 as well as , respectively:
[0024] In terms of flood control benefits, the simulated flood interception volume is used as the incentive function, which is defined as follows:
[0025] ;
[0026] ;
[0027] Where: represents the incentive function of flood control benefits, and Respectively Inbound and outbound traffic for each period, Indicates the flow threshold for starting to calculate the reservoir's flood control capacity. The full power generation flow in the current period; when the inflow of the target reservoir exceeds and greater than When the outflow is less than the inflow, the flood control capacity of this period is calculated; is the time period length;
[0028] In terms of power generation efficiency, the simulated power generation is used as the incentive function, which is defined as follows:
[0029] ;
[0030] Where: represents the incentive function of power generation efficiency, represents the comprehensive output coefficient, is the water head for the current period, is the power generation flow;
[0031] In terms of water resource benefits, the simulated abandoned water volume is used as the incentive function, which is defined as follows:
[0032] ;
[0033] ;
[0034] Where: represents the incentive function of water resource benefits, is the abandoned water flow;
[0035] In terms of water level, when the reservoir water level Falling in the water level safety zone ]When you are in the office, give positive incentives. 、 They represent the minimum and maximum values of the water level safety interval respectively. When the reservoir water level exceeds the water level safety interval, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is:
[0036] ;
[0037] Where: represents the activation function of the water level, is the reservoir water level at the current moment, is the positive excitation constant when the water level is in the safe range;
[0038] In terms of flow, when the reservoir outflow Falling in the traffic safety range[ ]When you are in the office, give positive incentives. 、 They represent the minimum and maximum values of the traffic safety interval respectively. When the outbound traffic exceeds the traffic safety interval, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is:
[0039] ;
[0040] Where: represents the incentive function of the flow, represents the traffic scaling factor, is the positive excitation constant when the outbound flow is in the safe range;
[0041] 2) Introduction of dynamic weights
[0042] First, define three intermediate variables 、 、 :
[0043] ;
[0044] ;
[0045] ;
[0046] Where: It is the maximum storage capacity of the reservoir. is the maximum inflow;
[0047] Define a linear combination concatenation index for each activation function as an unnormalized weight:
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] ;
[0053] Where, Reflects the critical water level threshold, It reflects the response sensitivity of the intelligent scheduling method to the current inflow and future forecast flow. It reflects the degree of influence of the current water storage ratio on the target weight. The three parameters are determined by the reservoir at different operating times and historical data;
[0054] use Softmax Function, converting the above weights into final weights:
[0055] ;
[0056] Where, i = 1, 2, 3, 4, 5 means , , The unnormalized weights of i = 1,2, 3, 4, 5 means The corresponding final weight;
[0057] 3) Total activation function The calculation formula is:
[0058] .
[0059] In the preferred scheme, the state transfer law function is that after the reservoir executes the discharge flow, the storage capacity is updated according to the water balance relationship to complete the state transfer. The water balance relationship is the change in water storage capacity of the target reservoir in the adjacent time period t to t+1, which is equal to the difference between the inflow and outflow water in the same period.
[0060] In a preferred solution, in step 4, a deep deterministic policy gradient algorithm is used to obtain the incentive signal and update the neural network parameters through continuous interactive training with the MDP environment.
[0061] In the preferred solution, the calculation process of the deep deterministic policy gradient algorithm is as follows:
[0062] 1) Initialization phase
[0063] The actor-critic model is used to build a neural network model, including an actor network and a critic network. The output layer in the actor network establishes a deterministic mapping from state to action. ,in, represents the network parameters, Indicates the current state, Actions generated for the actor network; the critic network is generated by The function implements the value evaluation of the state-action pair, where Represents network parameters;
[0064] The parameters of the actor-target network and the critic-target network are initially copied from the parameters of the actor-online network and the critic-online network, respectively, namely:
[0065] ;
[0066] ;
[0067] Where: The parameters of the actor-target network are gradually copied from the actor-online network parameters via soft updates; are the parameters of the critic target network, which are gradually copied from the critic online network parameters through soft updates;
[0068] 2) MDP environment interaction and data collection
[0069] In each round of training, the initial state is first obtained by resetting the MDP environment ; Then, at each time step Next, use the actor online network to output actions , the MDP environment executes this action and returns the incentive and the next state , construct transfer samples ( , , , , ), and store it in the experience replay buffer for subsequent updating of network parameters; Indicates the current state, actions generated for the actor network, It is the immediate incentive of environmental feedback. is the new state to which the action is transferred after execution. It is the termination status flag;
[0070] 3) Online network parameter update
[0071] When the number of samples in the experience replay buffer reaches a preset threshold, a small batch of samples is randomly sampled from it for neural network training and parameter update. The update process is divided into the following two parts:
[0072] Critic online network parameter update: For each sample ( , , , , ), calculate the target value :
[0073] ;
[0074] Where: is the discount factor;
[0075] The loss function is constructed using the mean square error:
[0076] ;
[0077] Where: represents the loss function of the critic online network, which is the predicted value ( ) and target value The mean square error between s is the number of samples in a mini-batch, that is, the number of data randomly sampled from the experience replay buffer for training each time;
[0078] Update the critic online network parameters using the gradient descent method in the Adam optimizer ;
[0079] Actor online network parameter update:
[0080] Define the actor loss as the negative mean of the values:
[0081] ;
[0082] Where: is the loss function of the actor online network;
[0083] Update the actor online network parameters via the gradient descent method in the Adam optimizer ;
[0084] 4) Target network soft update
[0085] The target network parameters are soft-updated to gradually approach the online network. The update formula is:
[0086] ;
[0087] ;
[0088] Where: is the soft update coefficient, which controls the progressive update speed of the target network.
[0089] In the preferred solution, in step 4, the actor network consists of three fully connected layers, the first two layers of which are Relu The activation function is used in the output layer. tanh The activation function constrains the action to the interval [-1,1], establishing a deterministic mapping from state to action. ,in Represents the online network parameters; in the synchronously constructed actor online network and its corresponding target network, the actor online network is responsible for generating actions in real time and interact with the environment, while the actor-target network provides a stable baseline for action evaluation;
[0090] The critic network adopts a three-layer fully connected layer structure, the first two layers use Relu The activation function performs nonlinear feature transformation, and the output layer directly outputs the value estimate of the state-action pair. Function implements the value evaluation of state-action pair, where Represents the online network parameters. In the synchronously constructed critic online network and its corresponding target network, the critic online network is responsible for immediate value prediction, while the critic target network provides the algorithm with a stable state-action pair value evaluation target.
[0091] The present invention provides a multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and a deterministic policy gradient algorithm, which has the following beneficial effects:
[0092] 1. The dynamic weighting mechanism proposed in this paper designs the weights of various scheduling objectives, such as flood control, power generation, and water resource utilization, as functions of environmental conditions. The relative weights of each objective in the incentive function are dynamically adjusted based on the current reservoir storage ratio, inflow, and future water inflow forecasts. Compared to traditional fixed-weight or stage-weighted approaches, this design enables adaptive optimization of reservoir scheduling strategies under non-stationary hydrological conditions. In particular, it dynamically switches scheduling priorities to more effectively balance different scheduling objectives, ensuring both safety and efficiency.
[0093] 2. By introducing a dynamic weight expression with state perception capabilities, the present invention realizes online collaborative scheduling of multi-objective trade-offs within the reinforcement learning framework, making up for the defect that traditional deep reinforcement learning cannot flexibly express dynamic preference changes in multi-objective problems. In particular, when faced with strongly nonlinear and highly uncertain water inflow conditions, the system can actively adjust the weight response to flood control or power generation based on the current water level trend of the reservoir and future water inflow forecasts, thereby improving the adaptability and robustness of the scheduling strategy. This mechanism not only improves the strategy convergence speed and training stability, but also significantly improves the generalization ability of the scheduling strategy in extreme scenarios, and has significant engineering applicability value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] The present invention will be further described below with reference to the accompanying drawings and examples:
[0095] Figure 1 is a flow chart of the method of the present invention;
[0096] Figure 2 This is a schematic diagram of the reservoir multi-objective optimization intelligent scheduling model;
[0097] Figure 3 This is a comparison chart of the actual operation and simulated scheduling water level and outflow flow process in the embodiment. DETAILED DESCRIPTION
[0098] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0099] This application example illustrates Reservoir A, which has a catchment area of approximately 1 million km² and falls within a monsoon climate zone, with precipitation concentrated between May and September. Reservoir A combines flood control, power generation, shipping, and water supply functions. Its normal storage capacity is 39.3 billion m³, and it has been in regular operation since 2010.
[0100] A multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm is implemented in reservoir A. Figure 1 As shown, the following steps are included:
[0101] Step 1: Collect the original hydrological data and reservoir characteristic curve of the target reservoir, and preprocess the original hydrological data.
[0102] The original hydrological data for Reservoir A includes inflow, reservoir water level, and forecast inflow. The characteristic curves for Reservoir A include the water level-discharge capacity curve, the head-maximum output curve, the head-output coefficient curve, the water level-storage capacity curve, and the outflow-downstream water level curve.
[0103] The preprocessing of the original hydrological data includes preprocessing missing values and outliers. For missing values in the original hydrological data, the missing values are filled by the nearest neighbor difference method. For outliers in the original hydrological data, the outliers are corrected by the nearest neighbor method.
[0104] Specifically, the missing values in the inflow data of Reservoir A are identified, and these missing values are filled by the nearest neighbor interpolation method. The outliers in the inflow data of Reservoir A are detected, and these outliers are corrected by the nearest neighbor interpolation method to obtain the complete inflow data sequence of Reservoir A.
[0105] Specifically, the missing values in the water level data of reservoir A are identified, these missing values are filled by the nearest neighbor interpolation method, the outliers in the water level data of reservoir A are detected, these outliers are corrected by the nearest neighbor interpolation method, and the complete water level data sequence of reservoir A is obtained.
[0106] Specifically, the missing values in the forecast inflow data of Reservoir A are identified, and these missing values are filled by the nearest neighbor interpolation method. The outliers in the forecast inflow data of Reservoir A are detected, and these outliers are corrected by the nearest neighbor interpolation method to obtain the complete forecast inflow data sequence of Reservoir A.
[0107] Among them, the nearest neighbor method refers to first determining the k samples closest to the sample with missing data based on the Euclidean distance, and then taking the weighted average of these k values to estimate the missing data of the sample.
[0108] The processed inflow data sequence of Reservoir A, the water level data sequence of Reservoir A and the forecast inflow data sequence of Reservoir A are unified to the same time scale using the average value method according to the forecast requirements, thereby obtaining the inflow data sequence, water level data sequence and forecast inflow data sequence with the same time scale.
[0109] Specifically, the embodiment of the present invention collects the daily inflow, daily water level and forecast inflow data for the next five days of Reservoir A during the flood season from June 10 to October 31, 2010 to 2021. After missing value interpolation and outlier correction, these data are obtained to obtain a set of daily inflow data series, a set of daily water level data series and a set of forecast inflow data series for the next five days.
[0110] Step 2: Based on the main functions and scheduling requirements of the target reservoir, a multi-objective optimization scheduling model including objective functions and constraints is established.
[0111] The objective function needs to cover scheduling objectives for multiple benefits such as flood control, power generation and water resources.
[0112] In the multi-objective optimization scheduling model, the objective function for flood control benefit is to maximize the annual flood control volume, the objective function for power generation benefit is to maximize the annual power generation, and the objective function for water resource benefit is to minimize the amount of abandoned water. The details are as follows:
[0113] (1) The flood control benefit adopts the maximum annual flood control volume as the objective function, and the calculation formula is as follows:
[0114] ;
[0115] ;
[0116] Where: is the annual flood control capacity, in 100 million m³, is the total number of simulation periods, and are the inbound and outbound flows of the current period, Indicates the flow threshold for starting to calculate the reservoir's flood control capacity. is the full traffic volume of the current period, in m³ / s. is the time period length, the unit is d. When the inflow of the target reservoir exceeds and greater than If the outflow is less than the inflow, the flood control capacity for that period is calculated.
[0117] Specifically, in the embodiment of reservoir A, through analysis of the unit operation data, it can be seen that: The value of will vary significantly under different head conditions due to the characteristics of the 32 units and actual operating conditions, and is generally around 30,000m³ / s.
[0118] (2) The power generation efficiency adopts maximizing annual power generation as the objective function, and the calculation formula is as follows:
[0119] ;
[0120] Where: E is the annual power generation, the unit is 100 million kW·h, represents the comprehensive output coefficient, is the water head, the unit is m, is the power generation flow rate, the unit is m³ / s.
[0121] (3) Water resource benefits take minimizing annual water abandonment as the objective function, and the calculation formula is as follows:
[0122] ;
[0123] ;
[0124] Where: The annual amount of water abandoned is expressed in 100 million m³. is the discarded water flow rate, the unit is m³ / s.
[0125] The constraints include outflow restrictions, water level restrictions, water balance principles, power station power generation capacity restrictions, water level and storage capacity relationships, and the relationship between outflow and downstream water levels.
[0126] Specifically, the restrictions are as follows:
[0127] The outbound flow restriction consists of three parts: the upper and lower limit constraints of the outbound flow, the outbound flow change rate constraint, and the outbound flow not exceeding the discharge capacity;
[0128] Water level restriction means that at any time, the water level of the target reservoir must be controlled between the minimum allowable water level and the maximum allowable water level;
[0129] The water balance principle states that the change in water storage capacity of the target reservoir from the adjacent period t to t+1 is equal to the difference between the water inflow and water outflow during the same period;
[0130] The power station output limit consists of two parts: the maximum output constraint and the output coefficient function. The maximum output constraint indicates that the actual output of the hydropower station does not exceed the maximum allowable output. The output coefficient function represents the functional relationship between the output coefficient of the hydropower station and the reservoir water level, as shown in the following formula:
[0131] ;
[0132] Where: is the output coefficient, is the reservoir water level in m.
[0133] The water level-capacity relationship is the amount of water stored in the target reservoir at any time, which is a function of the current water level, as shown in the following formula:
[0134] ;
[0135] Where: is the reservoir capacity, in m³;
[0136] The relationship between outflow and downstream water level shows that the downstream water level of the target reservoir has an obvious functional relationship with the outflow, as shown in the following formula:
[0137] ;
[0138] Where: It is the outbound flow.
[0139] Step 3: Map the above multi-objective optimization scheduling model into a Markov decision process to establish the corresponding reinforcement learning MDP environment.
[0140] The multi-objective optimization scheduling model is mapped into a Markov decision process, including building a complete environment, defining the state space, action space, incentive function and state transition law function.
[0141] In the reservoir operation problem, the decision variable is the outflow of the reservoir at the current moment, which is recorded as Since the variables are all continuous variables, the action space is a continuous space.
[0142] 1. State Space
[0143] Combined with the aforementioned objective function and constraints, the constructed state space needs to fully describe the operation of the target reservoir, including:
[0144] Reservoir capacity: The water storage capacity of reservoir A at the current moment, recorded as ;
[0145] Inflow: The inflow of reservoir A at the current moment, recorded as ;
[0146] Outflow: The outflow data of reservoir A at the previous moment is used, recorded as ;
[0147] Forecasted inflow: The inflow forecast value of reservoir A in the next N time periods at the current moment is recorded as:
[0148] ;
[0149] In the above formula: Respectively 、 2. 、 ,..., Inflow flow forecast value within the time period.
[0150] In this embodiment, the inflow forecast value of Reservoir A within the next 5 days from the current moment is recorded as:
[0151] ;
[0152] Then the state space can be expressed as:
[0153] .
[0154] 2. The calculation steps of the activation function are:
[0155] 1) Set incentive functions for flood control benefits, power generation benefits, water resource benefits, water level and flow respectively 、 、 、 as well as , respectively:
[0156] Flood control benefits: One of the goals of reservoir optimization is to maximize flood control capacity. To maximize flood control benefits, the simulated flood control capacity is directly used as the incentive function. The specific definition is as follows:
[0157] ;
[0158] ;
[0159] Where: represents the incentive function of flood control benefits, and Respectively Inbound and outbound traffic for each period, Indicates the flow threshold for starting to calculate the reservoir's flood control capacity. is the full power flow of the current period, in m³ / s. is the length of the period. When the inflow of the target reservoir exceeds and greater than If the outflow is less than the inflow, the flood control capacity for that period is calculated.
[0160] Specifically, in the embodiment of reservoir A, through analysis of the unit operation data, it can be seen that: The value of will vary under different head conditions due to the characteristics of the 32 units and actual operating conditions, and is generally around 30,000 m³ / s.
[0161] Power generation efficiency: One of the goals of reservoir optimization is to maximize power generation. To maximize power generation, simulated power generation is directly used as the incentive function. The specific definition is as follows:
[0162] ;
[0163] Where: represents the incentive function of power generation efficiency, represents the comprehensive output coefficient, is the water head in the current period, in meters, is the power generation flow rate, the unit is m³ / s.
[0164] Water resource efficiency: One of the goals of reservoir optimization is to reduce the amount of water abandoned. To minimize the amount of water abandoned, the simulated amount of water abandoned is directly used as the incentive function. The specific definition is as follows:
[0165] ;
[0166] ;
[0167] Where: represents the incentive function of water resource benefits, is the discarded water flow rate, the unit is m³ / s.
[0168] Water level: When the reservoir water level Falling in the water level safety zone ], a positive incentive is given, and when the water level exceeds the safe range, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is:
[0169] ;
[0170] Where: represents the activation function of the water level, is the reservoir water level at the current moment, in m, [ ] is the pre-set safe water level range, the unit is m. It is the positive incentive constant when the water level is in the safe range, and determines the penalty amplitude when the water level deviates from the safe range.
[0171] Flow: When the reservoir outflow Falling in the traffic safety range[ ], a positive incentive is given, and when the outbound flow exceeds the safe interval, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is:
[0172] ;
[0173] Where: represents the incentive function of the flow, [ ] is the pre-set safe outbound flow range, Represents the flow scaling factor, which is used to reduce the actual flow value to a smaller value range to reduce the impact of large values on the activation function. Because in actual operation, the flow rate is usually large, direct use may lead to numerical instability of the activation function or difficulty in gradient calculation. The unit is m³ / s. It is the positive incentive constant when the outbound flow is in the safe range, and determines the penalty amplitude when the flow deviates from the safe range.
[0174] Specifically, in the example of reservoir A, the actual outflow is analyzed and it is known that its value is usually large, which is significantly different from the magnitude of other incentive functions. Therefore, it is necessary to reasonably set To ensure that the activation function The output of is scaled to be comparable to other activation functions.
[0175] 2) Introduction of dynamic weights
[0176] According to the different operation stages of the reservoir in a year, and combined with the current storage capacity ratio, real-time inflow and forecast inflow, the weights of each incentive function are dynamically adjusted.
[0177] First, define three intermediate variables 、 、 :
[0178] ;
[0179] ;
[0180] ;
[0181] Where: is the current storage capacity of the reservoir, It is the maximum storage capacity of the reservoir, and the unit is 100 million m³. is the maximum inflow flow, in m³ / s.
[0182] Define a linear combination concatenation index for each activation function as an unnormalized weight:
[0183] ;
[0184] ;
[0185] ;
[0186] ;
[0187] ;
[0188] Where, , , As parameters, Is the critical water level threshold, generally set to ∈[0.80,0.90], determined according to the specific conditions of the reservoir. In this embodiment, 0.85; It is the response sensitivity of the intelligent dispatching method to the current inflow and future forecast flow. In the flood season, it is mainly used for flood control. In the dry season, the main focus is on power generation and water resource security. In this embodiment, 15; It reflects the degree of influence of the current water storage ratio on the target weight, especially when approaching high water level. , dry season is generally set .
[0189] use Softmax Function, converting the above weights into final weights:
[0190] ;
[0191] Where, i = 1, 2, 3, 4, 5 means , , The unnormalized weights of i = 1,2, 3, 4, 5 means The corresponding final weight;
[0192] 3) Total activation function The calculation formula is:
[0193] .
[0194] Specifically, in the example of Reservoir A, according to the scheduling regulations of Reservoir A, the water storage time of Reservoir A shall not be earlier than September 10. Based on this, in order to adapt to the operation requirements of different scheduling stages and optimize the weight distribution of scheduling targets, the model sets different parameter values in different time periods. During the main flood prevention period from June 10 to September 10, the parameter is set , , , in order to strengthen the weight of flood control benefits and appropriately consider the power generation benefits and water resource benefits goals; and during the water storage period from September 11 to October 31, set the parameters , , , in order to increase the weight of water storage and water resource efficiency goals.
[0195] 3. The state transition function states that after the reservoir discharges outflow, the storage capacity is updated according to the water balance relationship, completing the state transition. The water balance relationship is the change in water storage capacity of the target reservoir from the adjacent time period t to t+1, which is equal to the difference between the inflow and outflow during the same period.
[0196] When the environment receives the action generated by the algorithm, that is, the outbound traffic After executing this action, the reservoir environment state will change accordingly and transfer to the next moment. Specifically, after the reservoir executes the discharge flow, the storage capacity is updated according to the water balance relationship, thus completing the state transfer:
[0197] .
[0198] Step 4: Select the Deep Deterministic Policy Gradient (DDPG) algorithm to conduct interactive training with the environment and build a multi-objective optimization intelligent scheduling model for reservoirs, such as Figure 2 shown.
[0199] The Deep Deterministic Policy Gradient algorithm is a model-free deep reinforcement learning algorithm that combines the advantages of deterministic policy gradients with deep Q-networks. It is specifically designed to solve decision-making problems in continuous action spaces. Based on an actor-critic architecture, the algorithm employs a dual neural network, consisting of an online network and a target network, to approximate the policy function and action-value function, respectively. The network parameters are optimized using a stochastic gradient method.
[0200] The superiority of the deep deterministic policy gradient algorithm is mainly reflected in three aspects:
[0201] (1) Experience replay mechanism: The samples generated by the interaction between the actor and the environment are stored in the experience replay pool. Through batch sampling training, the data correlation is effectively reduced and the algorithm stability is improved;
[0202] (2) Target network soft update: By gradually updating the target network parameters, the oscillation problem during training is alleviated and convergence is accelerated;
[0203] (3) Dual network architecture: Both the actor and the critic adopt an online-target dual network structure to further reduce estimation bias and enhance learning robustness.
[0204] Using a deep deterministic policy gradient algorithm, the system continuously interacts with the environment to acquire incentive signals and update neural network parameters. Through continuous iteration and policy improvement, the training process gradually converges the neural network parameters, ultimately achieving an optimal or near-optimal strategy that can achieve multi-objective scheduling.
[0205] The calculation process of the deep deterministic policy gradient algorithm is as follows:
[0206] 1) Initialization phase
[0207] First, the neural network model of the algorithm is constructed using the actor-critic model. The neural network model constructed using the actor-critic model includes an actor network and a critic network. The actor network consists of three fully connected layers with a 256-256-1 neuron structure. The input layer and the middle layer of the actor network both use 256 neurons. The first two layers use Relu The activation function is used in the output layer. tanh The activation function constrains the action to the interval [-1,1], establishing a deterministic mapping from state to action. ,in, represents the network parameters, Indicates the current state, Actions generated for the actor network; in the synchronously constructed actor online network and its corresponding target network, the online network is responsible for generating actions in real time and interact with the environment, while the target network provides a stable baseline for action evaluation.
[0208] The critic network adopts a three-layer fully connected layer structure with a 256-256-1 neuron structure. The first two layers use Relu The activation function performs nonlinear feature transformation, and the output layer directly outputs the value estimate of the state-action pair. Function implements the value evaluation of state-action pair, where Represents the online network parameters, and in the synchronously constructed critic online network and its corresponding target network, the online network is responsible for immediate value prediction, while the target network provides the algorithm with a stable state-action pair value evaluation target.
[0209] The parameters of the target network are initially copied from the corresponding online network, namely:
[0210] , ;
[0211] Where: The parameters of the actor-target network are gradually copied from the actor-online network parameters via soft updates; are the parameters of the critic target network, which are gradually copied from the critic online network parameters through soft updates.
[0212] At the same time, a fixed-capacity experience replay buffer is established to store the transfer samples generated during the interaction between the algorithm and the environment ( , , , ,done). The buffer adopts a first-in-first-out storage strategy and automatically overwrites the earliest experience sample when the capacity limit is reached.
[0213] 2) Environmental interaction and data collection
[0214] In each round of training, the initial state is obtained by resetting the environment ; Then, at each time step Next, use the actor online network to output actions , the environment executes the action and returns the incentive and the next state , based on which the transfer sample is constructed ( , , , , ), and store it in the experience replay buffer for subsequent updating of network parameters; Indicates the current state, actions generated for the actor network, It is the immediate incentive of environmental feedback. is the new state to which the action is transferred after execution. It is the termination status flag.
[0215] 3) Online network parameter update
[0216] When the number of samples in the experience replay buffer reaches the preset threshold, a small batch of samples is randomly sampled from it for network training and network parameter update. The update process is divided into the following two parts:
[0217] Critic online network parameter update: For each sample ( , , , , ), calculate the target value :
[0218] ;
[0219] Where: is the discount factor;
[0220] The loss function is constructed using the mean square error:
[0221] ;
[0222] Where: represents the loss function of the critic online network, which is the predicted value ( ) and target value The mean square error between s is the number of samples in a mini-batch, that is, the number of data randomly sampled from the experience replay buffer for training each time;
[0223] Update the critic online network parameters using the gradient descent method in the Adam optimizer ;
[0224] Actor online network parameter update:
[0225] Define the actor loss as the negative mean of the values:
[0226] ;
[0227] Where: is the loss function of the actor online network.
[0228] Update the actor online network parameters via the gradient descent method in the Adam optimizer .
[0229] 4) Target network soft update
[0230] The target network parameters are soft-updated to gradually approach the online network. The update formula is:
[0231] ;
[0232] ;
[0233] Where: is the soft update coefficient, which controls the progressive update speed of the target network.
[0234] The deep deterministic policy gradient algorithm continuously iterates between interactive sampling in the environment and updating network parameters, gradually enabling the actor network to learn the optimal action that can output the maximum cumulative reward in each state, while the critic network continuously and accurately estimates the value of state-action, ultimately achieving algorithm convergence.
[0235] Specifically, in the A reservoir example, the total number of training cycles is set to 1000 to control the overall training scale; the neural network learning rate is set to 0.0001 to adjust the update step size of the network parameters; the discount factor , with a value of 0.9, used to balance the weight of immediate incentives and long-term returns; soft update coefficient , set to 0.01, controls the progressive update speed of the target network.
[0236] As shown in Tables 1–3, after intelligent optimization using the DDPG algorithm, Reservoir A outperformed the actual operation scenario in flood control, power generation, and water abandonment during the 2010–2021 flood season. Specifically, the simulated flood control capacity reached an average of 20.76 billion m³ per year, a 26.6% increase over the actual operation rate of 16.39 billion m³. This included increases of 8.845 billion m³ and 5.432 billion m³ in 2010 and 2021, respectively. Only in 2014 and 2018 did the actual flood control capacity slightly exceed or approach the simulated values, demonstrating that the optimized strategy significantly enhanced flood control capacity in most years. Regarding power generation performance, the simulated operation yielded an average annual power generation of 57.25 billion kWh, a 4.0% increase over the actual operation rate of 55.03 billion kWh. Furthermore, the simulated operation generated an additional 4.528 billion kWh and 1.731 billion kWh in 2019 and 2021, respectively, demonstrating a good balance between flood control and power generation. In terms of water abandonment, the simulated dispatching averaged 8.79 billion m³ per year, a decrease of 20.9% from the actual operation of 11.1 billion m³. Zero water abandonment was achieved in 2015, 2016, and 2019. Only in 2021 did the demand for water storage increase slightly by 6.498 billion m³, significantly reducing water waste overall. The above results fully demonstrate that the intelligent dispatching method proposed in this invention not only improves flood control benefits, but also takes into account power generation gains and water resource utilization. Figure 3 .
[0237]
[0238]
[0239]
[0240] Step 5: Deploy the trained reservoir multi-objective optimization intelligent scheduling model in the scheduling system of the target reservoir to perform multi-objective optimization intelligent scheduling on the reservoir.
[0241] The specific steps include:
[0242] S51. Integrate the multi-objective optimization intelligent scheduling model of the reservoir into the scheduling control platform, obtain information such as inflow, water level and predicted inflow in real time, input the trained model, and finally output the optimized scheduling decision in real time.
[0243] The multi-objective optimization intelligent scheduling model of Reservoir A is integrated into the scheduling control platform to obtain real-time information such as inflow, water level and predicted inflow, and input the trained model, and finally output the optimized scheduling decision of Reservoir A in real time.
[0244] S52. The model is continuously optimized through online learning and regular retraining to adapt to changing hydrological conditions and operational scheduling requirements.
[0245] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm, characterized by: include: Step 1: Collect the original hydrological data and reservoir characteristic curve of the target reservoir, and preprocess the original hydrological data; Step 2: According to the main functions and scheduling requirements of the target reservoir, the optimization objectives are determined from the three aspects of flood control benefits, power generation benefits and water resource benefits, and a multi-objective optimization scheduling model including objective functions and constraints is established; Step 3: Map the multi-objective optimization scheduling model into a Markov decision process to establish the corresponding reinforcement learning MDP environment. Specifically, this includes constructing a complete MDP environment, defining the state space, action space, incentive function, and state transition law function: The calculation steps of the activation function are: 1) Set incentive functions for flood control benefits, power generation benefits, water resource benefits, water level and flow respectively 、 、 、 as well as , respectively: In terms of flood control benefits, the simulated flood interception volume is used as the incentive function, which is defined as follows: ; ; Where: represents the incentive function of flood control benefits, and Respectively Inbound and outbound traffic for each period, Indicates the flow threshold for starting to calculate the reservoir's flood control capacity. The full power generation flow in the current period; when the inflow of the target reservoir exceeds and greater than When the outflow is less than the inflow, the flood control capacity of this period is calculated; is the time period length; In terms of power generation efficiency, the simulated power generation is used as the incentive function, which is defined as follows: ; Where: represents the incentive function of power generation efficiency, represents the comprehensive output coefficient, is the water head for the current period, is the power generation flow; In terms of water resource benefits, the simulated abandoned water volume is used as the incentive function, which is defined as follows: ; ; Where: represents the incentive function of water resource benefits, is the abandoned water flow; In terms of water level, when the reservoir water level Falling in the water level safety zone ]When you are in the office, give positive incentives. 、 They represent the minimum and maximum values of the water level safety interval respectively. When the reservoir water level exceeds the water level safety interval, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is: ; Where: represents the activation function of the water level, is the reservoir water level at the current moment, is the positive excitation constant when the water level is in the safe range; In terms of flow, when the reservoir outflow Falling in the traffic safety range[ ]When you are in the office, give positive incentives. 、 They represent the minimum and maximum values of the traffic safety interval respectively. When the outbound traffic exceeds the traffic safety interval, a negative incentive is given according to the degree of deviation. The greater the deviation, the more severe the penalty. The corresponding incentive function is: ; Where: represents the incentive function of the flow, represents the traffic scaling factor, is the positive excitation constant when the outbound flow is in the safe range; 2) Introduction of dynamic weights First, define three intermediate variables 、 、 : ; ; ; Where: It is the maximum storage capacity of the reservoir. is the maximum inflow; Define a linear combination concatenation index for each activation function as an unnormalized weight: ; ; ; ; ; Where, Reflects the critical water level threshold, It reflects the response sensitivity of the intelligent scheduling method to the current inflow and future forecast flow. It reflects the degree of influence of the current water storage ratio on the target weight. The three parameters are determined by the reservoir at different operating times and historical data; use Softmax Function, converting the above weights into final weights: ; Where, i = 1, 2, 3, 4, 5 means , , The unnormalized weights of i = 1, 2, 3, 4, 5 means The corresponding final weight; 3) Total activation function The calculation formula is: ; Step 4: Select a deep deterministic policy gradient algorithm to conduct interactive training with the environment and build a multi-objective optimization intelligent scheduling model for reservoirs; Step 5: Deploy the trained reservoir multi-objective optimization intelligent scheduling model in the scheduling system of the target reservoir to perform multi-objective optimization intelligent scheduling on the reservoir.
2. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 1 is characterized in that: In the step 1, preprocessing the original hydrological data includes preprocessing missing values and outliers. For missing values in the original hydrological data, the missing values are filled by interpolation using the nearest neighbor method. For outliers in the original hydrological data, the outliers are corrected by the nearest neighbor method.
3. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 1 is characterized in that: In the step 2, in the multi-objective optimization scheduling model, the flood control benefit adopts maximizing the annual flood interception volume as the objective function, the power generation benefit adopts maximizing the annual power generation as the objective function, and the water resource benefit adopts minimizing the abandoned water volume as the objective function; The constraints include outflow restrictions, water level restrictions, water balance principle, power station output restrictions, water level and storage capacity relationship, and the relationship between outflow and downstream water level.
4. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 1 is characterized in that: The state space is represented as: ; in: Indicates the target reservoir at the current moment t water storage capacity; represents the inflow of the target reservoir at the current moment; Indicates that the outflow data of the target reservoir at the previous moment is used; Indicates the target reservoir in the future at the current moment N Inflow flow forecast value within a period; ; In the above formula: Respectively 、 2. 、 ,..., Inflow flow forecast value within the time period.
5. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 1 is characterized in that: The state transfer law function is that after the reservoir executes the discharge flow, the storage capacity is updated according to the water balance relationship to complete the state transfer. The water balance relationship is the change in water storage capacity of the target reservoir in the adjacent time period t to t+1, which is equal to the difference between the inflow and outflow water in the same period.
6. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 1 is characterized in that: In step 4, a deep deterministic policy gradient algorithm is used to obtain the incentive signal and update the neural network parameters through continuous interactive training with the MDP environment.
7. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 6 is characterized in that: The calculation process of the deep deterministic policy gradient algorithm is as follows: 1) Initialization phase The actor-critic model is used to build a neural network model, including an actor network and a critic network. The output layer in the actor network establishes a deterministic mapping from state to action. ,in, represents the network parameters, Indicates the current state, Actions generated for the actor network; the critic network is generated by The function implements the value evaluation of the state-action pair, where Represents network parameters; The parameters of the actor-target network and the critic-target network are initially copied from the parameters of the actor-online network and the critic-online network, respectively, namely: ; ; Where: The parameters of the actor-target network are gradually copied from the actor-online network parameters via soft updates; are the parameters of the critic target network, which are gradually copied from the critic online network parameters through soft updates; 2) MDP environment interaction and data collection In each round of training, the initial state is first obtained by resetting the MDP environment ; Then, at each time step Next, use the actor online network to output actions , the MDP environment executes this action and returns the incentive and the next state , construct transfer samples ( , , , , ), and store it in the experience replay buffer for subsequent updating of network parameters; Indicates the current state, actions generated for the actor network, It is the immediate incentive of environmental feedback. is the new state to which the action is transferred after execution. It is the termination status flag; 3) Online network parameter update When the number of samples in the experience replay buffer reaches a preset threshold, a small batch of samples is randomly sampled from it for neural network training and parameter update. The update process is divided into the following two parts: Critic online network parameter update: For each sample ( , , , , ), calculate the target value : ; Where: is the discount factor; The loss function is constructed using the mean square error: ; Where: represents the loss function of the critic online network, which is the predicted value ( ) and target value The mean square error between N s is the number of samples in a mini-batch, that is, the number of data randomly sampled from the experience replay buffer for training each time; Update the critic online network parameters using the gradient descent method in the Adam optimizer ; Actor online network parameter update: Define the actor loss as the negative mean of the values: ;; Where: is the loss function of the actor online network; Update the actor online network parameters via the gradient descent method in the Adam optimizer ; 4) Target network soft update The target network parameters are soft-updated to gradually approach the online network. The update formula is: ; ; Where: is the soft update coefficient, which controls the progressive update speed of the target network.
8. The multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm according to claim 7 is characterized in that: In step 4, the actor network consists of three fully connected layers, the first two layers use Relu The activation function is used in the output layer. tanh The activation function constrains the action to the interval [-1,1], establishing a deterministic mapping from state to action. ,in Represents the online network parameters; in the synchronously constructed actor online network and its corresponding target network, the actor online network is responsible for generating actions in real time and interact with the environment, while the actor-target network provides a stable baseline for action evaluation; The critic network adopts a three-layer fully connected layer structure, the first two layers use Relu The activation function performs nonlinear feature transformation, and the output layer directly outputs the value estimate of the state-action pair. Function implements the value evaluation of state-action pair, where Represents the online network parameters. In the synchronously constructed critic online network and its corresponding target network, the critic online network is responsible for immediate value prediction, while the critic target network provides the algorithm with a stable state-action pair value evaluation target.
Citation Information
Patent Citations
Reservoir group joint optimization scheduling method based on MADDPG reinforcement learning
CN115952958A
Cascade hydropower station long-term scheduling decision-making method, system and equipment and storage medium
CN118691128A
Small reservoir intelligent flood discharge scheduling method based on reinforcement learning
CN120197889A