Wind-solar-pumped storage cascade reservoir depth reinforcement learning short-term stochastic optimization scheduling method
Patent Information
- Application Number
- CN202310427058.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-04-19
AI Technical Summary
[0003]本发明的目的在于克服上述不足,提供一种风-光-梯级水库深度强化学习短期随机优化调度方法,以解决风-光-梯级水库短期随机优化调度中,由多重随机变量大离散状态空间造成的“维数灾”导致计算速度慢、准确度不高的问题,并且提升发电效益、平滑互补系统出力波动
本发明提出的考虑径流、风电出力和光伏出力的随机性,以总发电量最大、剩余负荷最小为目标的基于深度强化学习DQN算法风-光-梯级水库多目标短期随机优化调度方法,技术效果如下:
Smart Images

Figure CN116720674B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reservoir optimization scheduling research, specifically to a deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs. Background Technology
[0002] In recent years, affected by the increasing scarcity of fossil fuels and the environmental pollution caused by the large-scale consumption of fossil fuels, countries have actively developed the construction and application of renewable energy. Research on renewable clean energy such as wind power, solar power, and hydropower has become a widely discussed topic. Conventional hydropower, as a clean energy source, is one of the most important forms of power system regulation and an effective way to improve the absorption of new energy. However, due to the strong fluctuations in wind and solar power output, how to utilize hydropower to suppress these fluctuations and ensure the safe and reliable operation of the power system is the primary issue in improving the absorption of new energy. Currently, the rapid development of wind and solar power has led to an increasing importance and difficulty in hydropower scheduling and operation. The difficulty mainly lies in the high dimensionality, randomness, non-convexity, multi-stage nature, and discretization of the long-term scheduling problem of cascade reservoirs. Compared to a single reservoir, the discreteness increases exponentially, and the addition of wind and solar power further increases the randomness of system output, leading to a further increase in the dimensionality of the problem and posing a significant challenge to optimal scheduling. Deep reinforcement learning algorithms, by combining the advantages of deep learning and reinforcement algorithms, have shown better effectiveness in addressing the optimal scheduling problem of cascade reservoirs combined with wind and solar power. Summary of the Invention
[0003] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs. This method addresses the problem of slow computation speed and low accuracy caused by the "curse of dimensionality" resulting from the large discrete state space of multiple random variables in the short-term stochastic optimization scheduling of wind-solar-cascade reservoirs. It also improves power generation efficiency and smooths out power output fluctuations in complementary systems.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by this invention is: a deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs, which includes the following steps: Step 1: Treat the short-term scheduling of the wind-solar-cascade reservoir system as a multi-stage decision problem, and divide the short-term scheduling into multiple stages to solve the scheduling strategy; Step 2: Based on the historical runoff data of the cascade reservoirs, the probability of inflow into the cascade reservoirs at each time period is described using the Pearson Type III distribution. Then, by comparing the correlation characteristics of the function in the Copula function with the correlation characteristics of the inflow into the cascade reservoirs, the function with the most similar correlation characteristics is selected to conduct spatial correlation analysis on the inflow into the cascade reservoirs, and the joint probability density function and joint probability distribution of the cascade reservoirs at each time period are obtained. Step 3: Solve the joint probability distribution of adjacent time periods of the cascade reservoirs using the function selected in Step 2 to describe their temporal correlation. Then, based on the conditional distribution formula, solve for the Markov state transition probability matrix of random runoff in adjacent time periods, and use the Markov Monte Carlo sampling method to randomly simulate the random scenario of inflow runoff into the cascade reservoirs. Step 4: Based on historical wind power output data, a Markov Monte Carlo sampling method based on kernel density estimation is proposed to solve the probability density function of wind power output and the Markov state transition probability matrix, and to randomly simulate future wind power output scenarios. Step 5: Based on historical photovoltaic power output data, a Markov Monte Carlo sampling method based on kernel density estimation is proposed to solve the probability density function of photovoltaic power output and the Markov state transition probability matrix, and to randomly simulate future photovoltaic power output scenarios. Step 6: The scenes obtained in Steps 3, 4 and 5 are reduced by Bisecting k-means clustering algorithm based on relevance distance, and then combined to form the scenes of each stage of the wind-solar cascade reservoir system. Step 7: Construct the environment for the deep reinforcement learning algorithm using the random scenario obtained in Step 6, and fine-tune the learning efficiency parameters of the deep reinforcement learning algorithm DQN (Deep Q-network) and the discretization accuracy of the reservoir water level. Determine the parameters of the DQN algorithm for the wind-solar-cascade reservoir system with the goal of maximizing system power generation and minimizing residual load. Finally, use the parameter-tuned algorithm to solve the optimal strategy for short-term scheduling of the wind-solar-cascade reservoir system.
[0005] Preferably, in step 1, the short-term scheduling is divided into 24 stages, that is, one hour is one stage to solve the scheduling strategy.
[0006] Preferably, step 2 involves obtaining the joint probability density function and joint probability distribution of the cascade reservoirs for each time period based on the Pearson Type III distribution. First, the Archimedean Copula family of Copula functions, widely used in hydrology, was selected. Then, by comparing the correlation characteristics of the main functions in the Archimedean Copula family with the correlation characteristics of inflow runoff from cascade reservoirs, the Clayton Copula function, with the most similar correlation characteristics, was ultimately chosen to describe the spatial correlation of the cascade reservoirs. The Pearson-III type probability distribution function of each reservoir was used as a marginal distribution to solve for the runoff state. Joint probability distribution of downstream reservoirs As shown in equation (1): (1) In the formula: Cascade reservoir runoff status Joint probability distribution of downstream reservoirs; for t Time period n Inflow of water from the first-level reservoir The Pearson-III type probability distribution function; for t The join parameters of the Clayton Copula function for the time period can be estimated using the Kendall correlation coefficient; Furthermore, the joint probability density function can be obtained by inverse transformation, i.e., by taking the partial derivatives separately.
[0007] Preferably, in step 3, the method for solving the Markov transition probability matrix of the cascade reservoir includes: Correlation analysis shows that the historical inflow data of the cascade reservoirs follows Markov characteristics. The specific steps for solving the Markov probability transition matrix of wind power output are as follows: S1: Obtained by solving the Clayton Copula function t Time and t Joint distribution function at time +1 , ; S2: Obtain the joint probability distribution of adjacent time periods of the cascade reservoirs using the Clayton Copula function. To describe time correlation; S3: Discretize the inflow runoff of each reservoir at each stage into the same number of states with the maximum and minimum inflow runoff as the range; S4: Based on the conditional distribution formula, calculate the Markov transition probability of the cascade reservoirs from runoff state in adjacent time periods. As shown in equation (2): (2) In the formula: The Markov transition probability of the runoff state of the cascade reservoirs in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, random variable inflow into the reservoir t The value of the time period The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Random variable inflow into the reservoir during a time period The value is And random variable inflow into the reservoir t The value of the time period The probability of; For random variables, inflow runoff t The value of the time period The probability of; S5: Solve for each discrete value of runoff to obtain the state transition probability matrix for this stage. ; S6: Repeat the above steps to obtain the state transition probability matrix for each stage, which is used to describe the Markov process of the inflow runoff of the cascade reservoirs in adjacent periods.
[0008] In step 3, the method for generating random scenarios of inflow into cascade reservoirs is as follows: The specific steps for generating random scenes using the MH algorithm in Markov Monte Carlo sampling are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =0, respectively from the transition matrix With uniform distribution Mid-sampling, obtained t+ Phase 1 sampled values and random values Accept sampled values The criterion is shown in equation (3): (3) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+1 time period t A set of random variables related to inflow into the reservoir over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0. Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e. t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage; thus obtaining z The samples that conform to the probability density function and state transition probability of each stage are used as the inflow scenarios of cascade reservoirs. Each scenario includes the inflow runoff value of each reservoir at each time period and the corresponding Markov state transition probability.
[0009] Preferably, in step 4, the method for solving the probability density function of wind power output includes: The choice between using a multidimensional or one-dimensional kernel density estimation function to fit the probability density distribution depends on the number of wind farms. If there is only one wind farm, a one-dimensional kernel density estimation function is used; if there are two or more, a multidimensional kernel density estimation function is used. Here, we will use a one-dimensional kernel density estimation function for explanation. Let... It is a probability density function For an unknown population of independent and identically distributed random variables, the kernel density estimate is defined as shown in equation (4): (4) In the formula: x It is a random variable; n For sample size; For the firsti One sample; h For bandwidth parameters; For a kernel function, it must satisfy the following conditions: , , ; As can be seen from the above formula, when the sample is known, kernel density estimation requires determining two important components: the bandwidth parameter. h and kernel function Since kernel density estimation averages the kernel function, the choice of bandwidth parameter has a much greater impact on the estimation result than the kernel function itself. A Gaussian kernel function is used, as shown in equation (5). (5) The optimal bandwidth parameter is solved using the least squares cross-validation test (LSCV) method, and the bandwidth parameter estimated by the multidimensional kernel function can be solved using the Gaussian kernel. The least squares cross-validation test (LSCV) method is a calculation method based on the minimum integral squared error (ISE) criterion. The ISE expression is shown in equation (6): (6) In the formula: It is a random variable; For kernel density estimation, It is a probability density distribution; The last term in the above three terms is independent of bandwidth. Therefore, the first two terms are used for least squares cross-validation test to minimize the bandwidth parameter. By substituting the one-dimensional kernel density estimation function, the one-dimensional least squares cross-validation test is shown in equation (7): (7) In the formula: For the first i One sample; For the first j One sample; h For bandwidth parameters; For kernel functions; This represents the kernel function convolving with itself; n is the sample size. In step 4, the method for solving the Markov state transition probability matrix of wind power output includes: Through correlation analysis, the random variable of wind power output follows Markov characteristics. The specific steps to solve the Markov probability transition matrix of wind power output are as follows: S1: Solve for the joint probability density function between two adjacent time periods. This involves integrating the probability density functions of wind power output at time t and time t+1 to obtain the probability distribution function. , As a marginal probability, the Copula function, which is most closely related to wind power output, is chosen to solve the joint probability function. ; S2: Discretize the wind power output into several states based on the maximum and minimum wind power output; S3: Based on the conditional distribution formula, the wind power output in adjacent time periods is obtained from the output state. Transition to the next state Markov transition probability As shown in equation (8): (8) In the formula: The Markov state transition probability of wind power output state in adjacent time periods; for t+ 1-period random variable inflow into the reservoir The value is hour, t Time-of-use random variable: wind power output The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Wind power output as a random variable over a time period The value is ,and t Time-of-use random variable: wind power output The value is The probability of; Wind power output is a random variable t The value of the time period The probability of; The state transition probability matrix is obtained based on the random runoff state transition probabilities of each adjacent time period. , used to describe Markov processes in adjacent periods of inflow into cascade reservoirs; In step 4, the method for generating random wind power output scenarios is as follows: The stochastic model is sampled using the MH algorithm in the Markov Monte Carlo sampling method. The specific steps are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t=1, respectively from the transition matrix With uniform distribution Mid-sampling, obtaining sampled values and random values Accept sampled values The criterion is shown in equation (9): (9) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0; Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e.t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage, thereby obtaining... z Samples that conform to the probability density function and state transition probability of each stage are used as wind power output scenarios; Each scenario includes the wind power output for each time period and the corresponding Markov state transition probability.
[0010] Preferably, in step 5, the method for solving the photovoltaic power output probability density function includes: The choice between using a multidimensional or one-dimensional kernel density estimation function to fit the probability density distribution is determined based on the number of photovoltaic power plants. If there is only one photovoltaic power plant, a one-dimensional kernel density estimation function is used; if there are two or more, a multidimensional kernel density estimation function is used. Here, we will use a one-dimensional kernel density estimation function for explanation. Let... It is a probability density function For an unknown population of independent and identically distributed random variables, the kernel density estimate is defined as shown in equation (10): (10) In the formula: x It is a random variable; n For sample size; For the first i One sample; h For bandwidth parameters; For a kernel function, it must satisfy the following conditions: , , ; As can be seen from the above formula, when the sample is known, kernel density estimation requires determining two important components: the bandwidth parameter. h and kernel function Since the kernel density averages the kernel function, the choice of bandwidth parameter has a much greater impact on the estimation result than the kernel function. Here, the Gaussian kernel function is chosen, as shown in equation (11): (11) The optimal bandwidth parameter is solved using the least squares cross-validation test (LSCV) method, and the bandwidth parameter estimated by the multidimensional kernel function can be solved using the Gaussian kernel. The least squares cross-validation test (LSCV) method is a calculation method based on the minimum integral squared error (ISE) criterion. The ISE expression is shown in equation (12): (12) In the formula: It is a random variable; For kernel density estimation, It is a probability density distribution; The last term in the above three terms is independent of bandwidth. Therefore, the first two terms are used for least squares cross-validation test to minimize the bandwidth parameter. By substituting the one-dimensional kernel density estimation function, the one-dimensional least squares cross-validation test is shown in equation (13): (13) In the formula: For the first i One sample; For the first j One sample; h For bandwidth parameters; For kernel functions; This represents the kernel function convolving with itself; n is the sample size. In step 5, the method for solving the photovoltaic output Markov state transition probability matrix includes: Through correlation analysis, the random variable of photovoltaic power output follows Markov characteristics. The specific steps to solve the Markov state probability transition matrix of photovoltaic power output are as follows: S1: Solve for the joint probability density function between the photovoltaic power output states of two adjacent time periods. This involves integrating the photovoltaic power output probability density functions at time t and t+1 to obtain the probability distribution function. , As the marginal probability, the Copula function, which is most closely related to photovoltaic output, is chosen to solve the joint probability function. ; S2: Discretize the photovoltaic output into several states based on the maximum and minimum photovoltaic output; S3: Based on the conditional distribution formula, the photovoltaic power output in adjacent time periods is obtained from the power output state. Transition to the next state Markov transition probability As shown in equation (14): (14) In the formula: The Markov state transition probability of photovoltaic power output states in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, the random variable photovoltaic output t Time period The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Photovoltaic power output as a random variable over a time period of 1 The value is And random variable photovoltaic output t The value of the time period The probability of; Photovoltaic output as a random variable t The value of the time period The probability of; The state transition probability matrix is obtained based on the random runoff state transition probabilities of each adjacent time period. , used to describe Markov processes in adjacent periods of inflow into cascade reservoirs; In step 5, the method for generating random photovoltaic output scenarios includes: The stochastic model is sampled using the MH algorithm in the Markov Monte Carlo sampling method. The specific steps are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =1, respectively from the transition matrix With uniform distribution Mid-sampling, obtaining sampled values and random values Accept sampled values The criterion is shown in equation (15): (15) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0. Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e. t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage; thus obtaining z Samples that conform to the probability density function and state transition probability of each stage are used as photovoltaic power output scenarios; Each scenario includes the photovoltaic power output for each time period and the corresponding Markov state transition probability.
[0011] Preferably, in step 6, the method for scene reduction and construction of scenes for each stage of the wind-solar cascade reservoir system based on the Bisecting k-means clustering algorithm of relevance distance includes: S1: Input random scene time series data; S2: Initialize the number of clusters k Optimal number of clusters in k-means clustering k Generally taken in Between, among N This represents the number of samples in the sample set, which is set in this invention. k It is 10, which is the square root of the sample size of 100; S3: Initialize all data into a single cluster and calculate its SSE value; S4: Select the cluster with the largest SSE and divide it into two clusters based on the k-means algorithm. The steps of clustering using the k-means algorithm are as follows: Divide the data into two groups and randomly select a cluster center for each group; calculate the correlation distance between each object in each group and the cluster center, and assign each object to the cluster center with the closest correlation distance; after each assignment, the cluster center will be recalculated based on the existing objects in the cluster; repeat the calculation and assignment until the termination condition is met; the termination conditions are mainly: no (or minimum number) objects are reassigned, no (or minimum number) cluster centers change, and the sum of squared errors is locally minimized. S5: Recalculate the error after partitioning, select the cluster with the largest SSE, and divide it into two clusters based on the k-means algorithm; S6: Repeat S4 and S5 until the number of clusters is reached. k =10 The algorithm terminates; S7: Output the cluster centers of the clusters after S6 has been divided; The sum of squared errors (SSE) is expressed as shown in equation (16): (16) In the formula: SSE represents the sum of squared errors; i Indicates the first in the cluster i One point; n This represents the total number of points in the cluster; Indicates the weight value; Indicates the first in the cluster i The value of each point; This represents the average value of all points in the cluster. The relevant distance expression is shown in equation (17): (17) In the formula: Represents the relevant distance; For sequence X and Y covariance; , Sequences X and Y covariance; Cluster centers of random scenarios for runoff, wind power output, and photovoltaic output are obtained separately and used as representative scenarios for permutation and combination to construct random scenarios of wind-solar-cascade complementary systems that consider multiple randomnesses.
[0012] Preferably, in step 7, the wind-solar-cascade reservoir system uses a reward function and its constraints that aim to maximize power generation and minimize residual load. The reward function with the objective of maximizing system power generation and minimizing residual load is shown in equation (18). (Since all reservoirs in the cascade reservoirs, as well as all wind farms and photovoltaic power stations, are operating simultaneously at the same time, the relevant symbols are represented by bold vectors.) (18) In the formula: express t Power generation benefits of wind-solar cascade complementary systems, taking into account time periods and residual load penalties; Cascade reservoirs t The output during a given time period reflects the production of hydropower energy; The penalty amount for the deviation between the output and load process of the complementary system; For each reservoir t Initial water level for the specified time period; for t The inflow of water into the headwater reservoir and the inflow of water into the downstream reservoirs during the specified time period are random variables. for t Power generation flow during a given time period; express t Wind power output during certain periods; Indicates photovoltaic power output; in, t Time-based cascade reservoir output As shown in equation (19): (19) In the formula: A This represents the overall output coefficient of each reservoir; For each reservoir t Average hydropower head over the period; Complementary system output deviation from load process penalty As shown in equation (20) (20) In the formula: for t Time-of-day load demand; for t Power output of cascade reservoirs during specific time periods; express t Wind power output during certain periods; Indicates photovoltaic power output; This is the penalty coefficient; β The penalty index; Optimal scheduling of complementary systems includes the following equality constraints and inequality constraints: The water balance constraints of the cascade reservoirs are shown in equation (21): (twenty one) In the formula: , Each reservoir t Initial and final storage capacity of the time period; , They are respectively t Inflow and outflow rates of each reservoir during the specified time period; For reservoir t The duration of power generation during a given period; The cascade water level constraint is shown in equation (22): (twenty two) In the formula: , Each reservoir t Minimum and maximum water level limits at the beginning of the time period; The power generation flow constraint is shown in equation (23): (twenty three) In the formula: This represents the maximum flow rate through each reservoir. The outbound flow constraint is shown in equation (24): (twenty four) In the formula: , Each reservoir t Minimum and maximum outbound flow allowed at the beginning of the time period; The maximum climbing output limit is shown in equation (25): (25) In the formula: for t Power output of cascade reservoirs during specific time periods; for t- Output of cascade reservoirs in one time period; This represents the maximum output amplitude in adjacent time periods; The transmission capacity of the tie line is shown in equation (26): (26) In the formula: for t Power output of cascade reservoirs during specific time periods; , Minimum and maximum limits for the power output of cascade hydropower to the power grid; The limitations on wind and solar power output are shown in equations (27, 28): (27) (28) In the formula: express t Wind power output during certain periods; Indicates photovoltaic power output; This is the maximum output limit for wind power. Maximum output limit for photovoltaic power; In step 7, the DQN is applied to the short-term optimization scheduling model of the wind-solar-cascade reservoir complementary system as follows: S1: Initialize the Q-value table and the initial state of the wind-solar-cascade complementary system; S2: Input the state transition matrix of cascade reservoir runoff, wind power output, and photovoltaic power output; S3: Based on the state transition matrix, the agent acquires knowledge samples by exploring and interacting with the environment using policies; S4: Update the neural network parameters based on the knowledge samples; S5: Repeat steps S3 to S4 until the agent reaches the final state to complete one training session; S6: Repeat S5 until the maximum number of training iterations is reached; S7: Output short-term scheduling strategy for wind-solar-cascade complementary systems.
[0013] The main steps of updating neural network parameters in the DQN algorithm are as follows: S1: The intelligent agent interacts with the environment to acquire a large number of knowledge samples, which are stored in the experience pool; S2: Extract knowledge samples from the experience pool through experience replay and input them into the neural network and loss function; S3: The neural network maps corresponding Q values based on knowledge samples; S4: This is accomplished by updating the main neural network parameters and Q-values through gradient descent of the loss function, and at each interval... β The next training iteration copies its parameters to the target neural network; The DQN algorithm trains agents mainly through two key technologies: experience replay and neural networks. In S1 to S2, the agent acquires a large number of knowledge samples by interacting with the environment and stores them as training data in the experience pool. When there are enough samples in the experience pool, a specified number of data are extracted in an unordered and random manner through experience replay for updating the neural network parameters. This increases the utilization rate of samples and breaks the correlation between samples, thus accelerating the convergence of the algorithm. To address the instability issue when using nonlinear functions to represent value functions, the DQN algorithm employs two networks with identical structures but different parameters to map Q values during process S3. Based on different input data, the main neural network is used to evaluate the action value corresponding to the current state of the cascade reservoir system, referred to as the Q estimate; the target neural network is used to evaluate the action value corresponding to the next state, referred to as the Q target value. In the S4 process, the parameters of the main neural network are based on the core idea of temporal difference. In each Q-value update, the gradient of the loss function is calculated to reduce the prediction error of the neural network. The temporal difference error is defined as the loss function as follows (29): (29) In the formula: For state Take action below The action value obtained by the main neural network is the Q estimate. For state Take action below The action value obtained by the target neural network, i.e., the Q target value; These are the network parameters of the main neural network; These are the network parameters of the target neural network; The discount rate is used to control the impact of future earnings on the present. The main neural network parameters and Q-value updates are shown in equation (30): (30) In the formula: This represents the gradient of the loss function; The learning rate is used to determine the extent to which the error is learned. The parameters of the target neural network are determined by intervals. β The training will adjust the parameters of the main neural network. Copy to the target neural network This parameter update method reduces the correlation between the Q estimate and the Q target value, improves algorithm stability, and ultimately enhances the agent's decision-making ability.
[0014] In step 7, the method for optimizing the DQN algorithm parameters of the wind-solar-cascade reservoir complementary system includes: The DQN algorithm has five learning efficiency parameters, including the learning rate. Greed rate Discount rate These are parameters related to reinforcement learning, and are the same as those in the Q-learning algorithm; the target neural network parameter update interval... β Parameters related to deep learning; water level discretization accuracy hThis parameter belongs to the "Other" category. Considering that water level adjustments in actual scheduling are often measured in meters, it is set to 1m. The Q-learning algorithm only needs to optimize the learning rate, greedy rate, and discount rate. To simplify the optimization process, the learning efficiency parameters of the DQN algorithm are divided into two parts: reinforcement learning-related parameters and deep learning-related parameters. First, the reinforcement learning-related parameters are optimized using the Q-learning algorithm. Then, based on its optimal parameters, the deep learning-related parameters are optimized using the DQN algorithm. That is: S1: By arranging and combining the learning rate and greedy rate in the Q-learning algorithm within a certain range, the parameter combination with the largest cumulative return value is selected as the optimal parameter for reinforcement learning. In addition, since the future returns of this reservoir scheduling model have no impact on the current returns, the discount rate is set to 1. S2: The optimal learning rate and greedy rate were applied to the DQN algorithm, and the neural network parameter update interval was tested and analyzed in the same way.
[0015] Beneficial effects of this invention: The present invention proposes a multi-objective short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs based on the deep reinforcement learning (DQN) algorithm, which considers the stochasticity of runoff, wind power output, and photovoltaic output, and aims to maximize total power generation and minimize surplus load. The technical effects are as follows: 1) This invention constructs a multi-objective short-term optimization scheduling model with the goal of maximizing total power generation and minimizing residual load. This model can meet load demand, smooth residual load fluctuations, and maximize the economic benefits of hydropower stations, wind power stations, and photovoltaic power stations.
[0016] 2) This invention considers multiple randomnesses when constructing the optimized scheduling model, including runoff randomness, wind power output randomness, and photovoltaic output randomness, making the model closer to reality. Methods for generating random scenarios corresponding to runoff, wind power output, and photovoltaic output are proposed respectively.
[0017] Since the reservoir runoff fluctuates relatively little, the spatial and temporal correlation of each stage of runoff is directly described by the Copula function based on the Pearson Type III distribution, and its randomness is described by its probability density function and Markov transition probability. The corresponding random scenario is generated based on the Markov Monte Carlo sampling method.
[0018] To address the significant fluctuations in wind and solar power output, this paper proposes a kernel density estimation method that does not require prior assumptions about the distribution of random variables. This method fits the probability density of wind and solar power output at each stage, solves for the probability density function and Markov state transition probability, describes the randomness of wind and solar power output, and generates corresponding random scenarios based on the Markov Monte Carlo sampling method.
[0019] 3) This invention constructs a large-scale scenario to describe the stochastic characteristics of the wind-solar-cascade reservoir complementary system. For the data characteristics of runoff, wind power output, and photovoltaic power output, the scenario is reduced by the Bisecting k-means algorithm based on the correlation distance. The cluster center is used as the representative scenario to describe the complex scenario characteristics, thereby improving the computational efficiency.
[0020] 4) The wind-solar-cascade reservoir complementary system suffers from a high degree of dimensionality and explosive growth in its state space due to its multiple random variables, resulting in a severe curse of dimensionality. The DQN algorithm combines the advantages of reinforcement learning and deep learning algorithms, effectively addressing the curse of dimensionality caused by the high dimensionality of state and decision variables in the short-term scheduling of wind-solar-cascade reservoir complementary systems, demonstrating good effectiveness. The DQN model shows a significant improvement in optimization performance compared to scheduling models based on traditional reinforcement learning algorithms. It employs artificial neural network fitting from deep learning algorithms instead of storing Q-values in a Q-table, greatly accelerating computation while maintaining accuracy. Furthermore, its experience replay technique breaks down the correlation between samples, increasing sample utilization, meaning more samples can be obtained with less historical data.
[0021] 5) This invention divides the DQN learning efficiency parameters into two parts: reinforcement learning-related parameters and deep learning-related parameters, and performs optimization step by step, which simplifies the optimization process and improves the solution performance of the DQN algorithm for short-term stochastic optimization scheduling of wind-solar-cascade reservoir systems. Attached Figure Description
[0022] Figure 1 A flowchart for short-term stochastic optimization scheduling of wind-solar-cascade reservoirs using deep reinforcement learning; Figure 2 The flowchart for the Bisecting k-means clustering algorithm based on relevance distance is shown below. Figure 3 Here is a flowchart of the DQN algorithm; Figure 4 Flowchart for updating neural network parameters in the DQN algorithm; Figure 5 For the selection of greedy rate and learning rate in the Q-learning algorithm; Figure 6 This is a comparison chart showing the convergence of the DQN algorithm and the Q-learning algorithm with the same number of iterations. Detailed Implementation
[0023] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0024] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0025] Step 1: Treat the short-term scheduling of the wind-solar-cascade reservoir system as a multi-stage decision problem, and divide the short-term scheduling into multiple stages to solve the scheduling strategy.
[0026] Step 2: Based on the historical runoff data of the cascade reservoirs, use the Pearson Type III distribution to describe the probability of inflow into each reservoir at each time period.
[0027] Step 3: By comparing the correlation characteristics of the function in the Copula function with the correlation characteristics of the inflow runoff of the cascade reservoirs, the function with the most similar correlation characteristics is selected to conduct spatial correlation analysis on the inflow runoff of the cascade reservoirs. Based on the Pearson Type III distribution obtained in Step 2, the joint probability density function and joint probability distribution of the cascade reservoirs at each time period are obtained.
[0028] Step 4: Solve the joint probability distribution of adjacent time periods of the cascade reservoirs using the function selected in Step 3 to describe their temporal correlation. Then, based on the conditional distribution formula, solve for the random runoff state transition probability matrix of adjacent time periods. Finally, based on the MH (Metropolis-Hasting) algorithm in the Markov Monte Carlo sampling method, randomly simulate 100 sets of cascade reservoir runoff scenarios.
[0029] Step 5: Based on historical wind power output data, solve for the probability density function and Markov state transition probability of wind power output using kernel density estimation. Then, use the Metropolis-Hasting (MH) algorithm in the Markov Monte Carlo sampling method to randomly simulate 100 sets of random scenarios for future wind power output. The bandwidth parameter of kernel density estimation is evaluated using the least squares cross-validation test (LSCV) method to evaluate the accuracy of kernel density estimation.
[0030] Step 6: Based on historical photovoltaic (PV) output data, solve for the probability density function of PV output and the Markov state transition probability using kernel density estimation. Then, use the Metropolis-Hasting (MH) algorithm in the Markov Monte Carlo sampling method to randomly simulate 100 sets of future PV output scenarios. The bandwidth parameter of kernel density estimation is evaluated using the least squares cross-validation test (LSCV) method to assess the accuracy of kernel density estimation.
[0031] Step 7: Based on the Bisecting k-means clustering algorithm, the scenes obtained in Steps 4, 5, and 6 are reduced to 10 groups of scenes respectively. By combining them, 1000 groups of wind and solar cascade reservoir system scenes are obtained.
[0032] Step 8: Construct the environment for the deep reinforcement learning algorithm using the random scenario obtained in Step 7. Divide the learning efficiency parameters of the DQN algorithm into two parts: reinforcement learning-related parameters and deep learning-related parameters. First, optimize the reinforcement learning-related parameters using the Q-learning algorithm. Then, optimize the deep learning-related parameters using the DQN algorithm based on the optimized parameters. Determine the parameters of the DQN algorithm for the wind-solar-cascade reservoir system with the goal of maximizing system power generation and minimizing residual load.
[0033] Step 9: Using the optimized algorithm with parameters obtained in Step 8, solve for the optimal short-term scheduling strategy of the wind-solar-cascade reservoir system.
[0034] In step 1, the wind-solar-cascade reservoir system refers to a complementary power generation system consisting of wind farms, photovoltaic power stations, and cascade reservoir hydropower stations. In step 1, short-term scheduling in this invention refers to the system's operational plan for the next day, with a daily scheduling cycle and hourly scheduling periods. In step 3, the spatial correlation of cascade reservoirs refers to the correlation of water inflows from cascade reservoirs distributed in different spaces at the same time. In step 3, the temporal correlation of cascade reservoirs refers to the correlation of water inflows from cascade reservoirs at different times.
[0035] In step 4, the relevant characteristics are obtained through the Pearson correlation test. The Markov correlation test examines whether the degree of correlation between variables satisfies the requirements of a finite Markov decision process. A finite Markov decision process is the classic formal expression of sequential decision-making, in which actions not only affect the current immediate payoff but also the subsequent state and future payoffs. The Pearson correlation test method is shown in equation (1): (1) In the formula: This represents the correlation coefficient between random variables in adjacent time periods; N The total number of years in the sample. , for t Time period and t +1 period i Historical data from [year], , for t Time period and t The average input value of all historical data for the +1 period, , They are respectively t Time period, t +1 period of historical data mean squared error.
[0036] In step 4, the Markov state transition probability matrix refers to the matrix describing the transition probabilities of each state of a variable in a Markov process. When a stochastic process... If the conditional probability satisfies the following equation, then the stochastic process is considered to have Markov properties, where... T For discrete time sets, for t The set of states of the random variable at time t is shown in equation (2): (2) In the formula, t represents discrete time; for t The set of states of a random variable at any given time; for t The state at any given moment.
[0037] Markov state transition probability matrix As shown in equation (3): (3) In the formula: M represents the total number of states composed of discrete runoff values, where , j for t -1 moment j A state, k for t Time of the first k One state; From -1 time period status j Transferred to t Markov state transition probabilities for state k in time period.
[0038] in The solution formula is shown in equation (4): (4) In the formula: From -1 time period status j Transferred to t Markov state transition probability for state k in time period; for t Time-based random variables The value is When, random variable t -1 time period The value is The probability of, where , They are respectively t Time period, t The set of random variable values in the -1 time period; for t random variables over time period The value of the random variable ,and t-1 time period random variable The value is The probability of; for t Random variables in the time interval -1 Value The probability of.
[0039] In step 4, the Markov Monte Carlo sampling method is a sampling method based on Markov transition probabilities combined with Monte Carlo sampling, aiming to obtain samples with the same distribution as the object to be sampled. Its basic steps are: constructing a Markov chain based on the distribution of the object to be sampled, and iterating until the stationary distribution of the Markov chain converges to the target distribution; starting from the initial state, performing state transitions based on the stationary Markov chain to obtain a sufficient number of samples. The stationary distribution of the Markov process is shown in equation (5): (5) In the formula: Transition matrix during MC process P The stable distribution; i , j Each of the two adjacent states.
[0040] In step 4, the cascade reservoir runoff scenario refers to a data sample that includes the inflow runoff from the cascade reservoirs and the corresponding Markov transition probabilities.
[0041] In step 5, kernel density estimation is a nonparametric estimation method used to fit the density function of an unknown random variable. Unlike other parametric estimation methods, this method has the advantage of not requiring prior knowledge of the random variable samples, i.e., it does not require assuming the sample data conforms to a certain distribution and avoids parametric estimation. However, due to the large fluctuations in wind and solar power output, if the kernel density is calculated by assuming it follows a certain distribution to predict wind and solar power output, the accuracy of the prediction results will be reduced. Therefore, this invention uses the nonparametric estimation method kernel density estimation to fit the probability density distribution of wind and solar power output. Let... It is a probability density function For an unknown population of independent and identically distributed random variables, the kernel density estimate is defined as shown in equation (6): (6) In the formula: x It is a random variable; n For sample size; For the first i One sample; h For bandwidth parameters; For a kernel function, it must satisfy the following conditions: , , .
[0042] Kernel function Here, we choose the most widely used Gaussian kernel function. The Gaussian kernel function is shown in equation (7): (7) And when random variable X When the dimension is multidimensional, the definition of multidimensional kernel density estimation is shown in equation (8): (8) In the formula: d For dimensions; X It is a multidimensional random variable; The dimension is j random variables; n For sample size; The dimension is j The first random variable i One sample; h d for d Bandwidth parameters under a given dimension; For a kernel function, it must satisfy the following conditions: , .
[0043] In steps 5 and 6, the wind power and photovoltaic output scenarios refer to data samples that include wind power output and corresponding Markov transition probabilities.
[0044] In step 7, since the number of scene combinations obtained in steps 4, 5 and 6 is huge, it causes great difficulty in calculation. Therefore, it is necessary to reduce the number of scenes in each stage of the wind-solar cascade reservoir system, that is, to remove some similar scenes, in order to shorten the solution time.
[0045] In step 7, the Bisecting k-means clustering algorithm based on correlation distance is an improved version of the k-means algorithm. The main improvements are twofold: Firstly, addressing the sensitivity of the traditional k-means algorithm to the selection of initial cluster centers, the Bisecting k-means algorithm incorporates hierarchical clustering, recursively splitting from top to bottom using a splitting method, resulting in greater distances between the centers and preventing the initial cluster centers from being selected into the same cluster. This improves computational speed and, to some extent, overcomes the algorithm's tendency to get trapped in local optima. Secondly, for long-series time-series data on runoff, wind power output, and photovoltaic power output, Euclidean distance clustering is ineffective; therefore, correlation distance is proposed to replace Euclidean distance in representing the correlation between two feature columns. The criterion for clustering is the sum of squared errors (SSE). This value evaluates the clustering results; a smaller SSE indicates that the data points are closer to their centroids, resulting in better clustering. The SSE is shown in equation (9). (9) In the formula: SSE represents the sum of squared errors; i Indicates the first in the cluster i One point; n This represents the total number of points in the cluster; Indicates the weight value; Indicates the first in the cluster i The value of each point; This represents the average value of all points in the cluster.
[0046] The relevant distance is shown in equation (10): (10) In the formula: Represents the relevant distance; For sequence X and Y covariance; , Sequences X and Y The covariance.
[0047] Relevant distance The smaller the value, the higher the correlation between the two feature columns.
[0048] In step 8, the deep reinforcement learning algorithm refers to an algorithm that, within the reinforcement learning framework, combines nonlinear fitting of neural networks in deep learning to approximate value functions and maximizes long-term gains through continuous trial and error. The main components of the reinforcement learning framework are the agent and the environment. The agent, based on a Markov decision process, interacts with the environment at discrete time steps to acquire knowledge, and is trained (i.e., updated) with the goal of maximizing value estimation, ultimately obtaining the optimal strategy. In this study, the cascade reservoir will act as the agent, interacting with the wind-solar cascade joint scheduling environment to acquire knowledge. The four-tuple knowledge samples acquired in each interaction are shown in equation (11): (11) In the formula: for t The state of the agent at any given time is determined by the scheduling time step. Cascade reservoir water level collection composition; for t The agent's state at time +1 is determined by the scheduling time step. Cascade reservoir water level collection composition; for t Real-time agent actions; for t The reward for each action at any given moment is the value of the power generation benefit generated by the scheduling. The solution is obtained based on random scenarios.
[0049] In the decision-making process, the expected value of an agent taking an action in a certain state is called the action value, also known as the Q-value. Reinforcement learning algorithms update the Q-value to approximate the optimal action value in order to find the optimal policy for multi-stage problems. The Bellman optimal action value function is shown in equation (12): (12) In the formula: Representing state Take action below The optimal action value obtained subsequently; The discount rate is used to control the impact of future earnings on the present. Indicates the system from state Transition to the next state The Markov transition probabilities are obtained by multiplying the Markov transition probabilities of each random variable in the system at each stage. They are used to describe the randomness of the inflow of water into the cascade reservoirs, wind power output, and photovoltaic power output in the system. S A set of states; for t Moment-based action rewards.
[0050] Regarding action selection, reinforcement learning determines it through an exploration-exploitation strategy. In this strategy, the agent learns from the environment by exploring and trying different actions, but the impact of actions on the outcome is uncertain. Exploring and exploiting the best action given the current information allows for high-value immediate gains, but it can easily lead to getting trapped in local optima. This paper uses a reinforcement learning framework... ε The -greedy exploration strategy has the advantage that it does not depend on any specific information about the environment. ε In the -greedy policy, the action selection probability is shown in equation (13): (13) In the formula: express t Moment State The probability of randomly selecting an action; express t Moment State The probability of the person with the highest assessed value choosing an action; The greed rate is the probability of choosing an action through exploration.
[0051] In step 8, the environment of the deep reinforcement learning algorithm refers to the constructed model containing sample knowledge. Step 8, aiming for maximum system power generation and minimum remaining load, means maximizing the total power generation of the wind-solar-cascade reservoir system and minimizing the difference between output and load, i.e., the output process most closely approximates the load process. Thus, the wind-solar-cascade reservoir system can maximize power generation efficiency and minimize remaining load fluctuations while ensuring power load demand. Unlike other methods, in deep reinforcement learning, the objective is solved through agent action rewards. In deep reinforcement learning, the environment sends a scalar value called a reward to the agent at each step of its action. The agent's sole objective is to maximize long-term total revenue; therefore, the reward determines the quality of the agent's decisions and is the main basis for determining the scheduling strategy.
[0052] The reward function with the objective of maximizing system power generation and minimizing residual load is shown in equation (14). (Since all reservoirs in the cascade reservoirs, as well as all wind farms and photovoltaic power stations, are operating simultaneously at the same time, the relevant symbols are represented by bold vectors.) (14) In the formula: express t Power generation benefits of wind-solar cascade complementary systems, taking into account time periods and residual load penalties; Cascade reservoirs t The output during a given time period reflects the production of hydropower energy; The penalty amount for the deviation between the output and load process of the complementary system; For each reservoir t Initial water level for the specified time period; for t The inflow of water into the headwater reservoir and the inflow of water into the downstream reservoirs during the specified time period are random variables. for t Power generation flow during a given time period; express t Wind power output during certain periods; This indicates the output of photovoltaic power.
[0053] in, t Time-based cascade reservoir output As shown in equation (15): (15) In the formula: A This represents the overall output coefficient of each reservoir; For each reservoir t Average hydropower head over a given period.
[0054] Complementary system output deviation from load process penalty As shown in equation (16) (16) In the formula: for t Time-of-day load demand; for t Power output of cascade reservoirs during specific time periods; express t Wind power output during certain periods; Indicates photovoltaic power output; This is the penalty coefficient; β This is the penalty index.
[0055] Optimal scheduling of complementary systems includes the following equality constraints and inequality constraints: The water balance constraints of the cascade reservoirs are shown in equation (17): (17) In the formula: , Each reservoir t Initial and final storage capacity of the time period; , They are respectively t Inflow and outflow rates of each reservoir during the specified time period; For reservoir t The duration of power generation during a given period.
[0056] The cascade water level constraint is shown in equation (18): (18) In the formula: , Each reservoir t Minimum and maximum water level limits at the beginning of the time period.
[0057] The power generation flow constraint is shown in equation (19): (19) In the formula: This represents the maximum flow rate through each reservoir.
[0058] The outbound flow constraint is shown in equation (20): (20) In the formula: , Each reservoir t Minimum and maximum outbound flow allowed at the beginning of the time period.
[0059] The maximum climbing output limit is shown in equation (21): (twenty one) In the formula: fort Power output of cascade reservoirs during specific time periods; for t- Output of cascade reservoirs in one time period; This represents the maximum output amplitude in adjacent time periods.
[0060] The transmission capacity of the tie line is shown in equation (22): (twenty two) In the formula: for t Power output of cascade reservoirs during specific time periods; , The minimum and maximum limits for the power output of cascade hydropower to the power grid.
[0061] The limitations on wind and solar power output are shown in equations (23, 24): (twenty three) (twenty four) In the formula: express t Wind power output during certain periods; Indicates photovoltaic power output; This is the maximum output limit for wind power. This is the maximum output limit for photovoltaic power.
[0062] In step 8, the DQN algorithm is a deep reinforcement learning algorithm. The DQN algorithm uses neural network fitting to achieve table approximation in high-dimensional scenarios. Through table approximation, the approximate Q-value can be obtained by inputting the agent's state and actions into the neural network. The Q-value is the expected value obtained by the agent taking a certain action in a certain state, called the action value. Therefore, deep reinforcement learning algorithms greatly improve the problem of the curse of dimensionality when using stochastic dynamic programming to solve the optimal scheduling of cascade reservoirs. The DQN algorithm completes agent training by updating neural network parameters based on the idea of temporal difference during the Q-value update process.
[0063] Since neural networks are nonlinear functions, it's theoretically impossible to prove whether the algorithm will converge. Therefore, the DQN algorithm uses two key techniques to optimize this problem. First, experience replay is used to break the correlation between samples and increase sample utilization. Second, a fixed target network is used to address the problem of unstable algorithm updates and improve training stability. The DQN algorithm uses two networks with identical structures but different parameters to map Q-values. Based on different input data, the main neural network is used to evaluate the action value corresponding to the current state of the cascade reservoir system, called the Q-estimate; the target neural network is used to evaluate the action value corresponding to the next state, called the Q-target value.
[0064] The main neural network parameters are based on the core idea of temporal difference, and are updated by calculating the gradient of the loss function in each Q-value update to reduce the prediction error of the neural network. The temporal difference error is defined as the loss function, as shown in equation (25): (25) In the formula: In the state action The action value, i.e., the Q estimate, is obtained from the main neural network. In the state action The action value, i.e., the Q target value, is obtained by the target neural network. The discount rate is used to control the impact of future earnings on the present. These are the network parameters of the main neural network; These are the network parameters of the target neural network; for t+ 1-moment action reward; This refers to timing difference error.
[0065] The parameters of the main neural network and the Q value are updated by gradient descent to address the temporal difference error, as shown in equation (26): (26) In the formula: This represents the gradient of the loss function; In the first k The next iteration and the state action Action value obtained by the lower main neural network; In the k+ 1 iteration and state action Action value obtained by the lower main neural network; These are the network parameters of the main neural network; It is expressed as the learning rate, which determines the extent to which the error is learned.
[0066] The parameters of the target neural network are determined by intervals. β The training will adjust the parameters of the main neural network. Copy to the target neural network This parameter update method reduces the correlation between the Q-estimate and the Q-target value, improves algorithm stability, and ultimately enhances the agent's decision-making ability. In step 8, the learning efficiency parameter directly affects the stability and search efficiency of deep reinforcement learning and is highly sensitive. The DQN algorithm has five learning efficiency parameters, including the learning rate. Greed rate Discount rate Target neural network parameter update interval β、 Water level discretization accuracy h In step 8, parameter tuning refers to adjusting the algorithm parameters to optimal values to achieve the best algorithm performance. DQN algorithm parameters include two main categories: control parameters and learning efficiency parameters. Because control parameters have low sensitivity, changes to them only cause gradual changes in the system state, while learning efficiency parameters have high sensitivity and directly affect the stability and search efficiency of deep reinforcement learning. Therefore, this invention focuses solely on parameter tuning for the learning efficiency parameters in the DQN algorithm. Specifically, the learning rate... Greed rate Discount rate Similar to the Q-learning algorithm, this belongs to the category of reinforcement learning parameters; the target neural network parameter update interval. β Parameters related to deep learning; water level discretization accuracy h This parameter is classified as "Other Parameters." Considering that water level adjustments in actual scheduling are often measured in meters, it is set to 1m. To simplify the optimization process, this invention divides the learning efficiency parameters of the DQN algorithm into two parts: reinforcement learning-related parameters and deep learning-related parameters. First, the reinforcement learning-related parameters are optimized using the Q-learning algorithm, and then the deep learning-related parameters are optimized using the DQN algorithm based on their optimal parameters.
[0067] In step 9, the optimal strategy refers to the short-term scheduling strategy of the wind-solar-cascade reservoir complementary system that maximizes total power generation and minimizes residual load.
[0068] The following is in conjunction with the appendix Figure 1 The specific implementation plan described above will be explained in detail below: 1. The short-term scheduling of the wind-solar-cascade reservoir system is divided into 12 stages, with each stage lasting one hour.
[0069] 2. The specific steps for describing the inflow probability of each reservoir during different time periods using the Pearson Type III distribution are as follows: The probability density function of Pearson type III is shown in equation (27): (27) In the formula: , , These represent the position, shape, and size of the Pearson-III type probability density curve in the probability grid paper, respectively. x For frequency; f ( x () represents the inflow runoff. e It is the base of the natural logarithm. for The function is shown in equation (28): (28) In the formula: Indicates the position of the Pearson-III probability density curve on the probability grid paper; e is the base of the natural logarithm.
[0070] Relevant parameter values were obtained statistically based on historical runoff data. These values were then substituted into Hessian probability graph paper, and adjustments were made using the fitting method. The value is set to maximize the fit between the scatter points and the probability density curve in the probability grid.
[0071] 3. The specific process of using the Copula function to solve for the joint probability density function and joint probability distribution of the cascade reservoirs at different time periods: Because the Archimedes Copula family of functions is relatively simple in construction and widely used in hydrology, it was initially selected as the first Copula family. By comparing the correlation characteristics of the main functions in the Archimedes Copula family with the correlation characteristics of inflow runoff from cascade reservoirs, this paper ultimately adopts the Clayton Copula function, which has the most similar correlation characteristics, to describe the spatial correlation of cascade reservoirs. The Pearson-III type probability distribution function of each reservoir is used as a marginal distribution to solve for the runoff state. Joint probability distribution of downstream reservoirs As shown in equation (29): (29) In the formula: Cascade reservoir runoff status Joint probability distribution of downstream reservoirs; for t Time period n Inflow of water from the first-level reservoir The Pearson-III type probability distribution function; for t The join parameters of the Clayton Copula function for the time period can be estimated using the Kendall correlation coefficient.
[0072] Furthermore, the joint probability density function can be obtained by inverse transformation, i.e., by taking the partial derivatives separately.
[0073] 4. See the appendix for the specific steps of solving the Markov transition probability matrix of the cascade reservoirs and generating random scenarios. Figure 2 : Through correlation analysis, the historical inflow data of the cascade reservoirs follows Markov characteristics. The specific steps for solving the Markov probability transition matrix of wind power output are as follows: S1: Obtained by solving the Clayton Copula function. t Time and tJoint distribution function at time +1 , S2: Obtain the joint probability distribution of adjacent time periods of the cascade reservoirs using the Clayton Copula function. S3: Discretize the inflow runoff of each reservoir at each stage into the same number of states, with the maximum and minimum inflow runoff as the ranges; S4: Calculate the Markov transition probability of the runoff state of the cascade reservoirs in adjacent time periods using the conditional distribution formula. As shown in equation (30): (30) In the formula: The Markov transition probability of the runoff state of the cascade reservoirs in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, random variable inflow into the reservoir t The value of the time period The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Random variable inflow into the reservoir during a time period The value is And random variable inflow into the reservoir t The value of the time period The probability of; For random variables, inflow runoff t The value of the time period The probability of.
[0074] S5: Solve for each discrete value of runoff to obtain the state transition probability matrix for this stage. S6: Repeat the above steps to obtain the state transition probability matrix for each stage, which is used to describe the Markov process of the inflow runoff of the cascade reservoirs in adjacent periods.
[0075] The specific steps for generating random scenarios using the MH algorithm in Markov Monte Carlo sampling are as follows: S1: Input the state transition probability matrix for each stage. Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =0, respectively from the transition matrix With uniform distribution Mid-sampling, obtained t+ Phase 1 sampled values and random values Accept sampled values The criterion is shown in equation (31): (31) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A set of random variables for inflow into the reservoir over a given period.
[0076] If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0. Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: The Markov chains constructed by repeating all stages of S2 converge to a stationary distribution; S4: In the first stage, i.e. t =1 Continue samplingz This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage; thus obtaining z The samples that conform to the probability density function and state transition probability of each stage are used as the inflow scenarios of cascade reservoirs. Each scenario includes the inflow runoff value of each reservoir at each time period and the corresponding Markov state transition probability.
[0077] 5. The specific steps for solving the probability density function of wind power output and the Markov state transition probability using the Markov Monte Carlo sampling method based on kernel density estimation, and for generating random scenarios, are detailed in the appendix. Figure 3 : The probability density distribution is fitted using either a multidimensional or one-dimensional kernel density estimation function based on the number of wind farms. If there is only one wind farm, a one-dimensional kernel density estimation function is used; if there are two or more, a multidimensional kernel density estimation function is used. Here, a one-dimensional kernel density estimation function is used for explanation. The optimal bandwidth parameter in the kernel density estimation function is solved using the least squares cross-validation test (LSCV) method, and its solution is relatively ideal and widely recognized. Furthermore, the bandwidth parameter of the multidimensional kernel function estimation can be solved using the Gaussian kernel. The least squares cross-validation test (LSCV) method is a calculation method based on the minimum integral squared error (ISE) criterion. ISE is shown in equation (32): (32) In the formula: It is a random variable; For kernel density estimation, It is a probability density distribution.
[0078] The last term in the above three terms is independent of bandwidth, so the first two terms are used for least squares cross-validation to minimize them and obtain the optimal bandwidth parameter. By substituting the one-dimensional kernel density estimation function, the one-dimensional least squares cross-validation test is shown in equation (33): (33) In the formula: For the first i One sample; For the first j One sample; h For bandwidth parameters; For kernel functions; This represents the kernel function convolved with itself; n is the sample size.
[0079] Through correlation analysis, the random variable of wind power output follows Markov characteristics. The specific steps to solve the Markov probability transition matrix of wind power output are as follows: S1: Solve for the joint probability density function between two adjacent time periods. This involves integrating the probability density functions of wind power output at time t and time t+1 to obtain the probability distribution function. , As a marginal probability, the Copula function, which is most closely related to wind power output, is chosen to solve the joint probability function. S2: Discretize the wind power output into several states based on the maximum and minimum wind power output; S3: Calculate the wind power output in adjacent time periods based on the conditional distribution formula, considering the output states. Transition to the next state Markov transition probability As shown in equation (34): (34) In the formula: The Markov state transition probability of wind power output state in adjacent time periods; for t+ 1-period random variable inflow into the reservoir The value is hour, t Time-of-use random variable: wind power output The value is The probability of, where , They are respectively t+ 1 time period t A set of random variables related to inflow into the reservoir over a specific time period; for t+ Wind power output as a random variable over a time period The value is ,and t Time-of-use random variable: wind power output The value is The probability of; Wind power output is a random variable t The value of the time period The probability of; The state transition probability matrix is obtained based on the random runoff state transition probabilities of each adjacent time period. This is used to describe the Markov process of inflow into a cascade reservoir during adjacent time periods. Then, the MH algorithm in the Markov Monte Carlo sampling method is used to sample this stochastic model. The specific steps are the same as those for generating a stochastic scenario using the MH algorithm in the Markov Monte Carlo sampling method for cascade reservoirs, only the objects differ. The sampled values are then accepted. The criterion is shown in equation (35), which differs from equation (31) in the ladder only in the object: (35) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period.
[0080] 6. The specific steps for solving the probability density function of photovoltaic power output and the Markov state transition probability using the Markov Monte Carlo sampling method based on kernel density estimation, and generating random scenarios, are the same as those for wind power, only the objects differ. See Appendix. Figure 4 Markov transition probability of photovoltaic output The only difference between this equation (34) and the one used in wind power is the object of the equation, as shown in equation (36): (36) In the formula: The Markov state transition probability of photovoltaic power output states in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, the random variable photovoltaic output t Time period The value is The probability of, where , They are respectively t+ 1 time period tA collection of inflow runoff that is a random variable over a specific time period; for t+ Photovoltaic power output as a random variable over a time period of 1 The value is And random variable photovoltaic output t The value of the time period The probability of; Photovoltaic output as a random variable t The value of the time period The probability of; Accept sampled values The criterion is shown in equation (37), which differs from that of the cascade (31) and wind power (35) only in the object: (37) In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period.
[0081] 7. The specific process of scene reduction and construction of each stage of the wind-solar cascade reservoir system using the Bisecting k-means clustering algorithm based on relevance distance is shown in the attached algorithm flowchart. Figure 2 : S1: Input random scene time series data; S2: Initialize cluster number k Optimal number of clusters in k-means clustering k Generally taken in Between, among N This represents the number of samples in the sample set, which is set in this invention. k S3: Initialize all data into a single cluster and calculate its SSE value; S4: Select the cluster with the largest SSE and divide it into two clusters based on the k-means algorithm, i.e., randomly select a cluster center for each group, calculate the correlation distance between each object in each group and the cluster center, and assign each object to the cluster center with the closest correlation distance; after each assignment, the cluster center is recalculated based on the existing objects in the cluster; repeat the calculation and assignment until the termination condition is met. The termination conditions are mainly: no (or minimum number) objects are reassigned, no (or minimum number) cluster centers change, and the sum of squared errors is locally minimized; S5: Recalculate the SSE value after partitioning, select the cluster with the largest SSE value, and divide the data into two clusters again; S6: Repeat S4 and S5 until the number of clusters is reached. k =10 The algorithm terminates; S7: Output the cluster centers of the clusters after S6 partitioning.
[0082] Cluster centers of random scenarios for runoff, wind power output, and photovoltaic output were obtained separately and used as representative scenarios for permutation and combination. A total of 1,000 random scenarios of the wind-solar-cascade complementary system considering multiple randomness were constructed as training data for the algorithm.
[0083] 8. From the appendix Figure 3 It can be seen that the specific steps of applying the deep reinforcement learning algorithm DQN to the short-term optimization scheduling model of the wind-solar-cascade reservoir complementary system are as follows: S1: Initialize neural network parameters; S2: Input random scenario of wind-solar cascade complementary system; S3: Agent explores and interacts with the environment to acquire knowledge samples; S4: Update neural network parameters based on knowledge samples; S5: Repeat steps S3 to S4 until the agent reaches the final state to complete one training cycle; S6: Repeat S5 to the maximum number of training cycles; S7: Output short-term scheduling strategy of wind-solar-cascade complementary system.
[0084] From the appendix Figure 4 As can be seen, the main steps of updating neural network parameters in the DQN algorithm are as follows: S1: The agent interacts with the environment to acquire a large number of knowledge samples, which are stored in an experience pool; S2: Knowledge samples are extracted from the experience pool through experience replay and input into the neural network and loss function; S3: The neural network maps corresponding Q values based on the knowledge samples; S4: The parameters of the main neural network are updated and the Q values are updated by performing gradient descent on the loss function, and this is completed at each interval. β The training iteration copies its parameters to the target neural network.
[0085] The specific process for parameter tuning of the DQN algorithm for wind-solar-cascade reservoir complementary systems: S1: By arranging and combining the learning rate and greedy rate within a certain range in the Q-learning algorithm, the parameter combination with the largest cumulative reward value is selected as the optimal parameter for reinforcement learning. (See appendix) Figure 6 Furthermore, since the future revenue of this reservoir scheduling model has no impact on the current revenue, the discount rate is set to 1; S2: The optimized learning rate and greedy rate are applied to the DQN algorithm, and the neural network parameter update interval is tested and analyzed in the same way, as shown in Table 1.
[0086] From the appendix Figure 5 It can be seen that the cumulative return is maximized when the learning rate is 0.03 and the greedy rate is 0.85. Therefore, both the Q-learning algorithm and the DQN algorithm select this parameter combination for training. Furthermore, according to Table 2, the update interval for the target neural network parameters in the DQN algorithm is set to 10. The final optimal parameters are shown in Table 2.
[0087] Table 1 Selection of target neural network parameter update interval
[0088] Table 2 Optimal parameters for the DQN algorithm
[0089] 9. Solve the optimal strategy for short-term scheduling of the wind-solar-cascade reservoir system based on the DQN algorithm with optimized parameters.
[0090] To demonstrate the advantages of the deep reinforcement learning algorithm DQN compared to traditional reinforcement learning algorithms such as Q-learning and SDP, DQN, Q-learning, and SDP were applied to a long-term stochastic optimization scheduling model of cascade reservoirs for comparative analysis and verification. The algorithm used a dataset consisting of 1000 random scenario samples from a wind-solar-cascade complementary system, with the training set, validation set, and test set ratios of 70%, 20%, and 10%, respectively. The number of iterations was set to 5000. Figure 6It can be seen that the Q-learning algorithm converges around 3000 iterations, while the DQN algorithm starts to accelerate its convergence around 800 iterations and converges around 2000 iterations, showing a significantly faster convergence speed compared to the Q-learning algorithm. In terms of solution quality, due to the traversal-based solution method of the SDP algorithm, using the SDP algorithm's solution as a standard, the solution results of the DQN, Q-learning, and SDP algorithms are almost identical. However, the solution times for DQN, Q-learning, and SDP algorithms are 81.23 minutes, 238.64 minutes, and 459.76 minutes, respectively. This means that the DQN algorithm is 65.96% faster than the Q-learning algorithm and 5.66 times faster than the SDP model, significantly improving learning efficiency.
Claims
1. A short-term stochastic optimization scheduling method based on deep reinforcement learning for wind-solar-cascade reservoirs, characterized in that: It includes the following steps: Step 1: Treat the short-term scheduling of the wind-solar-cascade reservoir system as a multi-stage decision problem, and divide the short-term scheduling into multiple stages to solve the scheduling strategy; Step 2: Based on the historical runoff data of the cascade reservoirs, the probability of inflow into the cascade reservoirs at each time period is described using the Pearson Type III distribution. Then, by comparing the correlation characteristics of the function in the Copula function with the correlation characteristics of the inflow into the cascade reservoirs, the function with the most similar correlation characteristics is selected to conduct spatial correlation analysis on the inflow into the cascade reservoirs, and the joint probability density function and joint probability distribution of the cascade reservoirs at each time period are obtained. Step 3: Solve the joint probability distribution of adjacent time periods of the cascade reservoirs using the function selected in Step 2 to describe their temporal correlation. Then, based on the conditional distribution formula, solve for the Markov state transition probability matrix of random runoff in adjacent time periods, and use the Markov Monte Carlo sampling method to randomly simulate the random scenario of inflow runoff into the cascade reservoirs. Step 4: Based on historical wind power output data, a Markov Monte Carlo sampling method based on kernel density estimation is proposed to solve the probability density function of wind power output and the Markov state transition probability matrix, and to randomly simulate future wind power output scenarios. Step 5: Based on historical photovoltaic power output data, a Markov Monte Carlo sampling method based on kernel density estimation is proposed to solve the probability density function of photovoltaic power output and the Markov state transition probability matrix, and to randomly simulate future photovoltaic power output scenarios. Step 6: The scenes obtained in Steps 3, 4 and 5 are reduced by Bisecting k-means clustering algorithm based on relevance distance, and then combined to form the scenes of each stage of the wind-solar cascade reservoir system. Step 7: Construct the environment for the deep reinforcement learning algorithm using the random scenario obtained in Step 6, and fine-tune the learning efficiency parameters of the deep reinforcement learning algorithm DQN and the discretization accuracy of the reservoir water level. Determine the parameters of the DQN algorithm for the wind-solar-cascade reservoir system with the goal of maximizing system power generation and minimizing residual load. Finally, use the parameter-tuned algorithm to solve the optimal strategy for short-term scheduling of the wind-solar-cascade reservoir system.
2. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 1, the short-term scheduling is divided into 24 stages, with one hour as one stage to solve the scheduling strategy.
3. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 2, the process of obtaining the joint probability density function and joint probability distribution of the cascade reservoirs for each time period based on the Pearson Type III distribution; First, the Archimedean Copula family of Copula functions, widely used in hydrology, was selected. Then, by comparing the correlation characteristics of the main functions in the Archimedean Copula family with the correlation characteristics of inflow runoff from cascade reservoirs, the Clayton Copula function, with the most similar correlation characteristics, was ultimately chosen to describe the spatial correlation of the cascade reservoirs. The Pearson-III type probability distribution function of each reservoir was used as a marginal distribution to solve for the runoff state. Joint probability distribution of downstream reservoirs As shown in equation (1): (1); In the formula: Cascade reservoir runoff status Joint probability distribution of downstream reservoirs; for t Time period n Inflow of water from the first-level reservoir The Pearson-III type probability distribution function; for t The join parameters of the Clayton Copula function for the time period can be estimated using the Kendall correlation coefficient; Furthermore, the joint probability density function can be obtained by inverse transformation, i.e., by taking the partial derivatives separately.
4. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 3, the method for solving the Markov transition probability matrix of the cascade reservoir includes: Correlation analysis shows that the historical inflow data of the cascade reservoirs follows Markov characteristics. The specific steps for solving the Markov probability transition matrix of wind power output are as follows: S1: Obtained by solving the Clayton Copula function t Time and t Joint distribution function at time +1 , ; S2: Obtain the joint probability distribution of adjacent time periods of the cascade reservoirs using the Clayton Copula function. To describe time correlation; S3: Discretize the inflow runoff of each reservoir at each stage into the same number of states with the maximum and minimum inflow runoff as the range; S4: Based on the conditional distribution formula, calculate the Markov transition probability of the cascade reservoirs from runoff state in adjacent time periods. As shown in equation (2): (2); In the formula: The Markov transition probability of the runoff state of the cascade reservoirs in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, random variable inflow into the reservoir t The value of the time period The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Random variable inflow into the reservoir during a time period The value is And random variable inflow into the reservoir t The value of the time period The probability of; For random variables, inflow runoff t The value of the time period The probability of; S5: Solve for each discrete value of runoff to obtain the state transition probability matrix for this stage. ; S6: Repeat the above steps to obtain the state transition probability matrix for each stage, which is used to describe the Markov process of the inflow runoff of the cascade reservoir in adjacent time periods; In step 3, the method for generating random scenarios of inflow into cascade reservoirs is as follows: The specific steps for generating random scenes using the MH algorithm in Markov Monte Carlo sampling are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =0, respectively from the transition matrix With uniform distribution Mid-sampling, obtained t+ Phase 1 sampled values and random values Accept sampled values The criterion is shown in equation (3): (3); In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0; Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e. t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage; thus obtaining z The samples that conform to the probability density function and state transition probability of each stage are used as the inflow scenarios of cascade reservoirs. Each scenario includes the inflow runoff value of each reservoir at each time period and the corresponding Markov state transition probability.
5. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 4, the method for solving the probability density function of wind power output includes: The choice between using a multidimensional or one-dimensional kernel density estimation function to fit the probability density distribution depends on the number of wind farms. If there is only one wind farm, a one-dimensional kernel density estimation function is used; if there are two or more, a multidimensional kernel density estimation function is used. Here, we will use a one-dimensional kernel density estimation function for explanation. Let... It is a probability density function For an unknown population of independent and identically distributed random variables, the kernel density estimate is defined as shown in equation (4): (4); In the formula: x It is a random variable; n For sample size; For the first i One sample; h For bandwidth parameters; For a kernel function, it must satisfy the following conditions: , , ; As can be seen from the above formula, when the sample is known, kernel density estimation requires determining two important components: the bandwidth parameter. h and kernel function Since kernel density estimation averages the kernel function, the choice of bandwidth parameter has a much greater impact on the estimation result than the kernel function itself. A Gaussian kernel function is used, as shown in equation (5). (5); The optimal bandwidth parameter is solved using the least squares cross-validation test (LSCV) method, and the bandwidth parameter estimated by the multidimensional kernel function can be solved using the Gaussian kernel. The least squares cross-validation test (LSCV) method is a calculation method based on the minimum integral squared error (ISE) criterion. The ISE expression is shown in equation (6): (6); In the formula: It is a random variable; For kernel density estimation, It is a probability density distribution; The last term in the above three terms is independent of bandwidth. Therefore, the first two terms are used for least squares cross-validation test to minimize the bandwidth parameter. By substituting the one-dimensional kernel density estimation function, the one-dimensional least squares cross-validation test is shown in equation (7): (7); In the formula: For the first i One sample; For the first j One sample; h For bandwidth parameters; For kernel functions; This represents the kernel function convolving with itself; n is the sample size. In step 4, the method for solving the Markov state transition probability matrix of wind power output includes: Through correlation analysis, the random variable of wind power output follows Markov characteristics. The specific steps to solve the Markov probability transition matrix of wind power output are as follows: S1: Solve for the joint density function between the wind power output states of two adjacent time periods; that is, integrate the wind power output probability density functions at time t and time t+1 to obtain the probability distribution function. , As a marginal probability, the Copula function, which is most closely related to wind power output, is chosen to solve the joint probability function. ; S2: Discretize the wind power output into several states based on the maximum and minimum wind power output; S3: Based on the conditional distribution formula, the wind power output in adjacent time periods is obtained from the output state. Transition to the next state Markov transition probability As shown in equation (8): (8); In the formula: The Markov state transition probability of wind power output state in adjacent time periods; for t+ 1-period random variable inflow into the reservoir The value is hour, t Time-of-use random variable: wind power output The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Wind power output as a random variable over a time period The value is ,and t Time-of-use random variable: wind power output The value is The probability of; Wind power output is a random variable t The value of the time period The probability of; The state transition probability matrix is obtained based on the random runoff state transition probabilities of each adjacent time period. , used to describe Markov processes in adjacent periods of inflow into cascade reservoirs; In step 4, the method for generating random wind power output scenarios is as follows: The stochastic model is sampled using the MH algorithm in the Markov Monte Carlo sampling method. The specific steps are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =1, respectively from the transition matrix With uniform distribution Mid-sampling, obtaining sampled values and random values Accept sampled values The criterion is shown in equation (9): (9); In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0; Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e. t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage, thereby obtaining... z Samples that conform to the probability density function and state transition probability of each stage are used as wind power output scenarios; Each scenario includes wind power output for each time period and the corresponding Markov state transition probability.
6. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 5, the method for solving the photovoltaic power output probability density function includes: The choice between using a multidimensional or one-dimensional kernel density estimation function to fit the probability density distribution is determined based on the number of photovoltaic power plants. If there is only one photovoltaic power plant, a one-dimensional kernel density estimation function is used; if there are two or more, a multidimensional kernel density estimation function is used. Here, we will use a one-dimensional kernel density estimation function for explanation. Let... It is a probability density function For an unknown population of independent and identically distributed random variables, the kernel density estimate is defined as shown in equation (10): (10); In the formula: x It is a random variable; n For sample size; For the first i One sample; h For bandwidth parameters; For a kernel function, it must satisfy the following conditions: , , ; As can be seen from the above formula, when the sample is known, kernel density estimation requires determining two important components: the bandwidth parameter. h and kernel function Since the kernel density averages the kernel function, the choice of bandwidth parameter has a much greater impact on the estimation result than the kernel function. Here, the Gaussian kernel function is chosen, as shown in equation (11): (11); The optimal bandwidth parameter is solved using the least squares cross-validation test (LSCV) method, and the bandwidth parameter estimated by the multidimensional kernel function can be solved using the Gaussian kernel. The least squares cross-validation test (LSCV) method is a calculation method based on the minimum integral squared error (ISE) criterion. The ISE expression is shown in equation (12): (12); In the formula: It is a random variable; For kernel density estimation, It is a probability density distribution; The last term in the above three terms is independent of bandwidth. Therefore, the first two terms are used for least squares cross-validation test to minimize the bandwidth parameter. By substituting the one-dimensional kernel density estimation function, the one-dimensional least squares cross-validation test is shown in equation (13): (13); In the formula: For the first i One sample; For the first j One sample; h For bandwidth parameters; For kernel functions; This represents the kernel function convolving with itself; n is the sample size. In step 5, the method for solving the photovoltaic output Markov state transition probability matrix includes: Through correlation analysis, the random variable of photovoltaic power output follows Markov characteristics. The specific steps to solve the Markov state probability transition matrix of photovoltaic power output are as follows: S1: Solve for the joint density function between the photovoltaic power output states of two adjacent time periods; that is, integrate the photovoltaic power output probability density functions at time t and time t+1 to obtain the probability distribution function. , As the marginal probability, the Copula function, which is most closely related to photovoltaic output, is chosen to solve the joint probability function. ; S2: Discretize the photovoltaic output into several states based on the maximum and minimum photovoltaic output; S3: Based on the conditional distribution formula, the photovoltaic power output in adjacent time periods is obtained from the power output state. Transition to the next state Markov transition probability As shown in equation (14): (14); In the formula: The Markov state transition probability of photovoltaic power output states in adjacent time periods; For random variables, inflow runoff t+ The value for time period 1 At that time, the random variable photovoltaic output t Time period The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ Photovoltaic power output as a random variable over a time period of 1 The value is And random variable photovoltaic output t The value of the time period The probability of; Photovoltaic output as a random variable t The value of the time period The probability of; The state transition probability matrix is obtained based on the random runoff state transition probabilities of each adjacent time period. , used to describe Markov processes in adjacent periods of inflow into cascade reservoirs; In step 5, the method for generating random photovoltaic output scenarios includes: The stochastic model is sampled using the MH algorithm in the Markov Monte Carlo sampling method. The specific steps are as follows: S1: Input the state transition probability matrix for each stage Initial state x 0 and the probability density distribution of each stage, and make the Markov stationary distribution of each stage. Equals the probability density distribution of each stage; S2: Let the number of transitions be... n =0, number of stages t =1, respectively from the transition matrix With uniform distribution Mid-sampling, obtaining sampled values and random values Accept sampled values The criterion is shown in equation (15): (15); In the formula: To meet Random values from a distribution; Based on the transition matrix The next stage of sampling values; for t+ Phase 1 sampled values The stable distribution; for t Phase State The stable distribution below; for t Time-based random variable inflow into the reservoir The value is hour, t+ 1-time random variable runoff The value of The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; for t+ 1-time random variable runoff The value is hour, t Time-based random variable inflow into the reservoir The value is The probability of, where , They are respectively t+ 1 time period t A collection of inflow runoff that is a random variable over a specific time period; If accepted, a state transition occurs. n = n+ 1, and let x 1= Otherwise, refuse sampling and let n = n+ 1, x 1= x 0; Continue the above steps until... n=N The Markov chain converges to a stationary distribution, meaning that the samples at this point all follow the target distribution; S3: Markov chains constructed by repeating all stages of S2 converge to a stationary distribution. S4: In the first stage, i.e. t =1 Continue sampling z This initial set is used as the initial state set for the runoff Markov process, and the state transition probability matrix for each stage is based on this initial state. Transfer and sampling are performed until the final stage; thus obtaining z Samples that conform to the probability density function and state transition probability of each stage are used as photovoltaic power output scenarios; Each scenario includes the photovoltaic power output for each time period and the corresponding Markov state transition probability.
7. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 6, the method for scene reduction and construction of scenes for each stage of the wind-solar cascade reservoir system based on the Bisecting k-means clustering algorithm of relevance distance includes: S1: Input random scene time series data; S2: Initialize the number of clusters k Optimal number of clusters in k-means clustering k Generally taken in Between, among N This represents the number of samples in the sample set, which is set in this invention. k It is 10, which is the square root of the sample size of 100; S3: Initialize all data into a single cluster and calculate its SSE value; S4: Select the cluster with the largest SSE and divide it into two clusters based on the k-means algorithm. The steps of clustering using the k-means algorithm are as follows: Divide the data into two groups and randomly select a cluster center for each group; calculate the correlation distance between each object in each group and the cluster center, and assign each object to the cluster center with the closest correlation distance; after each assignment, the cluster center will be recalculated based on the existing objects in the cluster; repeat the calculation and assignment until the termination condition is met; the termination conditions are mainly: no (or minimum number) objects are reassigned, no (or minimum number) cluster centers change, and the sum of squared errors is locally minimized. S5: Recalculate the error after partitioning, select the cluster with the largest SSE, and divide it into two clusters based on the k-means algorithm; S6: Repeat S4 and S5 until the number of clusters is reached. k =10 The algorithm terminates; S7: Output the cluster centers of the clusters after S6 has been divided; The sum of squared errors (SSE) is expressed as shown in equation (16): (16); In the formula: SSE represents the sum of squared errors; i Indicates the first in the cluster i One point; n This represents the total number of points in the cluster; Indicates the weight value; Indicates the first in the cluster i The value of each point; This represents the average value of all points in the cluster. The relevant distance expression is shown in equation (17): (17); In the formula: Represents the relevant distance; For sequence X and Y covariance; , Sequences X and Y covariance; Cluster centers of random scenarios for runoff, wind power output, and photovoltaic output are obtained separately and used as representative scenarios for permutation and combination to construct random scenarios of wind-solar-cascade complementary systems that consider multiple randomnesses.
8. The deep reinforcement learning short-term stochastic optimization scheduling method for wind-solar-cascade reservoirs according to claim 1, characterized in that: In step 7, the wind-solar-cascade reservoir system has a reward function and its constraints that aim to maximize power generation and minimize residual load. The reward function with the objective of maximizing system power generation and minimizing residual load is shown in equation (18): (18); In the formula: express t Power generation benefits of wind-solar cascade complementary systems, taking into account time periods and residual load penalties; Cascade reservoirs t The output during a given time period reflects the production of hydropower energy; The penalty amount for the deviation between the output and load process of the complementary system; For each reservoir t Initial water level for the specified time period; for t The inflow of water into the headwater reservoir and the inflow of water into the downstream reservoirs during the specified time period are random variables. for t Power generation flow during a given time period; express t Wind power output during certain periods; Indicates photovoltaic power output; in, t Time-based cascade reservoir output As shown in equation (19): (19); In the formula: A This represents the overall output coefficient of each reservoir; For each reservoir t Average hydropower head over the period; Complementary system output deviation from load process penalty As shown in equation (20) (20); In the formula: for t Time-of-day load demand; for t Power output of cascade reservoirs during specific time periods; express t Wind power output during certain periods; Indicates photovoltaic power output; This is the penalty coefficient; β The penalty index; Optimal scheduling of complementary systems includes the following equality constraints and inequality constraints: The water balance constraints of the cascade reservoirs are shown in equation (21): (21); In the formula: , Each reservoir t Initial and final storage capacity of the time period; , They are respectively t Inflow and outflow rates of each reservoir during the specified time period; For reservoir t The duration of power generation during a given period; The cascade water level constraint is shown in equation (22): (22); In the formula: , Each reservoir t Minimum and maximum water level limits at the beginning of the time period; The power generation flow constraint is shown in equation (23): (23); In the formula: This represents the maximum flow rate through each reservoir. The outbound flow constraint is shown in equation (24): (24); In the formula: , Each reservoir t Minimum and maximum outbound flow allowed at the beginning of the time period; The maximum climbing output limit is shown in equation (25): (25); In the formula: for t Power output of cascade reservoirs during specific time periods; for t- Output of cascade reservoirs in one time period; This represents the maximum output amplitude in adjacent time periods; The transmission capacity of the tie line is shown in equation (26): (26); In the formula: for t Power output of cascade reservoirs during specific time periods; , Minimum and maximum limits for the power output of cascade hydropower to the power grid; The limitations on wind and solar power output are shown in equations (27, 28): (27); (28); In the formula: express t Wind power output during certain periods; Indicates photovoltaic power output; This is the maximum output limit for wind power. Maximum output limit for photovoltaic power; In step 7, the DQN is applied to the short-term optimization scheduling model of the wind-solar-cascade reservoir complementary system as follows: S1: Initialize the Q-value table and the initial state of the wind-solar-cascade complementary system; S2: Input the state transition matrix of cascade reservoir runoff, wind power output, and photovoltaic power output; S3: Based on the state transition matrix, the agent acquires knowledge samples by exploring and interacting with the environment using policies; S4: Update the neural network parameters based on the knowledge samples; S5: Repeat steps S3 to S4 until the agent reaches the final state to complete one training session; S6: Repeat S5 until the maximum number of training iterations is reached; S7: Output short-term scheduling strategy for wind-solar-cascade complementary system; The main steps of updating neural network parameters in the DQN algorithm are as follows: S1: The intelligent agent interacts with the environment to acquire a large number of knowledge samples, which are stored in the experience pool; S2: Extract knowledge samples from the experience pool through experience replay and input them into the neural network and loss function; S3: The neural network maps corresponding Q values based on knowledge samples; S4: This is accomplished by updating the main neural network parameters and Q-values through gradient descent of the loss function, and at each interval... β The next training iteration copies its parameters to the target neural network; The DQN algorithm trains agents mainly through two key technologies: experience replay and neural networks. In S1 to S2, the agent acquires a large number of knowledge samples by interacting with the environment and stores them as training data in the experience pool. When there are enough samples in the experience pool, a specified number of data are extracted in an unordered and random manner through experience replay for updating the neural network parameters. This increases the utilization rate of samples and breaks the correlation between samples, thus accelerating the convergence of the algorithm. To address the instability issue when using nonlinear functions to represent value functions, the DQN algorithm employs two networks with identical structures but different parameters to map Q values during process S3. Based on different input data, the main neural network is used to evaluate the action value corresponding to the current state of the cascade reservoir system, referred to as the Q estimate; the target neural network is used to evaluate the action value corresponding to the next state, referred to as the Q target value. In the S4 process, the parameters of the main neural network are based on the core idea of temporal difference. In each Q-value update, the gradient of the loss function is calculated to reduce the prediction error of the neural network. The temporal difference error is defined as the loss function as follows (29): (29); In the formula: For state Take action below The action value obtained by the main neural network is the Q estimate. For state Take action below The action value obtained by the target neural network, i.e., the Q target value; These are the network parameters of the main neural network; These are the network parameters of the target neural network; The discount rate is used to control the impact of future earnings on the present. The main neural network parameters and Q-value updates are shown in equation (30): (30); In the formula: This represents the gradient of the loss function; The learning rate is used to determine the extent to which the error is learned. The parameters of the target neural network are determined by intervals. β The training will adjust the parameters of the main neural network. Copy to the target neural network This parameter update method reduces the correlation between the Q estimate and the Q target value, improves algorithm stability, and ultimately enhances the agent's decision-making ability. In step 7, the method for optimizing the DQN algorithm parameters of the wind-solar-cascade reservoir complementary system includes: The DQN algorithm has five learning efficiency parameters, including the learning rate. Greed rate Discount rate These are parameters related to reinforcement learning, and are the same as those in the Q-learning algorithm; the target neural network parameter update interval... β Parameters related to deep learning; water level discretization accuracy h This falls under the category of other parameters. Considering that water level adjustments in actual scheduling are often measured in meters, it is set to 1m. The only parameters that need to be tuned in the Q-learning algorithm are the learning rate, greedy rate, and discount rate. To simplify the tuning process, the learning efficiency parameters of the DQN algorithm are divided into two parts: reinforcement learning-related parameters and deep learning-related parameters. First, the reinforcement learning-related parameters are tuned using the Q-learning algorithm. Then, based on the optimized parameters, the deep learning-related parameters are tuned using the DQN algorithm. That is: S1: By arranging and combining the learning rate and greedy rate in the Q-learning algorithm within a certain range, the parameter combination with the largest cumulative return value is selected as the optimal parameter for reinforcement learning. In addition, since the future returns of this reservoir scheduling model have no impact on the current returns, the discount rate is set to 1. S2: The optimal learning rate and greedy rate were applied to the DQN algorithm, and the neural network parameter update interval was tested and analyzed in the same way.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
Medium-and-long-term hidden random scheduling method for cascade hydropower station of combined wind power photovoltaic power station
CN111476407A