A daily stochastic optimization scheduling method for wind-storage ecological power generation based on reinforcement learning
By optimizing the daily random optimization scheduling of wind-storage ecological power generation with the Q(λ)-learning algorithm based on reinforcement learning and the eligibility trace function, the problems of slow solution speed and easy falling into local optimality in the existing technology are solved, and more efficient scheduling decisions are achieved.
Patent Information
- Application Number
- CN202210318978.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-03-29
AI Technical Summary
The existing daily random optimization scheduling method for wind-storage ecological power generation has problems such as slow solution speed, easy to fall into local optimal solution or too long solution time during the solution process, making it difficult to be effectively applied in large-scale systems.
The Q(λ)-learning algorithm based on reinforcement learning is adopted, combined with the eligibility trace and heuristic function. By constructing the objective function to minimize the expected value of the square of the output deviation of the wind power-pumped storage combined system, the heuristic greedy strategy and eligibility trace function are used to extract action qualifications and optimize the scheduling strategy.
The solution speed and learning efficiency of the daily random optimization scheduling of wind-storage ecological power generation are improved, which avoids invalid exploration, shortens the solution time, and achieves faster scheduling decisions.
Smart Images

Figure CN114881404B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a daily random optimization scheduling method for wind-storage ecological power generation based on reinforcement learning, and belongs to the field of wind-storage optimization scheduling. Background Art
[0002] Vigorously developing a hydropower station dispatching system can meet the needs of hydrological station construction in various river basins, enhance practical capabilities in power generation and shipping, and improve the overall efficiency of hydropower stations. Among them, the stochastic optimization dispatching problem of a combined wind power and pumped storage system is a high-dimensional, multi-stage, nonlinear optimization problem with complex constraints and a wide range of factors to consider. Commonly used algorithms for solving such stochastic optimization dispatching problems include particle swarm optimization (PSO) and stochastic dynamic programming (SDP). However, as the scale of the system increases, these algorithms may encounter limitations during the solution process. For example, the PSO algorithm is prone to falling into local optimal solutions during the optimization process, requiring the algorithm to escape from local optimal solutions during the calculation process, making it difficult to find the theoretically optimal global solution. While the SDP algorithm can find the theoretically optimal solution, it is prone to the "curse of dimensionality" during the solution process, resulting in excessively long solution times and limited practical application.
[0003] Therefore, a daily stochastic optimization scheduling method for wind-storage ecological power generation with faster solution speed is needed. Summary of the Invention
[0004] In order to overcome the problems existing in the prior art, the present invention designs a daily random optimization scheduling method for wind-storage ecological power generation based on reinforcement learning.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A daily random optimization scheduling method for wind-storage ecological power generation based on reinforcement learning, characterized by comprising the following steps:
[0007] Pre-constructing an objective function, wherein the objective function takes minimizing the expected value of the square of the deviation between the actual output and the planned output of the wind power-pumped storage combined system as a goal;
[0008] Obtain the current actual wind power output value and the wind power output forecast value;
[0009] The Q(λ)-learning algorithm is used to iteratively solve the objective function and obtain the scheduling strategy:
[0010] Taking the current reservoir capacity as the initial state value, a heuristic greedy strategy is used to select actions from the reservoir inflow and outflow set. The eligibility trace function is used to extract the action qualifications, and the heuristic function is used to extract the heuristic information of the action. The reward value of executing the current action is calculated and the Q value is updated to obtain the Q value table. The scheduling strategy is determined based on the Q value table.
[0011] Furthermore, the method further includes: correcting the wind power output prediction value using a wind power prediction error probability density function that obeys Beta distribution.
[0012] Furthermore, an objective function is pre-constructed, specifically: within an optimization cycle, with the goal of minimizing the expected value of the square of the deviation between the actual output and the planned output of the wind power-pumped storage combined system, an objective function is constructed and constraints of the objective function are set.
[0013] Furthermore, the state is updated using the state transfer equation, which is:
[0014]
[0015] Where: V t 、V t+1 are the storage capacities of the reservoir at the beginning and end of period t; Q c is the pumping flow rate during period t, m 3 / s;Q fd is the power generation flow in period t, m 3 / s; ΔT is the time for power generation / pumping during period t.
[0016] Furthermore, the reward value is calculated based on the deviation between the actual wind power output value and the predicted value.
[0017] Furthermore, the heuristic greedy strategy is expressed as:
[0018]
[0019] Where: H t (s t ,a t ) is the heuristic function, ξ is the weight of the heuristic function, which is used to weight the influence of the heuristic function on action selection, ξH t (s t ,a t ) is the inspiration information for selecting an action.
[0020] Furthermore, the heuristic function is expressed as:
[0021]
[0022] Where: η is generally taken as about 0.01, the purpose of which is to prevent When the heuristic function H t (s t ,a t ) are all 0; π H (s t ) is a heuristic strategy.
[0023] Compared with the prior art, the present invention has the following characteristics and beneficial effects:
[0024] The present invention solves the wind-storage power generation optimization scheduling problem through the Q-learning (λ) algorithm. It can achieve the established long-term goals in a dynamic and stochastic wind-storage ecological power generation environment through repeated trial-and-error exploration. The eligibility trace and heuristic function are combined to improve the greedy strategy, avoid useless trial-and-error exploration, improve the learning efficiency and convergence speed of the objective function, and thus shorten the solution time. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of the heuristic Q(λ)-learning solution method of the present invention;
[0026] Figure 2 is the backward view of the TD(λ) algorithm;
[0027] Figure 3 This is the structure diagram of the Q-learning (λ) algorithm. DETAILED DESCRIPTION
[0028] The present invention will be described in more detail below with reference to the embodiments.
[0029] Example 1
[0030] In scenarios with large load fluctuations, it is necessary to adjust the output of the wind power-pumped storage combined system according to the grid dispatch requirements, and reasonably arrange its own output according to the daily output plan issued by the grid dispatch to reduce the frequency regulation pressure of other units in the system and ensure the stable operation of the power system. Based on this, a daily random optimization scheduling method for wind power generation based on reinforcement learning is proposed. Figure 1 As shown, the following steps are included:
[0031] Within one optimization cycle (one day), the goal is to minimize the expected value of the square of the deviation between the actual output and the planned output of the wind power-pumped storage combined system. The objective function is constructed as follows:
[0032]
[0033] Where: T is the number of time periods in a cycle; R t is the indicator function of period t; V t is the storage capacity of the upper reservoir of the pumped storage power station at the beginning of period t; P t gd is the power generated by the pumped storage power station during period t, less than 0 for pumping, greater than 0 for power generation; R t 、P t gd The expression is as follows:
[0034]
[0035] Where: After the wind power forecast error distribution function curve in period t is discretized into N values, its corresponding power is The corresponding probability is p t,i ; G t is the planned output of the wind power-pumped storage combined system during period t; when the pumped storage power station unit is in power generation state during period t, The value is 1, otherwise it is 0. t g is the power generation output corresponding to the unit during period t; when the pumped storage power station unit is in the motor state during period t, The value is 1, otherwise it is 0. t d is the pumping power corresponding to the unit during period t.
[0036] The objective function constraints are:
[0037] (1) Reservoir capacity constraints:
[0038] V min ≤V t ≤V max (3)
[0039] Where: V min 、V max They are the minimum and maximum available storage capacities of the reservoir in a pumped storage power station.
[0040] (2) Constraints on storage capacity changes during the first and last hours of each day:
[0041] V 24 -V0=0 (4)
[0042] Assuming that the pumped storage reservoir is regulated daily, the storage capacity of the reservoir is equal at the beginning and end of each day.
[0043] (3) Power generation and pumping constraints:
[0044]
[0045] Where: They are the upper and lower limits of power generation output of pumped storage power station units respectively.
[0046] The pumping power of a single unit of a pumped storage power station is treated as a fixed value:
[0047] P t d =P d k t (6)
[0048] Where: P d is the rated pumping power of a single unit in a pumped storage power station, k tis the total number of pumping units operating in period t.
[0049] (4) Mutually exclusive constraints:
[0050]
[0051] The units of a pumped-storage power station cannot be in the power generation and pumping states at the same time.
[0052] Example 2
[0053] The objective function is solved using the reinforcement learning heuristic Q(λ)-learning algorithm, which includes the following steps:
[0054] In this embodiment, the wind-storage system scheduling cycle is set to one day, and the scheduling cycle is divided into 24 time periods. The scheduling of wind power in each time period in the wind-storage system is uncertain. The storage capacity of the upper reservoir of the pumped storage power station is used as the state variable, and the power generation / pumping power of the pumped storage power station is used as the action variable. The state transition equation is constructed as follows:
[0055]
[0056] Where: V t 、V t+1 are the storage capacities of the reservoir at the beginning and end of period t; Q c is the pumping flow rate during period t, m 3 / s;Q fd is the power generation flow in period t, m 3 / s; ΔT is the time for power generation / pumping during period t.
[0057] According to the Bellman principle, in a certain state, maximizing future rewards is equivalent to maximizing the sum of immediate rewards and the maximum future rewards in the next state. When the wind power output in each period is an independent random variable, the recursive equation can be obtained:
[0058]
[0059] Where: It represents the expected value of the minimum square of the output deviation from the tth period to the 24th period; is the square value of the output deviation in the tth period; It is the expected value of the square of the output deviation in the period (t+1) to 24.
[0060] In the same state, different actions will result in different rewards. The wind-storage combined system aims to minimize the expected value of the square of the deviation between the planned output and the actual output. Therefore, in each state, when the actual output is not equal to the planned output, a penalty fee is required. The reward value at this time is:
[0061]
[0062] When the actual output is equal to the planned output, it means the expected effect is achieved. The reward value at this time is:
[0063] r t =G t (4)
[0064] The heuristic greedy strategy is used to select actions from the reservoir inflow and outflow sets. The heuristic greedy strategy is defined as follows:
[0065]
[0066] Where: H t (s t ,a t ) is the heuristic function, ξ is the weight of the heuristic function, which is used to weight the influence of the heuristic function on action selection, ξH t (s t ,a t ) is the inspiration information for selecting an action.
[0067] Obviously, the heuristic information ξH t (s t ,a t ) is greater than the action value function Q(s t ,a t ), the heuristic information will affect the action selection of the heuristic greedy strategy. And the larger the ξ, the greater the influence of the heuristic information on the heuristic greedy strategy. In order to avoid the heuristic information being too large and completely offsetting the action value function Q(s t ,a t ) on the heuristic greedy strategy, the heuristic function is constructed as follows
[0068]
[0069] Where: η is generally taken as about 0.01, the purpose of which is to prevent When the heuristic function H t (s t ,a t ) are all 0; π H (s t ) is a heuristic strategy, which is defined as the profit value R in this embodiment. t >0 action.
[0070] in, The upper limit of the heuristic function is limited, and η limits the lower limit of the heuristic function, ensuring that the heuristic information is slightly larger than the action value function Q(s t ,a t ).
[0071] As can be seen from Formula 6, after the eligibility trace extracts the eligibility of the action, the eligibility information is brought into the heuristic function and acts together on the greedy strategy.
[0072] Example 3
[0073] 3.1: Set the iteration statement N, state set S, action set A, learning rate α, discount factor γ, exploration rate ε, and decay factor λ; initialize the Q value table, number of iterations, and learning rate.
[0074] 3.2: Based on the wind power forecast value in the current period and the wind power forecast error probability density function that obeys the Beta distribution, calculate the wind power output in the current period.
[0075] 3.2: Assume that the storage capacity of the upper reservoir of the pumped storage power station in the current period is state S t ; Select action A from the outbound traffic set through a heuristic greedy strategy t (i.e., the pumping / power generation flow of the upper reservoir); using the eligibility trace function to extract the eligibility of the current action; using the heuristic function to extract the heuristic information of the current action;
[0076] According to the water balance equation, the water level of the upper reservoir in the next period is obtained as the new state S t+1 ; Calculate the reward r obtained by taking this action t , and update the action value function Q(S t ,A t )=Q(S t ,A t )+αδZ t (S,A).
[0077] 3.5: Calculate the wind power output for the next period using the same method as step 3.2.
[0078] 3.6: According to the deviation between the wind power output in the next period and the predicted output, the qualification information and heuristic information of the action are brought into the heuristic greedy strategy, and the new action A is selected through the heuristic greedy strategy. t+1 .
[0079] 3.7: Update the Q value table Q(S) according to the heuristic Q(λ)-learning algorithm t ,A t )=Q(S t ,A t )+αδZ t (S,A).
[0080] 3.8: Determine whether the algorithm has reached the specified number of iterations. If not, set time period t = t + 1 and return to step 3.2 to continue the iteration.
[0081] According to the Q value matrix after the iteration, the action sequence is selected as the scheduling decision.
[0082] Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
Claims
1. A daily random optimization scheduling method for wind-storage ecological power generation based on reinforcement learning, characterized in that: The following steps are involved: Pre-build the objective function; Obtain the current actual wind power output value and the wind power output forecast value; The Q(λ)-learning algorithm is used to iteratively solve the objective function and obtain the scheduling strategy: The current reservoir capacity is used as the initial state value. A heuristic greedy strategy is used to select actions from the reservoir inflow and outflow sets. The eligibility trace function is used to extract the action qualifications, and the heuristic function is used to extract the heuristic information of the action. The reward value of executing the current action is calculated and the Q value is updated to obtain the Q value table. The scheduling strategy is determined based on the Q value table, where: The heuristic greedy strategy is expressed as: Where: H t (s t ,a t ) is the heuristic function, ξ is the weight of the heuristic function, which is used to weight the influence of the heuristic function on action selection, ξH t (s t ,a t ) is the inspiration information for selecting an action.
2. The method for daily random optimization scheduling of wind-storage ecological power generation based on reinforcement learning according to claim 1 is characterized in that: Also includes: The wind power output forecast value is corrected using the wind power forecast error probability density function that obeys the Beta distribution.
3. The method for daily random optimization scheduling of wind-storage ecological power generation based on reinforcement learning according to claim 1 is characterized in that: The objective function is constructed in advance, specifically: within an optimization cycle, with the goal of minimizing the expected value of the square of the deviation between the actual output and the planned output of the wind power-pumped storage combined system, constructing the objective function and setting the constraint conditions of the objective function.
4. The method for daily random optimization scheduling of wind-storage ecological power generation based on reinforcement learning according to claim 1 is characterized in that: It also includes: using the state transition equation to update the state, which is expressed as: Where: V t 、V t+1 are the storage capacities of the reservoir at the beginning and end of period t; Q c is the pumping flow rate during period t, m 3 / s;Q fd is the power generation flow in period t, m 3 / s; ΔT is the time for power generation / pumping during period t.
5. According to the reinforcement learning-based daily random optimization scheduling method for wind-storage ecological power generation in claim 1, the reward value is calculated according to the deviation between the actual wind power output value and the predicted value.
6. According to the reinforcement learning-based daily random optimization scheduling method for wind storage ecological power generation in claim 1, the heuristic function is expressed as follows: Where: η>0; π H (s t ) is a heuristic strategy.
Citation Information
Patent Citations
Wind power-pumped storage combined system daily random dynamic scheduling method based on SARSA(lambda)algorithm
CN112054561A