Day-ahead and intraday joint dispatching method for regional power grid based on deep reinforcement learning

By introducing intraday rolling plans and deep reinforcement learning algorithms into the regional power grid, the problem of loose coordination in dispatching plans was solved, and efficient utilization of power grid resources and an increase in the new energy consumption rate were achieved.

CN115441437BActive Publication Date: 2025-09-16HEFEI UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211102713.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2025-09-16
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively utilize the multi-time-scale characteristics of elastic resources in regional power grids, resulting in a loose connection between the day-ahead dispatch plan and the AGC link, leading to problems such as wind curtailment or load loss.

Method used

A regional power grid day-ahead and intraday joint dispatching method based on deep reinforcement learning is adopted. By adding an intraday rolling plan between the day-ahead dispatching plan and AGC control, the intraday rolling dispatching model is optimized with the deep reinforcement learning algorithm, and the A2C algorithm is used for solving it, so as to achieve close connection and smooth transition of the dispatching plan.

Benefits of technology

It improves the real-time performance and computational efficiency of dispatching plans, ensures the rational use of flexible resources in the power grid, reduces wind power abandonment or load loss, and increases the absorption rate of new energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115441437B_ABST
    Figure CN115441437B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of power system dispatching optimization, and more specifically, relates to a regional power grid day-ahead-intraday joint dispatching method based on deep reinforcement learning, which establishes a regional power grid intraday rolling dispatching optimization model and proposes a dispatching strategy solution based on deep reinforcement learning. First, the day-ahead dispatching plan is formulated every day according to the day-ahead wind power and load forecast curve; then, an intraday rolling dispatching model is established for the regional power grid: objective function and constraints; finally, the intraday rolling model is solved using a deep reinforcement learning algorithm. This method adds an intraday rolling plan between the day-ahead dispatching plan and AGC control, so that the connection between the dispatching plans is closer and the transition is smoother. Compared with the traditional dispatching optimization method based on mathematical models and optimization solvers, the deep reinforcement learning algorithm is more real-time and greatly improves the solution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of regional power grid dispatch optimization, and more specifically, relates to a regional power grid day-ahead and intraday joint dispatching method based on deep reinforcement learning. Background Art

[0002] Since renewable energy generation is typically an intermittent power source, its output is volatile and uncertain. Traditional dispatch methods alone are unable to meet dispatch requirements, resulting in wind curtailment or load shedding. Therefore, it is necessary to further research a new dispatch method to rationally dispatch various resources in regional power grids and further improve the absorption rate of renewable energy.

[0003] Because errors in day-ahead forecasts of the output and load demand of renewable resources such as wind power are often difficult to avoid, if the next day's unit combination and unit output plan are formulated based solely on day-ahead wind power and load forecast data, a large power imbalance will occur in the AGC process, which is sometimes difficult to eliminate, resulting in wind curtailment or load loss. Generally, the forecast accuracy of renewable energy output and load demand such as wind power is directly related to the time scale. For example, intraday forecast accuracy is generally higher than day-ahead forecast accuracy. In addition, the response speeds of various dispatchable resources such as flexible loads in the power system may vary. The traditional day-ahead dispatch process, which directly connects with the AGC process, cannot fully utilize the multi-timescale characteristics of flexible resources in the regional power grid. However, current research has failed to fully utilize the multi-timescale characteristics of flexible resources in the regional power grid, resulting in a loose connection between dispatch plans and a less smooth transition.

[0004] Currently, the main approaches for solving power dispatch models include traditional solvers and deep reinforcement learning algorithms. Traditional mathematical model-based solvers can yield optimal solutions, but they are computationally inefficient for mixed integer programming problems and sometimes fail to meet real-time requirements. Deep reinforcement learning algorithms offer a new approach to solving such problems. The Advantage Actor-Critic (A2C) algorithm is a faster, simpler, and more robust parallel deep reinforcement learning algorithm that operates in a continuous action space. A2C utilizes synchronous learners for training, using multiple CPU threads on a single machine (each thread is referred to as a learner) for more efficient learning and significantly faster solution speeds than traditional methods. As elastic resources on both the source and load sides of the grid increase in number, deep reinforcement learning methods can better adapt to dispatching needs as the scale of the problem continues to expand. Therefore, research on power dispatch methods based on deep reinforcement learning has important theoretical and practical value. Summary of the Invention

[0005] In response to the problems existing in the prior art, the present invention proposes a regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning. This method adds an intraday rolling plan between the day-ahead scheduling plan and AGC control to make the connection between the scheduling plans closer and the transition smoother.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The regional power grid day-ahead and intraday joint dispatching method based on deep reinforcement learning includes the following steps:

[0008] Step 1: The day-ahead dispatch plan is formulated daily based on the day-ahead wind power and load forecast curves to obtain the start and stop plan of the thermal power units, the output plan of the thermal power units, the compensation price and reduction amount of Class A curtailable load, and the start time of the shiftable load operation;

[0009] Step 2: The objective function of the intraday rolling dispatch model is to minimize the sum of system operating costs and risk costs. The constraints are intraday power balance constraints, line transmission capacity constraints, upper and lower limits of thermal power unit output constraints, thermal power unit ramping constraints, and Class B curtailable load call constraints.

[0010] Step 3: Use deep reinforcement learning to solve the intraday rolling scheduling model and obtain the intraday scheduling plan.

[0011] This technical solution is further optimized by establishing the objective function of the intraday rolling scheduling model in step 2:

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019] Among them, k is the current period, and the wind power output and load demand in the future M*ΔT period are to be predicted; P i,t is the day-ahead output plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; ΔP i,t is the output adjustment of thermal power unit i in period t within a day, which is the decision variable of the model; and are the coal consumption cost, additional coal consumption cost and life loss cost of thermal power unit i after daily output adjustment; The load dispatching cost can be reduced for Class B during period t; is the risk cost of wind curtailment during period t; is the load loss risk cost of thermal power unit i during period t; δ i,t is the day-ahead start-stop plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; a i 、b i and c i is the coal consumption cost coefficient of unit i; is the coal consumption rate coefficient of unit i operating in deep peak regulation state; i is the coal consumption rate coefficient of unit i under the conventional minimum technical output state; z i,t It is used to indicate whether the thermal power unit is in a deep peak regulation state. When the unit is operating below the conventional minimum technical output, the value is 1; when the unit is operating above the conventional minimum technical output, the value is 0; ε i is the coal consumption rate of the thermal power unit at rated output; ρ coal is the unit coal price; N i,t (P i,t +ΔP i,t ) is the number of rotor cracking cycles of unit i, and its value is the same as (P i,t +ΔP i,t ) is closely related; i is the operating loss coefficient of the thermal power unit; is the purchase cost of unit i; ΔT is the length of time period t; ΔP t B It represents the load reduction amount of Class B load that can be reduced during period t; is the compensation price for the curtailable load of Class B during period t; cw is the risk cost coefficient of wind curtailment per unit electricity; N w is the number of wind farms in the regional power grid; is the wind power curtailment of the j-th wind farm under the extreme scenario of wind power output and load demand during period t; cl is the load loss risk cost coefficient per unit electricity; ΔP t cl is the load shedding power of the regional power grid under the extreme scenario of wind power output and load demand in period t.

[0020] This technical solution is further optimized by establishing the constraints of the intraday rolling scheduling model in step 2:

[0021] The constraints mainly include intraday power balance constraints, line flow constraints, upper and lower output limits of thermal power units, thermal power unit ramping constraints, and Class B curtailable load call constraints, as shown in the following formula:

[0022] The intraday power balance constraint is:

[0023]

[0024] Among them, N g is the number of thermal power units in the regional power grid, N w is the number of wind farms in the regional power grid, i and j represent the current thermal power unit i and wind power unit j respectively; and is the intraday ultra-short-term load forecast and wind power forecast, ΔP t B The reserve call amount of Class B load that can be reduced; ΔP t A The load reduction call amount of Class A; ΔP t cl is the load shedding amount during period t; P t sh P is the power consumption of the movable load during period t after scheduling; t sh* The power consumption of the movable load in the period t before dispatch;

[0025] The upper and lower limits of the thermal power unit output are:

[0026]

[0027] P i min ≤P i,t +ΔP i,t ≤P i max

[0028] Among them, P i min and P i max are the maximum and minimum output of thermal power unit i, respectively. For conventional thermal power units, P i min is the conventional minimum technical output. For the deep peak-shaving units after flexibility transformation, P i min The maximum peak-shaving depth after the unit transformation; and are the upward and downward reserve capacity values ​​of the regional power grid in the deep peak-shaving unit i during period t, respectively;

[0029] The thermal power unit climbing constraint:

[0030] -r i down ΔT≤(P i,t +ΔPi,t )-(P i,t-1 +ΔP i,t-1 )≤r i up ΔT

[0031] Among them, r i down and r i up are the downward and upward ramp rates of thermal power unit i, respectively, and ΔT is the time interval from t-1 to t;

[0032] The line power flow constraints are:

[0033]

[0034] Among them, T l,g 、T l,j and T l,b is the power transfer distribution coefficient, is the daily ultra-short-term forecast load value of the regional power grid at node k during period t after scheduling, and F l max is the upper limit of the power flow on line l;

[0035] The Class B load curtailment reserve call constraints are:

[0036] 0≤ΔP t B ≤P t B .

[0037] This technical solution is further optimized, and the step 3 is specifically as follows:

[0038] Based on the intraday rolling scheduling model established in step 2, a Markov decision model is established. The variables in the decision process include:

[0039] 1) State space construction: The state space includes the ultra-short-term load forecast value of the regional power grid, the ultra-short-term wind power forecast value, the unit output at the previous moment, and the day-ahead dispatch plan, namely:

[0040] S={P w ,P l ,P,P day-ahead}

[0041] Among them, P w is the set of ultra-short-term wind power forecast states within the regional power grid; P l is the set of ultra-short-term load power forecast states within the day; P is the set of output states of each thermal power unit at the previous moment; P day-ahead is the state set of the regional power grid's day-ahead dispatch plan;

[0042] 2) Action space structure: including the output adjustment range of thermal power units, the compensation price of Class B curtailable load and the curtailable load range, namely:

[0043] A={ΔP,ρ B ,ΔP B}

[0044] Where ΔP is the set of thermal power generation output adjustment actions within the regional power grid; ρ B is the compensation price action set for Class B curtailable load; ΔP B It is a set of actions for reducing the load reduction amount of Class B;

[0045] 3) Reward function construction: It includes three parts: the operating cost of the regional power grid's daily dispatch plan, the wind power curtailment / load loss penalty, and the safety constraint penalty. Among them, the operating cost of the regional power grid's daily rolling dispatch plan and the wind power curtailment / load loss penalty are the objective function described in claim 5. The safety constraint penalty is the system branch flow over-limit penalty, that is, the flow of the internal branch of the power grid exceeds the limit value it can withstand, which can be expressed as:

[0046]

[0047] in, Penalty for over-limit of current; ρ pf is the penalty coefficient for power flow exceeding the limit; μ l,t is a 0-1 variable, representing whether branch l exceeds the limit at time t, μ l,t =1 means the line current exceeds the limit, μ l,t =0 means the line power flow is within the limit; L is the total number of branches in the regional power grid;

[0048] Therefore, the agent reward function R can be expressed as:

[0049]

[0050] To maximize the reward, the sum of the grid's daily dispatch plan operating costs, wind curtailment / load loss penalties, and safety constraint penalties must be minimized.

[0051] This technical solution is further optimized, and the deep reinforcement learning algorithm in step 3 is an A2C algorithm.

[0052] This technical solution is further optimized, and the A2C algorithm is designed as follows:

[0053] The A2C algorithm consists of two deep networks, namely the Actor network and the Critic network. The Actor network inputs system state information and outputs the probability of action selection in the current state. The Critic network inputs system state information and outputs the value function of the current state. Based on the regional power grid dispatch environment information, the Actor network and the Critic network respectively output the dispatch plan for the next 4 hours and the state-value function of the current state. The dispatch plan is applied to the external environment to obtain the next state and reward, which are used as data for network training. After training is completed, the output of the Actor network is the intraday rolling dispatch plan of the regional power grid.

[0054] This technical solution is further optimized.

[0055] The Actor network needs to be updated based on the feedback from the Critic network, and the Critic network is updated based on the state transitions generated by the interaction between the agent and the environment; the Critic network uses the network parameter θ v Implement the state value function V(s;θ v ) and update the parameters according to the state value function, which can be expressed as:

[0056]

[0057] Where: L(θ v ) is the network loss function, r is the reward at this time, γ is the discount factor, For state s t+1 The value function when For state s t The value function when Critic network parameters when is i;

[0058] The Critic network inputs the system state information and outputs the value function of the current state. For the Actor network, it approximates the action strategy as a function expression, that is, π(s,a)≈π(a|s; θ π ), and further fitting approximation can be obtained as follows:

[0059]

[0060] Where: θ π is the weight parameter of the Actor network; different from the state transition probability P, p(a|s,θ π ) indicates that the network parameters are θ π The probability of taking action a in state s;

[0061] The objective function of the policy π can be expressed as

[0062]

[0063] Where R(a|s) represents the reward for executing action a in state s, Denotes the network parameters as θ π The probability of taking action a in state s, J(θ π ) indicates that the network parameters are θ π strategy at the time, Denotes the network parameters as θ π The expected reward obtained by taking action a in state s;

[0064] According to the gradient descent method, we know

[0065]

[0066]

[0067] Where, is the weight parameter of the Actor network at time t, is the weight parameter of the Actor network at time t+1, and α is the learning rate;

[0068] Furthermore, according to ▽f(x)=f(x)▽logf(x), we can deduce

[0069]

[0070] Using the action value function Q π (s,a) replaces R to get

[0071]

[0072] In order to make the feedback value both greater than zero and less than zero, the state value function V is added π (s) as the baseline value, we can get

[0073]

[0074] Define the advantage function A(s,a) as

[0075]

[0076] According to the above formula, we can get

[0077]

[0078] More generally, it can be expressed as

[0079]

[0080] The Actor network also inputs system state information and outputs the probability of action selection under the current state. Compared with the Critic network, the output layer of the Actor network is divided into a mean layer and a standard deviation layer. The output mean and variance form a normal distribution, and then the output value within the unit climbing constraint is sampled through the normal distribution to obtain the final scheduling action. In this way, continuous action evaluation is achieved, and at the same time, it is guaranteed that the thermal power unit will not exceed the climbing limit or output limit.

[0081] Different from the prior art, the beneficial effects of the present invention are mainly manifested in:

[0082] 1. The present invention adds a daily rolling plan between the day-ahead dispatch plan and AGC control. The traditional two-time-scale (day-ahead + AGC) dispatch mode is not sophisticated enough and lacks an intermediate transition link. The next day's unit combination and unit output plan are formulated only based on the day-ahead wind power and load forecast data. In the AGC link, a large power imbalance will occur, which is sometimes difficult to eliminate, resulting in wind curtailment or load loss. The addition of daily rolling dispatch makes the connection between dispatch plans tighter and the transition smoother.

[0083] 2. The present invention adopts a deep reinforcement learning algorithm to solve the intraday rolling scheduling model. Since the regional power grid dispatching center needs to interact with Class B load aggregators during the intraday rolling scheduling stage, and the intraday rolling scheduling time scale is relatively short, the system has certain real-time requirements for the formulation of scheduling plans. The use of deep reinforcement learning algorithms can improve computing efficiency and is more real-time than traditional scheduling optimization methods based on mathematical models and optimization solvers. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 This is a schematic diagram of the regional power grid architecture;

[0085] Figure 2 This is the flow chart of the day-ahead and intraday rolling scheduling;

[0086] Figure 3 This is a schematic diagram of the Critic network structure;

[0087] Figure 4 This is a diagram of the Actor network structure;

[0088] Figure 5 This is the A2C algorithm training framework diagram. DETAILED DESCRIPTION

[0089] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.

[0090] This paper discloses a method for regional power grid day-ahead and intraday joint scheduling based on deep reinforcement learning. This method incorporates intraday rolling plans between the day-ahead scheduling plan and AGC control, resulting in a tighter connection and smoother transitions between the scheduling plans. This deep reinforcement learning algorithm is more real-time than traditional scheduling optimization methods based on mathematical models and optimization solvers.

[0091] See also Figure 1 Figure 1 shows the architecture of a regional power grid. The regional power grid's power system includes conventional thermal power units, deep peak-shaving units, wind turbines, rigid loads, and flexible loads. Flexible loads include curtailable loads and shiftable loads. Curtailable loads include Class A curtailable loads and Class B curtailable loads. Class A curtailable loads are slow to respond and require long advance notice. The dispatch center plans and issues instructions for Class A curtailable loads before the day is up. Class B curtailable loads are short-term and quick-response loads. The dispatch center plans and issues instructions for Class B curtailable loads within a short period of time during the day.

[0092] See for example Figure 2 As shown in the figure, the process diagram of day-ahead and intraday rolling scheduling includes the following steps:

[0093] Step 1: The day-ahead dispatch plan is formulated daily based on the day-ahead wind power and load forecast curves to obtain the start and stop plan of the thermal power units, the output plan of the thermal power units, the compensation price and reduction amount of Class A curtailable load, and the start time of the shiftable load operation;

[0094] Step 2: Establish an intraday rolling dispatch model: objective function and constraints. The objective function is to minimize the sum of system operating costs and risk costs. The constraints are intraday power balance constraints, line flow constraints, thermal power unit output upper and lower limits constraints, thermal power unit ramping constraints, and Class B curtailable load call constraints:

[0095] Step 2.1: Establish the objective function of the intraday rolling scheduling model:

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103] Among them, k is the current period, and the wind power output and load demand in the future M*ΔT period are to be predicted; P i,t is the day-ahead output plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; ΔP i,t is the output adjustment of thermal power unit i in period t within a day, which is the decision variable of the model; and are the coal consumption cost, additional coal consumption cost and life loss cost of thermal power unit i after daily output adjustment; The load dispatching cost can be reduced for Class B during period t; is the risk cost of wind curtailment during period t; is the load loss risk cost of thermal power unit i during period t; δ i,t is the day-ahead start-stop plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; a i 、b i and c i is the coal consumption cost coefficient of unit i; is the coal consumption rate coefficient of unit i operating in deep peak regulation state; i is the coal consumption rate coefficient of unit i under the conventional minimum technical output state; z i,t It is used to indicate whether the thermal power unit is in a deep peak regulation state. When the unit is operating below the conventional minimum technical output, the value is 1; when the unit is operating above the conventional minimum technical output, the value is 0; ε i is the coal consumption rate of the thermal power unit at rated output; ρ coal is the unit coal price. N i,t (P i,t +ΔP i,t ) is the number of rotor cracking cycles of unit i, and its value is the same as (P i,t +ΔP i,t ) is closely related; i is the operating loss coefficient of the thermal power unit; is the purchase cost of unit i; ΔT is the length of time period t; ΔP t B It represents the load reduction amount of Class B load that can be reduced during period t; is the compensation price for the curtailable load of Class B during period t; cw is the risk cost coefficient of wind curtailment per unit electricity; N w is the number of wind farms in the regional power grid; is the wind power curtailment of the j-th wind farm under the extreme scenario of wind power output and load demand during period t; cl is the load loss risk cost coefficient per unit electricity; ΔP t clis the load shedding power of the regional power grid under the extreme scenario of wind power output and load demand during period t;

[0104] Step 2.2: Establish constraints for the intraday rolling scheduling model:

[0105] The constraints mainly include intraday power balance constraints, line transmission capacity constraints, upper and lower output limits of thermal power units, thermal power unit ramping constraints, and Class B curtailable load call constraints, as shown in the following formula:

[0106] The intraday power balance constraint is:

[0107]

[0108] Among them, N g is the number of thermal power units in the regional power grid, N w is the number of wind farms in the regional power grid, i and j represent the current thermal power unit i and wind power unit j respectively; P t loadl and It is the intraday ultra-short-term load forecast and wind power forecast. ΔP t B The reserve call amount of Class B load that can be reduced; ΔP t A The load reduction call amount of Class A; ΔP t cl is the load shedding amount during period t; P t sh P is the power consumption of the movable load during period t after scheduling; t sh* The power consumption of the movable load in the period t before dispatch;

[0109] The upper and lower limits of the thermal power unit output are:

[0110]

[0111] P i min ≤P i,t +ΔP i,t ≤P i max

[0112] Among them, P i min and P i max are the maximum and minimum output of thermal power unit i, respectively. For conventional thermal power units, P i min is the conventional minimum technical output. For the deep peak-shaving units after flexibility transformation, P i minThe maximum peak-shaving depth after the unit transformation; and are the upward and downward reserve capacity values ​​of the regional power grid in the deep peak-shaving unit i during period t.

[0113] The thermal power unit climbing constraint:

[0114] -r i down ΔT≤(P i,t +ΔP i,t )-(P i,t-1 +ΔP i,t-1 )≤r i up ΔT

[0115] Among them, r i down and r i up are the downward and upward climbing rates of thermal power unit i, respectively, and ΔT is the time interval from t-1 to t.

[0116] The line power flow constraints are:

[0117]

[0118] Among them, T l,g 、T l,j and T l,b is the power transfer distribution coefficient, is the daily ultra-short-term forecast load value of the regional power grid at node k during period t after scheduling, and F l max is the upper limit of the power flow on line 1.

[0119] The Class B load curtailment reserve call constraints are:

[0120] 0≤ΔP t B ≤P t B

[0121] Step 3: Use deep reinforcement learning to solve the intraday scheduling model:

[0122] Based on the intraday rolling scheduling model established in step 2, a Markov decision model is established. The variables in the decision process include:

[0123] 1) State space construction: The state space includes the ultra-short-term load forecast value of the regional power grid, the ultra-short-term wind power forecast value, the unit output at the previous moment, and the day-ahead dispatch plan, namely:

[0124] S={P w ,Pl ,P,P day-ahead}

[0125] Among them, P w is the set of ultra-short-term wind power forecast states within the regional power grid; P l is the set of ultra-short-term load power forecast states within the day; P is the set of output states of each thermal power unit at the previous moment; P day-ahead It is the set of regional power grid day-ahead dispatch plan states.

[0126] 2) Action space structure: including the output adjustment range of thermal power units, the compensation price of Class B curtailable load and the curtailable load range, namely:

[0127] A={ΔP,ρ B ,ΔP B}

[0128] Where ΔP is the set of thermal power generation output adjustment actions within the regional power grid; ρ B is the compensation price action set for Class B curtailable load; ΔP B It is a set of Class B load reduction actions.

[0129] 3) Reward function construction: This includes the cost of the regional grid's daily dispatch plan, the penalty for wind curtailment / load loss, and the safety constraint penalty. The cost of the regional grid's daily rolling dispatch plan and the penalty for wind curtailment / load loss are the objective functions of the daily rolling dispatch model established in step 2.1. The safety constraint penalty is the penalty for exceeding the limit of the system branch flow, i.e., the flow of the internal branch of the grid exceeds its limit value, which can be expressed as:

[0130]

[0131] in, Penalty for over-limit of current; ρ pf is the power flow over-limit penalty coefficient; μ l,t is a 0-1 variable, representing whether branch l exceeds the limit at time t, μ l,t =1 means the line current exceeds the limit, μ l,t =0 means the line flow is within the limit; L is the total number of branches in the regional power grid.

[0132] Therefore, the agent reward function R can be expressed as:

[0133]

[0134] To maximize the reward, the sum of the grid's daily dispatch plan operating costs, wind curtailment / load loss penalties, and safety constraint penalties must be minimized.

[0135] A2C algorithm design:

[0136] See for example Figure 3 Figure 1 shows the structure of a critic network. The critic network takes in the system state information and outputs the value function of the current state. The value function of the current state is obtained through the input layer, hidden layer, and output layer.

[0137] The A2C algorithm consists of two deep networks: an actor network and a critic network. The actor network inputs system state information and outputs the probability of action selection in the current state. The critic network inputs system state information and outputs a value function for the current state. Based on the regional power grid's dispatch environment, the actor and critic networks output a dispatch plan for the next four hours and a state-value function for the current state, respectively. The dispatch plan is applied to the external environment to determine the next state and reward, which serve as data for network training. After training is complete, the actor network outputs the daily rolling dispatch plan for the regional power grid.

[0138] The Actor network needs to be updated based on the feedback from the Critic network, and the Critic network is updated based on the state transitions generated by the interaction between the agent and the environment. The Critic network uses the network parameters θ v Implement the state value function V(s;θ v ) and update the parameters according to the state value function, which can be expressed as:

[0139]

[0140] Where: L(θ v ) is the network loss function, r is the reward at this time, γ is the discount factor, For state s t+1 The value function when For state s t The value function when Critic network parameters when i is i.

[0141] The Critic network inputs the system state information and outputs the value function of the current state. For the Actor network, it approximates the action strategy as a function expression, that is, π(s,a)≈π(a|s;θ π ), and further fitting approximation can be obtained as follows.

[0142]

[0143] Where: θ π is the weight parameter of the Actor network; different from the state transition probability P, p(a|s,θ π ) indicates that the network parameters are θ πThe probability of taking action a in state s.

[0144] The objective function of the policy π can be expressed as

[0145]

[0146] Where R(a|s) represents the reward for executing action a in state s, Denotes the network parameters as θ π The probability of taking action a in state s, J(θ π ) indicates that the network parameters are θ π strategy at the time, Denotes the network parameters as θ π The expected reward obtained by taking action a in state s.

[0147] According to the gradient descent method, we know

[0148]

[0149]

[0150] Where, is the weight parameter of the Actor network at time t, is the weight parameter of the Actor network at time t+1, and α is the learning rate.

[0151] Further, according to Can be launched

[0152]

[0153] Using the action value function Q π (s,a) replaces R to get

[0154]

[0155] In order to make the feedback value both greater than zero and less than zero, the state value function V is added π (s) as the baseline value, we can get

[0156]

[0157] Define the advantage function A(s,a) as

[0158]

[0159] According to the above formula, we can get

[0160]

[0161] More generally, it can be expressed as

[0162]

[0163] The Actor network also inputs system state information and outputs the probability of action selection under the current state. Compared to the Critic network, the Actor network's output layer is divided into a mean layer and a standard deviation layer. The output mean and variance form a normal distribution, which is then used to sample the output values ​​within the unit's ramping constraints to obtain the final scheduling action. This method ensures continuous action selection and prevents thermal power units from exceeding ramping limits or output limits.

[0164] Scheduling optimization framework of the A2C algorithm:

[0165] Based on the regional power grid's dispatch environment, the actor network and critic network output a dispatch plan for the next four hours and a state-value function for the current state, respectively. The dispatch plan is applied to the external environment to obtain the next state and reward, which serve as data for network training. After training is complete, the actor network's output becomes the regional power grid's daily rolling dispatch plan.

[0166] This invention incorporates a daily rolling schedule between the day-ahead scheduling plan and AGC control, making the connection between the scheduling plans tighter and the transition smoother. Compared to traditional scheduling optimization methods based on mathematical models and optimization solvers, the deep reinforcement learning algorithm is more real-time and significantly improves solution efficiency.

[0167] See for example Figure 4 Figure 2 shows the structure of an Actor network. The Actor network inputs system state information and outputs the probability of action selection under the current state. Compared to the Critic network, the Actor network's output layer consists of a mean layer and a standard deviation layer. The mean and variance of the output form a normal distribution, which is then used to sample the output values ​​within the unit ramping constraints to obtain the final scheduling action. This method ensures continuous action selection and prevents thermal power units from exceeding ramping limits or output limits.

[0168] Since the input information for both the actor and critic networks is the dispatching environment information of the regional power grid, their respective input and hidden layers extract features from this information. Therefore, this paper merges the input and hidden layers of the actor and critic networks, meaning that the actor and critic networks share the same input and hidden layers.

[0169] See for example Figure 5Figure 2 shows the A2C algorithm training framework. Based on the regional power grid's dispatch environment, the actor network and critic network output a dispatch plan for the next four hours and a state-value function for the current state, respectively. The dispatch plan is applied to the external environment to obtain the next state and reward, which serve as data for network training. After training is complete, the actor network's output becomes the regional power grid's daily rolling dispatch plan.

[0170] It should be noted that, in this document, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "include," "comprise," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, elements defined by the phrase "include..." or "comprising..." do not exclude the presence of additional elements in the process, method, article, or terminal device comprising the elements. Furthermore, in this document, "greater than," "less than," "exceeding," etc., are understood to exclude the number itself; "above," "below," "within," etc., are understood to include the number itself.

[0171] Although the above embodiments have been described, those skilled in the art may make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the above descriptions are merely embodiments of the present invention and do not limit the scope of patent protection of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of the present invention.

Claims

1. A regional power grid day-ahead and intraday joint dispatching method based on deep reinforcement learning, characterized by: The following steps are involved: Step 1: The day-ahead dispatch plan is formulated daily based on the day-ahead wind power and load forecast curves to obtain the start and stop plan of the thermal power units, the output plan of the thermal power units, the compensation price and reduction amount of Class A curtailable load, and the start time of the shiftable load operation; Step 2: The objective function of the intraday rolling dispatch model is to minimize the sum of system operating costs and risk costs. The constraints are intraday power balance constraints, line transmission capacity constraints, upper and lower limits of thermal power unit output constraints, thermal power unit ramping constraints, and Class B curtailable load call constraints. Step 3: Use deep reinforcement learning to solve the intraday rolling scheduling model and obtain the intraday scheduling plan. The step 3 is specifically as follows: Based on the intraday rolling scheduling model established in step 2, a Markov decision model is established. The variables in the decision process include: 1) State space construction: The state space includes the ultra-short-term load forecast value of the regional power grid, the ultra-short-term wind power forecast value, the unit output at the previous moment, and the day-ahead dispatch plan, namely: S={P w ,P l ,P,P day-ahead } Among them, P w is the set of ultra-short-term wind power forecast states within the regional power grid; P l is the set of ultra-short-term load power forecast states within the day; P is the set of output states of each thermal power unit at the previous moment; P day-ahead is the state set of the regional power grid's day-ahead dispatch plan; 2) Action space structure: including the output adjustment range of thermal power units, the compensation price of Class B curtailable load and the curtailable load range, namely: A={ΔP,ρ B ,ΔP B } Where ΔP is the set of thermal power generation output adjustment actions within the regional power grid; ρ B is the compensation price action set for Class B curtailable load; ΔP B It is a set of actions for reducing the load of type B; 3) Reward function construction: This includes the cost of the regional power grid's daily dispatch plan, the wind curtailment / load loss penalty, and the safety constraint penalty. The cost of the regional power grid's daily rolling dispatch plan and the wind curtailment / load loss penalty are the objective function. The safety constraint penalty is the system branch flow exceeding the limit penalty, i.e., the flow of the internal branch of the power grid exceeds the limit value it can withstand. It can be expressed as: in, Penalty for over-limit of current; ρ pf is the power flow over-limit penalty coefficient; μ l,t is a 0-1 variable, representing whether branch l exceeds the limit at time t, μ l,t =1 means the line current exceeds the limit, μ l,t =0 means the line power flow is within the limit; L is the total number of branches in the regional power grid; Therefore, the agent reward function R can be expressed as: To maximize the reward, the sum of the grid's daily dispatch plan operating costs, wind curtailment / load loss penalties, and safety constraint penalties must be minimized.

2. The regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning according to claim 1, characterized in that: In step 2, the objective function of the intraday rolling scheduling model is established: Among them, k is the current period, and the wind power output and load demand in the future M*ΔT period are to be predicted; P i,t is the day-ahead output plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; ΔP i,t is the output adjustment of thermal power unit i in period t within a day, which is the decision variable of the model; and are the coal consumption cost, additional coal consumption cost and life loss cost of thermal power unit i after daily output adjustment; The load dispatching cost can be reduced for Class B during period t; is the risk cost of wind curtailment during period t; is the load loss risk cost of thermal power unit i during period t; δ i,t is the day-ahead start-stop plan of thermal power unit i, which is a known quantity in the intraday rolling scheduling model; a i 、b i and c i is the coal consumption cost coefficient of unit i; is the coal consumption rate coefficient of unit i operating in deep peak regulation state; i is the coal consumption rate coefficient of unit i under the conventional minimum technical output state; z i,t It is used to indicate whether the thermal power unit is in a deep peak regulation state. When the unit is operating below the conventional minimum technical output, the value is 1; when the unit is operating above the conventional minimum technical output, the value is 0; ε i is the coal consumption rate of the thermal power unit at rated output; ρ coal is the unit coal price; N i,t (P i,t +ΔP i,t ) is the number of rotor cracking cycles of unit i, and its value is the same as (P i,t +ΔP i,t ) is closely related; i is the operating loss coefficient of the thermal power unit; is the purchase cost of unit i; ΔT is the length of time period t; ΔP t B It represents the load reduction amount of Class B load that can be reduced during period t; is the compensation price for the curtailable load of Class B during period t; cw is the risk cost coefficient of wind curtailment per unit electricity; N w is the number of wind farms in the regional power grid; is the wind power curtailment of the j-th wind farm under the extreme scenario of wind power output and load demand during period t; cl is the load loss risk cost coefficient per unit electricity; ΔP t cl is the load shedding power of the regional power grid under the extreme scenario of wind power output and load demand in period t.

3. The regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning according to claim 2 is characterized in that: In step 2, the constraints of the intraday rolling scheduling model are established: The constraints mainly include intraday power balance constraints, line flow constraints, upper and lower output limits of thermal power units, thermal power unit ramping constraints, and Class B curtailable load call constraints, as shown in the following formula: The intraday power balance constraint is: Among them, N g is the number of thermal power units in the regional power grid, N w is the number of wind farms in the regional power grid, i and j represent the current thermal power unit i and wind power unit j respectively; P t loadl and is the intraday ultra-short-term load forecast and wind power forecast, ΔP t B The reserve call amount of Class B load that can be reduced; ΔP t A The load reduction call amount of Class A; ΔP t cl is the load shedding amount during period t; P t sh P is the power consumption of the movable load during period t after scheduling; t sh* The power consumption of the movable load in the period t before dispatch; The upper and lower limits of the thermal power unit output are: P i min ≤P i,t +ΔP i,t ≤P i max Among them, P i min and P i max are the maximum and minimum output of thermal power unit i, respectively. For conventional thermal power units, P i min is the conventional minimum technical output. For the deep peak-shaving units after flexibility transformation, P i min The maximum peak-shaving depth after the unit transformation; and are the upward and downward reserve capacity values ​​of the regional power grid in the deep peak-shaving unit i during period t, respectively; The thermal power unit climbing constraint: -r i down ΔT≤(P i,t +ΔP i,t )-(P i,t-1 +ΔP i,t-1 )≤r i up ΔT Among them, r i down and r i up are the downward and upward ramp rates of thermal power unit i, respectively, and ΔT is the time interval from t-1 to t; The line power flow constraints are: Among them, T l,g 、T l,j and T l,b is the power transfer distribution coefficient, is the daily ultra-short-term forecast load value of the regional power grid at node k during period t after scheduling, and F l max is the upper limit of the power flow on line l; The Class B load curtailment reserve call constraints are: 0≤ΔP t B ≤P t B 。 4. The regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning according to claim 1, characterized in that: The deep reinforcement learning algorithm in step 3 is the A2C algorithm.

5. The regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning according to claim 4 is characterized in that: The A2C algorithm design: The A2C algorithm consists of two deep networks: the Actor network and the Critic network. The Actor network takes in system state information and outputs the probability of action selection in the current state. The Critic network takes in system state information and outputs the value function of the current state. Based on the regional power grid dispatch environment, the Actor network and the Critic network output a dispatch plan for the next four hours and a state-value function of the current state, respectively. The dispatch plan is applied to the external environment to obtain the next state and reward, which are used as data for network training. After training is completed, the output of the Actor network is the daily rolling dispatch plan of the regional power grid.

6. The regional power grid day-ahead and intraday joint scheduling method based on deep reinforcement learning according to claim 5, characterized in that: The Actor network needs to be updated based on the feedback from the Critic network, and the Critic network is updated based on the state transitions generated by the interaction between the agent and the environment; the Critic network uses the network parameter θ v Implement the state value function V(s;θ v ) and update the parameters according to the state value function, which can be expressed as: Where: L(θ v ) is the network loss function, r is the reward at this time, γ is the discount factor, For state s t+1 The value function when For state s t The value function when Critic network parameters when is i; The Critic network inputs the system state information and outputs the value function of the current state. For the Actor network, it approximates the action strategy as a function expression, that is, π(s,a)≈π(a|s; θ π ), and further fitting approximation can be obtained as follows: Where: θ π is the weight parameter of the Actor network; different from the state transition probability P, p(a|s,θ π ) indicates that the network parameters are θ π The probability of taking action a in state s; The objective function of the policy π can be expressed as Where R(a|s) represents the reward for executing action a in state s, Denotes the network parameters as θ π The probability of taking action a in state s, J(θ π ) indicates that the network parameters are θ π strategy at the time, Denotes the network parameters as θ π The expected reward obtained by taking action a in state s; According to the gradient descent method, we know Where, is the weight parameter of the Actor network at time t, is the weight parameter of the Actor network at time t+1, and α is the learning rate; Further, according to Can be launched Using the action value function Q π (s,a) replaces R to get In order to make the feedback value both greater than zero and less than zero, the state value function V is added π (s) as the baseline value, we can get Define the advantage function A(s,a) as According to the above formula, we can get More generally, it can be expressed as The Actor network also inputs system state information and outputs the probability of action selection under the current state. Compared with the Critic network, the output layer of the Actor network is divided into a mean layer and a standard deviation layer. The output mean and variance form a normal distribution, and then the output value within the unit climbing constraint is sampled through the normal distribution to obtain the final scheduling action. In this way, continuous action evaluation is achieved, and at the same time, it is guaranteed that the thermal power unit will not exceed the climbing limit or output limit.

Citation Information

Patent Citations

  • Grid multi-time-scale dispatching method

    CN109962499A

  • Electric power system day-ahead-intra-day cooperative scheduling method and system considering uncertainty of new energy and load intervals

    CN113193547A