A risk scheduling method considering N-1 safety constraints based on proximal policy optimization algorithm

Through the near-end strategy optimization algorithm and Markov reward process, a power system risk scheduling model considering N-1 safety constraints was established, which solved the shortcomings of traditional methods in simulating N-1 fault state and voltage risk, and achieved efficient and economical risk scheduling.

CN114142530BInactive Publication Date: 2025-05-13CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111116822.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional risk scheduling methods are difficult to accurately simulate line power in N-1 fault state, lack consideration of voltage risk, and have many decision variables, huge constraints, complex models, and low solution efficiency.

Method used

A near-end strategy optimization algorithm is used to establish a power system risk scheduling model that considers N-1 safety constraints, define the Markov reward process, and train it based on the reinforcement learning framework to obtain the power system risk scheduling optimization model.

Benefits of technology

It realizes rapid online application, avoids dimensional disasters, takes into account the consideration of voltage and line risks, and ensures the economic, safety and feasibility of the scheduling plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114142530B_ABST
    Figure CN114142530B_ABST
Patent Text Reader

Abstract

The present invention discloses a risk scheduling method considering N-1 safety constraints based on a proximal strategy optimization algorithm, and the steps are: 1) establishing a risk scheduling model that fully considers unit operation constraints, flow constraints and N-1 safety constraints; 2) defining the Markov reward process of the model under the reinforcement learning framework, determining the corresponding state space, action space and reward function, etc.; 3) based on the self-exploration learning mechanism of the proximal strategy optimization algorithm, through interaction with the power system environment, training a risk scheduling network that can ensure the safety and economy of power system operation under random changes in source and load and random grid failures. The present invention can be effectively applied in power systems containing new energy, and can quickly obtain an economical and safe optimal unit scheduling plan in any scenario to ensure the economy and safety of power system operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of risk dispatching of power systems containing a high proportion of renewable energy, and in particular to a risk dispatching method based on a proximal strategy optimization algorithm and taking into account N-1 safety constraints. Background Art

[0002] In the context of my country's carbon neutrality, vigorously developing renewable energy has become a national strategy and social consensus. Wind power and photovoltaics stand out among many renewable energy sources with their wide distribution, mature technology and low cost. In the new power system with wind and solar as the main body, the random volatility of wind and solar has brought huge challenges to the safe and stable operation of the power system.

[0003] Therefore, the risk scheduling problem of power systems with a high proportion of renewable energy has become a research hotspot in academia. However, the traditional probability model that uses a priori probability model to describe the uncertainty of wind and solar power is difficult to truly capture the uncertainty of source and load, and the traditional risk scheduling method based on DC power flow is difficult to accurately simulate the line power under the N-1 fault state, and lacks consideration of voltage risks. In addition, the traditional risk scheduling method has many decision variables, huge constraints, complex models, low solution efficiency and difficulty in solving the problem. Summary of the invention

[0004] The object of the present invention is to provide a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm, comprising the following steps:

[0005] 1) Establish a power system risk dispatch model considering N-1 security constraints.

[0006] The objective function of the power system risk dispatch model considering N-1 security constraints is as follows:

[0007]

[0008] Where T is the scheduling time period.

[0009] The unit power generation cost F at time t G,t , network loss cost F loss,t and risk penalty cost F risk,t They are as follows:

[0010]

[0011]

[0012]

[0013] Where N is the number of units. G,t,n is the actual output of n units at time t.n 、b n 、c n is the relevant cost coefficient of the nth unit. b is the number of network nodes. i,t , Q i,t and V i,t are the active power, reactive power and node voltage of the i-th node at the t-th moment. ij is the resistance between node i and node j. C sell N is the price of electricity sold by the power grid. c and N l are the number of expected accidents and the total number of power grid branches. risk is the risk penalty factor.

[0014] Among them, the line overload level Lsev of the lth line after the expected accident k occurs l,k As shown below:

[0015]

[0016] Where s1 and s2 represent the severity of overload when the line flow reaches the short-term emergency value LTE and the long-term emergency value STE, respectively. r is the percentage of flow on branch l under the expected fault k;

[0017] The severity of the voltage amplitude exceeding the limit at node i under the expected accident k is Vsev i,k As shown below:

[0018]

[0019] In the formula, K HV , A HV , B HV K is the coefficient of the high voltage severity function. LV , A LV , B LV is the coefficient of the low voltage severity function. V i is the voltage amplitude at node i.

[0020] The constraint conditions of the power system risk dispatch model considering N-1 safety constraints include power flow equation constraints, unit output constraints, ramp constraints, node voltage and line power constraints in base state and N-1 state.

[0021] The power flow equality constraints are as follows:

[0022]

[0023] Where P G,i,t , P wt,i,t , P pv,i,t and P D,i,tare the active outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. G,i,t , Q wt,i,t , Q pv,i,t and Q D,i,t are the reactive power outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. ij , B ij are the conductance and reactance between nodes i and j respectively. ij,t is the phase angle between node i and node j at the tth moment. V j,t is the voltage at node j.

[0024] The unit output constraints are as follows:

[0025]

[0026] In the formula, and are the lower and upper limits of the output of the nth unit respectively. G,n is the output of the nth unit.

[0027] The ramp constraints of the unit are as follows:

[0028]

[0029] Where U G,n and D G,n are the maximum upward and downward climbing powers of conventional unit n respectively. G,n,t , P G,n,t-1 They are the output of the nth unit at the tth moment and the t-1th moment respectively.

[0030] The node voltage constraints are as follows:

[0031]

[0032] Where V i,min and V i,max are the lower and upper limits of the voltage at node i, respectively. i 0 and V i k are the voltage amplitudes of node i in the base state and under the kth anticipated accident, respectively.

[0033] The line power constraints are as follows:

[0034]

[0035] Where PL l and PL l,maxare respectively the active power and the maximum transmission power of line l, and are the active power of line l in the base state and the kth expected accident respectively. Nl is the total number of lines. G ii is the self-conductance of node i. ij is the phase angle between nodes i and j;

[0036] 2) Determine the Markov reward process in the reinforcement learning framework.

[0037] The Markov reward process includes a four-element array {S, A, γ, R}. S, A, γ, and R represent the state space, action space, discount factor, and reward function, respectively.

[0038] The state space S is shown below:

[0039] S={S G ,S wt ,S pv ,S load} (12)

[0040] In the formula, S G is the output set of N conventional units, S wt and S pv Represent the predicted output of wind power and photovoltaic power, S load Output aggregation for power grid load forecasting.

[0041] The action space A is as follows:

[0042] A={A G,1 ,A G,2 ,…,A G,N} (13)

[0043] In the formula, A G,i Represents the output increment set of unit i.

[0044] The reward function R is as follows:

[0045]

[0046] In the formula, γ is the discount factor.

[0047] Among them, the immediate reward r t As shown below:

[0048] r t =-(F G,t +F loss,t +F risk,t ) (15)

[0049] In the formula, F G,t 、Floss,t and F risk,t They are the unit power generation cost, network loss cost and risk penalty cost at time t respectively.

[0050] The output increment set A of the unit i G,i It includes the output increment of unit i at different times. The output increment is used to calculate the dispatching output P of the computer group at the next moment. G,i,t+1 .

[0051] The dispatch output of the unit at the next moment P G,i,t+1 As shown below:

[0052] P G,i,t+1 =P G,i,t +ΔP G,i,t (16)

[0053] Where P G,i,t is the active output of the unit at the ith node at the tth moment. ΔP G,i,t The active output increment of the unit at the i-th node at the t-th moment.

[0054] 3) Based on the historical power data of the power system and the Markov reward process, the power system risk dispatch model is trained to obtain the power system risk dispatch optimization model.

[0055] The steps of training the power system risk dispatch model include:

[0056] 3.1) Randomly initialize the current strategy network parameters θ Q and the value network parameter θ π , determine the training period K, scheduling period T and target network update frequency C, and set k=0.

[0057] 3.2) Initialize the initial state s of the power system from the state space t , and let t=0.

[0058] 3.3) Based on state s in the current value network t Output action a t .

[0059] 3.4) Execute action a in the power system risk dispatch model t , and get the next state s t+1 , reward r t , let t=t+1, go to step 3).

[0060] 3.5) The learned time series samples {A T ,S T ,R T ,S T+1} is stored in the sample experience pool as the data set for training the network.

[0061] 3.6) Randomly collect m time series samples from the experience pool {A T ,S T ,R T ,S T+1}, and calculate the loss function L(θ Q ).

[0062] L(θ Q )=E(y t -Q(s t ,a t |θ Q )) 2 (17)

[0063] In the formula, Q(s t ,a t |θ Q ) is the Q value output by the current network at time t. E represents the mean function.

[0064] y t is the target Q value y t As shown below:

[0065] y t =r t +γQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) (18)

[0066] In the formula, r t is the instant reward at time t extracted from the experience pool. t+1 |θ π′ ) is the target policy network with parameters θ π′ The input state variable s t+1 The action variable output when Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the target network with parameters θ Q′ Next input state s t+1 and action variable π′(s t+1 |θ π′ ) under the input Q value.

[0067] 7) Update the gradient of the value network according to the following formula:

[0068]

[0069] In the formula, θ Qis the current value network parameter, μ Q is the learning rate of the current value network, is the loss function L(θ Q ) for the parameter θ Q The gradient of π θ (a t |s t ) and π θold (a t |s t ) are two different strategies of the strategy network, old and new. t is the advantage function. ε is the boundary value of the loss function. E t is the mean function.

[0070] 3.8) Update the parameters θ of the policy network π' ,Right now:

[0071]

[0072] In the formula, θ π is the current policy network parameter, μ π is the learning rate of the policy network, is the Q value Q(s,a|θ Q ) with respect to the gradient of action a. is the policy gradient.

[0073] 3.9) Determine whether k%C=1. If so, update the strategy network parameter θ based on formula (21) Q′ and the value network parameter θ π′ , and go to step 10), otherwise, go directly to step 3.10).

[0074]

[0075] In the formula, θ Q′ and θ π′ are the parameters of the target value network and the target policy network respectively, and τ is the soft update coefficient.

[0076] 3.10) Determine whether k>K holds true. If so, output the power system risk dispatch optimization model. Otherwise, set k=k+1 and return to step 3.2).

[0077] It is worth noting that the present invention establishes a risk scheduling model that fully considers the unit operation constraints, power flow constraints and N-1 safety constraints, defines the Markov reward process of the model under the reinforcement learning framework, determines the corresponding state space, action space and reward function, etc., and based on the self-exploration learning mechanism of the proximal strategy optimization algorithm, through interaction with the power system environment, trains a risk scheduling network that can ensure the safety and economy of power system operation under random changes in sources and loads and random grid failures.

[0078] The technical effect of the present invention is unquestionable. The present invention proposes a power system risk scheduling method based on a proximal strategy optimization algorithm and taking into account N-1 safety constraints. Compared with the traditional risk scheduling method, it avoids the dimensional disaster caused by too many decision variables, while taking into account the consideration of voltage risk and line risk, and can achieve rapid online application. The present invention can be widely used in power system risk scheduling in different regions and different scenarios. It can quickly obtain the scheduling plan of the power grid based on new energy and load input, and meet the N-1 safety constraints to ensure the economy, safety and feasibility of the scheduling plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is a schematic diagram of the method flow chart;

[0080] Figure 2 are the cost scheduling curves under two different schemes;

[0081] Figure 3 is the line power diagram in the base state;

[0082] Figure 4 is the node voltage diagram in the base state;

[0083] Figure 5 It is the line power diagram under N-1 fault state;

[0084] Figure 6 It is the node voltage diagram under N-1 fault state. DETAILED DESCRIPTION

[0085] The present invention is further described below in conjunction with the embodiments, but it should not be understood that the above subject matter of the present invention is limited to the following embodiments. Without departing from the above technical ideas of the present invention, various substitutions and changes are made according to the common technical knowledge and customary means in the art, which should all be included in the protection scope of the present invention.

[0086] Embodiment 1:

[0087] See also Figures 1 to 6 , a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm, comprising the following steps:

[0088] 1) Establish a power system risk dispatch model considering N-1 security constraints.

[0089] The objective function of the power system risk dispatch model considering N-1 security constraints is as follows:

[0090]

[0091] Where T is the scheduling time period.

[0092] The unit power generation cost F at time t G,t , network loss cost F loss,t and risk penalty cost F risk,t They are as follows:

[0093]

[0094]

[0095]

[0096] Where N is the number of units. G,t,n is the actual output of n units at time t. n 、b n 、c n is the relevant cost coefficient of the nth unit. b is the number of network nodes. i,t , Q i,t and V i,t are the active power, reactive power and node voltage of the i-th node at the t-th moment. ij is the resistance between node i and node j. C sell N is the price of electricity sold by the power grid. c and N l are the number of expected accidents and the total number of power grid branches. risk is the risk penalty factor.

[0097] Among them, the line overload level Lsev of the lth line after the expected accident k occurs l,k As shown below:

[0098]

[0099] Where s1 and s2 represent the severity of overload when the line flow reaches the short-term emergency value LTE and the long-term emergency value STE, respectively. r is the percentage of flow on branch l under the expected fault k;

[0100] The severity of the voltage amplitude exceeding the limit at node i under the expected accident k is Vsev i,k As shown below:

[0101]

[0102] In the formula, K HV , A HV , B HV K is the coefficient of the high voltage severity function. LV , A LV , B LV is the coefficient of the low voltage severity function. V i is the voltage amplitude at node i.

[0103] The constraint conditions of the power system risk dispatch model considering N-1 safety constraints include power flow equation constraints, unit output constraints, ramp constraints, node voltage and line power constraints in base state and N-1 state.

[0104] The power flow equality constraints are as follows:

[0105]

[0106] Where P G,i,t , P wt,i,t , P pv,i,t and P D,i,t are the active outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. G,i,t , Q wt,i,t , Q pv,i,t and Q D,i,t are the reactive power outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. ij , B ij are the conductance and reactance between nodes i and j respectively. ij,t is the phase angle between node i and node j at the tth time.

[0107] The unit output constraints are as follows:

[0108]

[0109] In the formula, and are the lower and upper limits of the output of the nth unit respectively. G,n is the output of the nth unit.

[0110] The ramp constraints of the unit are as follows:

[0111]

[0112] Where U G,n and D G,nare the maximum upward and downward climbing powers of conventional unit n respectively. G,n,t , P G,n,t-1 They are the output of the nth unit at the tth moment and the t-1th moment respectively.

[0113] The node voltage constraints are as follows:

[0114]

[0115] Where V i,min and V i,max are the lower and upper limits of the voltage at node i, respectively. i 0 and V i k are the voltage amplitudes of node i in the base state and under the kth anticipated accident, respectively.

[0116] The line power constraints are as follows:

[0117]

[0118] Where PL l and PL l,max are respectively the active power and the maximum transmission power of line l, and are the active power of line l in the base state and the kth expected accident respectively. Nl is the total number of lines. G ii is the self-conductance of node i.

[0119] 2) Determine the Markov reward process in the reinforcement learning framework.

[0120] The Markov reward process includes a four-element array {S, A, γ, R}. S, A, γ, and R represent the state space, action space, discount factor, and reward function, respectively.

[0121] The state space S is shown below:

[0122] S={S G ,S wt ,S pv ,S load} (12)

[0123] In the formula, S G is the output set of N conventional units, S wt and S pv Represent the predicted output of wind power and photovoltaic power, S load Output aggregation for power grid load forecasting.

[0124] The action space A is as follows:

[0125] A={AG,1 ,A G,2 ,…,A G,N} (13)

[0126] In the formula, A G,i Represents the output increment set of unit i.

[0127] The reward function R is as follows:

[0128]

[0129] In the formula, γ is the discount factor.

[0130] Among them, the immediate reward r t As shown below:

[0131] r t =-(F G,t +F loss,t +F risk,t ) (15)

[0132] In the formula, F G,t 、F loss,t and F risk,t They are the unit power generation cost, network loss cost and risk penalty cost at time t respectively.

[0133] The output increment set A of the unit i G,i It includes the output increment of unit i at different times. The output increment is used to calculate the dispatching output P of the computer group at the next moment. G,i,t+1 .

[0134] The dispatch output of the unit at the next moment P G,i,t+1 As shown below:

[0135] P G,i,t+1 =P G,i,t +ΔP G,i,t (16)

[0136] Where P G,i,t is the active output of the unit at the ith node at the tth moment. ΔP G,i,t The active output increment of the unit at the i-th node at the t-th moment.

[0137] 3) Based on the historical power data of the power system and the Markov reward process, the power system risk dispatch model is trained to obtain the power system risk dispatch optimization model.

[0138] The steps of training the power system risk dispatch model include:

[0139] 3.1) Randomly initialize the current strategy network parameters θ Q and the value network parameter θπ , determine the training period K, scheduling period T and target network update frequency C, and set k=0.

[0140] 3.2) Initialize the initial state s of the power system from the state space t , and let t=0.

[0141] 3.3) Based on state s in the current value network t Output action a t .

[0142] 3.4) Execute action a in the power system risk dispatch model t , and get the next state s t+1 , reward r t , let t=t+1, go to step 3).

[0143] 3.5) The learned time series samples {A T ,S T ,R T ,S T+1} is stored in the sample experience pool as the data set for training the network.

[0144] 3.6) Randomly collect m time series samples from the experience pool {A T ,S T ,R T ,S T+1}, and calculate the loss function L(θ Q ).

[0145] L(θ Q )=E(y t -Q(s t ,a t |θ Q )) 2 (17)

[0146] In the formula, Q(s t ,a t |θ Q ) is the Q value output by the current network at time t. The Q value is used to evaluate the quality of executing action A in state S, which is the cumulative reward function r in the scheduling cycle t The sum of . E represents the mean function.

[0147] y t is the target Q value y t As shown below:

[0148] y t =r t +γQ′(s t+1 ,π′(s t+1 |θπ′ )|θ Q′ ) (18)

[0149] In the formula, r t is the instant reward at time t extracted from the experience pool. t+1 |θ π′ ) is the target policy network with parameters θ π′ The input state variable s t+1 The action variable output when Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the target network with parameters θ Q′ Next input state s t+1 and action variable π′(s t+1 |θ π′ ) under the input Q value.

[0150] 7) Update the gradient of the value network according to the following formula:

[0151]

[0152] In the formula, θ Q is the current value network parameter, μ Q is the learning rate of the current value network, is the loss function L(θ Q ) for the parameter θ Q The gradient of π θ (a t |s t ) and π θold (a t |s t ) are two different strategies of the strategy network, old and new. t is the advantage function. ε is the boundary value of the loss function. E t is the mean function.

[0153] 3.8) Update the parameters θ of the policy network π' ,Right now:

[0154]

[0155] In the formula, θ π is the current policy network parameter, μ π is the learning rate of the policy network, is the Q value Q(s,a|θ Q ) with respect to the gradient of action a. is the policy gradient.

[0156] 3.9) Determine whether k%C=1. If so, update the strategy network parameter θ based on formula (21) Q′ and the value network parameter θ π′ , and go to step 10), otherwise, go directly to step 3.10).

[0157]

[0158] In the formula, θ Q′ and θ π′ are the parameters of the target value network and the target policy network respectively, and τ is the soft update coefficient.

[0159] 3.10) Determine whether k>K holds true. If so, output the power system risk dispatch optimization model. Otherwise, set k=k+1 and return to step 3.2).

[0160] Embodiment 2:

[0161] A risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm includes the following steps:

[0162] 1) Establish a power system risk dispatch model considering N-1 security constraints;

[0163] The objective function of the risk scheduling model is as follows:

[0164]

[0165] Where T is the scheduling time period, F G,t 、F loss,t and F risk,t They are the unit power generation cost, network loss cost and risk penalty cost at time t, and their calculation methods are as follows:

[0166]

[0167]

[0168]

[0169] Where N is the number of units, P G,t,n is the actual output of n units at time t, a n , b n , c n is the relevant cost coefficient of the nth unit. b is the number of network nodes, P i,t , Q i,t and V i,t are the active power, reactive power and node voltage of the i-th node at the t-th moment. ijis the resistance between node i and node j, C sell is the price of electricity sold by the power grid. c and N l are the number of expected accidents and the total number of power grid branches. l,k is the line overload level (severity) of the lth line after the expected accident k occurs, and its calculation formula is as follows:

[0170]

[0171] Where s1 and s2 represent the severity of overload when the line flow reaches short-term emergency (LTE, generally 1.1) and long-term emergency (STE, generally 1.2), respectively. Their values ​​can be determined based on the maximum time that the line can operate continuously under LTE and STE flows. r is the flow percentage on branch l under the expected fault k, and Vsev is the severity of voltage amplitude over-limit at node i under the expected accident k. i,k The calculation formula is as follows:

[0172]

[0173] In the formula, K HV , A HV , B HV K is the coefficient of the high voltage severity function; LV , A LV , B LV are the coefficients of the low voltage severity function, all of which are constants.

[0174] The power system risk dispatch model considering N-1 safety constraints considers power flow constraints, unit output constraints, ramp constraints, node voltage and line power constraints in base state and N-1 state;

[0175] The power flow equality constraints are as follows:

[0176]

[0177] Where P G,i,t , P wt,i,t , P pv,i,t and P D,i,t are the active output of thermal power unit, wind power, photovoltaic power and load of the i-th node at the t-th moment, Q G,i,t , Q wt,i,t , Q pv,i,t and Q D,i,t are the reactive power outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. ij , B ij and δ ijare the conductance, reactance and phase angle between node i and node j respectively. V j,t is the voltage at node j.

[0178] The unit output constraints are as follows:

[0179]

[0180] In the formula, and are the lower and upper limits of the output of the nth unit respectively.

[0181] The ramp constraints of the unit are as follows:

[0182]

[0183] Where U G,n and D G,n are the maximum upward and downward climbing powers of conventional unit n respectively.

[0184] The node voltage constraints are as follows:

[0185]

[0186] Where V i,min and V i,max are the lower and upper limits of the voltage at node i, respectively. i 0 and V i k are the voltage amplitudes of node i in the base state and under the kth anticipated accident, respectively.

[0187] The line power constraints are as follows:

[0188]

[0189] Where PL l and PL l,max are respectively the active power and the maximum transmission power of line l, and are the active powers of line l in the base state and under the kth anticipated accident respectively.

[0190] 2) Convert the risk scheduling model into a Markov reward process under the reinforcement learning framework;

[0191] The Markov reward process consists of a four-element array {S, A, γ, R} consisting of the state space S, the action space A, the discount factor γ and the reward function R.

[0192] In the power system risk dispatch model of this paper, the initial output value of conventional units, wind power, photovoltaic power and load output are selected as the state variables of the model, and the state space S of the model is obtained as follows:

[0193] S={S G ,S wt ,S pv ,S load} (12)

[0194] In the formula, S G is the output set of N conventional units, S wt and S pv Represent the predicted output of wind power and photovoltaic power, S load is the power grid load forecast output set. Therefore, in each scheduling period t, the corresponding state variable s can be uniquely determined from the state space S t =[s G,t ,s wt,t ,s pv,t ,s load,t ].

[0195] In the power system risk dispatch model of this embodiment, the incremental value of the unit output is set as the decision variable, so the action space A is as follows:

[0196] A={A G,1 ,A G,2 ,…,A G,N} (13)

[0197] In the formula, A G,i Represents the output increment set of conventional unit i.

[0198] In order to obtain the output increment ΔP of each unit G,n,t Afterwards, the dispatching output of the unit at the next moment can be obtained according to the following formula.

[0199] P G,i,t+1 =P G,i,t +ΔP G,i,t (14)

[0200] Where P G,i,t , is the active power output of the thermal power unit at the i-th node at the t-th moment.

[0201] Set the negative value of the objective function as the reward function, that is, the lower the cost, the greater the reward, thereby encouraging the agent to learn the optimal scheduling plan. Therefore, the immediate reward r t The calculation formula is as follows:

[0202] r t =-(F G,t +Floss,t +F risk,t ) (15)

[0203] In the formula, r t Indicates that the agent is in a certain state s t Next select action a t After that, you can get an immediate reward. For the entire scheduling period T, there is a cumulative reward function R as follows:

[0204]

[0205] Where R represents the cumulative reward obtained by the agent after obtaining the corresponding scheduling plan based on the external state variables of the system, and γ is called the discount factor.

[0206] 3) Based on the IEEE 30-node network and wind and solar data set, the risk scheduling model is trained using the proximal strategy optimization algorithm. The steps are as follows:

[0207] 3.1) Randomly initialize the current policy network and value network parameters θ Q and θ π , determine the training period K, scheduling period T and target network update frequency C, and set k = 0;

[0208] 3.2) Initialize the initial state s of the joint system from the state space t , and let t = 0;

[0209] 3.3) In the current value network based on s t Output action a t ;

[0210] 3.4) Execute action a in the risk scheduling model t , and get the next state s t+1 , reward r t , let t = t + 1, and go to step 3.3;

[0211] 3.5) The learned time series samples {A T ,S T ,R T ,S T+1} is stored in the sample experience pool as a data set for training the network;

[0212] 3.6) Randomly collect m time series samples from the experience pool {A T ,S T ,R T ,S T+1}, calculate the loss function of the value network through the following formula;

[0213] L(θ Q)=E(y t -Q(s t ,a t |θ Q )) 2 (17)

[0214] In the formula, Q(s t ,a t |θ Q ) is the Q value output by the current network at time t, y t is the target Q value, and its calculation formula is as follows:

[0215] y t =r t +γQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) (18)

[0216] In the formula, r t is the instantaneous reward at time t extracted from the experience pool, π′(s t+1 |θ π′ ) is the target policy network with parameters θ π′ The input state variable s t+1 The action variable output when Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the target network with parameters θ Q′ Next input state s t+1 and action variable π′(s t+1 |θ π′ ) under the input Q value.

[0217] 3.7) Update the gradient of the value network according to the following formula:

[0218]

[0219] In the formula, θ Q is the current value network parameter, μ Q is the learning rate of the current value network, is the loss function L(θ Q ) for the parameter θ Q The gradient of π θ (a t |s t ) and π θold (a t |s t ) are two different strategies of the strategy network, old and new. t is the advantage function, ε is the boundary value of the loss function, Et is the mean function.

[0220] 3.8) Then update the parameters of the policy network according to the following formula;

[0221]

[0222] In the formula, θ π is the current policy network parameter, μ π is the learning rate of the policy network.

[0223] is the policy gradient.

[0224] 3.9) If k%C=1, based on the soft update strategy, update the target network parameters according to the following formula

[0225]

[0226] In the formula, θ Q′ and θ π′ are the parameters of the target value network and the target policy network respectively, and τ is the soft update coefficient.

[0227] 3.10) Determine whether k>K holds. If so, output the training network. Otherwise, set k=k+1 and go to step 3.2.

[0228] Embodiment 3:

[0229] See also Figure 1-5 The purpose of the present invention is to provide a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm, comprising the following steps:

[0230] 1) Establish a power system risk dispatch model considering N-1 security constraints;

[0231] The objective function of the risk scheduling model is as follows:

[0232]

[0233] In the formula, the scheduling time period T = 96, F G,t 、F loss,t and F risk,t They are the unit power generation cost, network loss cost and risk penalty cost at time t, and their calculation methods are as follows:

[0234]

[0235]

[0236]

[0237] In the formula, the number of units N = 6, P G,t,n is the actual output of n units at time t, a n , b n , c n is the relevant cost coefficient of the nth unit, which is shown in the following table:

[0238] Table 1 Unit parameters

[0239]

[0240] Number of network nodes N b =30, P i,t , Q i,t and V i,t are the active power, reactive power and node voltage of the i-th node at the t-th moment. ij is the resistance between node i and node j, and the grid electricity price C sell is 0.5 yuan / kWh. The expected number of accidents N c =41, and the total number of power grid branches N l =41. Lsev l,k is the line overload level (severity) of the lth line after the expected accident k occurs, and its calculation formula is as follows:

[0241]

[0242] Where s1 and s2 represent the severity of overload when the line flow reaches short-term emergency (LTE, generally 1.1) and long-term emergency (STE, generally 1.2), respectively. Their values ​​can be determined based on the maximum time that the line can operate continuously under LTE and STE flows. r is the flow percentage on branch l under the expected fault k, and Vsev is the severity of voltage amplitude over-limit at node i under the expected accident k. i,k The calculation formula is as follows:

[0243]

[0244] In the formula, K HV , A HV , B HV K is the coefficient of the high voltage severity function; LV , A LV , B LV are the coefficients of the low voltage severity function, all of which are constants.

[0245] The power system risk dispatch model considering N-1 safety constraints considers power flow constraints, unit output constraints, ramp constraints, node voltage and line power constraints in base state and N-1 state;

[0246] The power flow equality constraints are as follows:

[0247]

[0248] Where P G,i,t , P wt,i,t , P pv,i,t and P D,i,t are the active output of thermal power unit, wind power, photovoltaic power and load of the i-th node at the t-th moment, Q G,i,t , Q wt,i,t , Q pv,i,t and Q D,i,t are the reactive power outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment. ij , B ij and δ ij are the conductance, reactance and phase angle between node i and node j respectively.

[0249] The unit output constraints are as follows:

[0250]

[0251] In the formula, and are the lower and upper limits of the output of the nth unit, respectively, and their values ​​are shown in Table 1.

[0252] The ramp constraints of the unit are as follows:

[0253]

[0254] Where U G,n and D G,n are the maximum upward and downward climbing powers of conventional unit n, respectively, and their values ​​are shown in Table 1.

[0255] The node voltage constraints are as follows:

[0256]

[0257] Where V i,min and V i,max are the lower and upper limits of the voltage at node i, which are 0.95 and 1.05 respectively. i 0 and V i k are the voltage amplitudes of node i in the base state and under the kth anticipated accident, respectively.

[0258] The line power constraints are as follows:

[0259]

[0260] Where PL l and PL l,max are respectively the active power and the maximum transmission power of line l, and are the active powers of line l in the base state and under the kth anticipated accident respectively.

[0261] 2) Convert the risk scheduling model into a Markov reward process under the reinforcement learning framework;

[0262] The Markov reward process consists of a four-element array {S, A, γ, R} consisting of the state space S, the action space A, the discount factor γ and the reward function R.

[0263] In the power system risk dispatch model of this paper, the initial output value of conventional units, wind power, photovoltaic power and load output are selected as the state variables of the model, and the state space S of the model is obtained as follows:

[0264] S={S G ,S wt ,S pv ,S load} (12)

[0265] In the formula, S G is the output set of N conventional units, S wt and S pv Represent the predicted output of wind power and photovoltaic power, S load is the power grid load forecast output set. Therefore, in each scheduling period t, the corresponding state variable s can be uniquely determined from the state space S t =[s G,t ,s wt,t ,s pv,t ,s load,t ].

[0266] In the power system risk dispatch model of this paper, the incremental value of the unit output is set as the decision variable, so the action space A is as follows:

[0267] A={A G,1 ,A G,2 ,…,A G,N} (13)

[0268] In the formula, A G,i Represents the output increment set of conventional unit i.

[0269] In order to obtain the output increment ΔP of each unit G,n,t Afterwards, the dispatching output of the unit at the next moment can be obtained according to the following formula.

[0270] P G,i,t+1 =PG,i,t +ΔP G,i,t (14)

[0271] Where P G,i,t , is the active power output of the thermal power unit at the i-th node at the t-th moment.

[0272] Set the negative value of the objective function as the reward function, that is, the lower the cost, the greater the reward, thereby encouraging the agent to learn the optimal scheduling plan. Therefore, the immediate reward r t The calculation formula is as follows:

[0273] r t =-(F G,t +F loss,t +F risk,t ) (15)

[0274] In the formula, r t Indicates that the agent is in a certain state s t Next select action a t After that, you can get an immediate reward. For the entire scheduling period T, there is a cumulative reward function R as follows:

[0275]

[0276] Where R represents the cumulative reward obtained by the agent after obtaining the corresponding scheduling plan based on the external state variables of the system, and γ is called the discount factor.

[0277] 3) Based on the IEEE 30-node network and wind and solar data set, the risk scheduling model is trained using the proximal strategy optimization algorithm. The steps are as follows:

[0278] 3.1) Randomly initialize the current policy network and value network parameters θ Q and θ π , determine the training period K = 100000, the scheduling period T = 96 and the target network update frequency C = 100, and set k = 0;

[0279] 3.2) Initialize the initial state s of the joint system from the state space t , and let t = 0;

[0280] 3.3) In the current value network based on s t Output action a t ;

[0281] 3.4) Execute action a in the risk scheduling model t , and get the next state s t+1 , reward r t , let t = t + 1, and go to step 3.3;

[0282] 3.5) The learned time series samples {A T ,S T ,R T ,S T+1} is stored in the sample experience pool as a data set for training the network;

[0283] 3.6) Randomly collect 128 time series samples from the experience pool {A T ,S T ,R T ,S T+1}, calculate the loss function of the value network through the following formula;

[0284] L(θ Q )=E(y t -Q(s t ,a t |θ Q )) 2 (17)

[0285] In the formula, Q(s t ,a t |θ Q ) is the Q value output by the current network at time t, y t is the target Q value, and its calculation formula is as follows:

[0286] y t =r t +γQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) (18)

[0287] In the formula, r t is the instantaneous reward at time t extracted from the experience pool, π′(s t+1 |θ π′ ) is the target policy network with parameters θ π′ The input state variable s t+1 The action variable output when Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the target network with parameters θ Q′ Next input state s t+1 and action variable π′(s t+1 |θ π′ ) under the input Q value.

[0288] 3.7) Update the gradient of the value network according to the following formula:

[0289]

[0290] In the formula, θ Q is the current value network parameter, the learning rate μ of the value network Q is 0.0001, is the loss function L(θ Q ) for the parameter θ Q The gradient of π θ (a t |s t ) and π θold (a t |s t ) are two different strategies of the strategy network, old and new. t is the advantage function, the loss function boundary value ε is 0.02, E t is the mean function.

[0291] 3.8) Then update the parameters of the policy network according to the following formula;

[0292]

[0293] In the formula, θ π is the current policy network parameter, the learning rate μ of the policy network π is 0.0001, is the policy gradient.

[0294] 3.9) If k%C=1, based on the soft update strategy, update the target network parameters according to the following formula

[0295]

[0296] In the formula, θ Q′ and θ π′ are the parameters of the target value network and the target policy network respectively, and the soft update coefficient τ is 0.01.

[0297] 3.10) Determine whether k>K holds. If so, output the training network. Otherwise, set k=k+1 and go to step 3.2.

[0298] 4) Application analysis

[0299] In order to verify the adaptability and superiority of the method proposed in this paper, this paper sets up the following comparative examples:

[0300] Case 1: Using particle swarm optimization to solve the above risk scheduling model

[0301] Case 2: Using the method proposed in this paper to train the power system risk dispatch network

[0302] The results of the two different cases are as follows Figure 1As shown, it can be seen that the economic cost of the method proposed in this paper is basically less than that of the particle swarm algorithm, and its performance indicators are shown in Table 1 below:

[0303] Table 2 Performance indicators under different schemes

[0304]

[0305] It can be seen from Table 2 above that the method proposed in this paper has very good economy and speed. Its overall economic cost is 0.57% lower than that of the particle swarm algorithm. It only takes 1.02s to give a scheduling plan, which is 15.87 times that of the conventional particle swarm algorithm. In addition, the network has good applicability and can adapt to the input of unwanted wind and load.

[0306] At the same time, after adopting the unit dispatch plan obtained by the power system risk dispatch network, the corresponding results under the base state and N-1 fault state are as follows: Figures 2 to 5 As shown, it can be seen that both the line power and the node voltage meet the normal safety requirements. Therefore, the method proposed in this paper can ensure the economy and safety of power system operation.

Claims

1. A risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm, characterized in that: The following steps are involved: 1) Establish a power system risk dispatch model considering N-1 security constraints; 2) Determine the Markov reward process in the reinforcement learning framework; 3) Based on the historical power data of the power system and the Markov reward process, the power system risk dispatch model is trained to obtain a power system risk dispatch optimization model; 4) Obtain real-time monitoring data of the power system and input it into the power system risk dispatch optimization model to obtain the power system dispatch plan; The objective function of the power system risk dispatch model considering N-1 security constraints is as follows: In the formula, T is the scheduling time period; The unit power generation cost F at time t G,t , network loss cost F loss,t and risk penalty cost F risk,t They are as follows: Where N is the number of units; P G,t,n is the actual output of n units at time t; a n 、b n 、c n is the relevant cost coefficient of the nth unit; N b is the number of network nodes; P i,t , Q i,t and V i,t are respectively the active power, reactive power and node voltage of the i-th node at the t-th moment; R ij is the resistance between node i and node j; C sell N is the price of electricity sold by the power grid; c and N l are the number of expected accidents and the total number of power grid branches; C risk is the risk penalty factor; Vsev i,k is the severity of voltage amplitude exceeding the limit at node i under the expected accident k; Among them, the line overload level Lsev of the lth line after the expected accident k occurs l,k As shown below: Where s1 and s2 represent the severity of overload when the line flow reaches the short-term emergency value LTE and the long-term emergency value STE respectively; r is the percentage of the flow on branch l under the expected fault k; The severity of the voltage amplitude exceeding the limit at node i under the expected accident k is Vsev i,k As shown below: In the formula, K HV , A HV , B HV K is the coefficient of the high voltage severity function; LV , A LV , B LV is the coefficient of the low voltage severity function; V i is the voltage amplitude at node i.

2. According to claim 1, a risk scheduling method based on a proximal policy optimization algorithm considering N-1 safety constraints is characterized by: The constraint conditions of the power system risk dispatch model considering N-1 safety constraints include power flow equation constraints, unit output constraints, ramp constraints, node voltage and line power constraints in base state and N-1 state.

3. According to claim 2, a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm is characterized in that: The power flow equality constraints are as follows: Where P G,i,t , P wt,i,t , P pv,i,t and P D,i,t are the active outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment; Q G,i,t , Q wt,i,t , Q pv,i,t and Q D,i,t are the reactive power outputs of thermal power units, wind power, photovoltaic power and loads at the i-th node at the t-th moment; G ij , B ij are the conductance and reactance between nodes i and j respectively; δ ij,t is the phase angle between node i and node j at the tth moment; V j,t is the voltage at node j.

4. According to claim 2, a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm is characterized in that: The unit output constraints are as follows: In the formula, and are the lower and upper limits of the output of the nth unit respectively; P G,n is the output of the nth unit.

5. According to claim 2, a risk scheduling method considering N-1 safety constraints based on a proximal policy optimization algorithm is characterized in that: The ramp constraints of the unit are as follows: Where U G,n and D G,n are the maximum upward and downward climbing powers of conventional unit n respectively; P G,n,t , P G,n,t-1 They are the output of the nth unit at the tth moment and the t-1th moment respectively.

6. The risk scheduling method considering N-1 safety constraints based on the proximal policy optimization algorithm according to claim 2 is characterized in that: The node voltage constraints are as follows: Where V i,min and V i,max are the lower and upper limits of the voltage at node i, respectively; V i 0 and V i k are the voltage amplitudes of node i in the base state and the kth anticipated accident respectively; The line power constraints are as follows: Where PL l and PL l,max are respectively the active power and the maximum transmission power of line l, and are the active power of line l in the base state and the kth expected accident respectively; Nl is the total number of lines; G ii is the self-conductance of node i; δ ij is the phase angle between node i and node j; G ij , B ij are the conductance and reactance between node i and node j respectively.

7. The risk scheduling method considering N-1 safety constraints based on the proximal policy optimization algorithm according to claim 1 is characterized in that: The Markov reward process includes a quaternion array {S, A, γ, R}; S, A, γ, and R represent the state space, action space, discount factor, and reward function, respectively.

8. The risk scheduling method considering N-1 safety constraints based on the proximal policy optimization algorithm according to claim 7 is characterized in that: The state space S is shown below: S={S G ,S wt ,S pv ,S load } (12) In the formula, S G is the output set of N conventional units, S wt and S pv Represent the predicted output of wind power and photovoltaic power, S load Output aggregation for grid load forecasting; The action space A is as follows: A={A G,1 ,A G,2 ,…,A G,N } (13) In the formula, A G,i represents the output increment set of unit i; The reward function R is as follows: In the formula, γ is the discount factor; Among them, the immediate reward r t As shown below: r t =-(F G,t +F loss,t +F risk,t ) (15) In the formula, F G,t 、F loss,t and F risk,t They are the unit power generation cost, network loss cost and risk penalty cost at time t respectively.

9. The risk scheduling method considering N-1 safety constraints based on the proximal policy optimization algorithm according to claim 8 is characterized in that: The output increment set A of the unit i G,i It includes the output increment of unit i at different times; the output increment is used to calculate the dispatching output P of the computer group at the next moment G,i,t+1 ; The dispatch output of the unit at the next moment P G,i,t+1 As shown below: P G,i,t+1 =P G,i,t +ΔP G,i,t (16) Where P G,i,t is the active output of the unit at the ith node at the tth moment; ΔP G,i,t The active output increment of the unit at the i-th node at the t-th moment.

10. The risk scheduling method considering N-1 safety constraints based on the proximal policy optimization algorithm according to claim 1 is characterized in that: The steps of training the power system risk dispatch model include: 1) Randomly initialize the current strategy network parameters θ Q and the value network parameter θ π , determine the training period K, scheduling period T and target network update frequency C, and set k = 0; 2) Initialize the initial state s of the power system from the state space t , and let t = 0; 3) Based on state s in the current value network t Output action a t ; 4) Execute action a in the power system risk dispatch model t , and get the next state s t+1 , reward r t , let t=t+1, go to step 3); 5) The learned time series samples {A T ,S T ,R T ,S T+1 } is stored in the sample experience pool as a data set for training the network; 6) Randomly collect m time series samples from the experience pool {A T ,S T ,R T ,S T+1 }, and calculate the loss function L(θ Q ); L(θ Q )=E(y t -Q(s t ,a t |θ Q )) 2 (17) In the formula, Q(s t ,a t |θ Q ) is the Q value output by the current network at time t; E represents the mean function; y t is the target Q value y t As shown below: y t =r t +γQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) (18) In the formula, r t is the instantaneous reward at time t extracted from the experience pool; π′(s t+1 |θ π′ ) is the target policy network with parameters θ π′ The input state variable s t+1 The action variable output when Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the target network with parameters θ Q′ Next input state s t+1 and action variable π′(s t+1 |θ π′ ) under the input Q value; 7) Update the gradient of the value network according to the following formula: In the formula, θ Q is the current value network parameter, μ Q is the learning rate of the current value network, is the loss function L(θ Q ) for the parameter θ Q The gradient of θ (a t |s t ) and π θold (a t |s t ) are two different strategies of the strategy network, old and new; A t is the advantage function; ε is the boundary value of the loss function; E t is the mean function; 8) Update the parameters θ of the policy network π ',Right now: In the formula, θ π is the current policy network parameter, μ π is the learning rate of the policy network; is the Q value Q(s,a|θ Q ) The gradient of action a; is the policy gradient; 9) Determine whether k%C=1. If so, update the strategy network parameter θ based on formula (21) Q′ and the value network parameter θ π′ , and go to step 10), otherwise, go directly to step 10); In the formula, θ Q′ and θ π′ are the parameters of the target value network and the target strategy network respectively, and τ is the soft update coefficient; 10) Determine whether k>K holds true. If so, output the power system risk dispatch optimization model. Otherwise, set k=k+1 and return to step 2).

Citation Information

Patent Citations

  • Photovoltaic-power-generation-containing power system emergency risk reduction control method

    CN111769560A

  • Dynamic power system economic dispatching method based on deep reinforcement learning

    CN112186743A