Adaptive Optimal AGC Control Method Based on Integral Reinforcement Learning

Through the adaptive optimal AGC control method of integral reinforcement learning, the frequency fluctuation problem caused by the instability of new energy generation in a single-region power system is solved, and fast and accurate frequency control is achieved, which extends the equipment life and reduces operating costs.

CN113346552BActive Publication Date: 2025-07-22STATE GRID CHONGQING ELECTRIC POWER COMPANY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110420781.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-19
Publication Date
2025-07-22
Estimated Expiration
2041-04-19

AI Technical Summary

Technical Problem

The existing AGC control method is difficult to quickly adapt to the instability and frequency fluctuations of new energy power generation in a single-region power system, and requires complete dynamic information of the system, resulting in difficulty in regulation and accelerated equipment aging.

Method used

Adaptive optimal AGC control method based on integral reinforcement learning is adopted, and the frequency response model of the single-region power system and the judge-executor neural network are established, and the strategy iterative algorithm in reinforcement learning is used to realize online learning and adaptive adjustment of the optimal control strategy.

Benefits of technology

In the case where the system dynamic model is unknown, the learning speed and accuracy are improved, frequency deviation is effectively suppressed, equipment aging is slowed down, and operation and maintenance costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113346552B_ABST
    Figure CN113346552B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive optimal AGC control method based on integral reinforcement learning. The steps are as follows: 1) Establish a single-area power system frequency response model and calculate the power system state space matrix; 2) Based on the policy iteration algorithm in reinforcement learning, establish a critic-actor neural network; the critic-actor neural network includes a critic network and an actor network; 3) Input the power system state space matrix into the critic-actor neural network and solve to obtain the optimal control strategy. The present invention uses an integral reinforcement learning strategy to learn the optimal cost function, enabling the learning process to be carried out under the condition that the system dynamic model is unknown, and improving the learning speed and accuracy from the perspective of weakening the persistent excitation condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems and their automation, and specifically to an adaptive optimal AGC control method based on integral reinforcement learning. Background Art

[0002] Currently, the structure of power systems is becoming increasingly complex, continuously expanding and extending to many remote areas. However, due to distance and natural conditions, the transmission costs in remote areas are high, the number of interconnection lines with other regions is limited or there are no interconnection lines. When a fault occurs in the interconnection line between regions, the local power system is prone to become an islanded single-region system. Therefore, the AGC control strategy for maintaining the stable operation of the single-region power grid is relatively important. At the same time, new energy generation often accounts for a large proportion in the power systems of these areas. Due to the instability of the output power of wind turbines, photovoltaic or tidal generating units, the frequency response of the power grid is prone to fluctuate. Coupled with the relatively small total inertia of the units in the single-region system, it is difficult to regulate the random fluctuations between the power generation side and the load side, resulting in a large frequency deviation. On the other hand, the system adjustment actions brought about by frequent frequency fluctuations also accelerate the aging of generator set components such as governors, increasing the operation and maintenance costs. The AGC control method based on the optimal control theory achieves the control purpose by minimizing the defined cost function related to the frequency deviation and the unit output. However, from the current research situation, the existing optimal control methods require complete dynamic information of the system, the optimal control strategy is difficult to solve, and it is easily affected by parameter changes and disturbance quantities. The adaptive optimal control methods proposed by some scholars can solve the optimal control strategy through online learning, but they face the problems of slow learning speed and inability to converge to the optimal, and still require the dynamic information of the system. If it is to be applied to the AGC control of a single-region power system, the adaptive optimal control strategy needs to solve the above problems to meet the requirements of actual operation. Summary of the Invention

[0003] The object of the present invention is to provide an adaptive optimal AGC control method based on integral reinforcement learning, including the following steps:

[0004] 1) Establish a frequency response model of a single-region power system and calculate the state space matrix of the power system;

[0005] The components of the power system include a governor, a turbine, a generator rotor, and a load.

[0006] The frequency response model of the single-region power system is as follows:

[0007]

[0008] In the formula, ΔX g (t) is the increment of the governor valve opening change; Is the increment ΔX g (t) differential; ΔP g (t) is the generator output change; Is the increment ΔP g (t) differential; Δf(t) is the frequency error increment; Is the differential of the increment Δf(t); ΔI(t) is the frequency error integral increment; Is the differential of the increment ΔI(t); ΔP d (t) is the load increment; T g 、T t 、T p Are the governor, turbine, and generator time constants respectively; K p 、K e Are the generator gain and integral control gain respectively; R d Is the governor speed droop rate; u(t) is the control strategy at time t;

[0009] Among them, the governor valve opening change increment ΔX g (t), the generator output change ΔP g (t), the frequency error increment Δf(t), and the frequency error integral increment ΔI(t) are the state variables of the single - area power system frequency response model; the load increment ΔP d (t) is the disturbance variable.

[0010] The power system state - space matrix is as follows:

[0011]

[0012] In the formula, x(t) represents the state variable; Represents the differential of the state variable;

[0013] Among them, the matrix A, matrix B, and matrix F are respectively as follows:

[0014]

[0015]

[0016] In the formula, R is the control variable weight.

[0017] 2) Based on the policy iteration algorithm in reinforcement learning, a critic - actor neural network is established; the critic - actor neural network includes a critic network and an actor network;

[0018] Both the critic network and the actor network include an input layer, a hidden layer, and an output layer;

[0019] The activation function of the critic network is χ(x) = [χ1(x), χ2(x),..., χ N (x)] T ; χ1(x), χ2(x),..., χ N (x) are neurons in the hidden layer of the critic network; the total number of neurons N ≥ n(n + 1) / 2; n is the number of state variables in the system.

[0020] The steps for the critic network to output the cost function V(x) include:

[0021] I) Establish an estimation expression for the cost function V(x), that is:

[0022] V(x) = w T χ(x) + ε a (x) (3)

[0023] In the formula, w = [w1, w2,..., w N T is the weight vector of the activation function vector χ(x); ε a (x) is the estimation error;

[0024] The partial derivative of the cost function V(x) with respect to the state variable x is as follows:

[0025]

[0026] II) The critic network learns the weight parameter vector through an adaptive parameter estimation method to obtain the estimated value of the weight vector At this time, the cost function is expressed in the form of the sum of the estimated value of the critic network, the estimation error, and the adaptive estimation error, as follows:

[0027]

[0028] In the formula, the adaptive estimation error

[0029] The Hamiltonian equation H(x(t, t + T), u) corresponding to the cost function is as follows:

[0030]

[0031] In the formula, V(x(t) is the cost function. Formula (6) is used to obtain the estimation error and the adaptive error of the critic network.

[0032] III) Calculate the Hamiltonian - Bellman equation error ε A , that is:

[0033] ​

[0034] In the formula, the Bellman equation error ε A = ε a (x(t + T)) - ε a (x(t)) is bounded; the enhanced signal term

[0035] IV) Calculate the adaptive estimation error ε on the time period [t, t + T] E = ε e (x(t + T)) – ε e (x(t)) and the total estimation error ε = ε A + ε E ;

[0036] Among them, the total estimation error ε satisfies the following formula:

[0037]

[0038] In the formula, the activation function equation Δχ(t) = χ(t + T) - χ(t);

[0039] V) Establish the adaptive estimation error cost function J of the critic network, that is:

[0040]

[0041] In the formula, J is the integral formula of the error quadratic term; β is the forgetting factor;

[0042] VI) The dynamic change of the estimated value of the weight vector is as follows:

[0043]

[0044] In the formula, Γ > 0 is the adaptive gain diagonal matrix; the normalization factor η = 1 + Δχ T Δχ;

[0045] VII) Define the integral term Ω(t) and the integral term Φ(t) as follows:

[0046]

[0047] In the formula, Ω is an N - order vector, and Φ is an N - order square matrix;

[0048] Substitute formula (10) into formula (9) to get:

[0049]

[0050] Among them, the dynamic processes of the vector Ω and the matrix Φ are as follows::

[0051]

[0052] In the formula, respectively represent the dynamic processes of the vector Ω and the matrix Φ;

[0053] VIII) Substitute formula (11) into formula (5) to obtain the cost function V(x).

[0054] The output of the executor network is as follows:

[0055]

[0056] In the formula, w is the weight vector; u is the control strategy; g is the dynamic characteristic of the system input, that is, the matrix B in the state space model.

[0057] 3) Input the power system state space matrix into the critic-executor neural network to solve and obtain the optimal control strategy.

[0058] The steps to solve and obtain the optimal control strategy include:

[0059] 3.1) Initialize the control strategy, denoted as u 0 ; Initialize the cost function, denoted as V 1 ;

[0060] 3.2) The critic network calculates the cost function V of the (i + 1)-th iteration according to the control strategy u of the i-th iteration i ; The initial value of i is 1; The cost function V i+1 is as follows: i+1 as follows:

[0061]

[0062] In the formula, V(x(t + T) is the cost function at time t + T; x(τ) is the state variable;

[0063] Among them, the utility parameter U(x(τ), u(x(τ)) is as follows:

[0064] U(x(τ), u(τ)) = x T (τ)Qx(τ) + u T (τ)Ru(τ) (16)

[0065] In the formula, Q is the state variable weight; R is the control variable weight;

[0066] 3.3) The executor network calculates the control strategy u of the (i + 1)-th iteration according to the cost function V i+1 , that is: i+1 Namely:

[0067]

[0068] 3.4) The executor network determines whether the cost function increment ΔV ≤ ε V and the control strategy increment Δu ≤ ε u hold. If so, the control strategy u i+1 is taken as the optimal control strategy; otherwise, let i = i + 1 and return to step 3.2); the cost function increment ΔV = V i+1 - V i ; the control strategy increment Δu = u i+1 - u i ε V and ε u are the cost function increment threshold and the control strategy increment threshold respectively.

[0069] It should be noted that the present invention establishes a frequency response model for a single - area power system. For a given system, by selecting appropriate state variables and linearizing them at the system equilibrium point, a corresponding frequency response model can be established and the system state - space matrix can be obtained.

[0070] Then, based on the policy iteration algorithm in reinforcement learning, a critic - executor neural network is established to implement learning and control. In reinforcement learning, the executor network (AGC controller) executes the control strategy and acts on the external environment (power system), and the critic network evaluates the current control action for policy evaluation, learns the return value (cost function) of the current policy, and feeds the system state variables and the return value back to the critic network. Among them, the critic network's learning of the cost function is based on the Weierstrass high - order approximation estimation method, approximating the unknown high - order polynomial as a combination of quadratic polynomials, establishing a Hamiltonian error equation based on the integral enhancement signal, solving the weight coefficient vector of the quadratic polynomial through the gradient method, and then obtaining the cost function. The executor network substitutes the learning result of the cost function into the Hamiltonian equation to solve the current control strategy.

[0071] Among them, when solving the weight vector of the approximate function of the cost equation by the gradient method, the present invention selects a quadratic error cost function, so that the persistent excitation condition of the recursive vector in parameter convergence can be weakened to persistent excitation within a finite time, and faster and more accurate learning of the cost function can be achieved.

[0072] Finally, through the simulation of MATLAB 2016 software, the effectiveness of the present invention is verified in a single - area power system model, proving that the present invention can achieve a better frequency - modulation effect.

[0073] The technical effect of the present invention is beyond doubt. The present invention uses the integral reinforcement learning strategy to learn the optimal cost function, enabling the learning process to be carried out when the system dynamic model is unknown, and improving the learning speed and accuracy from the perspective of weakening the persistent excitation condition. Description of the Drawings

[0074] Figure 1 It is a schematic diagram of adaptive optimal control;

[0075] Figure 2 It is a block diagram of the frequency response of a single - area power system;

[0076] Figure 3 It is a flowchart of the control algorithm;

[0077] Figure 4 It is the frequency error suppression effect of the present invention.

[0078] Figure 5 It is the frequency error suppression effect of traditional PI control. Specific implementation manners

[0079] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above - mentioned subject scope of the present invention is limited to the following embodiments. Without departing from the above - mentioned technical idea of the present invention, various substitutions and changes made according to ordinary technical knowledge and customary means in the art shall be included within the protection scope of the present invention.

[0080] Embodiment 1:

[0081] Refer to Figures 1 to 3 , the adaptive optimal AGC control method based on integral - type reinforcement learning includes the following steps:

[0082] 1) Establish a frequency response model of a single - area power system and calculate the state - space matrix of the power system;

[0083] The devices of the power system include a governor, a turbine, a generator rotor, and a load.

[0084] The frequency response model of the single - area power system is as follows:

[0085]

[0086] In the formula, ΔX g (t) is the increment of the governor valve opening; is the differential of the increment ΔX g (t); ΔP g (t) is the change in generator output; is the differential of the increment ΔP g (t); Δf(t) is the frequency error increment; is the differential of the increment Δf(t); ΔI(t) is the frequency error integral increment; is the differential of the increment ΔI(t); ΔP d (t) is the load increment; T g 、T t, T p are the time constants of the governor, turbine, and generator respectively; K p , K e are the generator gain and integral control gain respectively; R d is the governor speed droop rate; u(t) is the control strategy at time t;

[0087] Among them, the increment of the governor valve opening ΔX g (t), the change in generator output ΔP g (t), the increment of frequency error Δf(t), and the increment of frequency error integral ΔI(t) are the state variables of the single-area power system frequency response model; the load increment ΔP d (t) is the disturbance variable.

[0088] The state space matrix of the power system is as follows:

[0089]

[0090] In the formula, x(t) represents the state variable; represents the differential of the state variable;

[0091] Among them, matrices A, B, and F are as follows respectively:

[0092]

[0093]

[0094] In the formula, R is the control variable weight. When conducting model analysis, it is assumed that other state variables remain unchanged and only one variable changes, and this changing state variable is the control variable.

[0095] 2) Based on the policy iteration algorithm in reinforcement learning, a critic-actor neural network is established; the critic-actor neural network includes a critic network and an actor network;

[0096] Both the critic network and the actor network include an input layer, a hidden layer, and an output layer;

[0097] The activation function of the critic network is χ(x) = [χ1(x), χ2(x),..., χ N (x)] T ; χ1(x), χ2(x),..., χ N (x) are the neurons in the hidden layer of the critic network; the total number of neurons N ≥ n(n + 1) / 2; n is the number of state variables in the system.

[0098] The steps for the critic network to output the cost function V(x) include:

[0099] I) Establish the estimation expression of the cost function V(x), i.e.:

[0100] V(x) = w T χ(x) + ε a (x) (3)

[0101] In the formula, w = [w1, w2,..., w N T is the weight vector of the activation function vector χ(x); ε a (x) is the estimation error;

[0102] The partial derivative of the cost function V(x) with respect to the state variable x is as follows:

[0103]

[0104] II) The critic network learns the weight parameter vector through an adaptive parameter estimation method to obtain the estimated value of the weight vector At this time, the cost function is expressed in the form of the sum of the estimated value of the critic network, the estimation error, and the adaptive estimation error, as follows:

[0105]

[0106] In the formula, the adaptive estimation error

[0107] Equation 5 is the estimation form, and Equation 15 is the direct expression in the iterative process.

[0108] The Hamiltonian equation H(x(t, t + T), u) corresponding to the cost function is as follows:

[0109]

[0110] In the formula, V(x(t) is the cost function. Formula (6) is used to obtain the estimation error and the critic network adaptive error.

[0111] III) Calculate the Hamiltonian - Bellman equation error ε A , i.e.:

[0112]

[0113] In the formula, the Bellman equation error ε A = ε a (x(t + T)) - ε a (x(t)) is bounded; the enhancement signal term x(t + T), x(t) represent the state variables at time t + T and time t respectively. ​

[0114] IV) Calculate the adaptive estimation error ε over the time period [t, t + T]. E = ε e (x(t + T)) – ε e (x(t)) and the total estimation error ε = ε A + ε E ;

[0115] Wherein, the total estimation error ε satisfies the following formula:

[0116]

[0117] In the formula, the activation function equation is Δχ(t) = χ(t + T) - χ(t);

[0118] V) Establish the adaptive estimation error cost function J of the critic network, that is:

[0119]

[0120] In the formula, J is an integral formula of the quadratic term of the error; β is the forgetting factor;

[0121] VI) Dynamic change of the estimated value of the weight vector As follows:

[0122]

[0123] In the formula, Γ > 0 is the adaptive gain diagonal matrix; the normalization factor η = 1 + Δχ T Δχ;

[0124] VII) Define the integral term Ω(t) and the integral term Φ(t) as follows:

[0125]

[0126] In the formula, Ω is an N-order vector and Φ is an N-order square matrix;

[0127] Substitute formula (10) into formula (9) to get:

[0128]

[0129] Wherein, the dynamic processes of the vector Ω and the matrix Φ are as follows::

[0130]

[0131] In the formula, respectively represent the dynamic processes of the vector Ω and the matrix Φ;

[0132] VIII) Substitute formula (11) into formula (5) to obtain the cost function V(x).

[0133] The output of the executor network is as follows:

[0134]

[0135] Wherein, w is the weight vector; u is the control strategy; g is the dynamic characteristic of the system input, i.e., the matrix in the state space model

[0136] Equation (17) is a direct expression, and Equation (13) is the expression of the learning result of V by the neural network.

[0137] 3) Input the power system state space matrix into the critic-executor neural network to solve and obtain the optimal control strategy.

[0138] The steps to solve and obtain the optimal control strategy include:

[0139] 3.1) Initialize the control strategy, denoted as u 0 ; Initialize the cost function, denoted as V 1 ;

[0140] 3.2) The critic network calculates the cost function V of the (i + 1)-th iteration according to the control strategy u of the i-th iteration i ; The initial value of i is 1; The cost function V i+1 is as follows: i+1 as follows:

[0141]

[0142] Wherein, V(x(t + T) is the cost function at time t + T; x(τ) is the state variable;

[0143] Among them, the utility parameter U(x(τ), u(x(τ)) is as follows:

[0144] U(x(τ), u(τ)) = x T (τ)Qx(τ) + u T (τ)Ru(τ) (16)

[0145] Wherein, Q is the state variable weight; R is the control variable weight;

[0146] 3.3) The executor network calculates the control strategy u of the (i + 1)-th iteration according to the cost function V i+1 i.e.: i+1 That is:

[0147]

[0148] 3.4) The executor network judges that the cost function increment ΔV ≤ ε V and the control strategy increment Δu ≤ ε uWhether it holds. If so, use the control strategy u i+1 as the optimal control strategy. Otherwise, let i = i + 1 and return to step 3.2); the cost function increment ΔV = V i+1 -V i ; the control strategy increment Δu = u i+1 -u i ε V and ε u are the cost function increment threshold and the control strategy increment threshold respectively.

[0149] Example 2:

[0150] The adaptive optimal AGC control method based on integral reinforcement learning includes the following steps:

[0151] 1) Establish a power system frequency response model

[0152] This invention mainly studies the frequency control of a single-area power system. The typical devices therein include governors, turbines, generator rotors, and loads, and their dynamic models can all be approximated as first-order processes. The system state variables are selected as the governor valve opening change increment ΔX g (t), the generator output change ΔP g (t), the frequency error increment Δf(t), and the frequency error integral increment ΔI(t). The disturbance variable is the load increment ΔP d (t). The differential equations of this system are summarized as follows:

[0153]

[0154] The system state space model is expressed as:

[0155]

[0156]

[0157]

[0158] 2) Policy iteration of integral reinforcement learning

[0159] In the optimal control problem, a cost function V related to the system state x and input u is defined on an infinite time domain:

[0160]

[0161] where U(x,u) is an artificially defined utility equation, usually in quadratic form:

[0162] U(x(τ),u(τ)) = x T (τ)Qx(τ)+u T(τ)Ru(τ), (4)

[0163] Taking the partial derivative of the cost function with respect to time t, the Hamiltonian equation of this problem is obtained:

[0164]

[0165] Solving the equation H = 0 can obtain the optimal cost V * , and then substituting V * into to solve for the optimal control action u * . For continuous-time systems, the implementation of conventional reinforcement learning methods requires complete system dynamic information, which has a certain implementation difficulty. The integral-type reinforcement learning method can solve the optimal cost function by only using the input dynamic information of the model when solving this problem, avoiding the use of all system dynamic information. Considering the cost function containing an integral enhancement signal, for any time interval T > 0, the cost function is expressed in a new form:

[0166]

[0167] At this time, the Hamiltonian equation can be re-expressed as:

[0168]

[0169] When solving the cost function according to this formula, the dynamic information of the system is not required. The Policy iteration algorithm alternately implements two steps: Policy evaluation and Policy improvement. The algorithm initialization includes the initialization of the control policy u 0 and the initialization of the cost function V 1 , and the two steps are summarized as follows:

[0170] I) Policy evaluation

[0171] According to the control policy u i calculated in the i-th iteration, substitute it to solve the cost function V i+1 of the (i + 1)-th iteration:

[0172]

[0173] II) Policy improvement

[0174] According to the cost function V i+1 calculated in the (i + 1)-th iteration, calculate the control policy u i+1 of the (i + 1)-th iteration:

[0175]

[0176] The policy iteration algorithm alternates between the above two steps until the increments of the cost function and the control policy converge within a certain small threshold.

[0177] 3) Design of the Actor - critic Network

[0178] I) Critic Network for Policy Evaluation

[0179] The critic neural network approximately estimates the high - order cost function as a linear combination of low - order polynomials according to the Weierstrass high - order estimation method, and approximates the cost function by estimating the weight parameter vector corresponding to the low - order polynomials through the gradient method. Define the low - order polynomial vector χ(x) = [χ1(x), χ2(x),..., χ N (x)] T as the activation function vector, which serves as the neurons in the hidden layer of the neural network. If the low - order polynomial element χ i is in the quadratic form of the system state variables {x i (t)x j (t)}(i, j = 1, 2,..., n), assuming the number of state variables in the system is n, then the number of low - order polynomials N should satisfy N≥n(n + 1) / 2. At this time, the optimal control cost function can be estimated as:

[0180] V(x) = w T χ(x)+ε a (x), (10)

[0181] where w = [w1, w2,..., w N T is the weight vector of the activation function vector χ(x), ε a (x) is the estimation error. Considering that the partial derivative of the cost function with respect to the state variable x is used when calculating the control policy in equation (9), based on equation (10), the estimation expression of the partial derivative can be obtained:

[0182]

[0183] The estimation error ε a (x) and its partial derivative are both bounded. When the number of elements N in the activation function approaches ∞, ε a and both tend to 0. Therefore, as many activation elements as possible can be selected within the computing power. The critic network realizes the learning of the weight parameter vector through the adaptive parameter estimation method, and the estimated weight vector is expressed as The cost function can be further expressed as: ​

[0184]

[0185] Here is the adaptive estimation error. When the adaptive law and the signal excitation conditions can ensure the exponential stability of parameter estimation, ε e →0.

[0186] According to equation (7), the Bellman equation error ε A on the time interval [t, t + T] can be expressed as:

[0187]

[0188] where ε A = ε a (x(t + T)) - ε a (x(t)) is bounded. Here, the activation function equation is expressed as Δχ(t) = χ(t + T) - χ(t), and the enhancement signal term is denoted by μ(t) Define ε E = ε e (x(t + T)) – ε e (x(t)) as the adaptive estimation error on the time interval [t, t + T]. Then, use the total estimation error ε = ε A + ε E to represent the sum of the approximation error and the adaptive estimation error on the time interval [t, t + T]. Therefore, equation (13) can be rewritten as:

[0189]

[0190] Define the adaptive estimation error cost function J of the critic network:

[0191]

[0192] J is an integral of the error quadratic term. The exponential decay term avoids the unbounded cost caused by the integral effect. The forgetting factor β is related to the decay rate of historical dynamic information. Based on the gradient method, the dynamics of the estimated weights can be expressed as:

[0193]

[0194] where the constant Γ > 0 is the adaptive gain diagonal matrix, η = 1 + Δχ T Δχ is introduced as a normalization factor. For the convenience of representing the adaptive law, define the following integral terms:

[0195]

[0196] where Ω is an N-dimensional vector and Φ is an N×N square matrix. Therefore, equation (16) can be expressed as:

[0197]

[0198] The dynamic processes of vector Ω and matrix Φ can be expressed as:

[0199]

[0200] The selected error cost function preserves the historical information of the activation function Δχ(t). The adaptive process utilizes the dynamic information of the system at both the current and historical moments, weakening the persistent excitation condition of Δχ(t) necessary for exponentially stable parameter estimation to excitation within a finite time, which is easier to achieve. That is to say, the estimated parameters can converge to the true value in a shorter time, thereby achieving a better control effect. It is proven by the second Lyapunov method that when there is a bounded high-order estimation error ε a exists, the parameter estimation error can converge to a bounded value, and the cost function solved by the critic network is near the optimal value V * ; when the estimation error ε a = 0, the parameter estimation error can converge exponentially to 0. At this time, the critic network can solve the optimal cost function V * .

[0201] II) The actor network for policy update

[0202] The actor network calculates the control policy based on the learning result of the current critic network cost function:

[0203]

[0204] Assuming that the signal excitation condition for adaptive learning in the critic network can ensure the convergence of the parameter estimation result, according to the second Lyapunov method, it can be proven that when the high-order estimation error ε a of the neural network is a sufficiently small bounded value, the control policy solved by the actor network is a value within a bounded interval near the optimal policy u * , and the system state variables are bounded; when the estimation error ε a = 0, the actor network can solve the optimal policy u * .

[0205] Embodiment 3:

[0206] Referring to Figure 4 and Figure 5 , the adaptive optimal AGC control method based on integral reinforcement learning includes the following steps:

[0207] 1) System parameter setting

[0208] The controlled object is Figure 1The single - area power system shown, the governor time constant T g = 0.08, the turbine time constant T t = 0.1, the generator time constant T p = 20, the generator gain K p = 120, the governor droop rate R d = 2.5, the integral control gain K e = 1.

[0209] Define the optimal control cost function as in Equation (6), where the state variable weight Q = I of the utility equation U(x, u), the control variable weight R = 0.5, and the activation function χ(x) in the critic network is selected as a vector containing 10 quadratic - term elements The system state variable is initialized as x(0)=[0 0 0 0] T , and the initial value of the critic network weights is The adaptive gain matrix Γ = 10I, the adaptive forgetting factor β = 1.2, and the sampling period of the integral enhancement signal is T = 0.1 s.

[0210] 2) Algorithm performance and test results

[0211] The frequency deviation suppression effect of the control algorithm proposed by the present invention on the single - area power system is as Figure 4 shown, the control effect of the proportional - integral PI controller. There is an influence of a small load disturbance in the system. At 10 seconds, the system is subjected to a load disturbance of 0.25 p.u., and the disturbance disappears at 20 seconds. The frequency suppression effect of the control method proposed by the present invention for load disturbances is better than that of the classical proportional - integral method, which proves the effectiveness of the method.

[0212] To sum up, the present invention proposes a new method for AGC control of a single - area power system. This method is based on the policy iteration algorithm of integral - type reinforcement learning, and uses an actor - critic neural network to implement the two steps of policy evaluation and policy update in policy iteration. It can realize the learning of the cost function and the implementation of the optimal control strategy without knowing the system dynamic information, effectively improving the suppression effect of the frequency deviation of the power system and providing guidance for the parameter setting during the operation of the power system.

Claims

1. An adaptive optimal AGC control method based on integral reinforcement learning, characterized in that, It includes the following steps: 1) Establish a single - area power system frequency response model and calculate the power system state - space matrix; 2) Based on the policy iteration algorithm in reinforcement learning, establish a critic - actor neural network; The critic - actor neural network includes a critic network and an actor network; 3) Input the power system state - space matrix into the critic - actor neural network and solve to obtain the optimal control strategy; 4) The actor network executes the optimal control strategy in the power system; The single - area power system frequency response model is as follows: Where, ΔX g (t) is the increment of the governor valve opening change; is the differential of the increment ΔX g (t); ΔP g (t) is the change in the generator output; is the differential of the increment ΔP g (t); Δf(t) is the increment of the frequency error; is the differential of the increment Δf(t); ΔI(t) is the increment of the frequency error integral; is the differential of the increment ΔI(t); ΔP d (t) is the load increment; T g , T t , T p are the governor, turbine, and generator time constants respectively; K p , K e are the generator gain and the integral control gain respectively; R d is the governor speed droop rate; u(t) is the control strategy at time t; Among them, the increment of the governor valve opening ΔX g (t), the change in the generator output ΔP g (t), the frequency error increment Δf(t) and the frequency error integral increment ΔI(t) are the state variables of the single-area power system frequency response model; the load increment ΔP d (t) is the disturbance variable.

2. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 1, characterized in that: The devices of the power system include a governor, a turbine, a generator rotor, and a load.

3. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 1, characterized in that The power system state - space matrix is as follows: where \(x(t)\) represents the state variable; represents the differential of the state variable; Among them, matrix A, matrix B, and matrix F are respectively as follows: In the formula, R is the control variable weight.

4. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 1, characterized in that Both the critic network and the actor network include an input layer, a hidden layer, and an output layer; The activation function of the critic network is χ(x) = [χ1(x), χ2(x),..., χ N (x)] T ; χ1(x), χ2(x),..., χ N (x) are neurons in the hidden layer of the critic network; the total number of neurons N ≥ n(n + 1) / 2; n is the number of state variables in the system.

5. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 1, characterized in that The steps for the critic network to output the cost function V(x) include: 1) Establish an estimation expression for the cost function V(x), that is: V(x) = w T χ(x) + ε a (x) (3) where \(w = [w_1, w_2, \cdots, w N T is the weight vector of the activation function vector \(\chi(x)\); \(\varepsilon a (x)​ is the estimation error; The partial derivative of the cost function V(x) with respect to the state variable x is as follows: 2) The critic network learns the weight parameter vector through an adaptive parameter estimation method to obtain an estimated value of the weight vector. At this time, the cost function is expressed in the form of the sum of the estimated value of the critic network, the estimation error, and the adaptive estimation error, as follows: where the adaptive estimation error 3) Calculate the Hamiltonian-Bellman equation error ε over the time period [t, t+T] A , i.e.: where the Bellman equation error ε A = ε a (x(t + T)) - ε a (x(t)) is bounded; the enhanced signal term 4) Calculate the adaptive estimation error ε over the time period [t, t+T] E = ε e (x(t+T)) - ε e (x(t)) and the total estimation error ε = ε A + ε E ; Among them, the total estimation error ε satisfies the following formula: In the formula, the activation function equation is Δχ(t) = χ(t + T)-χ(t); 5) Establish the adaptive estimation error cost function J of the critic network, that is: In the formula, J is the integral of the error quadratic term; β is the forgetting factor; 6) Dynamic change of the estimated value of the weight vector As follows: where Γ>0 is an adaptive gain diagonal matrix; the normalization factor η = 1 + Δχ T Δχ; 7) Define the integral term Ω(t) and the integral term Φ(t) as follows: In the formula, Ω is an N - order vector, and Φ is an N - order square matrix; Substitute formula (10) into formula (9) to get: Among them, the dynamic processes of the vector Ω and the matrix Φ are as follows: In the formula, respectively represent the dynamic processes of the vector Ω and the matrix Φ; 8) Substitute formula (11) into formula (5) to obtain the cost function V(x).

6. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 5, wherein The output of the actor network is as follows: In the formula, w is the weight vector; u is the control strategy; g is the dynamic characteristic of the system input, that is, matrix B in the state - space model.

7. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 6, characterized in that The Hamiltonian equation H(x(t,t + T),u) corresponding to the cost function is as follows: In the formula, V(x(t)) is the cost function; formula (14) is used to calculate the estimation error and the critic network adaptive error.

8. The adaptive optimal AGC control method based on integral reinforcement learning according to claim 1, characterized in that The steps for solving the optimal control strategy include: 1) Initialize the control strategy, denoted as u 0 ; Initialize the cost function, denoted as V 1 ; 2) The critic network calculates the cost function V for the (i + 1)-th iteration according to the control strategy u for the i-th iteration i ; the initial value of i is 1; the cost function V i+1 is as follows: i+1 as shown below: In the formula, V(x(t + T)) is the cost function at time t + T; x(τ) is the state variable; Among them, the utility parameter U(x(τ),u(x(τ)) is as follows: U(x(τ),u(τ)) = x T (τ)Qx(τ)+u T (τ)Ru(τ) (16) In the formula, Q is the state variable weight; R is the control variable weight; 3) The executor network calculates the control policy u for the (i + 1)-th iteration according to the cost function V i+1 That is: i+1 ​ 4) The executor network determines whether the cost function increment ΔV ≤ ε V and the control strategy increment Δu ≤ ε u hold. If so, the control strategy u i+1 is taken as the optimal control strategy; otherwise, set i = i + 1 and return to step 2). The cost function increment ΔV = V i+1 - V i ; the control strategy increment Δu = u i+1 - u i ε V and ε u are the cost function increment threshold and the control strategy increment threshold, respectively.

Citation Information

Patent Citations

  • Circuit security constraint-considering provincial grid AGC (automatic generation control) unit dynamic optimization scheduling method

    CN104682392A

  • AGC real-time control strategy based on deep learning in big data environment

    CN111555363A