Single-machine infinite bus power system safety control method based on reinforcement learning

By applying reinforcement learning and adaptive control technology in a stand-alone infinity power system, the problem of safety control of power system under wrong data injection attacks is solved, and the system is stable, safe and high-performance is achieved.

CN120195982APending Publication Date: 2025-06-24HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing power system safety control methods have limitations when facing wrong data injection attacks, and it is difficult to effectively ensure the stability and security of the system.

Method used

The single-machine infinite power system safety control method based on reinforcement learning is adopted, and the virtual controller is designed through fractional-order mathematical model, adaptive inverse step control, dynamic surface control and reinforcement learning to achieve system stability and security.

Benefits of technology

In case of incorrect data injection attacks, ensure that the damaged system state does not violate constraints, all signals are bounded, the system is calm, and the overall performance is achieved, while reducing the communication burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120195982A_ABST
    Figure CN120195982A_ABST
Patent Text Reader

Abstract

The invention discloses a single-machine infinite-bus power system safety control method based on reinforcement learning, and the method comprises the following steps: building a damaged system expression after an attack signal is encountered based on a fractional-order single-machine infinite-bus power system mathematical model, and constructing a nonlinear function to restrain a damaged state variable; a single-machine infinite-bus power system is decomposed into two stages of subsystems; an equivalent auxiliary system is constructed by selecting a cost function and introducing an auxiliary variable, so that the optimization problem that a first-stage subsystem contains a fractional order dynamic system is solved; constructing a first Lyapunov function, and enabling a first-stage subsystem to tend to be stable; a second Lyapunov function is constructed, and a self-triggering control strategy is introduced to reduce the communication burden, so that a second-stage subsystem and the whole closed-loop system tend to be stable; all signals of the single-machine infinite-bus power system are bounded, the communication burden of the system can be effectively reduced, and under the condition that the system is attacked by wrong data injection, the state of the damaged system does not violate constraints, and the comprehensive performance is optimal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system control, and in particular, to a security control method for a single-machine infinite-bus power system based on reinforcement learning. Background Art

[0002] As one of the most important infrastructures in industrial society, the power system is crucial for both the national economy and people's livelihood. With the rapid development of the digitalization and intelligentization of the power system, the deep integration of the cyber-physical system not only improves the operation efficiency of the power grid but also makes it face increasingly severe cyber security challenges. Any potential malicious attack may pose a threat to the stability of the power network, thus affecting the safe operation of the regional power grid and even the entire power system. As a common cyber attack method, the false data injection attack has strong concealment and destructiveness. The attacker injects false data to change the integrity or authenticity of the data, so as to achieve the purpose of maliciously destroying the control performance of the system. Therefore, the research on its related control methods has received extensive attention.

[0003] The existing power system security control strategies mainly rely on classical control theories based on physical models, such as PID controllers, linear quadratic regulators, etc. Such methods often require accurate system parameters and linearized model assumptions. However, the actual power grid operation environment is complex and changeable. Especially when facing such complex cyber attacks with strong uncertainty and time-variation, traditional control methods have obvious limitations. Therefore, breakthrough technical means are urgently needed to solve the power system security control problem under false data injection attacks. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention provides a security control method for a single-machine infinite-bus power system based on reinforcement learning, realizing that the designed controller can ensure that all signals of the single-machine infinite-bus power system are bounded, can effectively reduce the system communication burden, and can ensure that under the condition of suffering from false data injection attacks, the damaged system state does not violate the constraints, the system is stabilized, and the comprehensive performance is optimal.

[0005] To achieve the above invention purpose, the present invention adopts the following technical solutions:

[0006] The first aspect of the present application provides a security control method for a single-machine infinite-bus power system based on reinforcement learning, including the following steps:

[0007] S101. Establish an expression of the damaged system after encountering an attack signal based on the fractional-order single-machine infinite-bus power system mathematical model, and construct a non-linear function λ(t) to constrain the damaged state variables;

[0008] S102. Based on adaptive backstepping control and dynamic surface control, construct a coordinate transformation equation to decompose the single-machine infinite bus power system into two-level subsystems;

[0009] S103. By selecting a cost function and introducing an auxiliary variable e to construct an equivalent auxiliary system, obtain the fractional-order Hamilton-Jacobi-Bellman equation and optimize the virtual controller α * , process the nonlinear terms in the fractional-order Hamilton-Jacobi-Bellman equation through a neural network and a fuzzy logic system, and design a virtual controller in combination with reinforcement learning Identification-Execution-Evaluation weight learning law and parameter adaptation law to solve the optimization problem of the fractional-order dynamic system in the first-level subsystem;

[0010] S104. Based on the coordinate transformation equation and the fractional-order Lyapunov stability theory, construct the first Lyapunov function V1, and confirm the designed virtual control law, weight learning law, and parameter adaptation law through the fractional-order derivative of V1, so that the first-level subsystem tends to be stable;

[0011] S105. By repeating the calculation processes of steps S103 and S104, construct the second Lyapunov function V2, and design the corresponding virtual controller Identification-Execution-Evaluation weight learning law and parameter adaptation law Introduce a self-triggered control strategy to reduce the communication burden and make the second-level subsystem and the entire closed-loop system tend to be stable.

[0012] Furthermore, the expression of the mathematical model of the fractional-order single-machine infinite bus power system is as follows:

[0013]

[0014] where D q represents the Caputo fractional-order differential symbol, 0 < q < 1; u represents the control input, and represent the relative angular velocity and angular frequency between two generators G1 and G2 respectively; and represent the moment of inertia, damping coefficient, maximum power, and mechanical power of the motor respectively;

[0015] Define the system state variables Let the output signal of the system be y = x1, then it has the following form:

[0016]

[0017] where \(q = 0.95\), \(f_1(x)=0\), \(f_2(x)=-0.5x^2-\sin(x_1)+2.75\sin(t)\)

[0018] Furthermore, step S101 specifically includes:

[0019] The damaged state after the system is attacked is where \(\omega(t)\) represents an unknown time-varying signal, and the expression of the damaged system is as follows:

[0020]

[0021] where \(p_1\) and \(p_2\) represent non-negative time-varying functions;

[0022] By constructing a non-linear function Furthermore, unilateral full-state constraint is achieved, and its expression is as follows:

[0023]

[0024] where, c i is a design parameter, and \(q_1\) and \(q_2\) represent non-negative time-varying functions. It should be noted that when the state variable approaches the set constraint boundary \(c\) i , then \(\lambda\) i \(\to\infty\), that is, as long as \(\lambda\) i is bounded, it also follows the constraint conditions.

[0025] Furthermore, in step S102, the expression of the coordinate transformation equation is as follows:

[0026]

[0027] where \(z_1\) and \(z_2\) represent error variables, and \(\alpha_1\), \(\eta_2\) and \(\xi_2\) represent the virtual control signal, the filtered output signal and the filtering error respectively;

[0028] The fractional-order filtering expression is as follows:

[0029] \(\eta_2(0)=\alpha_1(0)\)

[0030] where, is a design parameter, and \(\alpha_1(0)\) and \(\eta_2(0)\) represent their initial values.

[0031] Furthermore, in step S103, the expression of the cost function is where \(\Omega\) sDenote the admissible control set; the Hamilton-Jacobi-Bellman equation is Define the auxiliary variable e = D q-1 z, and we can get and z = D 1-q e; rewrite the cost function as Furthermore, the fractional-order Hamilton-Jacobi-Bellman equation can be obtained as where Using we can get the optimal controller Also, because J * (D 1-q e) = J * (z), equivalently, we can get where By introducing auxiliary variables and constructing an equivalent auxiliary system, the fractional-order Hamilton-Jacobi-Bellman equation and the optimal controller can be obtained, thus solving the optimization problem of the fractional-order dynamic system.

[0032] Furthermore, in step S103, according to the coordinate transformation equation in step S102, we can get

[0033]

[0034] Regarding the variable as the optimal virtual controller and defining the cost function we can obtain the optimal virtual controller

[0035] Furthermore, in step S103, to achieve the optimal control objective, construct the following auxiliary cost function where k1 > 0 is a design parameter; use the fuzzy logic system to approximate the unknown term where represents the optimal weight vector, ε1 is the approximation error, and there is a relationship and the input vector Since contains ideal parameter values that cannot be obtained, use the neural network to approximate the unknown function where, W1 * and represent the optimal weight vector and the approximation error respectively, represents the basis function and the input vector is By using the online training of the actor-critic network to obtain the optimal virtual controller, we can get

[0036]

[0037] where and respectively represent and the estimated values of;

[0038] Substitute the approximation terms and into the fractional - order Hamilton - Jacobi - Bellman equation, and the approximation error is

[0039]

[0040] To minimize the approximation error, it is necessary to satisfy Construct the Lyapunov function Then there is

[0041]

[0042] where

[0043] Design the weight learning law of the execution - evaluation neural network as

[0044]

[0045] where γ a1 , γ c1 > 0 are design parameters;

[0046] Substitute the above formula into D q S1, and we can get

[0047]

[0048] From this, it can be seen that can asymptotically converge to 0. Design the identification weight learning law and the parameter adaptation law as

[0049]

[0050] where ρ1, σ1 > 0 are design parameters.

[0051] Furthermore, step S104 specifically includes:

[0052] The expression of the Lyapunov function is Take the fractional - order derivative of V1 to get

[0053]

[0054] In view of and Then there is

[0055]

[0056] By using Young's inequality, we can obtain

[0057]

[0058]

[0059] Substituting the above equations and the designed weight learning law, we can obtain

[0060]

[0061] where is the parameter estimation error.

[0062] Furthermore, in step S105, according to the coordinate transformation equation in step S102, we can obtain

[0063]

[0064] Define the cost function The optimal virtual controller can be obtained

[0065] Furthermore, in step S105, in order to achieve the optimal control objective, construct the auxiliary cost function where k2 > 0 is the design parameter; in this process, use the fuzzy logic system to approximate the unknown term where represents the optimal weight vector, ε2 is the approximation error, and there is a relationship and the input vector Since there are ideal parameter values that cannot be obtained in, and use the neural network to approximate the unknown function where and respectively represent the optimal weight vector and the approximation error, represents the basis function and the input vector is Use the execution-evaluation network for online training to obtain the optimal virtual controller, and we can get

[0066]

[0067] where and respectively represent and the estimated values of;

[0068] Based on the above coordinate transformation equation and the fractional-order filter, we can obtain

[0069]

[0070] where and

[0071] Similarly, the design formulas of the weight learning law and the parameter adaptation law are as follows:

[0072]

[0073] where γ a2 , γ c2 , ρ2, σ2 and are all positive constants.

[0074] Advantages of the present application: It realizes a security control method against false data injection attacks by considering the impact of false data injection attacks on the power grid and provides a security control method against false data injection attacks with a single-machine infinite-bus power system as the research object; compared with general adaptive backstepping control methods, the present invention combines reinforcement learning, dynamic surface technology and fractional-order Lyapunov stability theory to achieve the optimal comprehensive performance of the control system while ensuring that the damaged state does not violate the constraints, all signals are bounded and the system is stabilized; in addition, the self-triggered control strategy adopted can reduce the update frequency of the control signal and eliminate Zeno behavior on the premise of ensuring the control performance, thereby effectively reducing the communication burden of the system.

[0075] The power system model is a complex non-linear model. If it is directly adopted in practice, it will bring a large number of problems and difficulties, which is not conducive to the development of research; while the single-machine infinite-bus power system, as a simplified mathematical model of the power grid, avoids some unnecessary influencing factors and helps the control research and stability analysis of the power system. Therefore, this patent mainly aims at the single-machine infinite-bus power system and designs an adaptive optimization controller based on reinforcement learning and self-triggered control strategy under the conditions of limited network communication bandwidth and false data injection attacks, so as to improve the stability, security and intelligence of the single-machine infinite-bus power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0077] Figure 1 is a schematic diagram of the steps of a security control method for a single-machine infinite-bus power system based on reinforcement learning according to the present invention;

[0078] Figure 2 is a schematic diagram of the single-machine infinite-bus power system according to the present invention;

[0079] Figure 3 are the original state trajectory and the damaged state trajectory diagrams of the single-machine infinite power system;

[0080] Figure 4 is the execution-evaluation weight curve diagram of the single-machine infinite power system;

[0081] Figure 5 is the adaptive law trajectory diagram of the single-machine infinite power system;

[0082] Figure 6 is the control input curve diagram of the single-machine infinite power system;

[0083] Figure 7 is the triggering interval diagram of the control input of the single-machine infinite power system. Specific implementation manners

[0084] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0085] The following illustrates the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0086] Embodiment 1:

[0087] A safety control method for a single-machine infinite power system based on reinforcement learning includes the following steps:

[0088] S101. Establish an expression of the damaged system after encountering an attack signal based on the fractional-order single-machine infinite power system mathematical model, and construct a nonlinear function λ(t) to constrain the damaged state variables;

[0089] It can be based on, as Figure 2 shown in the schematic diagram of the single-machine infinite power system, establish a fractional-order single-machine infinite power system mathematical model based on the Caputo fractional-order definition. The expression of the fractional-order single-machine infinite power system mathematical model is as follows:

[0090]

[0091] where Dq denotes the Caputo fractional differential symbol, where 0 < q < 1; u represents the control input, and represent the relative angular velocity and angular frequency between two generators G1 and G2, respectively; and represent the moment of inertia, damping coefficient, maximum power, and mechanical power of the motor, respectively;

[0092] Define the system state variables Let the output signal of the system be y = x1, then it has the following form:

[0093]

[0094] where q = 0.95, f1(x) = 0, f2(x) = -0.5x2 - sin(x1) + 2.75sin(t),

[0095] Based on the mathematical model of the fractional-order single-machine infinite-bus power system, establish the expression of the damaged system after encountering the attack signal, and construct the nonlinear function λ(t) to constrain the damaged state variables, specifically including:

[0096] The damaged state of the system after being attacked is where ω(t) represents an unknown time-varying signal, and the expression of the damaged system is as follows:

[0097]

[0098] where p1 and p2 represent non-negative time-varying functions;

[0099] By constructing the nonlinear function Furthermore, unilateral full-state constraint is achieved, and its expression is as follows:

[0100]

[0101] where, c i is a design parameter, and q1 and q2 represent non-negative time-varying functions. It should be noted that when the state variable approaches the set constraint boundary c i , then λ i →∞, that is, as long as λ i is bounded, also follows the constraint conditions.

[0102] S102. Based on adaptive backstepping control and dynamic surface control, construct the coordinate transformation equation to decompose the single-machine infinite-bus power system into two-level subsystems;

[0103] Through adaptive backstepping control and dynamic surface control, a coordinate transformation equation is constructed to decompose the single-machine infinite-bus power system into two-level subsystems, namely the first-level subsystem and the second-level subsystem.

[0104] Based on adaptive backstepping control and dynamic surface control, a coordinate transformation equation is constructed, and its expression is as follows:

[0105]

[0106] Among them, z1 and z2 represent error variables, and α1, η2, and ξ2 represent virtual control signals, filtered output signals, and filtered errors respectively;

[0107] The fractional-order filtering expression is as follows:

[0108] η2(0) = α1(0)

[0109] Among them, are design parameters, and α1(0) and η2(0) represent their initial values.

[0110] It should be noted that by introducing the fractional-order filter in the above form, this application can overcome the problem of repeated differentiation of the virtual control signal α i in the adaptive backstepping framework to reduce the corresponding computational burden. At the same time, considering the problem that the common first-order filter form βD q η2 + η2 = α1, β > 0 assumes that the upper bound value is known when dealing with the function term -D q α1, the introduction of and the parameter adaptation law related to can effectively avoid this conservative condition.

[0111] It should be noted that traditional control methods have obvious limitations, such as model dependence, static parameter design, and insufficient nonlinear processing ability, and are difficult to cope with the high uncertainty and strong nonlinearity problems commonly existing in power systems. Adaptive control methods can online estimate the uncertain parameters existing in the controlled object, so that the system can reach the control performance index under the adaptation law. For the unknown terms existing in the actual system, neural networks or fuzzy logic systems can be used to approximate the unknown functions. Considering the "complexity explosion" problem existing in the backstepping framework, this application can cleverly avoid this problem by introducing the dynamic surface technology.

[0112] S103. By selecting a cost function and introducing an auxiliary variable e to construct an equivalent auxiliary system, the fractional-order Hamilton-Jacobi-Bellman equation (HJB) and the optimized virtual controller α *, the neural network and the fuzzy logic system are used to process the non - linear terms contained in the fractional - order Hamilton - Jacobi - Bellman equation (HJB), and a virtual controller is designed by combining reinforcement learning. Identification - execution - evaluation weight learning law and parameter adaptation law to solve the optimization problem of the fractional - order dynamic system in the first - level subsystem;

[0113] For the first - level subsystem, a cost function is selected, the Hamilton - Jacobi - Bellman equation (HJB) is derived according to the Bellman optimization algorithm, and an equivalent auxiliary system is introduced by constructing an auxiliary variable e to obtain the fractional - order HJB equation and the optimized virtual controller α. * , the neural network and the fuzzy logic system are used to process the non - linear terms contained in the HJB equation, and a virtual controller is designed by combining reinforcement learning. Identification - execution - evaluation weight learning law and parameter adaptation law

[0114] Among them, the expression of the cost function is where Ω s represents the admissible control set; according to the Bellman optimization algorithm, the HJB equation can be obtained as Considering the existence of fractional - order differentiation, it is impossible to directly substitute it into the HJB equation for chain - rule differentiation. Define the auxiliary variable e = D q-1 z, and we can get and z = D 1- q e. At this time, the cost function is rewritten as Furthermore, the fractional - order HJB equation can be obtained as where Using the optimal controller can be obtained Also, because J * (D 1-q e) = J * (z), equivalently, we can get where By introducing the auxiliary variable and constructing the equivalent auxiliary system, the fractional - order HJB equation and the optimized controller can be obtained, thus solving the optimization problem of the fractional - order dynamic system.

[0115] According to the coordinate transformation equation in the above step S102, we can get

[0116]

[0117] Regarding the variable as the optimized virtual controller and defining the cost function An optimized virtual controller can be obtained Furthermore, to achieve the optimized control objective, the following auxiliary cost function is constructed

[0118] where k1 > 0 is a design parameter. A fuzzy logic system is used to approximate the unknown term where represents the optimal weight vector, ε1 is the approximation error, and there is a relationship and the input vector

[0119] Since contains ideal parameter values that cannot be obtained, a neural network is used to approximate the unknown function where and represent the optimal weight vector and the approximation error respectively represents the basis function and the input vector is Meanwhile, an execution-evaluation network is used for online training to obtain the optimized virtual controller, and we can get

[0120]

[0121] where and represent respectively and estimated values of

[0122] Substitute the above approximation terms and into the fractional-order HJB equation, and the approximation error is obtained as

[0123]

[0124] To minimize the approximation error, it is necessary to satisfy Construct the Lyapunov function Then there is

[0125]

[0126] where

[0127] Subsequently, the weight learning law of the execution-evaluation neural network is designed as

[0128]

[0129] where γ a1 , γ c1 > 0 are design parameters

[0130] Substitute the above equations into Dq From S1, it can be obtained that

[0131]

[0132] Thus, it can be known that it can asymptotically converge to 0. At the same time, design the identification weight learning law and parameter adaptation law as

[0133]

[0134] where ρ1, σ1 > 0 are design parameters.

[0135] S104. Based on the coordinate transformation equation and the fractional - order Lyapunov stability theory, construct the first Lyapunov function V1, and confirm the designed virtual control law, weight learning law and parameter adaptation law through the fractional - order derivative of V1, so that the first - level subsystem tends to be stable;

[0136] Based on the coordinate transformation equation and the fractional - order Lyapunov stability theory, construct the first Lyapunov function V1, and confirm whether the designed virtual control law, weight learning law and parameter adaptation law make the subsystem tend to be stable through the fractional - order derivative of V1;

[0137] Construct the Lyapunov function as where is the parameter estimation error. At the same time, taking the fractional - order derivative of V1, we can get

[0138]

[0139] In view of and then there is

[0140]

[0141] Using the Young's inequality to process, we can get

[0142]

[0143] Substitute the above formula and the designed weight learning law into it, we can get

[0144]

[0145] where is the parameter estimation error.

[0146] S105. By repeating the calculation process of step S103 and step S104, construct the second Lyapunov function V2, and design the corresponding virtual controller Identification - execution - evaluation weight learning law and parameter adaptation law Introduce a self-triggered control strategy to reduce the communication burden and make the second-level subsystem and the entire closed-loop system tend to be stable;

[0147] For the second-level subsystem, repeat the calculation processes of steps S103 and S104, construct a second Lyapunov function V2, and design a corresponding virtual controller Identification - execution - evaluation weight learning law And parameter adaptation law On this basis, introduce a self-triggered control strategy to reduce the communication burden and make the entire closed-loop system tend to be stable.

[0148] Based on the coordinate transformation equation in step S102, we can obtain

[0149]

[0150] Define the cost function An optimized virtual controller can be obtained Furthermore, to achieve the optimized control objective, construct the following auxiliary cost function. Among them, k2 > 0 is a design parameter.

[0151] During this process, use a fuzzy logic system to approximate the unknown term Where represents the optimal weight vector, ε2 is the approximation error, and there is a relationship And the input vector

[0152] Since contains ideal parameter values that cannot be obtained, a neural network is used to approximate the unknown function Where, and represent the optimal weight vector and approximation error respectively, represents the basis function and the input vector is At the same time, use the execution - evaluation network for online training to obtain an optimized virtual controller, and we can get

[0153]

[0154] Where, and represent and estimated values of

[0155] Based on the above coordinate transformation equation and fractional-order filter, we can obtain

[0156]

[0157] Among them, and

[0158] Similarly, the design formulas of the weight learning law and the parameter adaptation law are as follows:

[0159]

[0160] Among them, γ a2 , γ c2 , ρ2, σ2 and are all positive constants.

[0161] Most of the control methods for single-machine infinite bus power systems are designed based on time-triggered strategies. The disadvantage of this control method is that the continuous update of the system control signal will cause unnecessary waste of communication resources. In order to reduce the update frequency of the control signal and relieve the communication burden, the present invention introduces a self-triggered control strategy into the design of the controller, and its triggering rule is expressed as

[0162] u(t) = ν(t k )

[0163]

[0164] Among them, β, M, l, are all positive constants, and satisfy t k , t k+1 , k ∈ Z + represents the input update time, β|u(t)| + M is the time interval between two consecutive triggers, and l represent the change rate of the control signal interval. In the time interval [t k , t k+1 )], based on the above triggering rule, |υ(t) - u(t)| < β|u(t)| + M can be obtained, and the controller u is set as where ψ1(t), ψ2(t) ∈ [-1, 1].

[0165] Since 0 < 1 + βψ1 < 1 + β, and it can be obtained that Furthermore, construct a second Lyapunov function to ensure the stability of the system:

[0166]

[0167] Derive V2 and substitute the relevant design terms to obtain

[0168]

[0169] Among them, and and δ are normal constants.

[0170] By using the Young's inequality, it can be further obtained that

[0171]

[0172] Select and and define where λ min represents the minimum eigenvalue of the matrix , then

[0173] D q V2 ≤ -aV2 + b

[0174] According to the lemma: If V(t) is a positive definite function and has the following form:

[0175] D q V(t) ≤ -m1V(t) + m2

[0176] where q ∈ (0, 1], m1 > 0, m2 > 0 are constants. Then, V(t) is convergent and bounded.

[0177] Based on the above analysis, all signals of the closed-loop system are bounded. Therefore, the control input signal u is bounded, which further ensures is bounded, and the minimum time interval between two consecutive triggers satisfies T * = t k+1 - t k > 0, effectively avoiding the Zeno behavior.

[0178] To illustrate the control effect of the method of the present invention in detail, a simulation experiment will be carried out in MATLAB next. The initial conditions of the system are set as The control parameters are set as k1 = 10, k2 = 15, σ1 = 5, σ2 = 10, γ ai = 10, γ ci = 8, β = 0.3, l = 1.5, and the unknown signal related to the attack is selected as ω = -0.3 - 0.3cos(t), and the constraint parameters are set as c1 = 0.6, c2 = 0.5.

[0179] The following results are obtained through the MATLAB simulation experiment. Figure 3 For the original state trajectory and the damaged state trajectory diagram of the single-machine infinite bus power system, it can be seen that the damaged state and It can comply with the given constraint conditions, and even in the presence of adverse false data injection attacks, the system still has good stabilization performance. Figure 4 It is the execution-evaluation weight curve graph of the single-machine infinite power system; Figure 5 It is the trajectory graph of the adaptive law of the single-machine infinite power system, and all the signals in the graph are bounded; Figure 6 It is the control input curve graph of the single-machine infinite power system, and it can be concluded that the designed control method effectively avoids the continuous update of the control signal and reduces the communication burden; Figure 7 It is the trigger interval graph of the control input of the single-machine infinite power system, and Zeno behavior is effectively avoided. From the simulation experiment results, it can be obtained that the proposed security control method for the single-machine infinite power system based on reinforcement learning can ensure that the damaged state does not violate the constraints, all signals are bounded, and the control system is stabilized under false data injection attacks and limited communication bandwidth.

[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0181] The terms "first", "second", "third", etc. in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0182] As described above, the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A single-machine infinite power system safety control method based on reinforcement learning, characterized in that: The following steps are involved: S101. Based on the fractional-order single-machine infinite power system mathematical model, an expression of the damaged system after encountering an attack signal is established, and a nonlinear function λ(t) is constructed to constrain the damaged state variables; S102. Based on adaptive backstepping control and dynamic surface control, coordinate transformation equations are constructed to decompose the single-machine infinite power system into two-level subsystems; S103, construct an equivalent auxiliary system by selecting a cost function and introducing an auxiliary variable e to obtain the fractional-order Hamilton-Jacobi-Bellman equation and the optimized virtual controller α * , using neural networks and fuzzy logic systems to process the nonlinear terms contained in the fractional-order Hamilton-Jacobi-Bellman equations, and combining reinforcement learning to design a virtual controller Identification-Execution-Evaluation Weighted Learning Law and parameter adaptation law To solve the optimization problem of the first-level subsystem containing fractional-order dynamic systems; S104. Based on the coordinate transformation equation and fractional-order Lyapunov stability theory, construct the first Lyapunov function V1, and confirm the designed virtual control law, weight learning law and parameter adaptation law through the fractional-order derivative of V1, so that the first-level subsystem tends to be stable; S105, by repeating the calculation process of step S103 and step S104, construct the second Lyapunov function V2 and design the corresponding virtual controller Identification-Execution-Evaluation Weighted Learning Law and parameter adaptation law A self-triggering control strategy is introduced to reduce the communication burden and make the second-level subsystem and the entire closed-loop system stable.

2. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: The expression of the mathematical model of the fractional-order single-machine infinite power system is as follows: Among them, D q represents the Caputo fractional differential symbol, 0<q<1; u represents the control input, and Respectively represent the relative angular velocity and angular frequency between the two generators G1 and G2; and Respectively represent the motor's rotational inertia, damping coefficient, maximum power, and mechanical power; Define system state variables Let the system output signal y = x1, then it has the following form: Where, q = 0.95, f1(x) = 0, f2(x) = -0.5x2-sin(x1) + 2.75sin(t), 3. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: The step S101 specifically includes: The damaged state of the system after the attack is Where ω(t) represents the unknown time-varying signal, and the expression of the damaged system is as follows: Among them, p1 and p2 represent non-negative time-varying functions; By constructing a nonlinear function Then the unilateral full-state constraint is realized, and its expression is as follows: in, c i is the design parameter, q1 and q2 represent non-negative time-varying functions. It should be noted that when the state variable Approaching the set constraint boundary c i When λ i →∞, that is, as long as λ is guaranteed i is bounded, The constraints are also followed.

4. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S102, the coordinate transformation equation is expressed as follows: Among them, z1 and z2 represent error variables, α1, η2 and ξ2 represent virtual control signal, filter output signal and filter error respectively; The fractional order filter expression is as follows: in, are design parameters, α1(0) and η2(0) represent their initial values.

5. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S103, the cost function is expressed as Where Ω s represents the admissible control set; the Hamilton-Jacobi-Bellman equation is Define auxiliary variable e = D q- 1 z, we can get and z = D 1-q e; rewrite the cost function as Further, the fractional-order Hamilton-Jacobi-Bellman equation can be obtained as in use The optimal controller can be obtained Because J * (D 1-q e)=J * (z), equivalently, we can get in By introducing auxiliary variables and constructing equivalent auxiliary systems, we can obtain the fractional-order Hamilton-Jacobi-Bellman equations and the optimization controller, thereby solving the optimization problem of fractional-order dynamic systems.

6. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S103 of the single-machine infinite power system safety control method based on reinforcement learning, according to the coordinate transformation equation in step S102, it can be obtained The variable Optimized Virtual Controller And define the cost function You can get an optimized virtual controller 7. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S103, in order to achieve the optimization control target, the following auxiliary cost function is constructed: in, k1>0 is the design parameter; fuzzy logic system is used to approximate the unknown term in, represents the optimal weight vector, α1 is the approximation error, and there is a relationship And the input vector because There are unattainable ideal parameter values ​​in the , and neural networks are used to approximate unknown functions Among them, W1 * and denote the optimal weight vector and approximation error respectively, represents the basis function and the input vector is Using the execution-evaluation network online training to obtain the optimized virtual controller, we can get in, and Respectively and An estimated value of Substitute the approximation term and To the fractional-order Hamilton-Jacobi-Bellman equation, the approximation error is To minimize the approximation error, Constructing Lyapunov functions Then there is in, Design execution-evaluation neural network weight learning law Among them, γ a1 ,γ c1 >0 is the design parameter; Substituting the above formula into D q S1, can be obtained From this we can see that It can converge to 0 asymptotically, and the identification weight learning law and parameter adaptive law are designed as Among them, ρ1,σ1>0 are design parameters.

8. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: The step S104 specifically includes: The expression of Lyapunov function is Taking the fractional derivative of V1, we get Given that and Then there is Using Young's inequality, we can get Substituting the above formula and the designed weight learning law, we can get in, is the parameter estimation error.

9. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S105, according to the coordinate transformation equation in step S102, we can get Define the cost function You can get an optimized virtual controller In order to achieve the optimal control goal, an auxiliary cost function is constructed in, are the design parameters; in this process, a fuzzy logic system is used to approximate the unknown terms in represents the optimal weight vector, ε2 is the approximation error, and there is a relationship And the input vector because There are unattainable ideal parameter values ​​in the , and neural networks are used to approximate unknown functions in, and denote the optimal weight vector and approximation error respectively, represents the basis function and the input vector is Using the execution-evaluation network online training to obtain the optimized virtual controller, we can get in, and Respectively and An estimated value of Based on the above coordinate transformation equation and fractional-order filter, we can get in, and Similarly, the design formulas of weight learning law and parameter adaptation law are as follows: Among them, γ a2 ,γ c2 ,ρ2,σ2 and All are normal numbers.

10. The single-machine infinite power system safety control method based on reinforcement learning according to claim 1 is characterized in that: In step S105, the triggering rule of the self-triggering control strategy is expressed as u(t)=ν(t k ) Among them, β,M,l, are all normal numbers and satisfy represents the input update time, β|u(t)|+M is the time interval between two consecutive triggers, and l represents the rate of change of the control signal interval. k ,t k+1 ), based on the above triggering rules, we can get |υ(t)-u(t)|<β|u(t)|+M, and the controller u is set to where ψ1(t),ψ2(t)∈[-1,1]. Since 0<1+βψ1<1+β, and Available Construct a second Lyapunov function that ensures the stability of the system: Taking the derivative of V2 and substituting into the relevant design terms, we can get in, and and δ are positive constants; Using Young's inequality, we can get Select and And define Among them, λ min Representation Matrix The minimum eigenvalue of So IS q V2≤-aV2+b According to the lemma: Assume that V(t) is a positive definite function and has the following form: D q V(t)≤-m1V(t)+m2 Among them, q∈(0,1],m1>0,m2>0 are constants, then V(t) is convergent and bounded, all signals of the closed-loop system are bounded, so the control input signal u is bounded, thus ensuring The minimum time interval between two consecutive triggers satisfies T * =t k+1 -t k >0, effectively avoiding Zeno behavior.