A method for transient stability control of power system
By constructing a generative dynamic model and using PPO algorithm for strategy learning, the transient stability problem of high-permeability photovoltaic grid-connected power system is solved, and the stability and efficiency of the control strategy are improved.
Patent Information
- Application Number
- CN202510612494.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-13
AI Technical Summary
High permeability photovoltaic grid connection increases the difficulty of transient synchronization stability of the power system. The information delay of the wide-area measurement system affects the wide-area damping control effect, resulting in the failure of the transient stability control strategy.
Build a generative dynamic model, use PPO algorithm to learn strategy, generate control strategies and execute them, and optimize control strategies through reinforcement learning technology to reduce the impact of time lag.
It improves the stability and efficiency of the control strategy, reduces the impact of time lag on the control strategy, and achieves more stable transient stable control.
Smart Images

Figure CN120127706B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power system transient stability control, and in particular to a power system transient stability control method. Background Art
[0002] As the proportion of photovoltaic grid-connected power generation connected to the power grid continues to increase, the mutual influence between photovoltaic power generation units and the power system becomes more and more complex. The high penetration rate of photovoltaic grid connection will bring greater difficulties and challenges to the stable operation of the power system, especially the transient synchronization stability of the system.
[0003] On the other hand, the Wide Area Measurement System (WAMS), a crucial monitoring and control tool in power systems, aims to acquire and analyze power network status information in real time. It integrates and analyzes data from various measurement points to support system state estimation, fault detection, and control strategy formulation. However, due to the time lag in the transmission of WAMS information, errors in system state estimation are exacerbated, affecting the effectiveness of the wide-area damping control performed by the WAMS and phase measurement unit, rendering transient stability control strategies ineffective. Summary of the Invention
[0004] To solve the above technical problems existing in the prior art, the present invention provides a power system transient stability control method for an AC / DC system containing a high proportion of grid-connected photovoltaic power generation units, comprising the following steps:
[0005] Step S1: Constructing a generative dynamics model , and generate a dynamics model based on the hidden state and action at each time step Make predictions;
[0006] Step S2: Using the PPO algorithm, the predicted generative dynamics model Perform strategy learning, generate control strategies and execute them.
[0007] Preferably, step S1 specifically includes:
[0008] Step S101: Constructing a generative dynamics model , the formula is as follows:
[0009] ;
[0010] Where, for t The generation constraints of the moment, for t The generated result at the moment, for tThe hidden state of the moment, for t Moment-Generative Dynamics Model The action applied, R is the dimension symbol, b is the preset number of observations, d a is the length of the action tensor dimension, for t The terminator of time, for t The estimated benefit at the time, for t +1 moment generative dynamics model The observed values of the system state variables, d is a 1-dimensional Boolean termination tensor;
[0011] The action is the generator's PSS (Primary Synchronization Signa) signal; the benefit estimate is the absolute value of the difference between the generator speed and the rated speed; the observations include the generator voltage, the voltage amplitude of each node, the voltage phase of each node, the generator speed, and the time derivative of the generator speed.
[0012] Step S102: Calculate the specific distribution of the hidden state of the recognition network and calculate the specific distribution of the prior network hidden state. The formula is as follows:
[0013] ;
[0014] ;
[0015] Where, for t +1 moment identifies the hidden state of the network, for t +1 time prior to the hidden state of the network; To identify the network, To identify network parameters of the network; is the prior network, is the network parameter of the prior network; “~” is the variable distribution symbol;
[0016] Step 103: Hidden state and hidden state Input generation network Reconstructed, the formula is as follows:
[0017] ;
[0018] ;
[0019] Where, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, To generate the network parameters of the network;
[0020] Step 104: With the goal of maximizing the conditional likelihood logarithm, fit the reconstruction result and the original generated result. The formula is as follows:
[0021] ;
[0022] Where, x To generate constraints, y To generate the results, x and y are all input variables, z is the hidden state, To identify the probability distribution of network representation, for expectations;
[0023] Step S105: In hidden state Replace hidden state , predictive generative dynamics model .
[0024] Furthermore, after step S105, the method further includes:
[0025] Step S106: Constructing a generative dynamics model The objective function to be maintained , the formula is as follows:
[0026] ;
[0027] Where, Represents the KL divergence operator, The number of sampling times preset for the reconstruction process; is the sampling sequence number, for Moment Subsampled hidden states;
[0028] Step S107: Stacking generative modules of multiple time steps to optimize the objective function , the formula is as follows:
[0029] ;
[0030] Where, is the total number of time steps, is the time step number, for t Time has come t + T The hidden state of the prior network at any moment, for t Time has come t + T Moment-Generative Dynamics Model The action applied, for t Time has come t + T The estimated benefit at the time, for t Time has come t + T The terminator of time, for t +1 time to t +1+ T Moment-Generative Dynamics Model The observed values of the system state variables, for The hidden state of the prior network at any moment, for Moment-Generative Dynamics Model The action applied, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model Observed values of system state variables;
[0031] Step S108: Maximize the optimized objective function , by optimizing the network parameters , network parameters and network parameters Optimizing Generative Dynamics Models .
[0032] Preferably, step S2 specifically includes:
[0033] Step S201: Initialize the starting vector. The formula is as follows:
[0034] ;
[0035] Where, is the hidden state at the initial moment, Generate a dynamic model for the initial moment The action applied, is the estimated benefit at the initial moment, is the terminator of the initial moment;
[0036] Step S202: Identify the network , fault information and the starting vector to estimate the starting prior state ;
[0037] Step S203: Calculate the predicted transient stability sequence length using the following formula:
[0038] ;
[0039] Where, n is the trajectory length to be calculated, that is, the number of control cycles from the occurrence of the fault to the control signal reaching the controlled device; is the sequence length under transient control, is the transmission delay of fault information, is the computation time of the transient control strategy, Generative dynamics model The training time step of
[0040] Step S204: Calculate the real-time prior state, the formula is as follows:
[0041] ;
[0042] ;
[0043] Where, for Identify the hidden state of the network at all times, for Identify the hidden state of the network at all times, for Moment-Generative Dynamics Model The action applied, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for The number of steps predicted at each moment, for +1 moment identifies the hidden state of the network, for Identify the hidden state of the network at all times, for Identify the hidden state of the network at all times, for Moment-Generative Dynamics Model the action applied;
[0044] Step S205: Calculate the estimated benefit value and terminator. The formula is as follows:
[0045] ;
[0046] Where, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, H is the sequence difference length, for Moment-Generative Dynamics Model Observed values of system state variables;
[0047] Step S206: Based on the benefit estimate and the terminator, a control strategy is generated according to the PPO algorithm. The formula is as follows:
[0048] ;
[0049] Where, For the strategy network, are the current parameters of the policy network, is the timing difference error, is the discount factor, For the value network, is the current parameter of the value network;
[0050] Execute control strategies.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] The technical solution provided by the present invention can run the PPO algorithm on a dynamic generative dynamics model, search and optimize the optimal control trajectory and corresponding control strategy through reinforcement learning technology, reduce the impact of time lag, and improve the stability and efficiency of the control strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a schematic diagram of the principle structure of a conditional variational autoencoder in one embodiment of the present application.
[0054] Figure 2 This is a structural diagram of a production dynamics model under the PPO algorithm in one embodiment of the present application. DETAILED DESCRIPTION
[0055] The technical solution provided by the present invention will be further elaborated in combination with the embodiments and drawings.
[0056] Example 1:
[0057] In this embodiment, a Markov decision process is introduced into the generative dynamics system based on the conditional variational autoencoder to construct a generative dynamics model including multiple generative modules. , the formula is as follows:
[0058] ;
[0059] Where, for t The generation constraints of the moment, for t The generated result at the moment, for t The hidden state of the moment, for t Moment-Generative Dynamics Model The action applied, R is the dimension symbol, b is the preset number of observations, d a is the length of the action tensor dimension, for t The terminator of time, for t The estimated benefit at the time, for t +1 moment generative dynamics model The observed value of the system state variable. The principle structure of the conditional variational autoencoder is as follows Figure 1 As shown in the figure, X is the generation constraint, Y is the generation result, and Z is the corresponding hidden state.
[0060] The action is the PSS signal of the generator; the benefit estimate is the absolute value of the difference between the generator speed and the rated speed (per unit); the observation values include the generator voltage, the voltage amplitude of each node, the voltage phase of each node, the generator speed, and the time derivative of the generator speed.
[0061] Calculate the specific distribution of the hidden state of the recognition network and the specific distribution of the hidden state of the prior network. The formula is as follows:
[0062] ;
[0063] ;
[0064] Where, for t +1 moment identifies the hidden state of the network, for t +1 time prior to the hidden state of the network; To identify the network, To identify network parameters of the network; is the prior network, is the network parameter of the prior network; “~” is the variable distribution symbol, the symbol before it is the variable, and the symbol after it is the specific distribution.
[0065] Hidden state and hidden state Input generation network Reconstructed, the formula is as follows:
[0066] ;
[0067] ;
[0068] Where, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, is the network parameters for generating the network.
[0069] The goal is to maximize the logarithm of the conditional likelihood so that the reconstruction result is sufficiently similar to the original generated result. The formula is as follows:
[0070] ;
[0071] Where, x To generate constraints, y To generate the results, x and y are all input variables, z is the hidden state, To identify the probability distribution of network representation, for expectations;
[0072] In hidden state Replace hidden state , predictive generative dynamics model , the formula is as follows:
[0073] .
[0074] By utilizing t Hidden state at the moment Z t and actions performed a t Benefit evaluation value r t , terminator d t and subsequent states Make predictions to achieve state prediction of generative dynamics models, combined with the empirical lower bound of conditional variational autoencoders , and obtain the generative dynamics model The objective function to be maintained :
[0075] ;
[0076] ;
[0077] Where, Represents the KL divergence operator, The number of sampling times preset for the reconstruction process; c is the sampling sequence number, For the c Subsampled hidden states;
[0078] Stacking generative modules across multiple time steps to optimize the objective function .
[0079] To ensure the consistency of the physical meaning at each time step, the generative modules of multiple time steps are stacked. The formula is as follows:
[0080] ;
[0081] Where, T is the total number of time steps, is the time step number, for Time has come The hidden state of the prior network at any moment, for Time has come Moment-Generative Dynamics Model The action applied, for Time has come The estimated benefit at the time, for Time has come The terminator of time, for Time has come Moment-Generative Dynamics Model The observed values of the system state variables, for The hidden state of the prior network at any moment, for Moment-Generative Dynamics Model The action applied, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for Moment Subsampled hidden states.
[0082] Maximize the optimized objective function , by optimizing the network parameters , network parameters and network parameters Obtain an optimized generative dynamics model suitable for multi-step continuous state estimation , that is, the optimal generative dynamics model .
[0083] In the embodiment, for the transient stability control strategy problem studied by the present invention, the above content realizes the parameterization of the power system dynamics model. By running the reinforcement learning algorithm on the parameterized model, the transient stability control strategy and its evaluation results can be further parameterized, thereby directly optimizing the control strategy. Since the generative dynamics model is an approximation of the target dynamics system, it is necessary to continuously update its parameters by using the interaction data collected in the target system to ensure the long-term accuracy and stability of the model. This will cause the expression of the observed values of the system state variables in the latent space to change as the model parameters change. In order to adapt to such changes as quickly as possible and improve the utilization efficiency of samples, the present invention adopts the PPO algorithm (Proximal Policy Optimization) to perform policy learning in the latent space of the generative dynamics model, parameterize the transient stability control strategy and its evaluation results, and directly optimize the strategy.
[0084] The PPO algorithm is an efficient reinforcement learning algorithm that has been widely used in various control tasks. Its idea is to limit the instability of learning strategy updates by modifying the loss function during strategy updates, mainly through two methods: penalty and clipping. Figure 2 As shown, this embodiment adopts the shearing method to cut the length of MDP (Markov Decision Process) sampling trajectory , the PPO algorithm minimizes the following objective function J To learn:
[0085] ;
[0086] ;
[0087] Where, For the strategy network, are the current parameters of the policy network, E Express expectations, The sampling trajectory The network parameters of the corresponding policy network are, For input When the output is a t The probability density of V For the value network, are the parameters of the value network, is the generalized advantage function is the timing differential error, for The timing difference error at the moment, for tThe timing difference error at the moment, is the length of the MDP trajectory, is the discount factor, is the smoothing factor, is the shear coefficient, is the shear function, x 1 is the shear variable, 、 are the maximum and minimum values of the shear variable.
[0088] In this embodiment, the starting vector is first initialized, and the formula is as follows:
[0089] ;
[0090] Where, Z 0 is the hidden state at the initial moment, a 0 is the generative dynamics model at the initial moment The action applied, r 0 is the estimated benefit value at the initial moment, d 0 is the terminator of the initial moment.
[0091] By identifying the network , fault information and the starting vector to estimate the starting prior state , and infer the sequence length, the expression is:
[0092] ;
[0093] Where, n is the trajectory length to be calculated, that is, the number of control cycles from the occurrence of the fault to the control signal reaching the controlled device; h n is the sequence length under transient control, t a is the transmission delay of fault information (i.e. the information delay of fault information being sent to the dispatch control system after a fault occurs), t c is the calculation time of the transient control strategy (i.e. the time consumed by the control center to calculate the control strategy), Generative dynamics model The training time step.
[0094] Calculate the real-time prior state, the formula is as follows:
[0095] ;
[0096] ;
[0097] Where, for Identify the hidden state of the network at all times, for Identify the hidden state of the network at all times, for Moment-Generative Dynamics Model The action applied, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for The number of steps predicted at each moment, for +1 moment identifies the hidden state of the network, for Identify the hidden state of the network at all times, for Identify the hidden state of the network at all times, for Moment-Generative Dynamics Model Action applied.
[0098] Calculate the benefit estimate and terminator using the following formula:
[0099] ;
[0100] Where, for The estimated benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, H is the sequence difference length, for Moment-Generative Dynamics Model The observed values of the system state variables, for Moment-Generative Dynamics Model Action applied.
[0101] The system control trajectory is estimated and the corresponding MDP trajectory is obtained. The formula is as follows:
[0102] ;
[0103] Since the control sequence is given in advance and is not sufficient to support policy learning, the transient stability control strategy can be learned through the PPO algorithm in the latent space through the policy network and hidden state. The formula is as follows:
[0104] ;
[0105] Where, For the strategy network, are the current parameters of the policy network, is the timing difference error, is the discount factor, For the value network, The current parameters of the network are set to ; the control strategy is executed.
[0106] Figure 2 middle, a t,h for t + h The actions imposed by the moment-by-moment generative dynamics model.
[0107] In combination with the above embodiments, it can be concluded that the technical solution provided by the present invention can run the PPO algorithm on a dynamic generative dynamics model, search and optimize the optimal control trajectory and corresponding control strategy through reinforcement learning technology, reduce the impact of time lag, and improve the stability and efficiency of the control strategy.
[0108] Furthermore, by constructing an objective function to optimize the network parameters of the generative network, prior network, and identification network, the optimal generative dynamics model is obtained, which can further mitigate the impact of time lag and improve the stability and efficiency of the control strategy.
Claims
1. A method for controlling transient stability of a power system, characterized in that: The following steps are involved: Step S1: Construct a generative dynamics model D and make predictions based on the hidden state and action at each time step. Step S2: Using the PPO algorithm, perform strategy learning on the predicted generative dynamics model D, generate a control strategy, and execute it; Step S1 specifically includes: Step S101: Construct a generative dynamics model D, the formula is as follows: Where x t is the generation constraint at time t, y t is the generated result at time t, z t is the hidden state at time t, a t is the action imposed by the generative dynamics model D at time t, R is the dimension symbol, b is the preset number of observations, d a is the length of the action tensor dimension, d t is the terminator at time t, r t is the estimated benefit at time t, O t+1 is the observed value of the state variable of the generative dynamics model D system at time t+1, and d is a 1-dimensional Boolean termination tensor; The action is the PSS signal of the generator; the benefit estimate is the absolute value of the difference between the generator speed and the rated speed; the observations include the generator voltage, the voltage amplitude of each node, the voltage phase of each node, the generator speed, and the time derivative of the generator speed. Step S102: Calculate the specific distribution of the hidden state of the recognition network and calculate the specific distribution of the prior network hidden state. The formula is as follows: z t+1 ~p θ (z t+1 |[z t ,a t ]); Where, is the hidden state of the recognition network at time t+1, z t+1 is the hidden state of the prior network at time t+1; To identify the network, To identify the network parameters of the network; p θ is the prior network, θ is the network parameter of the prior network; "~" is the variable distribution symbol; Step 103: Hidden state and hidden state z t+1 Input generation network p ψ Reconstructed, the formula is as follows: Where, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, is the hidden state z t+1 The reconstructed benefit estimate, is the hidden state z t+1 The reconstructed terminator, is the hidden state z t+1 The reconstructed state observation variable value, ψ is the network parameter of the generated network; Step 104: With the goal of maximizing the conditional likelihood logarithm, fit the reconstruction result and the original generated result. The formula is as follows: In the formula, x is the generation constraint, y is the generation result, x and y are both input variables, z is the hidden state, To identify the probability distribution of network representation, for expectations; Step S105: Using hidden state z t+1 Replace hidden state Predictive generative dynamics model D.
2. A method for controlling transient stability of a power system according to claim 1, characterized in that: After step S105, the method further includes: Step S106: Construct the objective function that needs to be maintained by the generative dynamics model D The formula is as follows: In the formula, KL(·||·) represents the KL divergence operator, C is the number of samplings preset in the reconstruction process; c is the sampling order number, z τ+1 (c) is the hidden state of the cth sample at time τ+1; Step S107: Stacking generative modules of multiple time steps to optimize the objective function The formula is as follows: Where T is the total number of time steps, τ is the time step number, z t:t+T is the hidden state of the prior network from time t to time t+T, a t:t+T is the action imposed by the generative dynamics model D from time t to time t+T, r t:t+T is the estimated benefit from time t to time t+T, d t∶t+T is the terminator from time t to time t+T, o t+1∶t+1+T is the observed value of the state variable of the generative dynamics model D system from time t+1 to time t+1+T, z τ+1 is the hidden state of the prior network at time τ+1, a τ is the action imposed by the generative dynamics model D at time τ, r τ is the estimated benefit at time τ, d τ is the terminator at time τ, o τ+1 is the observed value of the state variables of the generative dynamics model D system at time τ+1; Step S108: Maximize the optimized objective function By optimizing the network parameters θ, network parameters and network parameters ψ to optimize the generative dynamics model D.
3. A method for controlling transient stability of a power system according to any one of claims 1 or 2, characterized in that: Step S2 specifically includes: Step S201: Initialize the starting vector. The formula is as follows: [z0,a0,r0,d0]=[0,0,0,0]; Where z0 is the hidden state at the initial moment, a0 is the action imposed by the generative dynamics model D at the initial moment, r0 is the estimated benefit at the initial moment, and d0 is the terminator at the initial moment; Step S202: Identify the network Fault information t0 and the starting vector to estimate the starting prior state Step S203: Calculate the predicted transient stability sequence length using the following formula: Where n is the length of the trajectory to be calculated, that is, the number of control cycles from the occurrence of the fault to the control signal reaching the controlled device; h n is the sequence length under transient control, t a is the transmission delay of fault information, t c is the computational time of the transient control strategy, Δt is the training time step of the generative dynamics model D; Step S204: Calculate the real-time prior state, the formula is as follows: Where, is the hidden state of the network at time τ, is the hidden state of the recognition network at time τ-1, a τ-1 is the action imposed by the generative dynamics model D at time τ-1, r τ-1 is the estimated benefit at time τ-1, d τ-1 is the terminator at time τ-1, o τ is the observed value of the state variable of the generative dynamics model D system at time τ, h is the number of steps predicted at time τ, z τ,h+1 is the hidden state of the recognition network at time τ+h+1, z τ,h is the hidden state of the recognition network at time τ+h, z τ,0 is the hidden state of the recognition network at time τ+0, a τ,h is the action imposed by the generative dynamics model D at time τ+h; Step S205: Calculate the estimated benefit value and terminator. The formula is as follows: Where r τ,h is the estimated benefit at time τ+h, d τ,h is the terminator at time τ+h, o τ,h is the observed value of the state variable of the generative dynamics model D system at time τ+h, H is the sequence difference length, o τ+1,h is the observed value of the state variable of the generative dynamics model D system at time τ+1+h; Step S206: Based on the benefit estimate and the terminator, a control strategy is generated according to the PPO algorithm. The formula is as follows: Where π is the policy network, ω π is the current parameter of the policy network, δ τ,h is the timing difference error, γ is the discount factor, V is the value network, ω v is the current parameter of the value network; Execute control strategies.
Citation Information
Patent Citations
Offline reinforcement learning method and device based on time reversal symmetry
CN119337960A
Interplanetary orbit transfer method based on hidden state and reinforcement learning
CN119861572A