Power system transient stability control method
By running the PPO algorithm on the generative dynamic model of the power system, the control strategy is optimized, and the impact of high permeability photovoltaic grid connection on the transient stability of the power system and the time delay of the wide-area measurement system is solved, and a more stable and efficient control strategy is achieved.
Patent Information
- Application Number
- CN202510612494.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
High permeability photovoltaic grid connection poses challenges to the transient synchronization stability of the power system, and the time delay phenomenon of the wide-area measurement system affects the effectiveness of the transient stability control strategy.
The generative dynamics model is adopted and the PPO algorithm is run on it, and the optimal control trajectory and corresponding control strategies are searched and optimized through reinforcement learning technology to reduce the impact of time lag and improve the stability and efficiency of the control strategy.
By running the PPO algorithm on a dynamic generative dynamic model, the control strategy can be optimized, the impact of time lag, and the stability and efficiency of the control strategy can be improved, solving the challenge of photovoltaic grid connection to the transient stability of the power system.
Smart Images

Figure CN120127706A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transient stability control of power systems, and in particular to a method for transient stability control of power systems. Background Art
[0002] With the increasing proportion of photovoltaic grid-connected power generation accessing the power grid, the mutual influence between photovoltaic power generation units and power systems has become more and more complex. High-penetration photovoltaic grid connection will bring greater difficulties and challenges to the stable operation of power systems, especially the transient synchronous stability of the system.
[0003] On the other hand, as an important monitoring and control tool in power systems, the Wide Area Measurement System (WAMS) aims to obtain and analyze the state information of the power network in real time, integrate and analyze the data from different measurement points, and provide support for system state estimation, fault detection, and control strategy formulation. However, due to the time-delay phenomenon in the transmission of WAMS information, it further exacerbates the system state estimation deviation, which will affect the wide-area damping control effect of the WAMS and the phase measurement unit, resulting in the failure of the transient stability control strategy. Summary of the Invention
[0004] To solve the above technical problems existing in the prior art, the present invention provides a method for transient stability control of a power system for an AC-DC system containing a high proportion of network-forming photovoltaic power generation units, including the following steps: Step S1: Construct a generative dynamics model , and predict the generative dynamics model according to the hidden state and action at each time step ; Step S2: Adopt the PPO algorithm to perform policy learning on the predicted generative dynamics model to generate a control strategy and execute it.
[0005] Preferably, step S1 specifically includes: Step S101: Construct a generative dynamics model , and the formula is as follows: ; In the formula, is the generation constraint condition at t time, is the generation result at t time, is the hidden state at t time, is the action applied by the generative dynamics model t at time,R is a dimension symbol, b is the preset number of observed values, d a is the length of the action tensor dimension, is t the terminator at time is t the benefit estimate value at time is t the generative dynamics model at time + 1 and the observed value of the system state variable, d is a 1D boolean termination tensor; Among them, the action is the PSS signal (Primary Synchronization Signa, main synchronization signal) of the generator; the benefit estimate value is the absolute value of the difference between the generator speed and the rated speed; the observed values include the generator voltage, the voltage amplitudes of each node, the voltage phases of each node, the generator speed, and the derivative of the generator speed with respect to time; Step S102: Calculate the specific distribution of the hidden state of the recognition network and the specific distribution of the hidden state of the prior network. The formula is as follows: ; ; In the formula, is t the hidden state of the recognition network at time is t the hidden state of the prior network at time is the recognition network, is the network parameter of the recognition network; is the prior network, is the network parameter of the prior network; "~" is the variable distribution symbol; Step 103: Input the hidden state and the hidden state into the generative network for reconstruction. The formula is as follows: ; ; In the formula, is the benefit estimate value after reconstructing the hidden state , is the terminator after reconstructing the hidden state , is the value of the state observation variable after reconstructing the hidden state , is the benefit estimate value after reconstructing the hidden state , is the hidden state The reconstructed terminator is the hidden state The value of the state observation variable after reconstruction is the network parameter of the generation network; Step 104: With the goal of maximizing the conditional likelihood logarithm, fit the reconstructed result and the original generation result. The formula is as follows: ; In the formula, x is the generation constraint condition, y is the generation result, x and y are both input variables, z is the hidden state, is the probability distribution represented by the recognition network, is the expectation of; Step S105: Use the hidden state to replace the hidden state and predict the generative dynamics model .
[0006] Furthermore, after Step S105, it also includes: Step S106: Construct the objective function that the generative dynamics model needs to maintain. The formula is as follows: ; In the formula, represents the KL divergence operator, is the preset number of sampling times in the reconstruction process; is the sampling sequence number, is at time the hidden state of the th sampling; Step S107: Stack the generative modules of multiple time steps and optimize the objective function ; In the formula, is the total number of time steps, is the time step sequence number, is t from time t + T to the hidden state of the prior network at time is t from time t + T to the action applied by the generative dynamics model at time ist Time to t + T Benefit estimate value at the time, is t Time to t + T Terminator at the time, is t Time to +1 t +1 + T Generative dynamics model at the time Observation value of the system state variable, is Hidden state of the prior network at the time, is Generative dynamics model at the time Applied action, is Benefit estimate value at the time, is Terminator at the time, is Generative dynamics model at the time Observation value of the system state variable; Step S108: Maximize the optimized objective function , by optimizing the network parameters , network parameter and network parameter to optimize the generative dynamics model .
[0007] Preferably, step S2 specifically includes: Step S201: Initialize the starting vector, the formula is as follows: ; In the formula, is the hidden state at the initial time, is the action applied by the generative dynamics model at the initial time , is the benefit estimate value at the initial time, is the terminator at the initial time; Step S202: Estimate the starting prior state through the recognition network , fault information and the starting vector; Step S203: Calculate the predicted transient stability sequence length, the formula is as follows: ; In the formula, nis the trajectory length to be calculated, i.e., the number of control cycles from the occurrence of the fault to the arrival of the control signal at the controlled device; is the sequence length under transient control, is the transmission delay of the fault information, is the computational time consumption of the transient control strategy, is the generative dynamics model of the training time step; Step S204: Calculate the real-time prior state, and the formula is as follows: ; ; In the formula, is the hidden state of the recognition network at time is the hidden state of the recognition network at time is the action applied by the generative dynamics model at time , is the benefit estimate value at time is the terminator at time is the observation value of the system state variable of the generative dynamics model at time , is the predicted number of steps at time is the hidden state of the recognition network at time +1, is the hidden state of the recognition network at time is the hidden state of the recognition network at time is the action applied by the generative dynamics model at time ; Step S205: Calculate the benefit estimate value and the terminator, and the formula is as follows: ; In the formula, is the benefit estimate value at time is the terminator at time is the observation value of the system state variable of the generative dynamics model at time , H is the sequence difference length, is the generative dynamics model at time Observed values of system state variables Step S206: Based on the benefit estimate value and the terminator, generate a control strategy according to the PPO algorithm. The formula is as follows: ; In the formula, is the policy network, is the current parameter of the policy network, is the temporal difference error, is the discount factor, is the value network, is the current parameter of the value network; Execute the control strategy.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: The technical solution provided by the present invention can run the PPO algorithm on a dynamic generative kinetic model, search and optimize the optimal control trajectory and the corresponding control strategy through reinforcement learning technology, reduce the influence of time delay, and improve the stability and efficiency of the control strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is the principle structure diagram of the conditional variational autoencoder in an embodiment of the present application.
[0010] Figure 2 is the structure diagram of the generative kinetic model under the PPO algorithm in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] The technical solution provided by the present invention will be further elaborated in detail in combination with the embodiments and the drawings.
[0012] Embodiment 1: In this embodiment, a Markov decision process is introduced into the generative kinetic system based on the conditional variational autoencoder, and a generative kinetic model including multiple generative modules is constructed , and the formula is as follows: ; In the formula, is t the generation constraint condition at time is t the generation result at time is t the hidden state at time is t the action applied by the generative kinetic model at time R is the dimension symbol, b is the preset number of observed valuesd a is the length of the dimension of the action tensor, is t the terminator at time is t the benefit estimate value at time is t the observed value of the system state variable of the generative dynamics model at time +1. The principle structure of the conditional variational autoencoder is as Figure 1 shown. In the figure, X is the generation constraint condition, Y is the generation result, and Z is the corresponding hidden state.
[0013] Among them, the action is the PSS signal of the generator; the benefit estimate value is the absolute value of the difference between the generator speed and the rated speed (per-unit value); the observed values include the generator voltage, the voltage amplitudes of each node, the voltage phases of each node, the generator speed, and the derivative of the generator speed with respect to time.
[0014] Calculate the specific distribution of the hidden state of the recognition network and the specific distribution of the hidden state of the prior network. The formula is as follows: ; ; In the formula, is t the hidden state of the recognition network at time is t the hidden state of the prior network at time is the recognition network, is the network parameter of the recognition network; is the prior network, is the network parameter of the prior network; "~" is the variable distribution symbol, the symbol before it is the variable, and the symbol after it is the specific distribution.
[0015] Input the hidden state and the hidden state into the generative network for reconstruction. The formula is as follows: ; ; In the formula, is the benefit estimate value after reconstructing the hidden state , is the terminator after reconstructing the hidden state , is the value of the state observation variable after reconstructing the hidden state , is the benefit estimate value after reconstructing the hidden state , is the hidden state The reconstructed terminator, is the hidden state The value of the state observation variable after reconstruction, is the network parameter of the generation network.
[0016] Aiming to maximize the conditional likelihood logarithm, making the reconstruction result similar enough to the original generation result, the formula is as follows: ; In the formula, x is the generation constraint condition, y is the generation result, x and y are both input variables, z is the hidden state, is the probability distribution represented by the recognition network, is the expectation of;
[0017] Using the hidden state to replace the hidden state , predict the generative dynamics model , the formula is as follows: .
[0018] By using t the hidden state at time Z t and the executed action a t to predict the benefit evaluation value r t , the terminator d t and the successor state to achieve the state prediction of the generative dynamics model, combined with the empirical lower bound of the conditional variational autoencoder, obtain the objective function that the generative dynamics model needs to maintain: ; ; In the formula, represents the KL divergence calculation operator, is the preset number of sampling times for the reconstruction process; c is the sampling sequence number, is the c hidden state of the Stack multiple time steps of the generative module to optimize the objective function .
[0019] To ensure the consistency of the physical meaning at each time step, the generative modules of multiple time steps are stacked. The formula is as follows: ; In the formula, T is the total number of time steps, is the time step number, for Time has come The hidden state of the prior network at any given moment, for Time has come Moment-Generative Dynamics Model The action applied, for Time has come The estimated value of benefit at the time, for Time has come The terminator of time, for Time has come Moment-Generative Dynamics Model The observed values of the system state variables, for The hidden state of the prior network at any given moment, for Moment-Generative Dynamics Model The action applied, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for Moment Subsampled hidden states.
[0020] Maximize the optimized objective function , by optimizing the network parameters , network parameters and network parameters Obtain an optimized generative dynamics model suitable for multi-step continuous state estimation , that is, the optimal generative dynamics model .
[0021] In the embodiment, for the transient stability control strategy problem studied by the present invention, the above content realizes the parameterization of the power system dynamics model. By running the reinforcement learning algorithm on the parameterized model, the transient stability control strategy and its evaluation results can be further parameterized, thereby directly optimizing the control strategy. Since the generative dynamics model is an approximation of the target dynamics system, it is necessary to continuously update its parameters by using the interaction data collected in the target system to ensure the lasting accuracy and stability of the model. This will cause the expression of the observed values of the system state variables in the latent space to change with the change of the model parameters. In order to adapt to this change as quickly as possible and improve the utilization efficiency of the samples, the present invention adopts the PPO algorithm (Proximal Policy Optimization) to perform strategy learning in the latent space of the generative dynamics model, parameterize the transient stability control strategy and its evaluation results, and directly optimize the strategy.
[0022] The PPO algorithm is an efficient reinforcement learning algorithm that has been widely used in various control tasks. Its idea is to limit the instability of learning strategy updates by modifying the loss function when updating the strategy, mainly through two methods: penalty and clipping. Figure 2 As shown, this embodiment adopts the shearing method to cut the length MDP (Markov Decision Process) sampling trajectory , the PPO algorithm minimizes the following objective function J To learn: ; ; In the formula, For the strategy network, are the current parameters of the policy network, E Express expectations, The sampling trajectory The network parameters of the corresponding policy network are For input When the output is a t The probability density of V For the value network, are the parameters of the value network, is the generalized advantage function is the timing difference error, for The timing difference error at time, for t The timing difference error at time, is the length of the MDP trajectory, is the discount factor, is the smoothing factor, is the shear coefficient, is the shear function, x 1 is the shear variable, , are the maximum and minimum values of the shear variable.
[0023] In this embodiment, the starting vector is first initialized, and the formula is as follows: ; In the formula, Z 0 is the hidden state at the initial moment, a 0 Generate a dynamic model for the initial moment The action applied, r 0 is the estimated benefit at the initial moment, d 0 The terminator for the initial moment.
[0024] By identifying the network , fault information and the starting vector to estimate the starting prior state , and infer the sequence length, the expression is: ; In the formula, n is the trajectory length to be calculated, that is, the number of control cycles from the occurrence of the fault to the control signal reaching the controlled device; h n is the sequence length under transient control, t a is the transmission delay of fault information (i.e. the information delay of fault information being sent to the dispatching control system after a fault occurs), t c is the calculation time of the transient control strategy (i.e. the time consumed by the control center to calculate the control strategy), Generative dynamics model The training time step.
[0025] Calculate the real-time prior state, the formula is as follows: ; ; In the formula, for Always identify the hidden state of the network, for Always identify the hidden state of the network, for Moment-Generative Dynamics Model The action applied, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for The number of steps predicted at each moment, for +1 moment identifies the hidden state of the network, for Always identify the hidden state of the network, for Always identify the hidden state of the network, for Moment-Generative Dynamics Model The action applied.
[0026] The benefit estimate and terminator are calculated as follows: ; In the formula, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, H is the sequence difference length, for Moment-Generative Dynamics Model The observed values of the system state variables, for Moment-Generative Dynamics Model The action applied.
[0027] The estimation of the system control trajectory is realized, and the corresponding MDP trajectory is obtained. The formula is as follows: ; Since the control sequence is given in advance and is not sufficient to support policy learning, the transient stability control strategy can be learned through the policy network and hidden state through the PPO algorithm in the hidden space. The formula is as follows: ; In the formula, For the strategy network, are the current parameters of the policy network, is the timing difference error, is the discount factor, For the value network, The current parameters of the network are set; the control strategy is executed.
[0028] Figure 2 middle, a t,h for t + h The actions imposed by the moment-by-moment generative dynamics model.
[0029] In combination with the above embodiments, it can be concluded that the technical solution provided by the present invention can run the PPO algorithm on a dynamic generative dynamics model, search and optimize the optimal control trajectory and the corresponding control strategy through reinforcement learning technology, reduce the impact of time lag, and improve the stability and efficiency of the control strategy.
[0030] Furthermore, by constructing the objective function to optimize the network parameters of the generative network, the prior network, and the discriminative network, the optimal generative dynamics model can be obtained, which can further alleviate the impact of time lag and improve the stability and efficiency of the control strategy.
Claims
1. A method for transient stability control of a power system, characterized in that: The following steps are involved: Step S1: Constructing a generative dynamics model , and generate a dynamics model based on the hidden state and action at each time step Make predictions; Step S2: Using the PPO algorithm, the predicted generative dynamics model Perform strategy learning, generate control strategies and execute them.
2. A method for controlling transient stability of a power system according to claim 1, characterized in that: Step S1 specifically includes: Step S101: Constructing a generative dynamics model , the formula is as follows: ; In the formula, for t The generation constraints at the moment, for t The generated result at the moment, for t The hidden state of the moment, for t Moment-Generative Dynamics Model The action applied, R is the dimension symbol, b is the preset number of observations, d a is the length of the action tensor dimension, for t The terminator of time, for t The estimated value of benefit at the time, for t +1 Moment Generative Dynamics Model The observed values of the system state variables, d is a 1-dimensional Boolean termination tensor; Among them, the action is the PSS signal of the generator; the benefit estimation value is the absolute value of the difference between the generator speed and the rated speed; the observation values include the generator voltage, the voltage amplitude of each node, the voltage phase of each node, the generator speed and the derivative of the generator speed with respect to time; Step S102: Calculate the specific distribution of the hidden state of the recognition network and calculate the specific distribution of the hidden state of the prior network. The formula is as follows: ; ; In the formula, for t +1 moment identifies the hidden state of the network, for t +1 time prior network hidden state; To identify the network, To identify network parameters of the network; is the prior network, is the network parameter of the prior network; "~" is the variable distribution symbol; Step 103: Hidden state and hidden state Input generation network Reconstructed, the formula is as follows: ; ; In the formula, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, Hidden state The reconstructed benefit estimate, Hidden state The reconstructed terminator, Hidden state The reconstructed state observation variable value, To generate the network parameters of the network; Step 104: With the goal of maximizing the conditional likelihood logarithm, fit the reconstruction result and the original generation result. The formula is as follows: ; In the formula, To generate constraints, To generate the results, and are all input variables, is the hidden state, To identify the probability distribution of network representations, for expectations; Step S105: In hidden state Replace hidden state , predictive generative dynamics model .
3. A method for controlling transient stability of a power system according to claim 2, characterized in that: After step S105, the method further includes: Step S106: Constructing a generative dynamics model The objective function to be maintained , the formula is as follows: ; In the formula, represents the KL divergence operator, The number of sampling times preset for the reconstruction process; is the sampling order number, for Moment Subsampled hidden states; Step S107: stacking generative modules of multiple time steps to optimize the objective function , the formula is as follows: ; In the formula, is the total number of time steps, is the time step number, for t Time has come t + T The hidden state of the prior network at any given moment, for t Time has come t + T Moment-Generative Dynamics Model The action applied, for t Time has come t + T The estimated value of benefit at the time, for t Time has come t + T The terminator of time, for t +1 Time to t +1+ T Moment-Generative Dynamics Model The observed values of the system state variables, for The hidden state of the prior network at any given moment, for Moment-Generative Dynamics Model The action applied, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model Observed values of system state variables; Step S108: Maximizing the optimized objective function , by optimizing the network parameters , network parameters and network parameters Optimizing Generative Dynamics Models .
4. A method for controlling transient stability of a power system according to any one of claims 2 or 3, characterized in that: Step S2 specifically includes: Step S201: Initialize the starting vector, the formula is as follows: ; In the formula, is the hidden state at the initial moment, Generate a dynamic model for the initial moment The action applied, is the estimated benefit at the initial moment, is the terminator of the initial moment; Step S202: Identify the network , fault information and the starting vector to estimate the starting prior state ; Step S203: Calculate the predicted transient stability sequence length, the formula is as follows: ; In the formula, n is the trajectory length to be calculated, that is, the number of control cycles from the occurrence of the fault to the control signal reaching the controlled device; is the sequence length under transient control, is the transmission delay of fault information, is the computation time of transient control strategy, Generative dynamics model The training time step of Step S204: Calculate the real-time prior state, the formula is as follows: ; ; In the formula, for Always identify the hidden state of the network, for Always identify the hidden state of the network, for Moment-Generative Dynamics Model The action applied, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, for The number of steps predicted at each moment, for +1 moment identifies the hidden state of the network, for Always identify the hidden state of the network, for Always identify the hidden state of the network, for Moment-Generative Dynamics Model the action applied; Step S205: Calculate the benefit estimate and terminator, the formula is as follows: ; In the formula, for The estimated value of benefit at the time, for The terminator of time, for Moment-Generative Dynamics Model The observed values of the system state variables, H is the sequence difference length, for Moment-Generative Dynamics Model Observed values of system state variables; Step S206: Based on the benefit estimate and the terminator, a control strategy is generated according to the PPO algorithm. The formula is as follows: ; In the formula, For the strategy network, are the current parameters of the policy network, is the timing difference error, is the discount factor, For the value network, is the current parameter of the value network; Execute control strategies.
Citation Information
Patent Citations
Wind power interval prediction method, system and storage medium
CN112365033A
Offline reinforcement learning method and device based on time reversal symmetry
CN119337960A
Interplanetary orbit transfer method based on hidden state and reinforcement learning
CN119861572A
Computational framework for modeling of physical process
US20200202057A1
Machine learning entity validation performance reporting
US20230368077A1