Power system emergency control method and system fusing DRL and STGNN

By integrating the DRL and STGNN methods, the problem that convolutional neural networks are difficult to capture the topological structure in the power system is solved, the real-time and accuracy of emergency control of the power system are improved, and the understanding of the power grid status and the dynamic adaptability of the control strategy are enhanced.

CN120810558APending Publication Date: 2025-10-17GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510700302.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Convolutional neural networks have difficulty in effectively capturing the topological information of the power grid when processing the complex topological structure of the power system, resulting in insufficient real-time performance of emergency control of the power system.

Method used

By integrating DRL and STGNN, the international standard example IEEE-39 model is constructed, and the spatiotemporal graph neural network is used to extract the topological structure characteristics of the power system. Deep reinforcement learning is then combined to make emergency control decisions, including feature extraction, action optimization and model training.

Benefits of technology

It improves the real-time and accuracy of emergency control of the power system, enhances the perception and generalization capabilities of the power grid status, and can better handle dynamically changing control problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120810558A_ABST
    Figure CN120810558A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power system emergency control method and system fusing DRL and STGNN. The method comprises the steps that an international standard example IEEE-39 model is built, electric power data are collected, and a data set is built; and performing feature extraction on the data in the data set, inputting the obtained features into the improved DDPG model fused with the DRL, and outputting a generator tripping action vector. And optimizing the generator tripping action vector, and updating network parameters in the model. According to the power system emergency control method and system fusing the DRL and the STGNN provided by the invention, the perception capability of the transient response time-space change trend is effectively improved through the time-space diagram neural network, and the real-time performance and the accuracy of the model are improved. On the basis, through the DRL model, learning of a complex control strategy is achieved, effective actions can be rapidly obtained, and the effectiveness of an emergency control scheme is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power system steady-state control, in particular to a power system emergency control method and system fusing DRL and STGNN. BACKGROUND

[0002] The operation uncertainty of power systems is significantly increased in the face of growing power grid interconnection and renewable energy proportion, leading to more complex and variable transient stability characteristics. This brings new challenges to the transient stability analysis and emergency control of power systems. Traditional methods such as stability defense system based on off-line generation and on-line update of stability control strategy table have the problem of insufficient real-time when facing rapidly changing grid states, which may delay emergency control measures such as generator tripping and load shedding, thereby affecting the stable operation of the system. Therefore, there is an urgent need for a real-time and accurate transient stability emergency control method to provide technical support for online transient stability defense control.

[0003] In the existing scheme, there is a model and method for extracting features of power systems using convolutional neural networks (CNN). This method is based on a CNN comprehensive model and a power system transient stability evaluation method based on steady-state characteristic quantities. Through the use of the results of multiple CNN models and the assistance of time domain simulation, the accuracy and reliability of power system transient stability evaluation are effectively improved. This method is especially suitable for handling sample imbalance problems and can overcome the neuron failure problem caused by activation functions.

[0004] However, there are obvious limitations to using convolutional neural networks for feature extraction in power systems. Convolutional neural network feature extraction is mainly used to process data with a clear grid structure, such as images, videos, and audio. It captures local features and hierarchical features in data through convolutional layers. The limitation of convolutional neural networks when processing power system data is that it assumes that the input data has a fixed grid structure, which is insufficient when processing topological information in power systems. Power systems are complex networks composed of nodes (such as generators, loads, and substations) and edges (such as transmission lines), with specific connection relationships between these nodes and edges. CNNs are difficult to directly capture this complex topological structure information because they are not designed to handle unordered and irregular graph data. SUMMARY

[0005] In view of the above problems, the present application is proposed.

[0006] Therefore, the technical problem solved by the present application is that there is obvious limitation in using convolutional neural network to extract features in the power system. CNN is difficult to directly capture the complex topological structure information in the power system because it is not designed to process unordered and irregular graph data.

[0007] To solve the above technical problems, the present application provides the following technical solutions: a power system emergency control method fusing DRL and STGNN, comprising: building an international standard example IEEE-39 model, collecting power data, and constructing a data set.

[0008] The data in the data set is subjected to feature extraction, and the obtained features are input into an improved DDPG model fusing DRL, and an action vector of generator tripping is output.

[0009] The action vector of generator tripping is optimized, and the network parameters in the model are updated.

[0010] As a preferred scheme of the power system emergency control method fusing DRL and STGNN, wherein: the building of the international standard example IEEE-39 model, the collection of power data, and the construction of the data set include that the IEEE-39 system contains 10 generators and 46 branches, the system reference capacity and the reference voltage are 100 MVA and 345 kV respectively, the generators adopt a 3-order model, the injection power of the power generation, the load and the reactive power compensation device is divided into 10 layers at 70% to 140% and ±10% random fluctuation is added, 3 types of data sets with different topological structures are set, including maintaining the original topological structure to construct data set A, randomly disconnecting one branch to construct data set B containing 34 topological structures, and randomly disconnecting two branches to construct data set C containing 550 topological structures.

[0011] The three-phase short-circuit fault of the simulation of all network branches is traversed, the fault lasts for 0.1s-0.2s, the operation data between the time after the fault and the time when the system stability coefficient η is less than or equal to 0.2 is collected and preprocessed.

[0012] The system stability coefficient is expressed as:

[0013]

[0014] Where, Δδ max represents the maximum value of the power angle difference of any two generators.

[0015] The collected data is mapped to a graph data structure, wherein the bus corresponds to the node in the graph, and the branch corresponds to the edge connecting the nodes. The correlation matrix A is constructed to describe the topology structure of the power grid containing N buses, and the element is expressed as:

[0016]

[0017] For the transient response characteristics of each node, a structure preserving energy function derived based on Kirchhoff's current law is adopted to construct, ensuring that the model can effectively capture and utilize transient information. The structure preserving energy function V(x) obtains the transient energy of the system by integrating the response information along the trajectory, expressed as:

[0018]

[0019] where P g,i represents the active power of the generator, P L,i represents the active power of the load, Q g,i represents the reactive power of the generator, Q L,i represents the reactive power of the load, V i represents the voltage amplitude of bus i, θ i represents the phase angle of bus i, P ij represents the active power of the branch between bus i and bus j, Q ij represents the reactive power of the branch between bus i and bus j, D i represents the generator damping, ω i represents the rotor angular velocity.

[0020] As a preferred scheme of the power system emergency control method combining DRL and STGNN, the structure preserving energy function includes constructing high-order features containing transient energy using system measurement responses, and constructing features to approximate the transient energy in the structure preserving energy function V(x).

[0021] The change values D Vi , D θi of the voltage Vi and the phase angle θi of bus i at each time step are used to approximate the dlnV i , dθ i terms, which are multiplied by the active power P g,i , the reactive power Q g,i of the generator, and the active power P L,i , the reactive power Q L,i of the load, to approximate the P g,i dθ i , Q g,i dlnV i , P L,i dθ i , Q L,i dlnV i terms in V(x).

[0022] Let the generator power angle at bus i be δ i , and the square term of the difference of δ i with respect to time is used to approximate the square term of the generator rotor angular velocity The dissipation term in V(x) is approximated by multiplying the generator damping

[0023] The transient energy features related to branches are stored at both ends of the bus. The bus i and bus j corresponding to each branch are determined according to the initial power direction of the branch. The branch reactive power Q ij , Q ji is multiplied by the voltage change value D Vi , D Vj at the source end bus i and bus j of the reactive power, respectively, to approximate the Q ij dlnV i , Q ji dlnV j term in V(x). The two features are stored in the node features corresponding to bus i and bus j, respectively. The branch active power P ij is multiplied by the change value D θij of the branch phase angle difference θ ij at each time step to approximate the P ij dθ ij term, and the term is stored in the node features corresponding to bus i and bus j, respectively.

[0024] The generated feature of the nth node at time t is represented as:

[0025]

[0026] wherein, represents the voltage amplitude of bus n at time t, represents the phase angle value of bus n at time t, represents the active power injected here, represents the reactive power injected here, represents the generator power angle value at bus n, and the feature value is set to 0 when the bus is not connected to a generator, and the remaining features are approximate quantities related to transient energy, and are represented as:

[0027]

[0028]

[0029] wherein, N(n) represents the adjacent bus set of bus n, represents the active power of the branch between bus n and bus j at time t, represents the phase angle difference of the branch between bus n and bus j at time t, represents the reactive power of the corresponding branch, represents the voltage amplitude difference of the corresponding branch.

[0030] In order to improve the perception efficiency of the model on the transient information, the element characteristic sequence in the whole transient process is modeled as a space-time graph data according to the topological correlation relationship, G={G(1), G(t), G(T)}, wherein G(t)=(X(t), A(t)), X(t) represents an element characteristic matrix at the tth moment, and A(t) represents an adjacency matrix at the tth moment.

[0031] As a preferred scheme of the power system emergency control method combining DRL and STGNN, wherein: the feature extraction on the data in the data set comprises: using a space-time graph neural network to extract features from the data set. The graph neural network is divided into three layers of input layer, middle layer and output layer. The input layer extracts features from the time domain data of the nodes. The input layer is composed of a convolutional neural network layer and an activation function layer. The input features are encoded by convolution operation and activation function, and are expressed as:

[0032] h=σ CNN (W CNN *x+b CNN )

[0033] wherein, W CNN represents the parameters of the convolution kernel, x represents the input feature matrix, h represents the output encoded feature matrix, b CNN represents the bias parameter matrix, W CNN *x represents the discrete convolution operation of x and W CNN , and σCNN(·) represents the activation function.

[0034] The time domain encoder is constructed by using the CNN convolution layer to extract the time sequence variation characteristics of the transient response characteristics in the space-time graph data, and the node time domain encoded features are generated, and are expressed as:

[0035]

[0036] wherein, f(·) represents a nonlinear activation function, represents a convolution operator, X∈R N×T×d represents the d-dimensional features of N nodes at T moments in the input space-time graph data, K i represents the convolution kernel group for the i-dimensional feature X i , which contains s convolution kernels in total, and the size of each convolution kernel is T*1. represents a matrix splicing operator, H CN ∈R N×k represents the k-dimensional time domain encoded features of N nodes output by the time domain encoder, and k=s*d.

[0037] For the node feature matrix X in the spatio-temporal graph data input to the model, firstly, X is convolved and nonlinearly mapped in the time dimension by the above formula, the convolution results in each feature dimension are spliced, and the time domain coding features HCN of each node are obtained. Secondly, the coding features HCN of each node and the node adjacency matrix A jointly constitute the graph data GCN=(HCN, A).

[0038] Subsequently, the graph data GCN is input into the intermediate layer as input, and the layer is a graph attention network GAT layer. The GAT layer mines the spatial distribution characteristics of the time domain coding features of GCN, and the attention mechanism operation form is represented as:

[0039]

[0040] Wherein, K i represents the key-value pair of the original feature information, V i represents the key-value pair of the original feature information, Q represents prior information, s(·) represents a correlation calculation function, and P represents updated features.

[0041] The GAT introduces a multi-head attention mechanism, obtains the output features h GT,i of the GAT layer node i by stacking K independent attention calculation units.

[0042]

[0043] Wherein, σ(·) represents an activation function, and a ij,k represents the normalized attention coefficient of the kth attention unit, and Wk represents the normalized attention coefficient and the corresponding parameter matrix of the kth attention unit.

[0044] After passing through the three-layer graph attention network layer, the obtained aggregated features are transmitted to the GAT output layer. The GAT output layer aggregates the information of the neighborhood nodes to improve the accuracy of the model. Finally, the output layer outputs the feature vector Z t that fuses time sequence and spatial information.

[0045] As a preferred scheme of the power system emergency control method fusing DRL and STGNN, wherein: the obtained features are input into the improved DDPG model fusing DRL, and the output machine action vector includes: the obtained features are input into the improved DDPG model for model training, and the DDPG model is composed of an action network and a value network. During the model training process, the batch size is set to 64, and the learning rates of the action network and the evaluation network are set to 0.001 and 0.005.

[0046] The DDPG model network is initialized, and the feature vector Z tThe action network inputted into the DDPG is used for action coding to output an action a t . The action a t includes a generator tripping amount a t and a load tripping amount b t .

[0047] The feature vector Z t is inputted again for state coding and the action vector a t is inputted into the value network of the DDPG model to obtain the Q value of the action. After the action a t is optimized, it is fed back to the transient stability simulation system.

[0048] As a preferred scheme of the power system emergency control method combining DRL and STGNN, the optimization of the generator tripping action vector includes that the action network outputs a generator tripping vector a t of the system in the current state, the generator tripping amount is distributed to all generators in the system in the execution link, and the action space of the agent is represented as:

[0049] A={P G1 ,P G2 ,…,P Gm ,P L1 ,P L2 ,…,P Ll}

[0050] Wherein, P Gi represents the tripping amount of the i-th generator, i=1, 2, …, m. P Lk represents the tripping amount of the k-th load, k=1, 2, …, l.

[0051] The generator tripping amount is distributed according to the size of the relative speed of the generator, and the load tripping amount is distributed according to the importance of the load.

[0052] The action of the agent is improved by introducing the knowledge of the power system, and the action amount outputted by the action network is distributed according to the size of the relative speed of the generator and the importance of the load.

[0053] As a preferred scheme of the power system emergency control method combining DRL and STGNN, the network parameters in the model are updated, including that the running state of the power system is adjusted according to the optimized action a t , the transient stability judgment flag of the power system is obtained through transient stability simulation, and the reward value of the action a t is calculated based on the transient stability judgment flag using the reward function.

[0054] The reward function is represented as:

[0055]

[0056] where T stable represents a set of stable states of the system, s t represents a set of agent states, R ST represents a large positive number, representing the reward of the system reaching stability after control. UST represents a large negative number, representing the penalty of instability after control.

[0057] In the control process, the short-term reward r t represents the transient stability coefficient, used to represent the transient stability of the system, and is represented as:

[0058]

[0059] where Δδ max represents the relative angle difference between any generator and the reference generator.

[0060] According to the reward value and the Q value, the network of the DDPG model is updated, the action network is μ(s∣θ μ ), the value network is the Q network, and the long-term return is represented as:

[0061] y i = r i + γQ'(s i+1 , μ'(s i+1 ∣θ μ′ )∣θ Q′ )

[0062] where γ represents the discount factor.

[0063] The loss function in the model training process is represented as:

[0064]

[0065] The action network μ(s∣θ μ ) is updated as follows:

[0066]

[0067] To improve the stability of the model training link, the target Q network Q' and the target action network μ' are used for training, and the parameter update method is represented as:

[0068]

[0069] where τ represents the update rate.

[0070] An emergency control system for power systems that combines DRL and STGNN, characterized by comprising,

[0071] A data collection module is used to build an international standard example IEEE-39 model, collect power data, and construct a data set.

[0072] A feature processing module is used to extract features from the data in the data set, and the obtained features are input into an improved DDPG model of the fusion DRL to output a machine tripping action vector.

[0073] An updating module is used to optimize the machine tripping action vector and update network parameters in the model.

[0074] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0075] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the method.

[0076] The beneficial effects of the present application are as follows: 1. Spatio-temporal feature fusion: CNN is good at capturing local features, but when dealing with data such as power grid system with complex topological structure, it cannot effectively consider the global structure information. The graph neural network not only retains the ability of CNN to extract time sequence features, but also extracts the topological structure features of the entire network through GAT, realizes the deep understanding of the topological structure of the power grid, and makes the model better perceive the state information of the power grid system.

[0077] 2. Dynamic strategy learning:

[0078] The DRL is introduced in the design scheme of the present application, so that the model can learn a series of control strategies, not just static feature mapping. In the power system, emergency control often needs to respond to the dynamic changes of the system state, and the sequential decision-making ability of DRL meets this demand, can better handle the control problem changing with time, and form a complex control strategy.

[0079] 3. Improved generalization ability:

[0080] The graph neural network enhances the generalization ability of the model by processing graph structure data. When facing an unknown power grid structure or operating state, the model can still make reasonable prediction and control by using the learned topological and time sequence feature relationship.

[0081] In summary, the method proposed in the present application combines spatio-temporal graph neural network (STGNN) and DRL, not only improves the understanding of the dynamic behavior of the power system, but also improves the real-time performance and accuracy of the emergency control, which is better than the local feature extraction method using only CNN. BRIEF DESCRIPTION OF DRAWINGS

[0082] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0083] Figure 1 The overall flowchart of the power system emergency control method combining DRL and STGNN provided for the first embodiment of the present application.

[0084] Figure 2 The CNN and GAT-based spatiotemporal graph neural network structure diagram of the power system emergency control method combining DRL and STGNN provided for the first embodiment of the present application.

[0085] Figure 3 The improved DDPG model structure diagram of the fusion spatiotemporal graph neural network of the power system emergency control method combining DRL and STGNN provided for the first embodiment of the present application.

[0086] Figure 4 The action space mapping relationship diagram of the power system emergency control method and system combining DRL and STGNN provided for the first embodiment of the present application.

[0087] Figure 5 The reward function curve diagram of the training process of the power system emergency control method and system combining DRL and STGNN provided for the second embodiment of the present application. DETAILED DESCRIPTION

[0088] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0089] Embodiment 1, refer to Figures 1-4 For an embodiment of the present application, a power system emergency control method combining DRL and STGNN is provided, which comprises:

[0090] S1: Build an international standard example IEEE-39 model, collect power data, and construct a data set.

[0091] It should be noted that the international standard example IEEE-39 model is built. The simulation platform is a portable computer with Intel Core i58300H 2.30GHz CPU and 16GB memory. The example IEEE-39 model is realized based on MATLAB R2018b and PSASP software environment. The IEEE-39 system contains 10 generators, 46 branches, and the system reference capacity and reference voltage are 100 MVA and 345 kV respectively. The generators adopt a 3-order model, and the injection power of the power generation, load and reactive power compensation device is divided into 10 layers between 70% and 140% and added with ±10% random fluctuation. Three types of data sets with different topological structures are set: data set A is constructed by maintaining the original topological structure, data set B is constructed by randomly breaking one branch to form 34 kinds of topological structures, and data set C is constructed by randomly breaking two branches to form 550 kinds of topological structures. Under the above operating mode, the three-phase short-circuit fault of the whole network branch is traversed, the fault lasts for 0.1s-0.2s, and the operating data between the fault and the time when the system stability coefficient η≤0.2 is collected and preprocessed.

[0092] The system stability coefficient is expressed as:

[0093]

[0094] where Δδ max represents the maximum value of the power angle difference of any two generators.

[0095] The collected data is mapped to a graph data structure, where the bus corresponds to the node in the graph, and the branch corresponds to the edge connecting the nodes. The correlation matrix A is constructed to describe the topology structure of the power grid containing N buses, and the element is expressed as:

[0096]

[0097] For the transient response characteristics of each node, a structure-preserving energy function derived based on Kirchhoff's current law is used to construct, which ensures that the model can effectively capture and utilize transient information.

[0098] It should be noted that for the transient response characteristics of each node, a structure-preserving energy function derived based on Kirchhoff's current law is used to construct, which ensures that the model can effectively capture and utilize transient information. The structure-preserving energy function V(x) is obtained by integrating the response information along the trajectory to obtain the transient energy of the system, which is expressed as:

[0099]

[0100] where P g,i represents the active power of the generator, P L,i represents the active power of the load, and Q g,iQ represents the reactive power of the generator L,i V represents the reactive power of the load i Vi represents the voltage amplitude of bus i i θi represents the phase angle of bus i ij Pi,j represents the active power of the branch between bus i and bus j ij Qi,j represents the reactive power of the branch between bus i and bus j i ω represents the generator damping i ω represents the rotor angular velocity

[0101] High-order features containing transient energy are constructed using system measurement responses, and the features are constructed to approximate the transient energy in the structure-preserving energy function V(x).

[0102] The change value D Vi of the voltage Vi and the phase angle θi of bus i at each time step θi approximates the dlnV i and dθ i terms, which are multiplied by the active power P g,i and the reactive power Q g,i of the generator and the active power P L,i and the reactive power Q L,i of the load, respectively, to approximate the P g,i dθ i , Q g,i dlnV i , P L,i dθ i , Q L,i dlnV i terms in V(x).

[0103] Let the generator power angle at bus i be δ i , and the square term of the difference of δ i with respect to time is used to approximate the square term of the rotor angular velocity of the generator , multiplied by the generator damping, to approximate the dissipation term in V(x).

[0104] The transient energy features related to the branch are stored at the two-end buses, and the corresponding bus i and bus j of each branch are determined according to the initial power direction of the branch. The reactive power Q ij of the branch is multiplied by the voltage change value D ji at the source bus i and bus j of the reactive power, respectively, to approximate the Q Vi dlnV Vj terms in V(x). ij dlnV i , Q ji dlnV jItem, store these two features into the node features corresponding to bus i and bus j respectively, and store the branch active power P ij Phase angle difference θ with the branch ij The change value D at each time step θij Multiply, approximate description P ij dθ ij Item, which is stored in the node features corresponding to bus i and bus j respectively.

[0105] The generated feature representation of the nth node at time t is:

[0106]

[0107] in, represents the voltage amplitude of bus n at time t, represents the phase angle value of bus n at time t, represents the active power injected here, represents the reactive power injected here, It represents the generator power angle value at bus n. When there is no generator connected to the bus, this characteristic value is set to 0. The other characteristics are approximate quantities related to transient energy, expressed as:

[0108]

[0109]

[0110] Where N(n) represents the set of adjacent buses of bus n, represents the active power of the branch between busbar n and busbar j at time t, It represents the phase angle difference between the branches of busbar n and busbar j at time t, represents the reactive power of the corresponding branch, Indicates the voltage amplitude difference of the corresponding branch.

[0111] In order to improve the model's perception efficiency of transient information, the component feature sequences in all transient processes are modeled as spatiotemporal graph data according to the topological association relationship, G = {G(1),…,G(t),…,G(T)}, where G(t) = (X(t),A(t)), X(t) represents the component feature matrix at the tth moment, and A(t) represents the adjacency matrix at the tth moment.

[0112] S2: Extract features from the data in the dataset, input the obtained features into the improved DDPG model integrated with DRL, and output the machine cutting action vector.

[0113] The spatiotemporal graph neural network is used to extract features from the dataset. Figure 2The graph neural network is divided into three layers of input layer, middle layer and output layer. The time domain data of the node is extracted by the input layer, and the input layer is composed of a convolutional neural network layer and an activation function layer. The input features are encoded by convolution operation and activation function, represented as:

[0114] h = σ CNN (W CNN *x + b CNN )

[0115] wherein W CNN represents the parameters of the convolution kernel, x represents the input feature matrix, h represents the output encoded feature matrix, b CNN represents the bias parameter matrix, W CNN *x represents the discrete convolution operation of x and W CNN , and σCNN(·) represents the activation function.

[0116] The time domain encoder is constructed by using the CNN convolution layer to extract the time sequence change characteristics of the transient response features in the space-time graph data, and the node time domain encoding features are generated, represented as:

[0117]

[0118] wherein f(·) represents a nonlinear activation function, represents a convolution operator, X ∈ R N×T×d represents the d-dimensional features of N nodes at T time points in the input space-time graph data, K i represents the convolution kernel group for the i-th dimensional feature X i , which contains s convolution kernels in total, and the size of each convolution kernel is T × 1. represents a matrix concatenation operator, H CN ∈ R N×k represents the k-dimensional time domain encoding features of N nodes output by the time domain encoder, k = s × d.

[0119] For the node feature matrix X in the input space-time graph data of the model, first, X is convolved and nonlinearly mapped according to the above formula according to the time dimension, the convolution results in each feature dimension are spliced, and then the time domain encoding features HCN of each node are obtained. Secondly, the encoding features HCN of each node and the node adjacency matrix A are jointly formed into graph data GCN = (HCN, A).

[0120] Subsequently, the graph data GCN is input into the middle layer as input, and the middle layer is a graph attention network GAT layer. The GAT layer mines the spatial distribution characteristics of the time domain encoding features of GCN, and the attention mechanism operation form is represented as:

[0121]

[0122] where K i represents the key-value pair of the original feature information, V i represents the key-value pair of the original feature information, Q represents the prior information, s(·) represents the correlation calculation function, and P represents the updated feature.

[0123] GAT introduces a multi-head attention mechanism, which obtains the output feature h GT,i of GAT layer node i by stacking K independent attention calculation units.

[0124]

[0125] where σ(·) represents the activation function, α ij,k represents the normalized attention coefficient of the kth attention unit, and Wk represents the normalized attention coefficient and the corresponding parameter matrix of the kth attention unit.

[0126] It should be noted that this mechanism can better handle complex relationships and patterns by processing multiple attention heads in parallel, thereby improving the model's generalization ability on unseen data. Each attention head focuses on different parts of the input data, which helps to reduce redundant information and improve the utilization of key information. In this way, the model can more efficiently capture the diversity and complexity of the input data, thereby improving the overall performance of the model.

[0127] Furthermore, after passing through the three-layer graph attention network layer, we input the obtained aggregated feature representation into the GAT output layer. The GAT output layer aims to more effectively aggregate the information of neighboring nodes and improve the accuracy of the model. It not only integrates information from the multi-head attention mechanism, but also optimizes the node feature representation to improve the performance of the model. This is crucial for ensuring that the model can understand and process graph data from a global perspective, and helps to improve the overall performance of the model, including improving accuracy and generalization ability, and speeding up the training process by reducing redundant information and optimizing feature representation. Finally, the output layer outputs a feature vector Z t that integrates temporal and spatial information.

[0128] The obtained features are input into the improved DDPG model for model training, as shown in Figure 3 . The DDPG model consists of an action network and a value network. During model training, the batch size is set to 64, and the learning rates of the action network and the evaluation network are set to 0.001 and 0.005, respectively.

[0129] The DDPG model network is initialized, and the feature vector Z t extracted by the spatio-temporal graph neural network is input into the action network of the DDPG, and the action encoding is performed to output the action a t . The action at including the generator tripping amount a t and the load tripping amount b t .

[0130] The feature vector Z t is input into the state encoding and the action vector a t is input into the value network of the DDPG model to obtain the Q value of the action. After action optimization of the action a t , it is fed back to the transient stability simulation system.

[0131] S3: Optimize the generator tripping action vector and update the network parameters in the model.

[0132] The action network outputs the generator tripping vector a t of the system in the current state, which is distributed to all generators in the system in the execution link, and the action space of the agent is represented as:

[0133] A = {P G1 , P G2 , …, P Gm , P L1 , P L2 , …, P Ll}

[0134] where P Gi represents the tripping amount of the i-th generator, i = 1, 2, …, m. P Lk represents the tripping amount of the k-th load, k = 1, 2, …, l.

[0135] The generator tripping amount is sorted according to the relative speed of the generator, and the load tripping amount is sorted according to the importance of the load.

[0136] It should be noted that in order to improve the performance of the model, the power system knowledge is introduced to improve the action of the agent, and the action amount output by the action network is sorted according to the relative speed of the generator and the importance of the load, as shown in Figure 4 Based on this method, the model can more accurately trip the generator with a larger relative speed, avoid tripping important loads, and more easily make the system return to stability and protect the supply of important loads.

[0137] According to the optimized action a t , the operating state of the power system is adjusted, the transient stability judgment flag of the power system is obtained through transient stability simulation, and the reward value of the current action a t is calculated based on the transient stability judgment flag using the reward function.

[0138] The reward function is represented as:

[0139]

[0140] wherein T stable represents a set of stable states of the system, s t represents a set of states of the agent, R ST represents a large positive number, representing the reward of the system reaching stability after control. UST represents a large negative number, representing the penalty of instability after control.

[0141] In the control process, the short-term reward r t represents a transient stability coefficient, used to represent the transient stability of the system, and is represented as:

[0142]

[0143] wherein Δδ max represents the relative angle difference between any generator and the reference generator.

[0144] According to the reward value and the Q value, the network of the DDPG model is updated, the action network is μ(s∣θ μ ), the value network is the Q network, and the long-term return is represented as:

[0145] y i = r i + γQ'(s i+1 , μ'(s i+1 ∣ θ μ′ )∣ θ Q′ )

[0146] wherein γ represents a discount factor.

[0147] The loss function in the model training process is represented as:

[0148]

[0149] The action network μ(s∣θ μ ) is updated in parameters:

[0150]

[0151] To improve the stability of the model training link, the target Q network Q' and the target action network μ' are used for training, and the parameter update method is represented as:

[0152]

[0153] wherein τ represents an update rate.

[0154] In the above embodiments, a power system emergency control system fusing DRL and STGNN is further included, specifically:

[0155] A data collection module is used to build an international standard example IEEE-39 model, collect power data, and construct a data set.

[0156] A feature processing module is used to extract features from the data in the data set, and the obtained features are input into an improved DDPG model of the fusion DRL to output a machine tripping action vector.

[0157] An updating module is used to optimize the machine tripping action vector and update network parameters in the model.

[0158] The computer device can be a server. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store data cluster data of a power monitoring system. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection.

[0159] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without limitation. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0160] As shown in Embodiment 2, Figure 5 An embodiment of the present application provides a power system emergency control method and system fusing DRL and STGNN, in order to verify the beneficial effects of the present application, scientific demonstration is carried out through simulation experiment.

[0161] The detailed steps of the embodiments have been given as examples in the technical solutions, and the following are the result images and effectiveness analysis:

[0162] The training cycle is set to 100000, Figure 5The change of the reward function during the training of the agent is shown. The training is first through about 9000 scenes of exploration process, and then through about 7000 scenes to reach convergence. Due to the consideration of the influence of random time delay, the agent selects more exploration actions in the initial training, and the internal strategy network is trained synchronously, so the reward function shows a certain range of fluctuations, and as the training process converges, the reward function also shows a convergent trend, indicating that a stable control strategy is finally learned. In the agent model training, a relatively stable control strategy can be obtained by calling less than 16000 scenes of simulation, and good convergence is also achieved in offline training.

[0163] Effectiveness analysis:

[0164] Under a certain operating mode, a three-phase short-circuit-to-ground fault occurs at the bus B16 side of the branch B16-B21 for 0.1s, and the fault branch is removed after 0.1s of fault duration. Based on the transient stability emergency control model, the current cut-off amount is 32% of the total power generation of the system, and this part of the power generation is distributed to G33, G34, G35 and G36 with the largest relative speed, and the generator power angle and bus voltage curves after emergency cut-off control are obtained, and the power angle is the value relative to the power angle of No. 32 generator. The transient stability control strategy formed by the model can effectively realize the transient stability emergency control of the power system.

[0165] If the same amount of cut-off is delayed to 0.7s to perform emergency cut-off operation, through time domain simulation of measured data, the system fails to recover stability, reflecting the importance of early implementation of emergency cut-off control measures, therefore, the emergency cut-off control scheme based on deep reinforcement learning in the present application has feasibility and timeliness.

[0166] In summary, in order to realize fast and accurate transient stability analysis and emergency control, the present application proposes a transient stability emergency control method based on improved deep reinforcement learning. In order to fully tap the time and space change trend of transient response, a multi-dimensional feature containing transient potential energy and other information is constructed, and a deep reinforcement learning model is improved based on a space-time graph neural network, and on this basis, an emergency control model is constructed, and power grid knowledge is integrated into the emergency control decision scheme, reducing the exploration of invalid decisions, and improving the performance of the model. The effectiveness of the proposed method is verified in the IEEE-39 system.

[0167] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A power system emergency control method integrating DRL and STGNN, characterized in that: include: Build the international standard IEEE-39 model, collect power data, and construct a data set; Perform feature extraction on the data in the dataset, input the obtained features into the improved DDPG model integrated with DRL, and output the machine cutting action vector; The cutting action vector is optimized and the network parameters in the model are updated.

2. The power system emergency control method integrating DRL and STGNN according to claim 1 is characterized in that: The IEEE-39 model of the international standard calculation example is constructed to collect power data and construct a data set. The IEEE-39 system includes 10 generators and 46 branches. The system base capacity and base voltage are 100 MVA and 345 kV respectively. The generator adopts a third-order model. The injected power of the power generation, load and reactive compensation device is evenly divided into 10 layers between 70% and 140% and a random fluctuation of ±10% is added. Three types of data sets with different topological structures are set, including maintaining the original topological structure to construct data set A, randomly disconnecting one branch to construct data set B containing 34 topological structures, and randomly disconnecting two branches to construct data set C containing 550 topological structures. Traverse and simulate a three-phase short-circuit fault in all branches of the network. The fault lasts for 0.1s-0.2s. Collect and pre-process the operating data from the time of the fault to the time when the system stability coefficient η≤0.

2. The system stability factor is expressed as: Among them, Δδ max Indicates the maximum value of the power angle difference between any two generators; The collected data is mapped into a graph data structure, where buses correspond to nodes in the graph and branches correspond to edges connecting nodes. An association matrix A is constructed to describe the topology of the power grid containing N buses, and the elements are represented as: For the transient response characteristics of each node, a structure-preserving energy function derived from Kirchhoff's current law is used to construct the model to ensure that the model can effectively capture and utilize transient information; the structure-preserving energy function V(x) obtains the transient energy of the system by integrating the response information along the trajectory, which is expressed as: Among them, P g,i Indicates the active power of the generator, P L,i Indicates the active power of the load, Q g,i Indicates the reactive power of the generator, Q L,i Represents the reactive power of the load, V i represents the voltage amplitude of bus i, θ i represents the phase angle of busbar i, P ij It represents the active power of the branch between busbar i and busbar j, Q ij It represents the reactive power of the branch between busbar i and busbar j, D i represents the generator damping, ω i represents the rotor angular velocity.

3. The power system emergency control method integrating DRL and STGNN as claimed in claim 2, characterized in that: The structure preserving energy function includes constructing a high-order feature containing transient energy using a system measurement response, and constructing the feature to approximate the transient energy in the structure preserving energy function V(x); The change value D of the voltage Vi and phase angle θi of bus i at each time step is used Vi 、D θi Approximate characterization of dlnV i , dθ i Item, and compare it with the active power P of the generator g,i , reactive power Q g,i And the active power P of the load L,i , reactive power Q L,i Multiply, approximate P in V(x) g,i dθ i , Q g,i dlnV i 、P L,i dθ i , Q L,i dlnV i item; Assume the generator power angle at busbar i is δ i , using δ i The square of the time difference is used to approximate the square of the generator rotor angular velocity Approximate the dissipative term in V(x) by multiplying it with the generator damping The transient energy characteristics related to the branch are stored at the busbars at both ends. The busbars i and j corresponding to each branch are specified according to the initial power direction of the branch. The branch reactive power Q ij , Q ji The voltage change value D at the source bus i and bus j of the reactive power is respectively Vi 、D Vj Multiply, approximate Q in V(x) ij dlnV i , Q ji dlnV j Item, store these two features into the node features corresponding to bus i and bus j respectively, and store the branch active power P ij Phase angle difference θ with the branch ij The change value D at each time step θij Multiply, approximate description P ij dθ ij Item, store the item in the node features corresponding to bus i and bus j respectively; The generated feature representation of the nth node at time t is: in, represents the voltage amplitude of bus n at time t, represents the phase angle value of bus n at time t, represents the active power injected here, represents the reactive power injected here, It represents the generator power angle value at bus n. When there is no generator connected to the bus, this characteristic value is set to 0. The other characteristics are approximate quantities related to transient energy, expressed as: Where N(n) represents the set of adjacent buses of bus n, represents the active power of the branch between busbar n and busbar j at time t, It represents the phase angle difference between the branches of busbar n and busbar j at time t, represents the reactive power of the corresponding branch, Indicates the voltage amplitude difference of the corresponding branch; In order to improve the model's perception efficiency of transient information, the component feature sequences in all transient processes are modeled as spatiotemporal graph data according to the topological association relationship, G = {G(1),…,G(t),…,G(T)}, where G(t) = (X(t),A(t)), X(t) represents the component feature matrix at the tth moment, and A(t) represents the adjacency matrix at the tth moment.

4. The power system emergency control method integrating DRL and STGNN as claimed in claim 3 is characterized by: The feature extraction of the data in the dataset includes extracting features from the dataset using a spatiotemporal graph neural network. The graph neural network is divided into three layers: an input layer, an intermediate layer, and an output layer. The input layer extracts features from the time domain data of the node. The input layer is composed of a convolutional neural network layer and an activation function layer. The input features are encoded through a convolution operation and an activation function, and are expressed as: h=σ CNN (W CNN *x+b CNN ) Among them, W CNN Represents the parameters of the convolution kernel, x represents the input feature matrix, h represents the output encoding feature matrix, b CNN represents the bias parameter matrix, W CNN *x represents x and W CNN The discrete convolution operation, σCNN(·) represents the activation function; The CNN convolutional layer is used to construct a time domain encoder to extract the temporal variation characteristics of the transient response features in the spatiotemporal graph data and generate the node time domain coding features, which can be expressed as: Where f(·) represents a nonlinear activation function, represents the convolution operator, X∈R N×T×d Represents the d-dimensional features of N nodes at T moments in the input spatiotemporal graph data, K i Represents the i-th dimension feature X i The convolution kernel group contains s convolution kernels, and the size of each convolution kernel is T×1; represents the matrix concatenation operator, H CN ∈R N×k represents the k-dimensional time-domain coding features of N nodes output by the time-domain encoder, k = s × d; For the node feature matrix X in the spatiotemporal graph data input to the model, first, convolve X according to the time dimension and perform nonlinear mapping, then concatenate the convolution results under each feature dimension to obtain the time domain coding feature HCN of each node. Secondly, the coding feature HCN of each node is combined with the node adjacency matrix A to form the graph data GCN = (HCN, A). The graph data GCN is then passed as input to the middle layer, which is the graph attention network GAT layer. The GAT layer mines the spatial distribution characteristics of the temporal encoding features of the GCN. The attention mechanism operation form is expressed as: Among them, K i Represents the key-value pair of the original feature information, V i represents the key-value pair of the original feature information, Q represents the prior information, s(·) represents the correlation calculation function, and P represents the updated feature; GAT introduces a multi-head attention mechanism, which obtains the output feature h of GAT layer node i by superimposing K independent attention calculation units. GT,i : Among them, σ(·) represents the activation function, α ij,k represents the normalized attention coefficient of the k-th attention unit, Wk represents the normalized attention coefficient of the k-th attention unit and the corresponding parameter matrix; After passing through the three-layer graph attention network layer, the obtained aggregated feature representation is passed to the GAT output layer; the GAT output layer aggregates the information of the neighboring nodes to improve the accuracy of the model; finally, the output layer outputs the feature vector Z that integrates temporal and spatial information t .

5. The power system emergency control method integrating DRL and STGNN according to claim 4 is characterized in that: The obtained features are input into the improved DDPG model integrated with DRL, and the output cutting action vector includes: the obtained features are input into the improved DDPG model for model training, and the DDPG model action network and value network are composed; During model training, the batch size was set to 64, and the learning rates of the action network and evaluation network were set to 0.001 and 0.005; Initialize the DDPG model network and use the spatiotemporal graph neural network to extract the feature vector Z of the sample data t Input into the action network of DDPG, perform action encoding to output action a t ;Action a t Including the amount of machine cut αt and the amount of load cut β t ; Again, the eigenvector Z t After state encoding and action vector a as input t Input into the value network of the DDPG model to obtain the Q value of the action; for action a t After the action optimization is performed, it is fed back to the transient stability simulation system.

6. The power system emergency control method integrating DRL and STGNN according to claim 5 is characterized in that: The optimization of the cutting action vector includes the action network outputting the cutting vector a of the system in the current state. t , in the execution phase, the power cut amount is distributed to all generators in the system, and the action space of the agent is expressed as: A={P G1 ,P G2 ,…,P Gm ,P L1 ,P L2 ,…,P Ll } Among them, P Gi represents the removal amount of the i-th generator, i = 1, 2, ..., m; P Lk represents the removal amount of the kth load, k = 1, 2, ..., l; The amount of generator shedding is allocated according to the relative speed of the generators, and the amount of load shedding is allocated according to the importance of the loads; The power system knowledge is introduced to improve the action of the intelligent agent, and the action quantity output by the action network is distributed according to the relative speed of the generator and the importance of the load.

7. The power system emergency control method integrating DRL and STGNN according to claim 6, characterized in that: The network parameters in the updated model include: t Adjust the operating state of the power system, obtain the transient stability judgment mark of the power system through transient stability simulation, and use the reward function to calculate the action a based on the transient stability judgment mark. t The reward value; The reward function is expressed as: Among them, T stable represents the set of stable states of the system, s t represents the state set of the agent, R ST Represents a large positive number, representing the reward for the system to reach stability after control; R UST represents a large negative number, representing the penalty for instability after control; During the control process, the short-term reward r t Expressed as transient stability coefficient, it is used to characterize the transient stability of the system and is expressed as: Among them, Δδ max It represents the relative power angle difference between any generator and the reference generator; The network of the DDPG model is updated according to the reward value and Q value, and the action network is μ(s|θ μ ), the value network is Q network, and the forward return is expressed as: y i =r i +γQ′(s i+1 ,μ′(s i+1 ∣θ μ′ )∣θ Q′ ) Where γ represents the discount factor; The loss function during model training is expressed as: Action network μ(s|θ μ ) to update the parameters: In order to improve the stability of the model training process, the target Q network Q' and the target action network μ' are used to participate in the training, and the parameter update method is expressed as: Here, τ represents the update rate.

8. A power system emergency control system integrating DRL and STGNN using the method according to any one of claims 1 to 7, characterized in that: Data collection module, builds the international standard IEEE-39 model, collects power data, and constructs a data set; The feature processing module extracts features from the data set, inputs the obtained features into the improved DDPG model integrated with DRL, and outputs the machine cutting action vector; The update module optimizes the cutting action vector and updates the network parameters in the model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-device parallel charging state identification method based on graph neural network

    CN121643146A

  • A multi-device parallel charging state recognition method based on a graph neural network

    CN121643146B