A method for locating spoofed data injection attacks in power grids based on adaptive spatiotemporal graph neural networks

By using an adaptive spatiotemporal graph neural network, and leveraging an improved multi-head self-attention and Transformer module to generate a sparse reconnected adjacency matrix, combined with spatiotemporal feature fusion, the problem of insufficient detection of low homogeneity graphs in existing technologies is solved, and more efficient localization of power grid spoofing data injection attacks is achieved.

CN119696841BActive Publication Date: 2025-10-31NORTH CHINA ELECTRIC POWER UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411749093.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-10-31
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing methods for detecting fake data injection attacks in power grids based on graph neural networks perform poorly in low-homogeneity graphs, struggle to fully utilize neighbor node information for effective feature representation, and lack robustness and adaptability in complex power grid environments.

Method used

An adaptive spatiotemporal graph neural network is adopted. By introducing a sliding window mechanism and an improved multi-head self-attention mechanism, a sparse reconnected adjacency matrix is ​​generated. An improved Transformer module is combined to perform global feature fusion and spatiotemporal feature extraction, identify reliable neighbor nodes, and use a temporal convolutional network to capture the temporal dependence and spatial features of power grid data.

Benefits of technology

It improves the accuracy and robustness of detecting fake data injection attacks on power grids, and can better identify attack locations in complex power grid environments, thus enhancing detection effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119696841B_ABST
    Figure CN119696841B_ABST
Patent Text Reader

Abstract

This invention discloses a method for locating power grid spoofing attacks based on an adaptive spatiotemporal graph neural network (GNN) within the field of smart grid security technology. The method includes: generating an optimized reconnection adjacency matrix to achieve adaptive neighborhood selection; performing global feature fusion using an improved Transformer module; inputting the self-embedding of nodes into a spatiotemporal feature fusion module, concatenating all obtained node embeddings, and inputting this concatenation into the improved Transformer module to fuse multi-source features, obtaining the final total node embedding, which is then input into a classification layer to determine node anomalies, thereby locating power grid spoofing attacks. This method considers both the temporal dependence and spatial correlation of power grid data while improving the performance of GNNs in low-homogeneity graph data; and combines adaptive neighborhood selection and spatiotemporal feature fusion to enhance the detection and location of attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart grid security technology, and specifically relates to a method for locating fake data injection attacks in the power grid based on an adaptive spatiotemporal graph neural network. Background Technology

[0002] Against the backdrop of global informatization and intelligentization, modern smart grids are gradually transforming into power cyber-physical systems. Smart grids integrate information and communication technologies into large-scale power networks to achieve more efficient power generation, transmission, and distribution. Terminal acquisition units and vector measurement units acquire physical measurement values ​​and transmit them to monitoring and data acquisition systems. State estimation, after analyzing these measurements, provides the control center with the optimal estimate of the system state; its accuracy is crucial for the stable operation of the power grid. The interaction of large amounts of data between the physical and network layers brings efficiency but also increases the risk of network attacks. False Data Injection Attacks (FDIA) compromise the integrity of power grid data, posing a significant threat to power system security. By constructing specific attack vectors, FDIA bypasses the system's malicious data detection mechanisms, affecting the results of power system state estimation. Incorrect estimates lead the control center to make erroneous decisions, impacting the normal operation of the power grid and causing substantial economic losses.

[0003] To detect data injection attacks, artificial intelligence has rapidly developed with the advent of the big data era. Simultaneously, the power grid generates an ever-increasing volume of data. To address the data security issues brought about by this massive data, advanced neural network algorithms are used to analyze the characteristics of power grid operation data, thereby enabling data-driven methods for FDIA detection. Existing anomaly detection methods based on Graph Neural Networks (GNNs) consider the topological structure information of the power grid and have already shown certain advantages in power grid FDIA detection. However, traditional GNN models propagate and aggregate features through a fixed adjacency matrix, making them more suitable for graphs with highly homogeneous nodes. Homogeneity refers to the similar distribution of certain features or attributes among members of a population. In graph-structured data, homogeneity often manifests as connected nodes having similar features and the same labels.

[0004] Most existing graph neural networks assume that the graph has high homomorphism and perform poorly in low homomorphic graphs. Due to the dynamic complexity of the power grid topology, the load characteristics of each bus are different, and the power characteristics of adjacent buses with physical connections may also differ significantly. Existing detection methods based on graph neural networks are difficult to fully utilize the information of neighboring nodes in the power grid for effective feature representation.

[0005] Therefore, there is an urgent need for a method for locating fake data injection attacks in the power grid based on an Adaptive Spatial Temporal Graph Neural Network (ADSTGNN). This method would address the problem of poor performance of GNNs in low-homogeneity graphs in existing technologies, and would also have better robustness and adaptability in complex power grid environments, thereby improving the effectiveness of attack detection. Summary of the Invention

[0006] The purpose of this invention is to provide a method for locating spoofed data injection attacks in power grids based on an adaptive spatiotemporal graph neural network, characterized by the following steps:

[0007] Step S1: Obtain the power grid operation data, model the power grid as a graph, obtain the graph structure data, introduce the sliding window mechanism, and generate the graph structure time series X(Input) of historical data.

[0008] Step S2: Based on the graph structure time series X (Input), combine graph structure learning and node embedding to generate an optimized reconnection adjacency matrix A. * To achieve adaptive neighborhood selection;

[0009] Step S3: Based on the graph structure time series X (Input), perform global feature fusion using the improved Transformer module to obtain the fused feature vector X. fusion The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego ;

[0010] Step S4: Self-embedding of the node X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C The neighbor embedding output by the spatiotemporal feature fusion module includes: reconnection neighbor embedding. and the neighbor embedding of nodes and

[0011] Step S5: Set the total embedding of the connected nodes X C The input is fed into the improved Transformer module, which fuses multi-source features to obtain the final node embedding X. out Embed the final node into X out Input is fed into the classification layer to determine node anomalies, thereby enabling the location of fake data injection attacks on the power grid.

[0012] The acquisition of power grid operation data in step S1 includes: selectively reading historical data from the SCADA system of the power system; selecting the bus phase angle, voltage amplitude, active and reactive power injection of the bus, and active and reactive power injection of each branch to form the original dataset Z; and performing mean removal and normalization processing on the original dataset Z.

[0013] The step of modeling the power grid as a graph and obtaining graph structure data includes: taking the bus in the power grid as nodes and the connection relationship between the bus in the power grid as edges, and converting the original dataset Z into graph structure data according to the number of nodes and the feature number of each node;

[0014] The introduction of the sliding window mechanism includes setting the sliding step size of the sliding window mechanism to 1 time step.

[0015] In step S2, graph structure learning and node embedding are combined to generate an optimized reconnection adjacency matrix A. * include:

[0016] Given a graph-structured time series X (Input) of historical data, an improved multi-head self-attention mechanism is used to calculate the attention scores between nodes. A K-nearest neighbor and minimum threshold method are jointly applied to select the node with the highest attention score as the new neighbor of the target node, generating an optimized reconnection adjacency matrix A. * ;

[0017] The optimized reconnection adjacency matrix A * It is a sparse and reliable reconnection adjacency matrix.

[0018] The improved multi-head self-attention mechanism is as follows:

[0019] Q i =XW i Q (1)

[0020]

[0021] S=Mean(head1, head2,..., head n (3)

[0022] In the formula, Q i The query matrix is ​​obtained by linear transformation of X, where X represents the node feature matrix, and W... i Q It is the learnable weight matrix of the i-th head (i = 1, ..., n); This represents the normalized Qi; for The transpose of the matrix; Equation (2) calculates matrix Q. iThe cosine score between each pair of row vectors is used to generate the cosine score matrix head. i ∈R n×n Mean represents the average of the score matrices of all attention heads, resulting in a symmetric attention score matrix S∈R. n×n , where element S ij The value of is in the range of [-1, 1].

[0023] The K-nearest neighbor and minimum threshold method includes:

[0024] Define a positive integer γ and a non-negative threshold ∈ . Retain the attention scores of those points that are greater than ∈ or belong to the top γ largest elements in the corresponding row, and set the remaining attention scores to 0. Reconnect the adjacency matrix A. * ∈R n×n Represented as:

[0025]

[0026] In the formula, Let S represent the set of the first γ largest elements in the i-th row of matrix S.

[0027] The specific process of global feature fusion using the improved Transformer module is as follows:

[0028] X o =ReLU(XW0+b0) (5)

[0029] Q i =X o W i Q (6)

[0030] K i =X o W i K (7)

[0031] V i =X o W i V (8)

[0032] X MHSA =Concat(head1, head2,..., head n (9)

[0033]

[0034] In the formula, X oThe feature representation obtained after linear transformation and ReLU activation of the input node feature matrix X, where W0 is the weight matrix, b0 is the bias vector, and Q is the weight vector. i K i and V i They are respectively made by X o The query, key, and value matrix W obtained by linear transformation i Q W i K and W i V Learnable weight matrices corresponding to the query, key, and value matrices, respectively; X MHSA The output is the result through a multi-head self-attention mechanism. For K i The transpose of d k The vector dimension of the query and key matrix;

[0035] To overcome the vanishing gradient problem and alleviate overfitting, residual connections and Dropout operations are introduced:

[0036] X fusion =Dropout(X o +X MHSA )+X o (11)

[0037] In the formula, X fusion For the improved output of the Transformer module;

[0038] The output X of the improved Transformer module fusion Simplified representation: FusionBlock(X):

[0039] X fusion =FusionBlock(X) (12)

[0040] The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego for:

[0041] X ego =X fusion ||X o (13).

[0042] The self-embedding of nodes X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C include:

[0043] Temporal convolutional networks are used to extract features across time steps from the input time series data;

[0044] By employing extended causal convolution with gating mechanism, long-term dependencies in the data can be effectively obtained;

[0045] By using extended causal convolution and introducing an expansion factor, for an extended causal convolutional network with dilation = [1, 2, 4], the number of convolutions in each layer remains unchanged, and convolutional expansion is performed in the next layer;

[0046] Introducing a gating mechanism, we define a gated temporal convolutional layer:

[0047]

[0048] In the formula, the initial input is the self-embedded X ego σ represents the sigmoid activation function. and These are the convolution kernel parameters used to generate the main feature signal and the gate signal, respectively. l+1 and c l+1 The corresponding bias term is represented by ⊙, which indicates element-wise dot product, and ★ indicates extended causal convolution. Let be the input feature tensor of the l-th layer at time step t;

[0049] The learnable parameters of the convolution kernel at time step t and The convolution operation is represented as:

[0050]

[0051] In the formula, M is the kernel length; is the s-th parameter in the convolution kernel; d is the expansion factor, whose size controls the jump distance, that is, an input is selected every d steps;

[0052] The final output X of the temporal convolution module t As input to the spatial feature extraction module, the initial adjacency matrix A is used for feature aggregation:

[0053]

[0054] In the formula, k = 1, 2, ..., K, represents the number of feature aggregations. Let D represent the normalized adjacency matrix with self-loops, D = diag((A+I)1 n Let be the degree matrix of A with self-loops, where I∈R n×n It is the identity matrix; and Defined as the neighbor embedding of a node;

[0055] Using the generated reconnection adjacency matrix A * To perform feature aggregation:

[0056]

[0057] In the formula, It is a standardized A * D * =diag(A * 1 n ) is A * The degree matrix, the result of feature aggregation This is called reconnected neighbor embedding;

[0058] Connect the self-embedding, reconnected neighbor embedding, and node neighbor embedding to form the total embedding X of the node. C :

[0059]

[0060] For X C Perform Dropout to generate a learnable weight vector w. The dimension of the weight vector w is the same as the output dimension. Then, compare the weight vector w with X. C Perform the Hadamard product operation to highlight X. C Key feature dimensions:

[0061] X out =ReLU(w⊙X) C (19).

[0062] Step S5 specifically includes:

[0063] Again, using the improved Transformer module for X out Processing is performed; based on the multi-label classification results output by the model, and the probability of the node output at the last time step, the nodes with higher anomaly probabilities are identified as targets of attack.

[0064] The embedding of the final node into X out The following are examples of nodes that are anomalies when input to the classification layer:

[0065] The Softmax layer outputs the probability of each node being abnormal, determining potentially attacked nodes and their corresponding locations. Specifically, the process involves a classification layer that determines the method of node anomaly.

[0066] output=softmax(FusionBlock(X out W1)) (20)

[0067] In the formula, W1 is a trainable weight matrix; X outThe final total node embedding is used; the output is the multi-label classification result.

[0068] Another object of the present invention is to provide a power grid spoofing attack localization device based on the power grid spoofing attack localization method based on the adaptive spatiotemporal graph neural network according to the present invention, characterized in that it comprises:

[0069] Data acquisition module: used to acquire power grid operation data and generate historical data in a graph structure time series.

[0070] Preprocessing module: Used to preprocess the acquired power grid operation data, including normalizing the data, converting it into graph structure data, and introducing a sliding window mechanism to obtain the processed power grid operation data;

[0071] Adaptive Spatiotemporal Graph Neural Network Model Module: Inputs the processed data into the trained model to obtain the probability of each node being attacked, thereby locating the attacked main line and taking further protective measures;

[0072] The adaptive spatiotemporal graph neural network model includes: a graph structure learning module, a global feature fusion module, a spatiotemporal feature fusion module, and a Softmax classification module; the graph structure learning module is used to mine the intrinsic relationships between nodes and identify new reliable neighbors for each node; the global feature fusion module fuses more valuable information from the entire graph; the spatiotemporal feature fusion module extracts the temporal and spatial features of the data; and the Softmax classification module is used to classify the input data and output the final detection result.

[0073] The beneficial effects of this invention are as follows:

[0074] This invention discloses a method for locating and detecting spoofing data injection attacks based on an adaptive spatiotemporal graph neural network. Collected power grid operation data is input into a pre-trained model to obtain labels for determining the possible locations of spoofing data injection attacks. Compared to existing technologies, this invention establishes a reconnected graph structure through a graph structure learning module, a global feature fusion module, and a spatiotemporal feature fusion module. This identifies more reliable neighbors for each node in the graph and integrates valuable information from the entire graph. Specifically, it includes the following beneficial effects:

[0075] (1) By using the improved multi-head self-attention mechanism, the intrinsic relationship between nodes is mined from the original graph data. Combining the ideas of K-nearest neighbors and minimum threshold method, reliable neighbors are identified for each node, thereby generating a reconnected adjacency matrix.

[0076] (2) Using the improved Transformer module, more valuable information is fused from the whole graph, and the fused feature vector is combined with the initial feature vector to form node self-embedding.

[0077] (3) The spatiotemporal feature fusion module fully captures the time dependence and spatial features of the data, and the original adjacency matrix and the reconnected adjacency matrix are respectively subjected to feature aggregation to form the neighbor embedding of the node. The generated node embeddings are connected to each other, so that the model can better identify local and global information, which is conducive to more accurate detection of power grid FDIA and determination of the location of the attacked bus. Attached Figure Description

[0078] Figure 1 This is a flowchart illustrating a method for locating spoofed data injection attacks in power grids based on an adaptive spatiotemporal graph neural network, as described in this invention.

[0079] Figure 2 This is a schematic diagram of the fake data injection attack detection process according to an embodiment of the present invention;

[0080] Figure 3(a) is a schematic diagram of the graph structure learning module and the global feature fusion module in the overall architecture of the adaptive spatiotemporal graph neural network according to an embodiment of the present invention;

[0081] Figure 3(b) is a schematic diagram of the spatiotemporal feature fusion module and classifier in the overall architecture of the adaptive spatiotemporal graph neural network according to an embodiment of the present invention;

[0082] Figure 4 A schematic diagram illustrating the introduction of a sliding window mechanism for graph-structured data in an embodiment of the present invention;

[0083] Figure 5 This is a schematic diagram of the extended causal convolution structure according to an embodiment of the present invention. Detailed Implementation

[0084] This invention provides a method for locating fake data injection attacks in power grids based on an adaptive spatiotemporal graph neural network. The invention will be further described in detail below with reference to the accompanying drawings.

[0085] like Figure 1 The embodiment of the present invention disclosed herein discloses a method for locating spoofed data injection attacks in power grids based on an adaptive spatiotemporal graph neural network, comprising:

[0086] Step S1: Obtain the power grid operation data, model the power grid as a graph, obtain the graph structure data, introduce the sliding window mechanism, and generate the graph structure time series X(Input) of historical data.

[0087] Step S2: Based on the graph structure time series X (Input), combine graph structure learning and node embedding to generate an optimized reconnection adjacency matrix A. *To achieve adaptive neighborhood selection;

[0088] Step S3: Based on the graph structure time series X (Input), perform global feature fusion using the improved Transformer module to obtain the fused feature vector X. fusion The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego ;

[0089] Step S4: Self-embedding of the node X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C The neighbor embedding output by the spatiotemporal feature fusion module includes: reconnection neighbor embedding. and the neighbor embedding of nodes and

[0090] Step S5: Set the total embedding of the connected nodes X C The input is fed into the improved Transformer module, which fuses multi-source features to obtain the final node embedding X. out Embed the final node into X out Input is fed into the classification layer to determine node anomalies, thereby enabling the location of fake data injection attacks on the power grid.

[0091] In this embodiment, the fake data injection attack detection process is implemented as follows: Figure 2 As shown, the purpose of applying the detection process is to input the collected power grid operation data into a pre-trained model to obtain labels for determining the possible locations where spoofing attacks may exist. Compared with existing technologies, this invention establishes a reconnected graph structure through a graph structure learning module, a global feature fusion module, and a spatiotemporal feature fusion module, identifying more reliable neighbors for each node in the graph and integrating valuable information from the entire graph.

[0092] In this embodiment, the following technical solution is specifically adopted to achieve the purpose of the detection process:

[0093] Acquire normal operation data of the power grid and generate a graph-structured time series X(Input) of historical data;

[0094] Historical data from the Supervisory Control and Data Acquisition (SCADA) system in the power system are selectively read. The original dataset Z is formed by selecting the bus phase angle, voltage amplitude, active and reactive power injected into the bus, and active and reactive power injected into each branch. The dataset is then subjected to mean removal and normalization.

[0095] The power grid is modeled as a graph, with the busbars in the power grid as nodes and the connections between the busbars as edges. Based on the number of nodes and the feature number of each node, the original dataset Z is transformed into graph structure data.

[0096] To fully explore the time dependencies between historical data and capture short-term dynamics and local patterns, a sliding window mechanism is introduced.

[0097] By combining graph structure learning and node embedding, an optimized reconnection adjacency matrix A is generated. * This enables adaptive neighborhood selection and provides more reliable neighbor information for each node.

[0098] The processed graph structure data is input, and an improved multi-head self-attention mechanism is used to calculate the attention score between nodes. Then, the K-nearest neighbor and minimum threshold methods are jointly applied to select the node with the highest attention score as the new neighbor of the target node, resulting in a sparse, connected and reliable reconnected adjacency matrix.

[0099] An improved Transformer module is used to help the model identify more valuable features in the entire graph and to focus more on these features. The identified features are then fused with the initial node features as self-embeddings of the nodes, which are then input into the subsequent spatiotemporal feature fusion module.

[0100] A spatiotemporal feature fusion module is constructed. First, a temporal correlation layer is used, employing extended causal convolutions with gating mechanisms to learn the temporal dependencies between input data. Then, a spatial correlation layer, composed of graph convolutional neural networks, is used to extract spatial features. The initial adjacency matrix and the established reconnected adjacency matrix work together. The initial adjacency matrix is ​​input to obtain the neighbor embeddings of nodes, and the reconnected adjacency matrix is ​​input to obtain the reconnected neighbor embeddings. Finally, the two types of node embeddings are concatenated with the node's self-embedding, incorporating both node features and domain information. This integrated approach, combining local and global information, facilitates the identification of subsequent attacks.

[0101] A multi-label classification layer is constructed, with the total embedding of the connected nodes as input. After the features are extracted by the linear transformation of the fully connected layer, the Softmax layer outputs the probability of each node being abnormal, thereby identifying the nodes that may be attacked and their corresponding locations, so that decision-makers can take further remedial measures.

[0102] This invention discloses a method for locating spoofed data injection (FDIA) attacks in power grids. This method uses graph structure learning to capture the internal correlations between nodes, enabling nodes to identify more reliable neighbors. Furthermore, it integrates more valuable information from the graph for detection tasks through node feature fusion. Finally, it introduces an extended causal convolutional network with a gating mechanism combined with a graph convolutional network to capture the temporal dependencies of node information and further extract spatial features, thereby improving the detection and localization capabilities of FDIA attacks in power grids. This invention addresses the problem of poor performance of GNNs in low-homogeneity graphs and exhibits better robustness and adaptability in complex power grid environments.

[0103] The overall architecture of the adaptive spatiotemporal graph neural network in this embodiment of the invention is as follows: Figure 3(a) and 3(b) As shown, Figure 3(a) is a schematic diagram of the graph structure learning module and the global feature fusion module; Figure 3(b) is a schematic diagram of the spatiotemporal feature fusion module and the classifier.

[0104] The following section, in conjunction with the accompanying diagrams, provides a detailed description of the implementation process for each step.

[0105] Step S1: Obtain the power grid operation data, model the power grid as a graph, obtain the graph structure data, introduce the sliding window mechanism, and generate the graph structure time series X(Input) of historical data.

[0106] The acquisition of power grid operating data in step S1 includes:

[0107] Historical data from the SCADA system in the power system are selectively read. The original dataset Z is formed by selecting the bus phase angle, voltage amplitude, active and reactive power injected into the bus, and active and reactive power injected into each branch. The dataset is then subjected to mean removal and normalization.

[0108] The power grid is modeled as a graph, with the busbars in the power grid as nodes and the connections between the busbars as edges. Based on the number of nodes and the feature number of each node, the original dataset Z is transformed into graph structure data.

[0109] To fully explore the time dependencies between historical data and capture short-term dynamics and local patterns, a sliding window mechanism is introduced, with a sliding step size set to one time step. The processing mechanism is as follows: Figure 4 As shown.

[0110] Step S2: Based on the graph structure time series X (Input), combine graph structure learning and node embedding.

[0111] Generate an optimized reconnection adjacency matrix A * To achieve adaptive neighborhood selection;

[0112] In this embodiment, a reconnection adjacency matrix is ​​generated. Graph neural networks designed based on the assumption of high homogeneity are not entirely applicable to graphs with low homogeneity such as power grids. To address this issue, a reconnection adjacency matrix is ​​constructed through graph structure learning, providing reliable neighbor information for each node.

[0113] The attention scores between nodes are calculated using the Multi-Head Self-Attention (MHSA) mechanism to uncover their intrinsic connections. The MHSA mechanism is modified to focus on the generation of undirected graphs.

[0114] The improved multi-head self-attention mechanism is as follows:

[0115] Q i =XW i Q (1)

[0116]

[0117] S=Mean(head1, head2,..., head n (3)

[0118] In the formula, Q i The query matrix is ​​obtained by linear transformation of X, where X represents the node feature matrix, and W... i Q It is the learnable weight matrix of the i-th head (i = 1, ..., n); Represents the normalized Qx; for The transpose of the matrix; Equation (2) calculates matrix Q. i The cosine score between each pair of row vectors is used to generate the cosine score matrix head. i ∈R n×n Mean represents the average of the score matrices of all attention heads, resulting in a symmetric attention score matrix S∈R. n×n , where element S ij The value of is in the range of [-1, 1].

[0119] In this embodiment, equation (2) calculates matrix Q. i The cosine score between each pair of row vectors is used to generate the cosine score matrix head. i ∈R n×n To ensure that the generated reconnection adjacency matrix conforms to the properties of an undirected graph and better describes the intrinsic relationships between nodes, a method different from the multi-head self-attention mechanism in existing technologies is used. Calculate weighted similarity, using This is used to calculate and ensure that the resulting attention score matrix is ​​symmetric; the mean average of the score matrices of all attention heads is then performed to obtain a symmetric attention score matrix S∈R. n×n , where element S ij The value of is in the range of [-1, 1].

[0120] The MHSA mechanism can construct multiple fully connected graphs, allowing each node to connect to all other nodes. This capability enables each node to aggregate features from any other node, but it also introduces excessive computational burden and noise. To address this issue, the ideas of K-Nearest Neighbors (KNN) and the minimum threshold method are combined to select the most reliable neighbor for each node. The K-Nearest Neighbors and minimum threshold methods include:

[0121] Define a positive integer γ and a non-negative threshold ∈ . Retain the attention scores of those points that are greater than ∈ or belong to the top γ largest elements in the corresponding row, and set the remaining attention scores to 0. The new adjacency matrix A is... * ∈R n×n Represented as:

[0122]

[0123] In the formula, Let S represent the set of the first γ largest elements in the i-th row of matrix S.

[0124] Thus, we obtain the sparse, non-negative, and reliable reconnection adjacency matrix A at each time step. * Ensure that the adjacency matrix A is reconnected. * Its sparsity and high reliability enable adaptive neighborhood selection of nodes in the power grid.

[0125] Those skilled in the art will know that the regenerated reconnection adjacency matrix A in this embodiment * Sparsity and high reliability are essential for better reflecting real-world physical topologies. In actual power grids, due to cost and other considerations, each busbar cannot be connected by a single branch. The goal is to ensure grid connectivity with as few branches as possible. Therefore, the adjacency matrix of a power grid modeled as a graph is necessarily sparse. Calculating the similarity from a node to all other nodes using a multi-head self-attention mechanism, without subsequent minimum thresholding, results in a fully connected graph, which is unrealistic and increases computational complexity. Therefore, sparsity is an essential characteristic of power grid modeling as a graph, while high reliability is achieved by selecting the top few nodes with the highest similarity among all calculated nodes, identifying them as new neighbors.

[0126] Step S3: Based on the graph structure time series X (Input), perform global feature fusion using the improved Transformer module to obtain the fused feature vector X. fusion The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego ;

[0127] In this embodiment, step S3 performs global feature fusion of the graph, including: using the improved Transformer module to identify more valuable features in the entire graph; fusing the identified features with the initial node features to generate richer initial feature representations for the nodes, and using this as the self-embedding of the nodes, inputting it into subsequent modules;

[0128] In this embodiment, based on the graph-structured time series X (Input), a modified Transformer module is used to perform global feature fusion to obtain the fused feature vector X. fusion Considering the difficulty in capturing long-range dependencies in graph neural networks, simply aggregating neighborhood information may not capture the relationships between distant nodes. The improved Transformer module helps the model identify important features of the entire graph and focuses more on these important features.

[0129] The improved Transformer module includes: a two-layer Dropout and a residual connection layer;

[0130] Compared to the original Transformer architecture in existing technologies, the improved Transformer module omits regularization and the Feed Forward fully connected layer. Specifically, in graph data, the number of nodes and edges is unbalanced. Regularization can easily impose excessive weight penalties on low-degree nodes. Using Dropout, which allows for random discarding, avoids this situation and effectively reduces the risk of overfitting. Furthermore, by omitting the Feed Forward layer, the focus is placed on the multi-head self-attention mechanism to capture useful information from the entire graph, reducing computational complexity.

[0131] The specific process of global feature fusion using the improved Transformer module is as follows:

[0132] X o =ReLU(XW0+b0) (5)

[0133] Q i =X o W i Q (6)

[0134] K i =X o Wi K (7)

[0135] V i =X o W i V (8)

[0136] X MHSA =Concat(head1, head2,..., head n (9)

[0137]

[0138] In the formula, X o The feature representation obtained after linear transformation and ReLU activation of the input node feature matrix X, where W0 is the weight matrix, b0 is the bias vector, and Q is the weight vector. i K i and V i They are respectively made by X o The query, key, and value matrix W obtained by linear transformation. i Q W i K and W i V Learnable weight matrices corresponding to the query, key, and value matrices, respectively; X MHSA The output is the result through a multi-head self-attention mechanism. For K i The transpose of d k The vector dimension of the query and key matrix;

[0139] To overcome the vanishing gradient problem and alleviate overfitting, residual connections and Dropout operations are introduced:

[0140] X fusion =Dropout(X o +X MHSA )+X o (11)

[0141] In the formula, X fusion For the improved output of the Transformer module;

[0142] For ease of subsequent description, the output X of the improved Transformer module will be... fusion Simplified representation: FusionBlock(X):

[0143] Xfusion =FusionBlock(X) (12)

[0144] In this embodiment, global feature fusion of the graph is performed, including: using the improved Transformer module to identify more valuable features in the entire graph; fusing the identified features with the initial node features to generate richer initial feature representations for the nodes, and using this as the self-embedding of the nodes, which is then input into subsequent modules.

[0145] The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego for:

[0146] X ego =X fusion ||X o (13)

[0147] Step S4: Self-embedding of the node X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C The neighbor embedding output by the spatiotemporal feature fusion module includes: reconnection neighbor embedding. and the neighbor embedding of nodes and

[0148] In this embodiment, step S4 concatenates the fused feature vector with the original feature vector as the node's self-embedding, inputs it into the spatiotemporal feature fusion module, and obtains the total embedding of the connected nodes. This includes: passing through a temporal correlation layer, where extended causal convolution with a gating mechanism is used to learn the temporal dependencies between input data; passing through a spatial correlation layer, composed of a graph convolutional neural network, used to extract spatial features; combining the initial adjacency matrix and the established reconnected adjacency matrix; inputting the initial adjacency matrix to obtain the node's neighbor embedding, and inputting the reconnected adjacency matrix to obtain the reconnected neighbor embedding; concatenating the two types of node embeddings with the node's self-embedding, which simultaneously includes the node's own features and domain information, and combining local and global information to obtain the total embedding of the connected nodes.

[0149] The self-embedding of nodes X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C include:

[0150] To extract temporal features, a temporal convolutional network is used to extract features across time steps from the input time series data;

[0151] By employing extended causal convolution with gating mechanism, long-term dependencies in the data can be effectively obtained;

[0152] Specifically, this includes: using extended causal convolution with gating mechanism to effectively obtain long-term dependencies in the data; TCN is a convolutional neural network architecture that uses causal convolutional layers to process time-series data, and in this embodiment, extended causal convolution with gating mechanism is selected.

[0153] By utilizing extended causal convolution and introducing a dilation factor, the structure of a temporal convolutional network with dilation = [1, 2, 4] is as follows: Figure 5 As shown, the number of convolutions in each layer remains unchanged, and convolution expansion is performed in the next layer.

[0154] To fully model the nonlinear relationship in the time dimension, a gating mechanism is introduced, defining a gated temporal convolutional layer as follows:

[0155]

[0156] In the formula, the initial input is the self-embedded X ego σ represents the sigmoid activation function. and These are the convolution kernel parameters used to generate the main feature signal and the gate signal, respectively. l+1 and c l+1 The corresponding bias term is represented by ⊙, which indicates element-wise dot product, and ★ indicates extended causal convolution. Let be the input feature tensor of the l-th layer at time step t;

[0157] Given The learnable parameters of the convolution kernel at time step t and The convolution operation is represented as:

[0158]

[0159] In the formula, M is the kernel length; is the s-th parameter in the convolution kernel; d is the expansion factor, whose size controls the jump distance, that is, an input is selected every d steps;

[0160] Subsequently, the final output X of the temporal convolution module t As input to the spatial feature extraction module, the initial adjacency matrix A is used for feature aggregation:

[0161]

[0162] In the formula, k = 1, 2, ..., K, represents the number of feature aggregations. Let D represent the normalized adjacency matrix with self-loops, D = diag((A+I)1 n Let be the degree matrix of A with self-loops, where I∈R n×n It is the identity matrix; and Defined as the neighbor embedding of a node;

[0163] Similarly, using the generated reconnection adjacency matrix A * To perform feature aggregation:

[0164]

[0165] In the formula, It is a standardized A * D * =diag(A * 1 n ) is A * The degree matrix, the result of feature aggregation This is called reconnected neighbor embedding;

[0166] The total embedding of a node is obtained by concatenating self-embedding, reconnected neighbor embedding, and node neighbor embedding:

[0167]

[0168] For X C Perform Dropout to generate a learnable weight vector w. The dimension of the weight vector w is the same as the output dimension. Then, compare the weight vector w with X. C Perform the Hadamard product operation to highlight X. C Key feature dimensions:

[0169] X out =ReLU(w⊙X) C (19)

[0170] Step S5: Set the total embedding of the connected nodes X C The input is fed into the improved Transformer module, which fuses multi-source features to obtain the final node embedding X. out Embed the final node into X out Input is fed into the classification layer to determine node anomalies, thereby enabling the location of fake data injection attacks on the power grid.

[0171] In this embodiment, step S5 again utilizes the improved Transformer module to process X. out The processing aims to optimize the final total node embedding by integrating information from the node itself, the original neighbor relationships, and the reconnected neighbor relationships, fusing multi-source features to generate a more robust, high-quality node embedding.

[0172] Based on the multi-label classification results output by the model, and the probability of the node output at the last time step, the nodes with higher anomaly probabilities are identified as targets of attack.

[0173] The embedding of the final node into X out The following are examples of nodes that are anomalies when input to the classification layer:

[0174] The Softmax layer outputs the probability of each node being abnormal, determining potentially attacked nodes and their corresponding locations. Specifically, the process involves a classification layer that determines the method of node anomaly.

[0175] output=softmax(FusionBlock(X out W1)) (20)

[0176] In the formula, W1 is a trainable weight matrix; X out The final total node embedding is used; the output is the multi-label classification result.

[0177] Another embodiment of the present invention discloses a power grid spoofing attack localization device based on the power grid spoofing attack localization method according to the present invention, comprising:

[0178] Data acquisition module: used to acquire power grid operation data and generate historical data in a graph structure time series.

[0179] Preprocessing module: Used to preprocess the acquired power grid operation data, including normalizing the data, converting it into graph structure data, and introducing a sliding window mechanism to obtain the processed power grid operation data;

[0180] Adaptive Spatiotemporal Graph Neural Network Model Module: Inputs the processed data into the trained model to obtain the probability of each node being attacked, thereby locating the attacked main line and taking further protective measures;

[0181] The adaptive spatiotemporal graph neural network model includes: a graph structure learning module, a global feature fusion module, a spatiotemporal feature fusion module, and a Softmax classification module; the graph structure learning module is used to mine the intrinsic relationships between nodes and identify new reliable neighbors for each node; the global feature fusion module fuses more valuable information from the entire graph; the spatiotemporal feature fusion module extracts the temporal and spatial features of the data; and the Softmax classification module is used to classify the input data and output the final detection result.

[0182] In this embodiment, the power grid fake data injection attack localization device is used to implement the power grid fake data injection attack localization method based on adaptive spatiotemporal graph neural network disclosed in this invention.

[0183] To verify the effectiveness of the power grid spoofing data injection attack localization method based on adaptive spatiotemporal graph neural network disclosed in this invention, the following comparative verification experiment was conducted.

[0184] The evaluation metrics used in the comparative verification experiment include:

[0185] accuracy

[0186] Accuracy

[0187] Recall rate

[0188] F1 score

[0189] The experimental results of the comparative verification experiment are shown in Table 1. As can be seen from Table 1, the experimental results of the power grid false data injection attack localization method based on adaptive spatiotemporal graph neural network disclosed in this invention, represented by Ours, have achieved better experimental results in terms of accuracy, precision, recall and F1 score compared with other network structures in the prior art.

[0190] Table 1. Statistical table of comparative experimental results

[0191]

[0192] Note: Table 1 contains model definitions.

[0193] LSTM (Long Short-Term Memory): Long Short-Term Memory Network

[0194] GRU (Gated Recurrent Unit): Gated Recurrent Unit

[0195] Gated-TCN (Gated Temporal Convolutional Network): A gated temporal convolutional network

[0196] Transformer: Transformer model

[0197] GCN (Graph Convolutional Networks): Graph Convolutional Networks

[0198] GAT (Graph Attention Network): Graph Attention Network

[0199] GGNN (Gated Graph Neural Network): A gated graph neural network

[0200] STGCN (Spatial-Temporal Graph Convolutional Network): A spatiotemporal graph neural network

[0201] To evaluate the importance of each module in the power grid spoofing attack localization method based on adaptive spatiotemporal graph neural network disclosed in this invention, ablation experiments were conducted from the following four aspects. In this embodiment, the ablation experiments included the following four aspects:

[0202] 1. Remove reconnected neighbor embeddings from the final total embedding;

[0203] 2. Remove node self-embeds from the final total embedding;

[0204] 3. Remove the improved Transformer module after the final total embedding;

[0205] 4. Remove the temporal feature extraction module, i.e., the extended causal convolution module with gating mechanism;

[0206] The results of the ablation experiment are shown in Table 2. As can be seen from Table 2, the complete method for locating power grid spoofing data injection attacks based on adaptive spatiotemporal graph neural networks disclosed in this invention, represented by Ours, achieves better experimental results in terms of accuracy, precision, recall, and F1 score compared to removing some modules.

[0207] Table 2 Statistical table of ablation test results

[0208]

[0209] In this embodiment, through comparative verification experiments against other network structures in the prior art and ablation experiments against the method itself, it is demonstrated that the power grid fake data injection attack localization method based on adaptive spatiotemporal graph neural network disclosed in this invention can solve the problem of poor performance of GNN in low homogeneity graphs in the prior art, and has better robustness and adaptability in complex power grid environments, thereby improving the attack detection effect.

Claims

1. A method for locating spoofed data injection attacks in power grids based on an adaptive spatiotemporal graph neural network, characterized in that, Includes the following steps: Step S1: Obtain the power grid operation data, model the power grid as a graph, obtain the graph structure data, introduce the sliding window mechanism, and generate the graph structure time series X(Input) of historical data. Step S2: Based on the graph structure time series X (Input), combine graph structure learning and node embedding to generate an optimized reconnection adjacency matrix A. * To achieve adaptive neighborhood selection; Step S3: Based on the graph structure time series X (Input), perform global feature fusion using the improved Transformer module to obtain the fused feature vector X. fusion The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego ; Step S4: Self-embedding of the node X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C The neighbor embedding output by the spatiotemporal feature fusion module includes: reconnection neighbor embedding. and the neighbor embedding of nodes and Step S5: Set the total embedding of the connected nodes X C The input is fed into the improved Transformer module, which fuses multi-source features to obtain the final node embedding X. out Embed the final node into X out The data is input to the classification layer to determine node anomalies, thus enabling the location of fake data injection attacks on the power grid. In step S2, graph structure learning and node embedding are combined to generate an optimized reconnection adjacency matrix A. * include: Given a graph-structured time series X (Input) of historical data, an improved multi-head self-attention mechanism is used to calculate the attention scores between nodes. A combined K-nearest neighbor and minimum threshold method is applied to select the node with the highest attention score as the new neighbor of the target node, generating an optimized reconnection adjacency matrix A. * ; The optimized reconnection adjacency matrix A * It is a sparse and reliable reconnection adjacency matrix; The improved multi-head self-attention mechanism is as follows: Q i =XW i Q (1) S=Mean(head1,head2,…,head n ) (3) In the formula, Q i The query matrix is ​​obtained by linear transformation of X, where X represents the node feature matrix, and W... i Q It is the learnable weight matrix of the i-th head (i = 1, ..., n); Represents the normalized Q i ; for The transpose of the matrix; Equation (2) calculates matrix Q. i The cosine score between each pair of row vectors is used to generate the cosine score matrix head. i ∈R n×n Mean represents the average of the score matrices of all attention heads, resulting in a symmetric attention score matrix S∈R. n×n , where element S ij The value of is in the range of [-1, 1]; The self-embedding of nodes X ego By concatenating the neighbor embeddings output by the spatiotemporal feature fusion module, the total embedding X of the connected nodes is obtained. C include: Temporal convolutional networks are used to extract features across time steps from the input time series data; By employing extended causal convolution with gating mechanism, long-term dependencies in the data can be effectively obtained; By using extended causal convolution and introducing an expansion factor, for an extended causal convolutional network with dilation = [1,2,4], the number of convolutions in each layer remains unchanged, and convolutional expansion is performed in the next layer. Introducing a gating mechanism, we define a gated temporal convolutional layer: In the formula, the initial input is the self-embedded X ego σ represents the sigmoid activation function. and These are the convolution kernel parameters used to generate the main feature signal and the gate signal, respectively. l+1 and c l+1 The corresponding bias term is represented by ⊙, which indicates element-wise dot product, and ★ indicates extended causal convolution. Let be the input feature tensor of the l-th layer at time step t; The learnable parameters of the convolution kernel at time step t and The convolution operation is represented as: In the formula, M is the kernel length; is the s-th parameter in the convolution kernel; d is the expansion factor, whose size controls the jump distance, that is, an input is selected every d steps; The final output X of the temporal convolution module t As input to the spatial feature extraction module, the initial adjacency matrix A is used for feature aggregation: In the formula, k = 1, 2, ..., K, represents the number of feature aggregations. Let D represent the normalized adjacency matrix with self-loops, D = diag((A+I)1 n Let be the degree matrix of A with self-loops, where I∈R n×n It is the identity matrix; and Defined as the neighbor embedding of a node; Using the generated reconnection adjacency matrix A * To perform feature aggregation: In the formula, It is a standardized A * D * =diag(A * 1 n ) is A * The degree matrix, the result of feature aggregation This is called reconnected neighbor embedding; Connect the self-embedding, reconnected neighbor embedding, and node neighbor embedding to form the total embedding X of the node. C : For X C Perform Dropout to generate a learnable weight vector w, where the dimension of the weight vector w is the same as that of X. C The output dimensions are the same, so the weight vector w is the same as X. C Perform the Hadamard product operation to highlight X. C Key feature dimensions: X out =ReLU(w⊙X C ) (19)。 2. The method for locating power grid spoofing data injection attacks based on adaptive spatiotemporal graph neural networks according to claim 1, characterized in that, The acquisition of power grid operation data in step S1 includes: selectively reading historical data from the SCADA system of the power system; selecting the bus phase angle, voltage amplitude, active and reactive power injection of the bus, and active and reactive power injection of each branch to form the original dataset Z; and performing mean removal and normalization processing on the original dataset Z. The step of modeling the power grid as a graph and obtaining graph structure data includes: taking the bus in the power grid as nodes and the connection relationship between the bus in the power grid as edges, and converting the original dataset Z into graph structure data according to the number of nodes and the feature number of each node; The introduction of the sliding window mechanism includes setting the sliding step size of the sliding window mechanism to 1 time step.

3. The method for locating power grid spoofing data injection attacks based on adaptive spatiotemporal graph neural networks according to claim 1, characterized in that, The K-nearest neighbor and minimum threshold method includes: Define a positive integer γ and a non-negative threshold ∈ . Retain the attention scores of those points that are greater than ∈ or belong to the top γ largest elements in the corresponding row, and set the remaining attention scores to 0. Reconnect the adjacency matrix A. * ∈R n×n Represented as: In the formula, Let S represent the set of the first γ largest elements in the i-th row of matrix S.

4. The method for locating power grid spoofing data injection attacks based on adaptive spatiotemporal graph neural networks according to claim 1, characterized in that, The specific process of global feature fusion using the improved Transformer module is as follows: X o =ReLU(XW0+b0) (5) Q i =X o W i Q (6) K i =X o W i K (7) V i =X o W i V (8) X MHSA =Concat(head1,head2,…,head n ) (9) In the formula, X o The feature representation obtained after linear transformation and ReLU activation of the input node feature matrix X, where W0 is the weight matrix, b0 is the bias vector, and Q is the weight vector. i K i and V i They are respectively made by X o The query, key, and value matrix W obtained by linear transformation i Q W i K and W i V Learnable weight matrices corresponding to the query, key, and value matrices, respectively; X MHSA The output is the result through a multi-head self-attention mechanism. For K i The transpose of d k The vector dimension of the query and key matrix; To overcome the vanishing gradient problem and alleviate overfitting, residual connections and Dropout operations are introduced: X fusion =Dropout(X o +X MHSA )+X o (11) In the formula, X fusion For the improved output of the Transformer module; The output X of the improved Transformer module fusion Simplified representation: FusionBlock(X): X fusion =FusionBlock(X) (12) The fused feature vector X fusion With the original feature vector X o Connection as a self-embedded X node ego for: X ego =X fusion ||X o (13)。 5. The method for locating power grid spoofing data injection attacks based on adaptive spatiotemporal graph neural networks according to claim 1, characterized in that, Step S5 specifically includes: Again, using the improved Transformer module for X out Processing is performed; based on the multi-label classification results output by the model, and the probability of the node output at the last time step, the nodes with higher anomaly probabilities are identified as targets of attack. The embedding of the final node into X out The following are examples of nodes that are anomalies when input to the classification layer: The Softmax layer outputs the probability of each node being abnormal, identifying potentially attacked nodes and their corresponding locations. The specific process involves a classification layer that determines the method of node anomaly. output=softmax(FusionBlock(X out W1)) (20) In the formula, W1 is a trainable weight matrix; X out The final node is embedded; the output is the multi-label classification result.

6. A device for locating power grid spoofing data injection attacks, used to implement the power grid spoofing data injection attack location method based on adaptive spatiotemporal graph neural network as described in any one of claims 1-5, characterized in that, include: Data acquisition module: used to acquire power grid operation data and generate historical data in a graph structure time series. Preprocessing module: Used to preprocess the acquired power grid operation data, including normalizing the data, converting it into graph structure data, and introducing a sliding window mechanism to obtain the processed power grid operation data; Adaptive Spatiotemporal Graph Neural Network Model Module: Inputs the processed data into the trained model to obtain the probability of each node being attacked, thereby locating the attacked main line and taking further protective measures; The adaptive spatiotemporal graph neural network model includes: a graph structure learning module, a global feature fusion module, a spatiotemporal feature fusion module, and a Softmax classification module; the graph structure learning module is used to mine the intrinsic relationships between nodes and identify new reliable neighbors for each node; the global feature fusion module fuses more valuable information from the entire graph; the spatiotemporal feature fusion module extracts the temporal and spatial features of the data; and the Softmax classification module is used to classify the input data and output the final detection result.

Citation Information

Patent Citations

  • Machine room anomaly detection method and device based on graph structure and abnormal attention mechanism

    CN115018021A

  • Motion recognition method and system based on fusion graph convolutional network and Transform network

    CN115100574A