Electric power system abnormal remote signaling detection method based on graph auto-encoder model

Through the graph autoencoder model and deep learning technology, combined with residual enhancement graph attention encoder and physical prior, the problem of high false alarm rate of remote signal data abnormal detection in the power system is solved, efficient and intelligent detection of remote signal abnormality is achieved, and the grid data quality and system operation reliability are improved.

CN120597196AActive Publication Date: 2025-09-05SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510672861.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the topological changes of dynamic graphs in power systems, resulting in a high false alarm rate for abnormal detection of remote signal data, and traditional methods are difficult to capture abnormal patterns in dynamic processes, affecting the accuracy of state estimation and the safe and stable operation of the power system.

Method used

The remote signal detection method of power system anomaly based on the graph autoencoder model is adopted. By combining the graph autoencoder model and deep learning technology, combining residual-enhancing graph attention encoder, node-edge joint attention mechanism, bidirectional long and short-term memory network and physical prior, sensitive perception and accurate recognition of remote signal abnormalities are achieved.

Benefits of technology

It achieves efficient and intelligent detection of telesignaling anomalies, reduces the false alarm rate, improves the quality of power grid data and the reliability of system operation, and has physical explainability and efficient anomaly detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597196A_ABST
    Figure CN120597196A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power system abnormal remote signaling detection method based on a graph auto-encoder model, and the method comprises the steps: collecting measurement data of an electric power system, carrying out the preprocessing, obtaining a graph data set with an abnormal label, and dividing the graph data set into a training set and a test set; inputting the graph data in the training set into the graph auto-encoder model for training; after training is completed, abnormal score distribution is counted based on normal edge samples in a training set, a threshold value is set to serve as a follow-up judgment basis, and threshold value selection takes the accuracy rate and the recall rate on a test set as an adjustment and optimization target; in a test stage, image data in a test set are input to carry out edge feature reconstruction and anomaly scoring, anomaly judgment is carried out on edges in combination with a set threshold value, and a preliminary abnormal edge detection result is output; and the output abnormal edge detection result is input into the graph restoration module, the restored edge structure and edge features are output, the damaged remote signaling state in the power grid is restored, and the integrity of the graph structure and the operation credibility of the power system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal remote signaling detection of power systems, and in particular to a method for detecting abnormal remote signaling of power systems based on a graph autoencoder model. Background Art

[0002] With the rapid development of new power systems, the scale and complexity of data processed by power systems are rapidly increasing, posing new challenges for information collection, processing, and analysis. To ensure the safe and stable operation of power systems, power system dispatch centers require a comprehensive understanding of the system's real-time operating status and the ability to predict its operating trends. State estimation is a key technology in situational awareness systems, aiming to accurately perceive the system's operating status. Accurate state estimation results are crucial for power system applications such as optimal power flow calculations, load forecasting, and economic dispatch. However, the accuracy of state estimation is highly dependent on the accuracy and timeliness of collected data. In actual operation, monitoring systems are prone to misjudgment due to switch (also known as telesignaling) status misreporting (e.g., a closed switch mistakenly reported as open), node measurement noise (e.g., abnormal voltage and current fluctuations), and dynamic grid topology changes (e.g., fault isolation or load switching). This can lead to erroneous dispatch decisions and even cascading failures. Furthermore, SCADA and other measurement equipment can generate topological errors or poor measurement data during information collection due to factors such as device failures, data transmission and communication anomalies, device aging, environmental interference, and sudden changes in system operating status. These abnormal data not only seriously affect the accuracy of state estimation, but may also lead to a decline in the convergence performance of power system information processing, thereby increasing the risk of grid operation. Therefore, effective grid topology anomaly detection and timely correction of measurement data are of great practical significance for improving the reliability of state estimation results, ensuring data quality, and enhancing the safe and stable operation of power systems.

[0003] Information collection in power systems typically relies on SCADA systems, which collect data every 15 minutes. Consequently, the data exhibits significant dynamic characteristics. Traditional static graph modeling methods cannot fully reflect the dynamic changes in the power system's operating status. In contrast, dynamic graph modeling can more accurately capture the topology of the power grid and the evolution of measurement data over time, more closely resonating with actual operating conditions. However, the dynamic evolution of nodes and edges in dynamic graphs poses significant challenges in data analysis, especially anomaly detection. On the one hand, nodes and edges in dynamic graphs may change frequently, and anomaly characteristics may be dynamic, local, or global. On the other hand, this variability increases the complexity of the data, making it difficult for traditional anomaly detection methods to effectively capture abnormal patterns in dynamic processes.

[0004] Anomaly detection, a critical task for ensuring power system operational safety and data integrity, aims to identify and locate unusual events in data sets that significantly deviate from normal patterns. Timely and effective detection of anomalies is crucial for improving system reliability and mitigating potential safety risks. However, traditional anomaly detection methods typically rely on static feature analysis, making it difficult to effectively handle the dynamic characteristics of data and achieving efficient and accurate detection in highly dynamic environments.

[0005] In recent years, deep learning-based anomaly detection methods for dynamic graphs have made significant progress and have been successfully applied to a variety of fields, including social networks, knowledge graphs, and network security. These methods typically leverage the inherent temporal characteristics and relational structure of dynamic graphs to effectively identify anomalous events in data, thereby improving the security and integrity of various network systems. For example, NetWalk proposed a node vector representation method based on a deep autoencoder, combining clique embedding and reservoir sampling techniques to achieve efficient updates of dynamic graphs and effectively identify structural anomalies in the network through clustering algorithms. AddGraph constructed a gated graph convolutional network (GCN) framework based on an attention mechanism, which can simultaneously capture short-term and long-term patterns in dynamic graphs and introduce negative sampling and edge loss functions to detect anomalous edges. Furthermore, the TADDY method introduced a Transformer network architecture to simultaneously model the spatiotemporal characteristics of dynamic graphs, effectively improving the performance of dynamic graph anomaly detection. Although dynamic graph anomaly detection has achieved promising results in these fields, research in the power system field is still in its early stages, and existing research results are relatively limited. Current research on anomaly detection based on dynamic graphs in the power system field primarily focuses on detecting false data injection attacks (FDIAs). For example, the DynWatch-Local model utilizes a graph distance metric driven by neighborhood knowledge to dynamically estimate the distribution of measurement data of the current power grid state based on historical data for anomaly detection. The TGNN method combines graph neural networks and gated recurrent networks (GRUs) to effectively detect and locate false data injection attacks in power systems. The GGNN method, on the other hand, uses an attention mechanism to fuse power system operating data with topological information, extracting spatiotemporal features of nodes to improve detection effectiveness. However, these methods primarily focus on telemetry data, and research on anomaly detection in power system topological connectivity anomalies, namely telesignaling data, has yet to be fully explored.

[0006] Because power system measurement data is highly dependent on the dynamically changing grid topology and load conditions, its contextual dependence is strong, posing a significant challenge to accurately modeling the conditional distribution of measurement data. Existing methods that ignore background knowledge and physical laws of the power grid can easily lead to high false positive rates, reducing the reliability of practical applications. Furthermore, most existing methods are limited to a single scale when extracting features, making it difficult to simultaneously and efficiently capture the complex spatiotemporal features at both local and global scales in dynamic graphs. This leads to a high probability of missed or false positives, especially in complex and changing scenarios involving abnormal events. Summary of the Invention

[0007] The purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method for detecting abnormal telesignaling in power systems based on a graph autoencoder model. By integrating the graph autoencoder model with deep learning technology, it can effectively mine the potential structural patterns in power system data, comprehensively integrate dynamic spatiotemporal characteristics and grid physical knowledge, and achieve sensitive perception and accurate identification of telesignaling anomalies, providing a new solution with high efficiency, intelligence and physical interpretability for telesignaling anomaly detection and structural repair in power systems.

[0008] To achieve the above objectives, the present invention provides a technical solution: a method for detecting abnormal remote signaling in a power system based on a graph autoencoder model, comprising the following steps:

[0009] S1. Collect measurement data from the power system and construct it into a graph with node and edge relationships. Simulate two types of telesignaling anomalies in the graph, called edge anomalies, namely false disconnection anomalies and false connection anomalies. Through anomaly injection and annotation, generate a labeled graph dataset containing normal edge and abnormal edge samples, and finally divide it into training and test sets.

[0010] S2. Input the graph data in the training set into the designed graph autoencoder model for training. The encoder part of the model uses a residual-enhanced graph attention mechanism and a node-edge joint attention mechanism to achieve joint modeling of graph structure and temporal information, maintain the stability of node representation and extract key structural features. It also combines a bidirectional long short-term memory network to capture the dynamic temporal changes of nodes, fuse multi-scale temporal features, and generate a robust edge embedding representation. The graph decoder part of the above edge embedding representation input model performs edge feature reconstruction and anomaly score prediction. The graph decoder combines the physical prior of the power system to construct a physical constraint embedding space, realizes edge attribute reconstruction through the inter-edge self-attention mechanism, and generates anomaly scores reflecting the degree of structural or attribute deviation. A joint loss function is used for optimization during training, including reconstruction loss, Kirchhoff current balancing loss, and classification loss for anomaly detection.

[0011] In the training phase, the anomaly score distribution is calculated based on normal edge samples and the threshold is optimized, with the precision and recall of the test set as the tuning target. In the testing phase, edge features are reconstructed and anomaly scores are performed on the input graph data. Abnormal edges are determined based on the threshold, and preliminary anomaly detection results are output.

[0012] S4. The detected abnormal edges are input into the graph convolution repair module, which integrates graph topology, node features and edge attribute information. Through the neighborhood information propagation mechanism, the connection status and electrical parameters of the abnormal edges are corrected. The repaired edges can accurately restore the true topological relationship of the power grid, effectively eliminate telesignaling errors, and improve the quality of power grid data and the operational reliability of the power system.

[0013] Further, in step S1, the measurement data of the power system is collected, and the collected measurement data is constructed into graph data with node and edge relationships, and two typical types of telesignaling anomalies, called edge anomalies, are simulated or identified in the graph data, including false disconnection anomalies and false connection anomalies, and the graph data is injected, labeled and normalized for preprocessing, and finally a graph data set with anomaly labels is formed, that is, the graph data set contains normal edge samples and abnormal edge samples; the graph data set is divided into a training set and a test set; the characteristics of the nodes include node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power, and the characteristics of the edges include starting side active power PF, starting side reactive power QF, terminal side active power PT, terminal side reactive power QT and power system state estimation measurement data in the graph adjacency matrix; the graph data is represented as a time series graph set T′ represents the total time, where the graph structure at each moment is defined as graph G t ={V t ,E t}, where V t represents the set of nodes in the power system at time t, E t Represents the edge set at time t; the node represents the bus, and the edge represents the connection relationship of the transmission branch, reflecting the operating topology and state of the power system at time t; each node v∈V t Contains the time series characteristics of node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power; each edge e∈E t Includes active power PF measured at the starting point and end point t PT t and reactive power QF t , QT t The topological connection state of the power grid is represented by the adjacency matrix express, represents the topology of the power grid at time t, that is, whether there is a connection between node i and node j, Indicates that there is a connection line between node i and node j at time t, Indicates that there is no connection between node i and node j at time t, and time t indicates that the adjacency matrix is ​​graph G t The structural description reflects the dynamic changes of the graph structure over time.

[0014] Furthermore, in step S2, during the training process, the graph autoencoder model first performs structural and temporal joint modeling on the graph data through the encoder, adopts a residual-enhanced graph attention encoder, and combines the node-edge joint attention mechanism based on the graph attention network to extract key feature information in the graph structure and maintain the stability of the original node representation. Subsequently, by introducing a bidirectional long short-term memory network, the dynamic change characteristics of the node in the time dimension are captured, and the temporal information of different scales is integrated to obtain a more robust edge embedding representation. The residual-enhanced graph attention encoder is a graph neural network layer designed for power system graph structure data. By fusing the node-edge joint feature representation and historical state memory, it effectively captures abnormal information. It includes: a node-edge joint attention mechanism based on the graph attention network, dynamic residual feature enhancement, a cross-attention historical state retention mechanism and a temporal information extraction module, which are specifically introduced as follows:

[0015] ①Node-edge joint attention mechanism based on graph attention network:

[0016] To capture the correlation between power line anomalies and the status of nodes at both ends, an improved graph attention mechanism is designed, namely the node-edge joint attention mechanism based on the graph attention network. Through the node-edge joint attention mechanism, node attributes and topological relationships are explicitly integrated, as follows:

[0017] Given the feature vector x of the i-th node i ∈R d and the feature vector e of the edge connecting node i and node j ij ∈R e , where R is the real number field, d is the dimension of each node feature vector, and e is the feature dimension of each edge; first, feature alignment is performed through linear projection:

[0018]

[0019]

[0020] Where W h is the weight matrix of node features; b h is the bias term of the node feature; W e is the edge feature weight matrix; b e is the edge feature bias; is the normalized feature of node i; is the edge feature after nonlinear activation; LayerNorm normalizes the node features to eliminate dimensional differences; ReLU nonlinearly activates the edge features;

[0021] Use the graph attention mechanism for feature interaction and calculate the edge attention weight:

[0022]

[0023] Where, represents the normalized attention weight of node j to node i in the lth attention head; a is a learnable attention vector used to measure the importance of the node-edge joint feature; l represents the lth attention head, || represents the feature splicing operation; represents the normalized features of node i in the lth attention head; represents the normalized features of node j in the lth attention head; Represents the edge features after nonlinear activation in the lth attention head;

[0024] Finally, the node representation h is obtained through multi-head aggregation attn :

[0025]

[0026] Where L is the total number of attention heads, N(i) is the neighbor set of node i;

[0027] ②Dynamic residual feature enhancement:

[0028] The topology of the power system changes dynamically. To avoid the degradation of deep network characteristics, an adaptive residual connection mechanism is designed as follows:

[0029] Define the residual mapping function:

[0030]

[0031] Where, d in is the dimension of the input feature, i.e. x i Dimension; d out Represents the dimension of the output feature, that is, W after mapping res x i Dimension; W res is the weight matrix of the residual mapping; F res It is a dynamic residual mapping function that can solve the key problems of dimension mismatch and gradient instability in deep learning models;

[0032] Dynamically fuse attention features and residual features through a gating mechanism:

[0033] g=σ(W g [hattn ||F res (x i )]);

[0034] h fusion =g⊙h attn +(1-g)⊙F res (x i );

[0035] Where σ is the Sigmoid function, which smooths the gate control changes and alleviates the gradient vanishing problem of deep networks; ⊙ represents element-by-element multiplication; g is the gating vector, whose value is between [0, 1]. The fusion ratio of attention and residual information obtained by the gating mechanism is used to control the dynamic fusion ratio of the two types of information. The larger the g, the more reliable the current attention feature is, and the smaller the g, the more dependent on the historical state; h fusion is the fused feature vector, which combines attention information and residual information; W g is the weight matrix of the gating mechanism; this gating mechanism can adaptively adjust the fusion ratio of original features and attention features, showing better adaptability in the scenario of dynamic topology changes in the power system;

[0036] ③ Cross-attention history state retention mechanism:

[0037] In order to preserve the state evolution characteristics of the node itself, a cross-layer attention mechanism is introduced. The historical residual features are used as the knowledge base, and key historical information is adaptively selected and integrated with the current state. The current layer output is used as the query, and the historical residual features are used as the key-value:

[0038] Q=h fusion W q ;

[0039] K=F res (x i )W k ;

[0040] V=F res (x i )W v ;

[0041] Calculate cross-layer attention weight β i :

[0042]

[0043] Through cross-layer attention and nonlinear changes, historical information is fused with current features to generate the final output h out :

[0044] h out =LayerNorm(ELU(β iV+h fusion ));

[0045] ELU activation function:

[0046]

[0047] In the formula, ELU has a non-zero gradient in the negative interval, which prevents node features from not being updated for a long time and alleviates the Dead ReLU problem; the exponential term e x Smooths negative value changes, suitable for scenarios with positive and negative values ​​in power characteristics; W q 、W k 、W v is the cross-attention query vector Q, key vector K, and value vector V projection weight; x is the input eigenvalue; α is the scaling factor of the ELU function in the negative interval, which controls the amplitude of the negative value mapping; the cross-attention history state retention mechanism enables nodes to adaptively retain key historical state features when aggregating neighborhood information, enhancing the modeling ability of the continuous evolution of nodes and edges;

[0048] ④Time series information extraction module:

[0049] When detecting abnormal edges in power systems, it is found that the abnormal state of edges is not only affected by the node state at the historical moment, but can also be triggered by contextual factors at the future moment. To this end, a bidirectional long short-term memory network (Bi-LSTM) is introduced to enhance the modeling capability of time series context. Bi-LSTM considers both forward (historical) and backward (future) information in the sequence, and can capture more comprehensive time-dependent characteristics.

[0050] For each target edge, its state at multiple time steps is fed into the Bi-LSTM as an input sequence. The Bi-LSTM consists of two parallel LSTM layers, which transmit information in the forward and reverse directions of time. The forward LSTM layer extracts historical context features, while the backward LSTM layer models the potential impact of future moments. The hidden states output by the two are fused at each time step to form an edge temporal representation containing bidirectional temporal information.

[0051] The fused representation not only reflects the state of the edge at the current moment, but also implicitly integrates the state change trend related to the adjacent nodes in the previous and next time windows, thereby providing more discriminative temporal feature support for subsequent anomaly scores.

[0052] Furthermore, in step S2, the extracted edge embedding representation is input into the graph decoder for edge feature reconstruction and anomaly score prediction. A physical prior-based graph decoder is used to construct a physical constraint embedding space in combination with the physical information of the power system. The edge attribute features are reconstructed using the inter-edge self-attention mechanism. The physical prior-based graph decoder includes: a dynamic edge feature generation mechanism, an inter-edge self-attention mechanism, and a physical prior fusion mechanism. The details are as follows:

[0053] ① Dynamic edge feature generation mechanism:

[0054] We choose to dynamically generate edge features based on nodes rather than directly relying on original edge features. That is, we implicitly model edge features through the evolution of node states to solve the problem that historical edge features are not fully recorded or are interfered with by noise. We propose a dynamic edge feature generation module that combines a graph attention mechanism with an edge feature generator to dynamically generate edge features from the node's potential representation. The graph attention network generates a high-level semantic representation of the node by dynamically aggregating neighborhood information, while the edge feature generator dynamically generates edge features using the fused features of the source and target nodes.

[0055] Using a multi-head GAT layer, each head independently learns different attention weights, and finally integrates multi-view features by splicing:

[0056]

[0057] Where L is the total number of attention heads; l is the number of the lth attention head; σ is the Sigmoid function, the smooth gate controls the change and alleviates the gradient vanishing problem of deep networks; W l Represents the learnable linear transformation weight matrix corresponding to the lth attention head, which is used to transform the neighbor node feature h j Mapped to the attention space for subsequent weighted summation; represents the normalized attention weight of node j to node i in the lth attention head; h i is the multi-head attention output feature of node i; || represents the feature concatenation operation; N(i) is the neighbor set of node i; j∈N(i) means that node j is a neighbor node of node i; h j The feature vector representing node j is the neighbor node feature of node i;

[0058] The edge feature generator considers both the high-level abstract features of nodes extracted by the GAT layer and the original latent features, thereby improving the model's tolerance to sparse or noisy data:

[0059] e f =[h src ||h dst ||x src ||x dst];

[0060] Where, e f is the fusion edge feature, which is h src ||h dst ||x src ||x dst The concatenation result of the four features is used to comprehensively represent the structure and attribute information; h src is the node starting feature output by the GAT layer; h dst is the node termination feature output by the GAT layer; x src is the original starting feature of the node; x dst is the original termination feature of the node;

[0061] ②Inter-edge self-attention mechanism:

[0062] To effectively model edge representation, we designed an edge feature decoding module based on a self-attention mechanism. In the abnormal edge detection task, the goal is to identify edges with potential faults or measurement errors in the power system. Because edge power flow measurements may be missing or falsified, we construct edge representations based on more robust node states.

[0063] The designed edge feature decoding module takes the state features of the nodes at both ends of the edge as input, explicitly models the interaction relationship between nodes through the self-attention mechanism, and obtains the edge representation containing contextual information;

[0064] The starting node feature e connected by a given edge fs With the terminal node feature e ft First, the two are spliced ​​together to form the basic feature representation h of the edge base :

[0065] h base =[e fs ||e ft ];

[0066] In the edge feature decoding module, the self-attention mechanism is used to model the edge representation. The input of the self-attention mechanism is the basic feature representation h base , the output of the self-attention mechanism is the edge representation h attn :

[0067] MultiheadAttn(Q,K,V)=Concat(head1,head2,…,head L )W o ;

[0068] Among them, the attention mechanism adopts a multi-head attention mechanism, that is, the original input is decomposed into multiple heads, the attention weight corresponding to each head is calculated separately, and then the attention weights obtained by each head are spliced ​​to obtain the final attention weight; the Lth head W o is the output projection matrix, W L is the weight matrix of the L-th head, Attn is the attention function, Q, K, V are query, key and value vectors respectively;

[0069] Then, h base With h attn Splicing and inputting a multi-layer perceptron to obtain the final edge feature representation E pred :

[0070] E pred =MLP([h base ||h attn ]);

[0071] ③Physical prior fusion mechanism:

[0072] To enhance the model's ability to perceive power physical information, a physical feature adapter is designed to map the original physical features into a latent space representation compatible with graph neural networks. The adapter is then introduced into the model as part of the node or edge input, thereby providing prior information on the power grid's operating status.

[0073] First, the physical features are used as input and mapped to a latent space representation compatible with the graph neural network through a linear transformation. The physical feature adapter receives the original physical features of the node, including: voltage amplitude V m , the node active power P of the generator and load, the node reactive power Q of the generator and load, the starting side active power PF, the starting side reactive power QF, the terminal side active power PT and the terminal side reactive power QT, and generate the latent space physical feature h through single-layer linear projection and nonlinear activation function Tanh phy , expressed as:

[0074] h phy =Tanh(W adapter ·[V m ,P,Q,PT,PF,QT,QF]+b adapter );

[0075] Where W adapter and b adapter To be a learnable parameter, the Tanh function constrains the feature range to [-1, 1] to avoid numerical divergence of physical quantities;

[0076] The physical feature adapter transforms the latent space physical feature h phy and edge feature representation Epred Fusion is performed to explicitly inject the neighborhood knowledge constraint H:

[0077] H=MLP([E pred ,h phy ]);

[0078] During the execution of the physical prior fusion mechanism, prior information is implicitly injected through the joint optimization of the feature space. The loss function imposes penalties on predictions that violate physical laws. Backpropagation forces the physical feature adapter to learn feature representations that conform to domain knowledge, fuse graph features with physical latent features, and generate predictions that conform to domain laws.

[0079] Furthermore, in step S2, the edge feature reconstruction result output by the graph decoder and the anomaly score are input into the joint loss function for optimization. The joint optimization objective includes three types of loss terms: reconstruction loss, which is used to improve the reconstruction accuracy of edge attributes; Kirchhoff current balance loss, which is used to ensure the physical consistency of the system state; classification loss, which is used to optimize the anomaly detection performance according to the anomaly score;

[0080] The total loss function L′ consists of three terms:

[0081] L′=αL cls +βL recon +γL kcl ;

[0082] Where α, β and γ are hyperparameters, which control the weights of the three constraints respectively; L cls 、L recon and L kcl They correspond to classification loss, reconstruction loss and Kirchhoff current balancing loss respectively; the specific introductions of classification loss, reconstruction loss and Kirchhoff current balancing loss are as follows:

[0083] ①Classification loss L cls Focal Loss is used to address the serious imbalance between abnormal and normal edge samples. By focusing more on abnormal edge samples, the model is prevented from being biased towards the majority class.

[0084] L cls =-Σ(1-p t ) γ log(p t );

[0085] Where p t is the model's predicted probability for the true category. If the sample is abnormal, it is called a positive sample, then p t is the probability that the model predicts an abnormality; if the sample is normal, it is called a negative sample, then p tis the probability that the model predicts normal; γ is the focusing parameter that adjusts the weight of difficult and easy samples. When γ>0, the weight of easy-to-classify samples is reduced, forcing the model to focus on difficult-to-classify samples. When γ=0, it degenerates to the standard cross entropy loss.

[0086] ②Reconstruction loss L recon Physically constrained MSE is used to ensure that the power flow results predicted by the model conform to the actual physical laws of the power grid, preventing the model from fitting data only based on statistical characteristics, and improving generalization performance;

[0087]

[0088] Among them, the power flow is predicted by minimizing the model With the real power flow e ij The mean square error between them forces the model output to conform to the actual power system state;

[0089] ③Kirchhoff current balancing loss L kcl By explicitly introducing node power conservation constraints, the physical consistency and credibility of model predictions are further strengthened, reducing false positives caused by non-physical predictions. The goal of this approach is to transform the power system's energy conservation law into a differentiable mathematical form and embed it into the training process of the machine learning model, enabling the model to not only learn the statistical laws in the data but also follow the underlying physical rules. To achieve this goal, the following core concepts must be clearly defined: edge power flow, node power conservation, and KCL constraints.

[0090] ① Edge power flow definition:

[0091] For the edge connecting the nodes v i -v j , define the power flow characteristics:

[0092]

[0093]

[0094] Where, e ij Represents the slave node v i Points to node v j The edge eigenvector of ij [0]、e ji [1], e ij [2], e ji [3], Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input; each edge e contains bidirectional active and reactive power flow information i→j and j→i;

[0095] ②Definition of node power conservation:

[0096] ΔP i =P i gen -P i load -P i in -P i out ;

[0097]

[0098] Where ΔQ i For node v i The reactive power imbalance, ΔP i For node v i The active power imbalance, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input, P i gen For node v i The active power of the generator, P i load For node v i The load active power, For node v i The reactive power of the generator, For node v i The reactive power of the load;

[0099] ③KCL constraint definition:

[0100] By balancing node power, the model output is forced to follow physical laws, thus reducing misjudgment of anomalies:

[0101]

[0102] Where V is the node set.

[0103] Furthermore, in step S3, after training is completed, the anomaly score distribution is statistically calculated based on the normal edge samples in the training set, and a threshold is set as the basis for subsequent judgment. The threshold selection is optimized based on the precision and recall rate on the test set. In the testing phase, the graph data in the test set is input for edge feature reconstruction and anomaly scoring. The edges are judged to be abnormal based on the set threshold, and preliminary abnormal edge detection results are output. The calculation of the anomaly score is based on the reconstructed features of the edge by the graph decoder of the graph autoencoder model, which is used to measure the degree of deviation of the edge in structure or attribute. The larger the reconstruction error, the less the edge conforms to the normal pattern in the graph structure, and is therefore regarded as a potential abnormal edge.

[0104] Let edge e = (s, v), then the anomaly score score(e) of the edge is the reconstruction error of the graph autoencoder model on the edge, that is:

[0105]

[0106] Where, Edge features reconstructed from the physical prior-based graph decoder part of the graph autoencoder model; f is the original edge feature; ||·||2 is the L2 norm; s is the starting node of edge e, v is the ending node of edge e, and e=(s,v) is the edge connecting s and v. The anomaly score score(e) is used to measure the ability of the graph autoencoder model to fit the edge structure. The larger the error, the more likely the edge is an anomaly.

[0107] After obtaining the anomaly score of each edge, a threshold τ is set. When the anomaly score score(e) of an edge satisfies score(e)>τ, the edge is judged to be an anomaly edge. The selection of the threshold is tuned according to the performance indicators on the test set to achieve a balance between precision and recall.

[0108] Furthermore, in step S4, for edges judged to be abnormal, the graph repair model uses a graph convolutional neural network to repair them. The graph convolutional neural network propagates the node representation on the updated graph structure containing the abnormal edges, so that the node embedding integrates the semantic information of the surrounding context. Subsequently, the graph repair model takes the repaired node representation as input, and reconstructs the properties of the abnormal edge using a multi-layer perceptron by splicing the features of the nodes at both ends of the edge. For structurally abnormal edges, if the connection is disconnected, the graph repair model introduces candidate edges and uses the updated node representation of the graph convolutional neural network to determine the rationality of its connection, and restores or removes these connections when necessary. Finally, the graph repair model outputs the repaired graph structure and edge features to achieve joint repair of telemetry data errors and structural anomalies, thereby enhancing the accuracy of power system measurement data.

[0109] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0110] 1. This paper proposes a physically constrained graph autoencoder model. By converting Kirchhoff's current law into a physically constrained loss, this model, along with the classification and reconstruction losses of the graph autoencoder, forms a joint optimization objective, achieving a synergistic improvement in anomaly detection accuracy and compliance with physical laws.

[0111] 2. The present invention designs a residual-enhanced graph attention encoder, constructs a coding structure coupled with a node-edge joint attention mechanism based on a graph attention network, dynamic residual feature enhancement, a cross-attention history state retention mechanism, and a temporal information extraction module. The node-edge joint attention mechanism is used to fuse topology and features, and the gating mechanism is combined to effectively alleviate the problem of deep feature degradation.

[0112] 3. The present invention constructs a graph decoder based on physical priors, introduces a physical adapter, a dynamic edge feature generation mechanism, and an edge self-attention mechanism into the graph decoder, realizes the embedding of physical feature information into the model latent space, and improves the reconstruction capability and sensitivity to topological anomalies.

[0113] 4. The present invention integrates multimodal and multi-scale feature representation capabilities, and simultaneously utilizes graph structure features, node electrical quantities and time series information to perform dynamic multi-scale modeling, thereby improving the generalization ability and robustness of the model in complex scenarios.

[0114] 5. The present invention achieves end-to-end high-performance anomaly detection. By introducing reconstruction error as the anomaly score and combining it with physical constraint optimization objectives, the accuracy of edge anomaly detection is significantly improved and the false alarm rate is significantly reduced.

[0115] 6. Experiments based on the IEEE118 standard power grid test set demonstrate that the graph autoencoder model proposed in this paper exhibits excellent performance under various anomaly patterns. In experiments with a 5% anomaly ratio, the model maintained a high accuracy of 98.7% and a precision of 96.5%, while achieving a recall rate of 92.3%, demonstrating that it can accurately distinguish between normal and abnormal edges in low anomaly ratio scenarios while maintaining comprehensive coverage of abnormal situations. When the anomaly ratio increases to 10%, the model performance remains stable, with only slight fluctuations in accuracy (98.5%) and precision (96.2%). The recall rate of 91.0% verifies that it can maintain excellent detection coverage when the number of anomalies increases. Even under the high anomaly ratio of 15%, the model still demonstrates significant robustness, with an accuracy of 98.0%, a precision of 95.5%, and a recall rate of 89.8%, demonstrating its reliable detection capability in high-noise environments. These experimental results fully verify that the model has accuracy, physical consistency and engineering adaptability in the task of abnormal edge telesignaling detection in power systems. Its performance shows gradient stability as the abnormality ratio increases, providing high reliability support for the stable operation of smart grids. BRIEF DESCRIPTION OF THE DRAWINGS

[0116] Figure 1 Flowchart of the method of the present invention.

[0117] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0118] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0119] like Figure 1 and Figure 2 As shown, this embodiment discloses a method for detecting abnormal remote signaling of a power system based on a graph autoencoder model, the specific details of which are as follows:

[0120] S1. Collect measurement data from the power system and construct the collected measurement data into graph data with node and edge relationships. Simulate or identify two typical telesignaling anomalies in the graph data, called edge anomalies, including false disconnection anomalies and false connection anomalies. Perform anomaly injection, annotation, and normalization preprocessing on the graph data to ultimately form a graph dataset with anomaly labels, that is, the graph dataset contains normal edge samples and abnormal edge samples; divide the graph dataset into a training set and a test set.

[0121] The characteristics of the nodes include node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power; the characteristics of the edges include starting side active power PF, starting side reactive power QF, end side active power PT, end side reactive power QT and power system state estimation measurement data in the graph adjacency matrix; the graph data is represented as a time series graph set T′ represents the total time, where the graph structure at each moment is defined as graph G t ={V t ,E t}, where V t represents the set of nodes in the power system at time t, E t Represents the edge set at time t; the node represents the bus, and the edge represents the connection relationship of the transmission branch, reflecting the operating topology and state of the power system at time t; each node v∈V t Contains the time series characteristics of node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power; each edge e∈E t Includes active power PF measured at the starting point and end point t PT t and reactive power QF t , QT t The topological connection state of the power grid is represented by the adjacency matrix express, represents the topology of the power grid at time t, that is, whether there is a connection between node i and node j, Indicates that there is a connection line between node i and node j at time t, Indicates that there is no connection between node i and node j at time t, and time t indicates that the adjacency matrix is ​​graph G t The structural description reflects the dynamic changes of the graph structure over time.

[0122] In this embodiment, the collected measurement data is specifically the NREL-118 test system dataset, which uses the NREL-118 test system dataset released by the National Renewable Energy Laboratory (NREL) of the United States. This dataset is reconstructed based on the IEEE 118 node standard test system and integrates the power generation capacity and load characteristics of the 2024 public case database of the Western Interconnection (WECC) of the United States. The dataset spans from 0:00 on January 1, 2024 to 24:00 on December 31, 2024, covering 8760 hours of synchronized time series data throughout the year. The dataset includes 3 interconnected regions, 118 buses, 186 transmission lines, and 327 generators, covering 9 types of power generation technologies (including thermal power, wind power, photovoltaic power, etc.), and includes hourly actual values ​​(Real-Time, RT) and day-ahead forecast values ​​(Day-Ahead, DA) of load, wind power output, and photovoltaic output time series. This dataset supports deep coupling of synchronized time series data with the physical model of the power grid, and supports dynamic topology anomaly detection and operation status deduction.

[0123] For the NREL-118 test system dataset, we simulated three types of abnormal situations:

[0124] The first type involves disconnected lines that should be connected. This involves randomly selecting edges from the original topology and removing them to create disconnection anomalies. During this removal process, the overall network topology remains connected to avoid isolated nodes.

[0125] The second type involves connecting a line that should be disconnected. New edges are randomly inserted between previously unconnected pairs of nodes, and their features are randomly initialized. The active power features of the new edges follow a Gaussian distribution, while the reactive power features are randomly sampled from the features of existing healthy edges.

[0126] The third type is abnormal line characteristics: some edges are randomly selected from the remaining normally operating lines, and their active power characteristics are adjusted to ensure that the total input / output power of the edges remains unchanged.

[0127] For this data set, the overall anomaly proportions were set at 5%, 10%, and 15%, respectively, of which line connection anomalies (should be connected but disconnected) accounted for 30%, line misconnection anomalies (should be disconnected but connected) accounted for 30%, and line feature anomalies accounted for 40%.

[0128] S2. Input the graph data in the training set into the designed graph autoencoder model for training. During the training process, the graph autoencoder model first performs structural and temporal joint modeling on the graph data through the encoder, adopts the residual enhanced graph attention encoder, and combines the node-edge joint attention mechanism based on the graph attention network to extract the key feature information in the graph structure and maintain the stability of the original node representation. Subsequently, by introducing the bidirectional long short-term memory network, the dynamic change characteristics of the node in the time dimension are captured, and the temporal information of different scales is integrated to obtain a more robust edge embedding representation. The extracted edge embedding representation is input into the graph decoder for edge embedding. Feature reconstruction and anomaly score prediction use a physical prior-based graph decoder to build a physical constraint embedding space based on the physical information of the power system. The edge attention mechanism is used to reconstruct the attribute features of the edge, and the anomaly score of the edge is generated based on the reconstructed attribute features. It is used to measure the degree of deviation of the edge in structure or attribute. The edge feature reconstruction results and the anomaly score output by the graph decoder are input into the joint loss function for optimization. The joint optimization objective includes three types of loss terms: reconstruction loss, used to improve the reconstruction accuracy of edge attributes; Kirchhoff current balancing loss, used to ensure the physical consistency of the system state; classification loss, used to optimize the anomaly detection performance according to the anomaly score.

[0129] The residual-enhanced graph attention encoder is a graph neural network layer designed for power system graph structure data. By fusing node-edge joint feature representation with historical state memory, it effectively captures abnormal information. It includes: a node-edge joint attention mechanism based on the graph attention network, dynamic residual feature enhancement, a cross-attention historical state retention mechanism, and a temporal information extraction module. The details are as follows:

[0130] ①Node-edge joint attention mechanism based on graph attention network:

[0131] To capture the correlation between power line anomalies and the status of nodes at both ends, this paper designs an improved graph attention mechanism, namely a node-edge joint attention mechanism based on a graph attention network. Traditional GNNs only model the interactions between nodes and ignore the direct impact of edge features. Through the node-edge joint attention mechanism, node attributes and topological relationships are explicitly integrated, as follows:

[0132] Given the feature vector x of the i-th node i ∈R d and the feature vector e of the edge connecting node i and node j ij ∈R e , where R is the real number field, d is the dimension of each node feature vector, and e is the feature dimension of each edge; first, feature alignment is performed through linear projection:

[0133]

[0134]

[0135] Where W h is the weight matrix of node features; b h is the bias term of the node feature; W e is the edge feature weight matrix; b e is the edge feature bias; is the normalized feature of node i; is the edge feature after nonlinear activation; LayerNorm normalizes the node features to eliminate dimensional differences; ReLU nonlinearly activates the edge features;

[0136] Use the graph attention mechanism for feature interaction and calculate the edge attention weight:

[0137]

[0138] Where, represents the normalized attention weight of node j to node i in the lth attention head; a is a learnable attention vector used to measure the importance of the node-edge joint feature; l represents the lth attention head, || represents the feature splicing operation; represents the normalized features of node i in the lth attention head; represents the normalized features of node j in the lth attention head; Represents the edge features after nonlinear activation in the lth attention head;

[0139] Finally, the node representation h is obtained through multi-head aggregation attn :

[0140]

[0141] Where L is the total number of attention heads, N(i) is the neighbor set of node i;

[0142] ②Dynamic residual feature enhancement:

[0143] The topology of the power system changes dynamically, and traditional fixed residual connections are difficult to adapt to the feature fusion requirements under different operating conditions. To avoid the degradation of deep network features, an adaptive residual connection mechanism is designed, as follows:

[0144] Define the residual mapping function:

[0145]

[0146] Where, d in is the dimension of the input feature, i.e. x i Dimension; d outRepresents the dimension of the output feature, that is, W after mapping res x i Dimension; W res is the weight matrix of the residual mapping; F res It is a dynamic residual mapping function that can solve the key problems of dimension mismatch and gradient instability in deep learning models;

[0147] Dynamically fuse attention features and residual features through a gating mechanism:

[0148] g=σ(W g [h attn ||F res (x i )]);

[0149] h fusion =g⊙h attn +(1-g)⊙F res (x i );

[0150] Where σ is the Sigmoid function, which smooths the gate control changes and alleviates the gradient vanishing problem of deep networks; ⊙ represents element-by-element multiplication; g is the gating vector, whose value is between [0, 1]. The fusion ratio of attention and residual information obtained by the gating mechanism is used to control the dynamic fusion ratio of the two types of information. The larger the g, the more reliable the current attention feature is, and the smaller the g, the more dependent on the historical state; h fusion is the fused feature vector, which combines attention information and residual information; W g is the weight matrix of the gating mechanism; this gating mechanism can adaptively adjust the fusion ratio of original features and attention features. Compared with fixed residual connections, it shows better adaptability in scenarios with dynamic topological changes in power systems.

[0151] ③ Cross-attention history state retention mechanism:

[0152] In order to preserve the state evolution characteristics of the node itself, a cross-layer attention mechanism is introduced. The historical residual features are used as the knowledge base, and key historical information is adaptively selected and integrated with the current state. The current layer output is used as the query, and the historical residual features are used as the key-value:

[0153] Q=h fusion W q ;

[0154] K=F res (x i )W k ;

[0155] V=F res (x i )W v ;

[0156] Calculate cross-layer attention weight β i :

[0157]

[0158] Through cross-layer attention and nonlinear changes, historical information is fused with current features to generate the final output h out :

[0159] h out =LayerNorm(ELU(β i V+h fusion ));

[0160] ELU activation function:

[0161]

[0162] In the formula, ELU has a non-zero gradient in the negative interval, which prevents node features from not being updated for a long time and alleviates the Dead ReLU problem; the exponential term e x Smooths negative value changes, which is suitable for scenarios with positive and negative values ​​in power characteristics (such as bidirectional power flow); W q 、W k 、W v is the cross-attention query vector Q, key vector K, and value vector V projection weight; x is the input eigenvalue; α is the scaling factor of the ELU function in the negative interval, which controls the amplitude of the negative value mapping; the cross-attention history state retention mechanism enables nodes to adaptively retain key historical state features when aggregating neighborhood information, enhancing the modeling ability of the continuous evolution of nodes and edges;

[0163] ④Time series information extraction module:

[0164] When detecting abnormal edges in power systems, it is found that the abnormal state of edges is not only affected by the node state at the historical moment, but can also be triggered by contextual factors at the future moment. To this end, a bidirectional long short-term memory network (Bi-LSTM) is introduced to enhance the modeling capability of time series context. Bi-LSTM considers both forward (historical) and backward (future) information in the sequence, and can capture more comprehensive time-dependent characteristics.

[0165] For each target edge, its state at multiple time steps is fed into the Bi-LSTM as an input sequence. The Bi-LSTM consists of two parallel LSTM layers, which transmit information in the forward and reverse directions of time. The forward LSTM layer extracts historical context features, while the backward LSTM layer models the potential impact of future moments. The hidden states output by the two are fused at each time step to form an edge temporal representation containing bidirectional temporal information.

[0166] The fused representation not only reflects the state of the edge at the current moment, but also implicitly integrates the state change trend related to the adjacent nodes in the previous and next time windows, thereby providing more discriminative temporal feature support for subsequent anomaly scores.

[0167] The physical prior-based graph decoder includes a dynamic edge feature generation mechanism, an edge self-attention mechanism, and a physical prior fusion mechanism. The details are as follows:

[0168] ① Dynamic edge feature generation mechanism:

[0169] In the task of detecting abnormal edges in power systems, edge features usually change dynamically over time. If the original features corresponding to the edges are used directly, there may be problems with historical edge features not being fully recorded or being interfered with by noise. Therefore, we choose to dynamically generate edge features based on nodes rather than relying directly on the original edge features. That is, we implicitly model the edge features through the evolution of node states to solve the problem of historical edge features not being fully recorded or being interfered with by noise. We propose a dynamic edge feature generation module that integrates a graph attention mechanism and an edge feature generator to dynamically generate edge features from the node's potential representation. The graph attention network generates a high-level semantic representation of the node by dynamically aggregating neighborhood information, while the edge feature generator dynamically generates edge features using the fused features of the source and target nodes.

[0170] Using a multi-head GAT layer, each head independently learns different attention weights, and finally integrates multi-view features by splicing:

[0171]

[0172] Where L is the total number of attention heads; l is the number of the lth attention head; σ is the Sigmoid function, the smooth gate controls the change and alleviates the gradient vanishing problem of deep networks; W l Represents the learnable linear transformation weight matrix corresponding to the lth attention head, which is used to transform the neighbor node feature h j Mapped to the attention space for subsequent weighted summation; represents the normalized attention weight of node j to node i in the lth attention head; h i is the multi-head attention output feature of node i; || represents the feature concatenation operation; N(i) is the neighbor set of node i; j∈N(i) means that node j is a neighbor node of node i; h j The feature vector representing node j is the neighbor node feature of node i;

[0173] The edge feature generator considers both the high-level abstract features of nodes extracted by the GAT layer and the original latent features, thereby improving the model's tolerance to sparse or noisy data:

[0174] e f =[h src ||h dst ||x src ||x dst ];

[0175] Where, e f is the fusion edge feature, which is h src ||h dst ||x src ||x dst The concatenation result of the four features is used to comprehensively represent the structure and attribute information; h src is the node starting feature output by the GAT layer; h dst is the node termination feature output by the GAT layer; x src is the original starting feature of the node; x dst is the original termination feature of the node;

[0176] ②Inter-edge self-attention mechanism:

[0177] To effectively model edge representation, we designed an edge feature decoding module based on a self-attention mechanism. In the abnormal edge detection task, the goal is to identify edges with potential faults or measurement errors in the power system. Because edge power flow measurements may be missing or falsified, we construct edge representations based on more robust node states.

[0178] Specifically, the designed edge feature decoding module takes the state features of the nodes at both ends of the edge as input, explicitly models the interaction relationship between nodes through the self-attention mechanism, and obtains the edge representation containing contextual information;

[0179] The starting node feature e connected by a given edge fs With the terminal node feature e ft First, the two are spliced ​​together to form the basic feature representation h of the edge base :

[0180] h base =[e fs ||e ft ];

[0181] In the edge feature decoding module, the self-attention mechanism is used to model the edge representation. The input of the self-attention mechanism is the basic feature representation h base , the output of the self-attention mechanism is the edge representation h attn :

[0182] MultiheadAttn(Q,K,V)=Concat(head1,head2,...,headL )W o ;

[0183] Among them, the attention mechanism adopts a multi-head attention mechanism, that is, the original input is decomposed into multiple heads, the attention weight corresponding to each head is calculated separately, and then the attention weights obtained by each head are spliced ​​to obtain the final attention weight; the Lth head W o is the output projection matrix, W L is the weight matrix of the L-th head, Attn is the attention function, Q, K, V are query, key and value vectors respectively;

[0184] Then, h base With h attn Splicing and inputting a multi-layer perceptron to obtain the final edge feature representation E pred :

[0185] E pred =MLP([h base ||h attn ]);

[0186] ③Physical prior fusion mechanism:

[0187] To enhance the model's ability to perceive power physical information, a physical feature adapter is designed to map the original physical features into a latent space representation compatible with graph neural networks. The adapter is then introduced into the model as part of the node or edge input, thereby providing prior information on the power grid's operating status.

[0188] First, the physical features are used as input and mapped to a latent space representation compatible with the graph neural network through a linear transformation. The physical feature adapter receives the original physical features of the node, including: voltage amplitude V m , node active power P (including generators and loads), node reactive power Q (including generators and loads), branch flow characteristics (active power PF at the starting point, reactive power QF at the starting point, active power PT at the end point, and reactive power QT at the end point), and generate latent space physical features h through single-layer linear projection and nonlinear activation function (Tanh) phy , expressed as:

[0189] h phy =Tanh(W adapter ·[V m ,P,Q,PT,PF,QT,QF]+b adapter );

[0190] Where W adapter and b adapterTo be a learnable parameter, the Tanh function constrains the feature range to [-1, 1] to avoid numerical divergence of physical quantities;

[0191] The physical feature adapter transforms the latent space physical feature h phy and edge feature representation E pred Fusion is performed to explicitly inject the neighborhood knowledge constraint H:

[0192] H=MLP([E pred ,h phy ]);

[0193] During the execution of the physical prior fusion mechanism, prior information is implicitly injected through the joint optimization of the feature space. The loss function imposes penalties on predictions that violate physical laws. Backpropagation forces the physical feature adapter to learn feature representations that conform to domain knowledge, fuse graph features with physical latent features, and generate predictions that conform to domain laws.

[0194] To address the multi-constraint problem of abnormal edge detection in power systems, a joint optimization objective is designed, which contains three constraints: reconstruction loss, used to improve the reconstruction accuracy of edge attributes; Kirchhoff current leveling loss (KCL), used to ensure the physical consistency of system states; classification loss, used to optimize the anomaly detection performance according to the anomaly score;

[0195] The total loss function L′ consists of three terms:

[0196] L′=αL cls +βL recon +γL kcl ;

[0197] Where α, β and γ are hyperparameters, which control the weights of the three constraints respectively; L cls 、L recon and L kcl They correspond to classification loss, reconstruction loss and Kirchhoff current balancing loss respectively; the specific introductions of classification loss, reconstruction loss and Kirchhoff current balancing loss are as follows:

[0198] ①Classification loss L cls Focal Loss is used to address the serious imbalance between abnormal and normal edge samples. By focusing more on abnormal edge samples, the model is prevented from being biased towards the majority class.

[0199] L cls =-∑1-p t ) γ log(p t );

[0200] Where p t is the model's predicted probability for the true category. If the sample is abnormal, it is called a positive sample, then pt is the probability that the model predicts an abnormality; if the sample is normal, it is called a negative sample, then p t is the probability that the model predicts normal; γ is the focusing parameter that adjusts the weight of difficult and easy samples. When γ>0, the weight of easy-to-classify samples is reduced, forcing the model to focus on difficult-to-classify samples. When γ=0, it degenerates to the standard cross entropy loss.

[0201] ②Reconstruction loss L recon Physically constrained MSE is used to ensure that the power flow results predicted by the model conform to the actual physical laws of the power grid, preventing the model from fitting data only based on statistical characteristics, and improving generalization performance;

[0202]

[0203] Among them, the power flow is predicted by minimizing the model With the real power flow e ij The mean square error between them forces the model output to conform to the actual power system state;

[0204] ③Kirchhoff current balancing loss L kcl By explicitly introducing node power conservation constraints, the physical consistency and credibility of model predictions are further strengthened, reducing false positives caused by non-physical predictions. The goal of this approach is to transform the power system's energy conservation law into a differentiable mathematical form and embed it into the training process of the machine learning model, enabling the model to not only learn the statistical laws in the data but also follow the underlying physical rules. To achieve this goal, the following core concepts must be clearly defined: edge power flow, node power conservation, and KCL constraints.

[0205] ① Edge power flow definition:

[0206] For the edge connecting the nodes v i -v j , define the power flow characteristics:

[0207]

[0208]

[0209] Where, e ij Represents the slave node v i Points to node v j The edge eigenvector of ij [0]、e ji [1], e ij [2], e ji [3], Represents the slave node v i Flow direction v j The active power output, Represents the slave node vj Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input; each edge e contains bidirectional active and reactive power flow information i→j and j→i;

[0210] ②Definition of node power conservation:

[0211] ΔP i =P i gen -P i load -P i in -P i out ;

[0212]

[0213] Where ΔQ i For node v i The reactive power imbalance, ΔP i For node v i The active power imbalance, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input, P i gen For node v i The active power of the generator, P i load For node v i The load active power, For node v i The reactive power of the generator, For node v i The reactive power of the load;

[0214] ③KCL constraint definition:

[0215] By balancing node power, the model output is forced to follow physical laws, thus reducing misjudgment of anomalies:

[0216]

[0217] Where V is the node set.

[0218] S3. After training is completed, the anomaly score distribution is statistically calculated based on the normal edge samples in the training set, and a threshold is set as the basis for subsequent judgment. The threshold selection is optimized based on the precision and recall rate on the test set. In the testing phase, the graph data in the test set is input for edge feature reconstruction and anomaly scoring. The edges are judged as abnormal based on the set threshold, and the preliminary abnormal edge detection results are output.

[0219] The calculation of the anomaly score is based on the edge reconstruction features of the graph decoder of the graph autoencoder model. The larger the reconstruction error, the more the edge does not conform to the normal pattern in the graph structure, and is therefore considered a potential anomaly edge.

[0220] Let edge e = (s, v), then the anomaly score score(e) of the edge is the reconstruction error of the graph autoencoder model on the edge, that is:

[0221]

[0222] Where, Edge features reconstructed from the physical prior-based graph decoder part of the graph autoencoder model; f is the original edge feature; ||·||2 is the L2 norm; s is the starting node of edge e, v is the ending node of edge e, and e=(s,v) is the edge connecting s and v. The anomaly score score(e) is used to measure the ability of the graph autoencoder model to fit the edge structure. The larger the error, the more likely the edge is an anomaly.

[0223] After obtaining the anomaly score of each edge, a threshold τ is set. When the anomaly score score(e) of an edge satisfies score(e)>τ, the edge is judged to be an anomaly edge. The selection of the threshold is tuned according to the performance indicators on the test set to achieve a balance between precision and recall.

[0224] S4. The output abnormal edge detection result is input into the graph repair module. The graph repair module is designed based on the graph convolutional neural network (GCN) and receives the graph structure, node embedding and reconstructed edge features as input. In the case of deviation in abnormal edge connection or attributes, the node representation is fused with contextual information through the graph convolution propagation mechanism, and the attributes or connection status of the abnormal edge are reconstructed based on the representation of the nodes at both ends of the edge. Finally, the repaired edge structure and edge features are output to restore the damaged telesignaling status in the power grid, thereby improving the integrity of the graph structure and the operational credibility of the power system.

[0225] For edges judged to be abnormal, the graph repair model uses a graph convolutional neural network to repair them. The graph convolutional neural network propagates the node representation on the updated graph structure containing the abnormal edges, so that the node embedding integrates the semantic information of the surrounding context. Subsequently, the graph repair model takes the repaired node representation as input, and reconstructs the properties of the abnormal edge using a multi-layer perceptron by splicing the features of the nodes at both ends of the edge. For structurally abnormal edges, if the connection is disconnected, the graph repair model introduces candidate edges and uses the updated node representation of the graph convolutional neural network to determine the rationality of its connection, and restores or removes these connections when necessary. Finally, the graph repair model outputs the repaired graph structure and edge features to achieve joint repair of telemetry data errors and structural anomalies, thereby enhancing the accuracy of power system measurement data.

[0226] The following describes in detail the experimental results of the power system abnormality remote signaling detection method based on the graph autoencoder model in this embodiment:

[0227] Based on the final test results, the performance is evaluated using three indicators: accuracy, precision, and recall. To evaluate the performance of the present invention, an ablation experiment was conducted. By comparing with models that remove the temporal information extraction module, the node-edge joint attention mechanism based on the graph attention network, the dynamic residual feature enhancement and cross-attention mechanism, and the physical prior-based graph decoder, the impact of each module on the model's accuracy, precision, and recall is analyzed. Corresponding comparative tests were conducted on the three abnormality ratios of 5%, 10%, and 15% divided by the NREL-118 test system dataset. The comparison results for 5% are shown in Table 1.

[0228] Table 1

[0229]

[0230] The results in the table above show that in the experiment with a 5% anomaly ratio, the present invention still maintained a high accuracy of 98.7% and precision of 96.5%, demonstrating that the present invention can effectively distinguish normal and abnormal edges at low anomaly ratios, and the recall rate of 92.3% indicates that it can well cover all abnormal situations. In contrast, the model after removing the module showed a significant performance degradation, especially in the recall rate. After removing the time series information extraction module, the recall rate dropped significantly by 85.6%, indicating that time series features are crucial for power system anomaly detection. After removing the physical prior graph decoder, the precision and recall rates dropped further, indicating that the introduction of physical constraints is crucial to the accuracy of anomaly detection. The 10% comparison results are shown in Table 2.

[0231] Table 2

[0232]

[0233] The results in the table above show that in the experiment with a 10% anomaly ratio, the present invention still shows significant advantages. The stability of the accuracy rate of 98.5% and the precision rate of 96.2% has been verified, and the recall rate of 91.0% is relatively high, indicating that the present invention can still detect anomalies more comprehensively in more abnormal situations. In contrast, removing the time series information extraction module causes the recall rate to drop to 82.3%, demonstrating the importance of time series features in dynamic power grid environments. The reduction in precision and recall rate caused by removing the node-edge joint attention mechanism and the physical prior decoder emphasizes the necessity of these mechanisms in the model. After removing the dynamic residual feature enhancement module, the model accuracy decreases slightly, but the recall rate is relatively stable, indicating that it has little impact on overall performance.

[0234] The comparison results of 15% are shown in Table 3.

[0235] Table 3

[0236]

[0237] The results in the table above show that in experiments with a 15% anomaly ratio, the performance of our method is gradually affected by the increasing number of anomalies. Despite this, our method still demonstrates strong robustness and superiority, achieving an accuracy of 98.0%, a precision of 95.5%, and a recall of 89.8%. This demonstrates that even with a high anomaly ratio, our method can effectively identify anomalous edges with a low false positive rate. Removing the temporal information extraction module significantly reduces the recall rate to 78.9%, demonstrating that temporal information is even more important for anomaly detection at high anomaly ratios. Removing the node-edge joint attention mechanism and the physical constraint decoder significantly reduces both precision and recall, highlighting the essential role of these modules in high anomaly detection. Removing the dynamic residual feature enhancement module has little impact on recall, but reduces precision.

[0238] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A power system abnormality telesignaling detection method based on a graph autoencoder model is characterized by: The following steps are involved: S1. Collect measurement data from the power system and construct it into a graph with node and edge relationships. Simulate two types of telesignaling anomalies in the graph, called edge anomalies, namely false disconnection anomalies and false connection anomalies. Through anomaly injection and annotation, generate a labeled graph dataset containing normal edge and abnormal edge samples, and finally divide it into training and test sets. S2. Input the graph data in the training set into the designed graph autoencoder model for training. The encoder part of the model uses a residual-enhanced graph attention mechanism and a node-edge joint attention mechanism to achieve joint modeling of graph structure and temporal information, maintain the stability of node representation and extract key structural features. It also combines a bidirectional long short-term memory network to capture the dynamic temporal changes of nodes, fuse multi-scale temporal features, and generate a robust edge embedding representation. The graph decoder part of the above edge embedding representation input model performs edge feature reconstruction and anomaly score prediction. The graph decoder combines the physical prior of the power system to construct a physical constraint embedding space, realizes edge attribute reconstruction through the inter-edge self-attention mechanism, and generates anomaly scores reflecting the degree of structural or attribute deviation. A joint loss function is used for optimization during training, including reconstruction loss, Kirchhoff current balancing loss, and classification loss for anomaly detection. In the training phase, the anomaly score distribution is calculated based on normal edge samples and the threshold is optimized, with the precision and recall of the test set as the tuning target. In the testing phase, edge features are reconstructed and anomaly scores are performed on the input graph data. Abnormal edges are determined based on the threshold, and preliminary anomaly detection results are output. S4. The detected abnormal edges are input into the graph convolution repair module, which integrates graph topology, node features and edge attribute information. Through the neighborhood information propagation mechanism, the connection status and electrical parameters of the abnormal edges are corrected. The repaired edges can accurately restore the true topological relationship of the power grid, effectively eliminate telesignaling errors, and improve the quality of power grid data and the operational reliability of the power system.

2. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: In step S1, the measurement data of the power system is collected, and the collected measurement data is constructed into graph data with node and edge relationships, and two typical types of telesignaling anomalies, called edge anomalies, are simulated or identified in the graph data, including false disconnection anomalies and false connection anomalies, and the graph data is injected, labeled and normalized for preprocessing, and finally a graph data set with anomaly labels is formed, that is, the graph data set contains normal edge samples and abnormal edge samples; the graph data set is divided into a training set and a test set; the characteristics of the nodes include node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power, and the characteristics of the edges include starting side active power PF, starting side reactive power QF, terminal side active power PT, terminal side reactive power QT and power system state estimation measurement data in the graph adjacency matrix; the graph data is represented as a time series graph set T ′ Represents the total time, where the graph structure at each moment is defined as graph G t ={V t ,E t }, where V t represents the set of nodes in the power system at time t, E t Represents the edge set at time t; the node represents the bus, and the edge represents the connection relationship of the transmission branch, reflecting the operating topology and state of the power system at time t; each node v∈V t Contains the time series characteristics of node voltage amplitude, generator active power, generator reactive power, load active power and load reactive power; each edge e∈E t Includes active power PF measured at the starting point and end point t PT t and reactive power QF t , QT t The topological connection state of the power grid is represented by the adjacency matrix express, represents the topology of the power grid at time t, that is, whether there is a connection between node i and node j, Indicates that there is a connection line between node i and node j at time t, Indicates that there is no connection between node i and node j at time t, and time t indicates that the adjacency matrix is ​​graph G t The structural description reflects the dynamic changes of the graph structure over time.

3. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: In step S2, during the training process, the graph autoencoder model first performs structural and temporal joint modeling on the graph data through the encoder, adopts the residual-enhanced graph attention encoder, and combines the node-edge joint attention mechanism based on the graph attention network to extract key feature information in the graph structure and maintain the stability of the original node representation. Subsequently, by introducing a bidirectional long short-term memory network, the dynamic change characteristics of the node in the time dimension are captured, and the temporal information of different scales is integrated to obtain a more robust edge embedding representation. The residual-enhanced graph attention encoder is a graph neural network layer designed for power system graph structure data. By fusing the node-edge joint feature representation and historical state memory, it effectively captures abnormal information. It includes: a node-edge joint attention mechanism based on the graph attention network, dynamic residual feature enhancement, a cross-attention historical state retention mechanism and a temporal information extraction module, which are specifically introduced as follows: ①Node-edge joint attention mechanism based on graph attention network: To capture the correlation between power line anomalies and the status of nodes at both ends, an improved graph attention mechanism is designed, namely the node-edge joint attention mechanism based on the graph attention network. Through the node-edge joint attention mechanism, node attributes and topological relationships are explicitly integrated, as follows: Given the feature vector x of the i-th node i ∈R d and the feature vector e of the edge connecting node i and node j ij ∈R e , where R is the real number field, d is the dimension of each node feature vector, and e is the feature dimension of each edge; first, feature alignment is performed through linear projection: Where W h is the weight matrix of node features; b h is the bias term of the node feature; W e is the edge feature weight matrix; b e is the edge feature bias; is the normalized feature of node i; is the edge feature after nonlinear activation; LayerNorm normalizes the node features to eliminate dimensional differences; ReLU nonlinearly activates the edge features; Use the graph attention mechanism for feature interaction and calculate the edge attention weight: Where, represents the normalized attention weight of node j to node i in the lth attention head; a is a learnable attention vector used to measure the importance of the node-edge joint feature; l represents the lth attention head, || represents the feature splicing operation; represents the normalized features of node i in the lth attention head; represents the normalized features of node j in the lth attention head; Represents the edge features after nonlinear activation in the lth attention head; Finally, the node representation h is obtained through multi-head aggregation attn : Where L is the total number of attention heads, N(i) is the neighbor set of node i; ②Dynamic residual feature enhancement: The topology of the power system changes dynamically. To avoid the degradation of deep network characteristics, an adaptive residual connection mechanism is designed as follows: Define the residual mapping function: Where, d in is the dimension of the input feature, i.e. x i Dimension; d out Represents the dimension of the output feature, that is, W after mapping res x i Dimension; W res is the weight matrix of the residual mapping; F res It is a dynamic residual mapping function that can solve the key problems of dimension mismatch and gradient instability in deep learning models; Dynamically fuse attention features and residual features through a gating mechanism: g=σ(W g [h attn ||F res (x i )]); h fusion =g⊙h attn +(1-g)⊙F res (x i ); Where σ is the Sigmoid function, which smooths the gate control changes and alleviates the gradient vanishing problem of deep networks; ⊙ represents element-by-element multiplication; g is the gating vector, whose value is between [0, 1]. The fusion ratio of attention and residual information obtained by the gating mechanism is used to control the dynamic fusion ratio of the two types of information. The larger the g, the more reliable the current attention feature is, and the smaller the g, the more dependent on the historical state; h fusion is the fused feature vector, which combines attention information and residual information; W g is the weight matrix of the gating mechanism; this gating mechanism can adaptively adjust the fusion ratio of original features and attention features, showing better adaptability in the scenario of dynamic topology changes in the power system; ③ Cross-attention history state retention mechanism: In order to preserve the state evolution characteristics of the node itself, a cross-layer attention mechanism is introduced. The historical residual features are used as the knowledge base, and key historical information is adaptively selected and integrated with the current state. The current layer output is used as the query, and the historical residual features are used as the key-value: Q=h fusion W q ; K=F res (x i )W k ; V=F res (x i )W v ; Calculate cross-layer attention weight β i : Through cross-layer attention and nonlinear changes, historical information is fused with current features to generate the final output h out : h out =LayerNorm(ELU(β i V+h fusion )); ELU activation function: In the formula, ELU has a non-zero gradient in the negative interval, which prevents node features from not being updated for a long time and alleviates the Dead ReLU problem; the exponential term e x Smooths negative value changes, suitable for scenarios with positive and negative values ​​in power characteristics; W q 、W k 、W v is the cross-attention query vector Q, key vector K, and value vector V projection weight; x is the input eigenvalue; α is the scaling factor of the ELU function in the negative interval, which controls the amplitude of the negative value mapping; the cross-attention history state retention mechanism enables nodes to adaptively retain key historical state features when aggregating neighborhood information, enhancing the modeling ability of the continuous evolution of nodes and edges; ④Time series information extraction module: When detecting abnormal edges in power systems, it is found that the abnormal state of edges is not only affected by the node state at the historical moment, but can also be triggered by contextual factors at the future moment. To this end, a bidirectional long short-term memory network (Bi-LSTM) is introduced to enhance the modeling capability of time series context. Bi-LSTM considers both forward (historical) and backward (future) information in the sequence, and can capture more comprehensive time-dependent characteristics. For each target edge, its state at multiple time steps is fed into the Bi-LSTM as an input sequence. The Bi-LSTM consists of two parallel LSTM layers, which transmit information in the forward and reverse directions of time. The forward LSTM layer extracts historical context features, while the backward LSTM layer models the potential impact of future moments. The hidden states output by the two are fused at each time step to form an edge temporal representation containing bidirectional temporal information. The fused representation not only reflects the state of the edge at the current moment, but also implicitly integrates the state change trend related to the adjacent nodes in the previous and next time windows, thereby providing more discriminative temporal feature support for subsequent anomaly scores.

4. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: In step S2, the extracted edge embedding representation is input into the graph decoder for edge feature reconstruction and anomaly score prediction. A physical prior-based graph decoder is used to construct a physical constraint embedding space in combination with the physical information of the power system. The edge attribute features are reconstructed using the edge self-attention mechanism. The physical prior-based graph decoder includes: a dynamic edge feature generation mechanism, an edge self-attention mechanism, and a physical prior fusion mechanism. The details are as follows: ① Dynamic edge feature generation mechanism: We choose to dynamically generate edge features based on nodes rather than directly relying on original edge features. That is, we implicitly model edge features through the evolution of node states to solve the problem that historical edge features are not fully recorded or are interfered with by noise. We propose a dynamic edge feature generation module that combines a graph attention mechanism with an edge feature generator to dynamically generate edge features from the node's potential representation. The graph attention network generates a high-level semantic representation of the node by dynamically aggregating neighborhood information, while the edge feature generator dynamically generates edge features using the fused features of the source and target nodes. Using a multi-head GAT layer, each head independently learns different attention weights, and finally integrates multi-view features by splicing: Where L is the total number of attention heads; l is the number of the lth attention head; σ is the Sigmoid function, the smooth gate controls the change and alleviates the gradient vanishing problem of deep networks; W l Represents the learnable linear transformation weight matrix corresponding to the lth attention head, which is used to transform the neighbor node feature h j Mapped to the attention space for subsequent weighted summation; represents the normalized attention weight of node j to node i in the lth attention head; h i is the multi-head attention output feature of node i; || represents the feature concatenation operation; N(i) is the neighbor set of node i; j∈N(i) means that node j is a neighbor node of node i; h j The feature vector representing node j is the neighbor node feature of node i; The edge feature generator considers both the high-level abstract features of nodes extracted by the GAT layer and the original latent features, thereby improving the model's tolerance to sparse or noisy data: e f =[h src ||h dst ||x src ||x dst ]; Where, e f is the fusion edge feature, which is h src ||h dst ||x src ||x dst The concatenation result of the four features is used to comprehensively represent the structure and attribute information; h src is the node starting feature output by the GAT layer; h dst is the node termination feature output by the GAT layer; x src is the original starting feature of the node; x dst is the original termination feature of the node; ②Inter-edge self-attention mechanism: To effectively model edge representation, we designed an edge feature decoding module based on a self-attention mechanism. In the abnormal edge detection task, the goal is to identify edges with potential faults or measurement errors in the power system. Because edge power flow measurements may be missing or falsified, we construct edge representations based on more robust node states. The designed edge feature decoding module takes the state features of the nodes at both ends of the edge as input, explicitly models the interaction relationship between nodes through the self-attention mechanism, and obtains the edge representation containing contextual information; The starting node feature e connected by a given edge fs With the terminal node feature e ft First, the two are spliced ​​together to form the basic feature representation h of the edge base : h base =[e fs ||e ft ]; In the edge feature decoding module, the self-attention mechanism is used to model the edge representation. The input of the self-attention mechanism is the basic feature representation h base , the output of the self-attention mechanism is the edge representation h attn : MultiheadAttn(Q,K,V)=Concat(head1,head2,…,head L )W o ; Among them, the attention mechanism adopts a multi-head attention mechanism, that is, the original input is decomposed into multiple heads, the attention weight corresponding to each head is calculated separately, and then the attention weights obtained by each head are spliced ​​to obtain the final attention weight; the Lth head W o is the output projection matrix, W L is the weight matrix of the L-th head, Attn is the attention function, Q, K, V are query, key and value vectors respectively; Then, h base With h attn Splicing and inputting a multi-layer perceptron to obtain the final edge feature representation E pred : E pred =MLP([h base ||h attn ]); ③Physical prior fusion mechanism: To enhance the model's ability to perceive power physical information, a physical feature adapter is designed to map the original physical features into a latent space representation compatible with graph neural networks. The adapter is then introduced into the model as part of the node or edge input, thereby providing prior information on the power grid's operating status. First, the physical features are used as input and mapped to a latent space representation compatible with the graph neural network through a linear transformation. The physical feature adapter receives the original physical features of the node, including: voltage amplitude V m , the node active power P of the generator and load, the node reactive power Q of the generator and load, the starting side active power PF, the starting side reactive power QF, the terminal side active power PT and the terminal side reactive power QT, and generate the latent space physical feature h through single-layer linear projection and nonlinear activation function Tanh phy , expressed as: h phy =Tanh(W adapter ·[V m ,P,Q,PT,PF,QT,QF]+b adapter ); Where W adapter and b adapter To be a learnable parameter, the Tanh function constrains the feature range to [-1, 1] to avoid numerical divergence of physical quantities; The physical feature adapter transforms the latent space physical feature h phy and edge feature representation E pred Fusion is performed to explicitly inject the neighborhood knowledge constraint H: H=MLP([E pred ,h phy ]); During the execution of the physical prior fusion mechanism, prior information is implicitly injected through the joint optimization of the feature space. The loss function imposes penalties on predictions that violate physical laws. Backpropagation forces the physical feature adapter to learn feature representations that conform to domain knowledge, fuse graph features with physical latent features, and generate predictions that conform to domain laws.

5. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: In step S2, the edge feature reconstruction result output by the graph decoder and the anomaly score are input into the joint loss function for optimization. The joint optimization objective includes three types of loss terms: reconstruction loss, which is used to improve the reconstruction accuracy of edge attributes; Kirchhoff current balance loss, which is used to ensure the physical consistency of the system state; classification loss, which is used to optimize the anomaly detection performance based on the anomaly score; Total loss function L ′ Contains three items: L′=αL cls +βL recon +γL kcl ; Where α, β and γ are hyperparameters, which control the weights of the three constraints respectively; L cls , L recon and L kcl They correspond to classification loss, reconstruction loss and Kirchhoff current balancing loss respectively; the specific introductions of classification loss, reconstruction loss and Kirchhoff current balancing loss are as follows: ①Classification loss L cls Focal Loss is used to address the serious imbalance between abnormal and normal edge samples. By focusing more on abnormal edge samples, the model is prevented from being biased towards the majority class. THE cls =-∑(1-p t )γlog(p t ); Where p t is the model's predicted probability for the true category. If the sample is abnormal, it is called a positive sample, then p t is the probability that the model predicts it to be an anomaly; If the sample is normal, it is called a negative sample, then p t is the probability that the model predicts normal; γ is the focusing parameter that adjusts the weight of difficult and easy samples. When γ>0, the weight of easy-to-classify samples is reduced, forcing the model to focus on difficult-to-classify samples. When γ=0, it degenerates to the standard cross entropy loss. ②Reconstruction loss L recon Physically constrained MSE is used to ensure that the power flow results predicted by the model conform to the actual physical laws of the power grid, preventing the model from fitting data only based on statistical characteristics, and improving generalization performance; Among them, the power flow is predicted by minimizing the model With the real power flow e ij The mean square error between them forces the model output to conform to the actual power system state; ③Kirchhoff current balancing loss L kcl By explicitly introducing node power conservation constraints, the physical consistency and credibility of model predictions are further strengthened, reducing false positives caused by non-physical predictions. The goal of this approach is to transform the power system's energy conservation law into a differentiable mathematical form and embed it into the training process of the machine learning model, enabling the model to not only learn the statistical laws in the data but also follow the underlying physical rules. To achieve this goal, the following core concepts must be clearly defined: edge power flow, node power conservation, and KCL constraints. ① Edge power flow definition: For the edge connecting the nodes v i -v j , define the power flow characteristics: Where, e ij Represents the slave node v i Points to node v j The edge eigenvector of ij [0]、e ji [1], e ij [2], e ji [3], Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input; each edge e contains bidirectional active and reactive power flow information i→j and j→i; ②Definition of node power conservation: ΔP i =P i gen -P i load -P i in -P i out ; Where ΔQ i For node v i The reactive power imbalance, ΔP i For node v i The active power imbalance, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Active power input, Represents the slave node v i Flow direction v j The active power output, Represents the slave node v j Flow direction v i Reactive power input, P i gen For node v i The active power of the generator, P i load For node v i The load active power, For node v i The reactive power of the generator, For node v i The reactive power of the load; ③KCL constraint definition: By balancing node power, the model output is forced to follow physical laws, thus reducing misjudgment of anomalies: Where V is the node set.

6. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: After the training is completed in step S3, the abnormal score distribution is calculated based on the normal edge samples in the training set, and a threshold is set as the basis for subsequent judgment. The threshold is selected based on the precision and recall rate on the test set as the tuning target; During the testing phase, graph data from the test set is input for edge feature reconstruction and anomaly scoring. Edge anomalies are determined based on a set threshold, and preliminary abnormal edge detection results are output. The anomaly score is calculated based on the edge reconstruction features of the graph decoder of the graph autoencoder model. It is used to measure the degree of deviation of the edge in structure or attributes. The larger the reconstruction error, the less consistent the edge is with the normal pattern in the graph structure, and thus it is considered a potential abnormal edge. Let edge e = (s, v), then the anomaly score score(e) of the edge is the reconstruction error of the graph autoencoder model on the edge, that is: Where, Edge features reconstructed from the physical prior-based graph decoder part of the graph autoencoder model; f is the original edge feature; ||·||2 is the L2 norm; s is the starting node of edge e, v is the ending node of edge e, and e=(s,v) is the edge connecting s and v. The anomaly score score(e) is used to measure the ability of the graph autoencoder model to fit the edge structure. The larger the error, the more likely the edge is an anomaly. After obtaining the anomaly score of each edge, a threshold τ is set. When the anomaly score score(e) of an edge satisfies score(e)>τ, the edge is judged to be an anomaly edge. The selection of the threshold is tuned according to the performance indicators on the test set to achieve a balance between precision and recall.

7. The power system abnormality remote signaling detection method based on the graph autoencoder model according to claim 1 is characterized in that: In step S4, for edges judged to be abnormal, the graph repair model uses a graph convolutional neural network to repair them. The graph convolutional neural network propagates the node representation on the updated graph structure containing the abnormal edge, so that the node embedding integrates the semantic information of the surrounding context. Subsequently, the graph repair model takes the repaired node representation as input, and reconstructs the properties of the abnormal edge using a multi-layer perceptron by splicing the features of the nodes at both ends of the edge. For structurally abnormal edges, if the connection is disconnected, the graph repair model introduces candidate edges and uses the updated node representation of the graph convolutional neural network to determine the rationality of its connection, and restores or removes these connections when necessary. Finally, the graph repair model outputs the repaired graph structure and edge features to achieve joint repair of telemetry data errors and structural anomalies, thereby enhancing the accuracy of power system measurement data.

Citation Information

Patent Citations

  • Aero-engine gas path performance anomaly detection system based on depth auto-encoder

    CN114742165A

  • Database index data anomaly prediction method based on gated convolution and graph attention

    CN118585936A

  • Intelligent power grid anomaly detection method and system based on graph neural network

    CN119622562A

  • Image anomaly detection method for latent space auto-regression based on memory enhancement

    WO2022095645A1

  • Multivariate time series anomaly detection method for intelligent internet of things system

    WO2024207627A1

Cited By

  • Power distribution network topology correction method fusing topology anomaly recognition and structure credible reasoning

    CN120833077A

  • Anomaly detection method and detection system for electric energy meter

    CN121051662A

  • APT attack detection method and device based on graph attention learning and electronic equipment

    CN121283670A

  • APT attack detection method and device based on graph attention learning and electronic equipment

    CN121283670B

  • Power system anomaly detection method based on voltage-current joint attention network

    CN121347921A