Graph model-based power grid anomaly detection method

Through the grid abnormality detection method based on graph model, the graph distance is calculated using branch break distribution factors, combined with time-weighted calculation and weighted distribution model generation, the problems of high false alarm rate, serious missed alarm and low detection efficiency in complex topological associations and dynamic topological changes are solved, and higher detection accuracy and robustness are achieved.

CN120177944APending Publication Date: 2025-06-20QINGYUAN INFORMATION TECHNOLOGY (QINGHAI) CO LTD

Patent Information

Application Number
CN202510446488.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When traditional power grid abnormality detection methods face complex topological associations and dynamic topological changes, they are prone to problems such as high false alarm rate, serious missed alarms and low detection efficiency.

Method used

The grid abnormality detection method based on graph model is adopted, and the grid graph model is constructed, and the graph distance is calculated using branch break distribution factors, combined with time-weighted calculation, weighted distribution model generation, abnormality determination and pseudo-label allocation, detection of multiple types of abnormalities is achieved.

Benefits of technology

It improves the accuracy and robustness of grid abnormal detection, can identify isolated random anomalies, local line failures and global topological anomalies, and adapts to the real-time monitoring needs of large-scale power grids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120177944A_ABST
    Figure CN120177944A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent monitoring and control of a power system, and discloses a power grid anomaly detection method based on a graph model, and the method comprises the steps: building a power grid graph model, measuring the influence of line disconnection on the power flow distribution of other lines through employing a branch disconnection distribution factor, and calculating the graph distance between different time topological states. On the basis, a time weighting strategy based on a graph distance is provided to screen historical observation data more similar to the current topological state, and a robust reference distribution model is generated in combination with a weighted median and a weighted quartile distance; deviation degree calculation is carried out on a real-time measured value and the distribution model, when the deviation degree exceeds a preset threshold value, it is judged that abnormity exists, false label marking is carried out, and dynamic adjustment is carried out on the abnormal weight according to data state changes in continuous detection. By means of the method, high-precision and low-false-alarm fault warning and positioning can be achieved under the condition that topology dynamics and various types of abnormal characteristics are considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent monitoring and control of power systems, and particularly to a power grid anomaly detection method based on a graph model. Background Art

[0002] With the continuous expansion of the scale of the power grid and the continuous improvement of the intelligent level, the number and types of sensors in the power system have shown exponential growth. Most traditional anomaly detection methods are based on fixed thresholds or simple statistical models. These methods can play a certain role when the power grid structure is relatively stable and the scale is small. However, as the power grid topology and operating conditions become increasingly complex, their limitations gradually emerge: Lack of utilization of topological coupling relationships: Traditional methods often only focus on the data information of single-point sensors, ignoring the electrical coupling between sensors and the relevance between lines. In a real power grid, the change in the operating state of any line may affect the power flow distribution of adjacent lines, and traditional methods are difficult to accurately capture this global linkage characteristic. Insufficient adaptability to dynamic topological changes: In the daily operation of the power system, situations such as line addition and deletion, load transfer, equipment failure, and maintenance switching often occur, causing dynamic changes in the topological structure and power flow distribution. If only static or single statistical thresholds are used, it is difficult to adjust the parameters in a timely manner according to the changing power grid state, often resulting in a high false alarm rate or missed detection phenomenon. Lack of pertinence to the diversity of anomaly patterns: The sources of power grid anomalies are diverse, including both local events such as random interference, sensor failures, and line failures, and may also occur global topological changes or complex anomalies caused by the superposition of multiple factors. Traditional methods usually cannot take into account various types of anomaly patterns and are prone to failure when dealing with new types or comprehensive anomalies. It is difficult to balance the recognition accuracy and real-time performance: With the growth of the sensor data volume, overly complex detection algorithms will lead to excessive computational overhead and are difficult to meet the requirements of the power system for real-time monitoring; while relying only on simple and fast detection models may cause a decrease in recognition accuracy.

[0003] In recent years, graph models have received attention in power grid anomaly detection. The physical and logical structures of the power grid naturally possess graph characteristics: each node can represent a sensor or other power equipment, and the edge can represent a line or a power flow connection. Through the graph model, the local and global structure information of the power grid can be organically combined, so as to more precisely depict the electrical coupling relationship between sensors. However, how to effectively quantify the impact of topological changes on power distribution in a dynamic topological environment and integrate it with sensor measurement data is still one of the challenges faced by graph models when applied to power grid anomaly detection.

[0004] Based on this, in order to improve the detection accuracy and robustness of multi-type anomalies in a complex and changing power grid environment, an anomaly detection method is needed that can naturally integrate topological information during the modeling process, quantify the impact of topological changes on power flow distribution, and comprehensively analyze multi-source measurement data. Summary of the Invention

[0005] The purpose of the present invention is to overcome the problems of high false alarm rate, serious missed reports, and low detection efficiency that easily occur in traditional power grid anomaly detection methods when facing complex topological associations and dynamic topological changes, and to provide an anomaly detection method based on a graph model combined with the line outage distribution factor.

[0006] To achieve the above object, the present invention adopts the following technical means: A power grid anomaly detection method based on a graph model, comprising the following steps:

[0007] 1) Construct a power grid graph model: Various devices in the power grid are used as nodes, and lines and power flows are used as edges to form a power grid graph model with a node set and an edge set;

[0008] 2) Graph distance calculation based on the line outage distribution factor: By comparing the power grid topological states at different times, the line outage distribution factor (LODF) is used to measure the influence intensity of newly added or deleted lines on the power flow distribution of other lines, and the graph distance between the topological states at two times is calculated by combining the edge weights;

[0009] 3) Anomaly detection algorithm based on graph distance calculation: It includes time-weighted calculation based on graph distance, generation of a weighted distribution model, and anomaly determination and pseudo-label assignment.

[0010] Further, the time-weighted calculation based on graph distance: Obtain the power grid topological state at the current moment, compare it with the topological states at each moment within the historical window, and calculate the graph distance between each pair of moments; The graph distance is determined by the line outage distribution factor and is used to quantify the influence intensity of a certain line on the power flow of other lines when it is disconnected; According to the magnitude of the graph distance, the weights of historical data with too large a difference from the current topology are reduced to zero, more representative historical data is screened out, and a distance-decreasing function is used to assign a time-weighting factor to it.

[0011] Further, generation of a weighted distribution model: Perform weighted statistics on the historical sensor measurement values that are retained after screening, focusing on key features such as edge anomaly indicators, average anomaly indicators, and dispersion indicators; Calculate the weighted median and weighted interquartile range respectively to obtain a reference distribution model that is more robust to outliers; This reference distribution model not only considers the similarity between the current topological state and the historical topological state, but also effectively suppresses the deviation caused by extreme values at the distribution statistics level.

[0012] Furthermore, anomaly determination and pseudo-label assignment: compare the measured value at the current moment with the weighted distribution model, and determine whether it is abnormal based on the degree of deviation (the ratio of the difference between the current value and the weighted median to the weighted interquartile range); when the degree of deviation exceeds the preset threshold, the sensor measurement data is marked as abnormal and assigned a pseudo-label; for data marked as pseudo-labels, their weights for the distribution model will be set to zero in subsequent tests to avoid abnormal data from continuously interfering with the update of the model; if the abnormal value returns to a normal level at a subsequent moment, its pseudo-label can be removed and re-included in the distribution statistics.

[0013] Furthermore, the graph distance calculation based on the branch disconnection distribution factor includes the following steps:

[0014] a) Obtain the power grid topology at the current time and historical time, and identify the newly added or deleted lines;

[0015] b) Calculate the branch disconnection distribution factor to measure the impact of a single line disconnection on the power flow distribution of other lines;

[0016] c) Combine the edge weights (such as the power flow of the line) and the impact value of the branch disconnection distribution factor, and accumulate all the newly added or deleted lines to obtain the graph distance between the topological states at two moments. This distance can be used to quantify the physical impact of topological changes into a unified metric.

[0017] Furthermore, the edge weight is obtained by normalizing the power flow of the line or its importance in the grid topology, and the contribution value of the high-weight line is amplified when calculating the graph distance.

[0018] Furthermore, the time-weighted calculation strategy based on graph distance in step 3) includes:

[0019] a) Calculate the graph distance between the current topological state and the topological state at the historical moment;

[0020] b) if the graph distance is greater than a preset threshold, the historical data is discarded; if the graph distance does not exceed the threshold, a time weighting factor is assigned to the historical data according to a distance decreasing rule;

[0021] c) Complete the screening and weighting of historical measurement data based on the above weighting factors.

[0022] Furthermore, the calculation of the deviation of the sensor measurement value in the abnormality determination and pseudo-label assignment step includes at least the following abnormality indicators:

[0023] a) Edge anomaly index: the maximum power flow change in the sensor connection line;

[0024] b) Average abnormal index: the average value of the power flow change of the sensor connection line;

[0025] c) Dispersion anomaly index: The standard deviation of the change in power flow of the sensor connection line.

[0026] Furthermore, when performing time weighting, historical data is further screened according to the geographical location of the sensor or its electrical distance from the current node to improve the positioning accuracy of local anomalies or regional faults.

[0027] Furthermore, after completing the anomaly determination and pseudo-label assignment, it also includes an iterative update step for the weighted distribution model. When the data with the assigned pseudo-labels remains in an abnormal state for more than a preset duration or the deviation continuously increases within consecutive detection cycles, the data of this sensor is excluded from the current weighted statistical window to prevent long-term outliers from causing continuous interference to the weighted distribution model.

[0028] Through the above process, the algorithm can not only identify isolated random anomalies but also effectively detect complex situations such as local line faults and global topology anomalies. The time-weighted distribution model constructed based on graph distance can adaptively reflect the operating characteristics of the power grid at different times and provide reliable anomaly alarms when significant deviations occur in multi-source measurement data, significantly improving the accuracy and robustness of power grid anomaly detection.

[0029] Furthermore, the power grid structure is intuitively represented by a graph model, and the impact of topological changes on power distribution is quantified using branch outage distribution factors; by introducing context information of topological changes in the graph model and performing weighted statistics on historical data, accurate identification of multiple types of anomalies (such as random anomalies, local line faults, and global topological changes, etc.) is achieved; statistical quantities such as weighted median and weighted interquartile range are used in the distribution model to reduce the interference of extreme values on detection and meet the real-time monitoring requirements of large-scale power grids; the modeling and operation processes can be adjusted according to the scale and topological complexity of the power grid, and it is applicable to the operation scenarios of intelligent power grids at different levels and scales. Brief Description of the Drawings

[0030] Figure 1 It is a graph of the F1-score index result of a specific embodiment of the present invention.

[0031] Figure 2 It is a graph of the AUC index result of a specific embodiment of the present invention. Detailed Embodiments

[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the detailed embodiments of the present invention will be described below.

[0033] In the smart grid, the dynamic changes in the grid topology are not only reflected in the addition and deletion of physical lines but also significantly affect the power flow distribution among various lines, thus having a profound impact on the sensor measurement values. Therefore, if only the sensor measurement values themselves are concerned while ignoring the evolution of the grid topology, it is very easy to lead to misjudgment or missed judgment of abnormal events. The Line Outage Distribution Factor (LODF), as a sensitivity index to measure the impact of line outages on the power flow of other lines in the system, can be quickly calculated by assuming a lossless DC power flow model or a linearized AC power flow model and is often used to estimate the linear influence range and degree of line outages. Now, we explicitly embed the impact of topological changes on the power flow into the graph model, enabling anomaly detection to not only consider the deviation of the sensor data itself but also comprehensively evaluate the system-level cascading effects brought about by topological structure changes, thereby enhancing the ability to locate and identify abnormal events in a dynamic environment.

[0034] The Line Outage Distribution Factor describes the impact of the disconnection of a specific line on the power flow changes of other lines, and its definition is the ratio between the initial power flow of the disconnected line and the power flow changes of other lines caused by it. For a disconnected line k, its impact on the power flow of another line l can be expressed by the following formula:

[0035]

[0036] where d kl represents the ratio between the power flow change of line k after disconnection and the original power flow of line k, Δf l is the power flow change amount of line l, and f k is the power flow value of line k before disconnection. The calculation of the Line Outage Distribution Factor is based on the power flow equation of the power system and is usually linearized under the premise of assuming a DC power flow model. This index has a clear physical meaning and can intuitively quantify the influence range and intensity of the disconnection of a certain line on the rest of the system.

[0037] The core idea of the graph distance calculation method proposed based on the Line Outage Distribution Factor is to embed the impact of topological changes on the power flow distribution into the graph model and measure the distance between two topological states based on this. Specifically, assume that the smart grid corresponds to two graphs G i =(V i , E i ) and G j =(V j , E j ) at times i and j respectively, where V i and E i represent the node set and edge set at time i respectively. Let G union be Gi and G j union state diagram, that is, G union in E union = E i ∪E j and V union = V i ∪V j (as shown in the figure below). When the graph changes from G i to G j , its topological change can be represented by the difference in the edge sets of the two graphs E i \E j and E j \E i , that is, the set of newly added or deleted lines. To calculate the impact of this change on the system operation, this algorithm takes each line in the topological change as a separate perturbation source and quantifies its impact on the power flow of other lines through the branch outage distribution factor.

[0038] Specifically, for each newly added or deleted line p, the average impact on the power flow change of all other lines can be calculated based on its branch outage distribution factor. Let d pl be the impact factor of the fault of line p on line l, and the set of all lines in the system is L, then the impact of line p on the entire system can be initially expressed as:

[0039]

[0040] where, |E i ∪E j | represents the union of the edge sets at two time points in the system, and |d pl | represents the absolute value of the impact of line p on the power flow of line l. In this way, the impact intensity of the topological change of a single line on the power distribution of the entire power grid can be quantified. On this basis, this application defines the graph distance D(G i , G j ) to measure the global change degree between two topological states.

[0041] To further enhance the role of significant changes and at the same time reduce the sensitivity to small changes, we introduce a non-linear amplification technique for the contribution value of the branch outage distribution factor. The contribution value x p of edge p is defined as:

[0042]

[0043] where, w pl is the weight of edge l in the network, reflecting its importance, and is defined as:

[0044]

[0045] Edge weight v pl Edges with a greater impact of the embedding of pl on the power distribution have higher weights when calculating the graph distance, thus more accurately reflecting the actual impact of topological changes on the power grid operation. The branch outage distribution factor essentially reflects the sensitivity of the removal of an edge to the power distribution of other edges. Combining the edge weight v pl , on the one hand, we can assign higher influence to edges with larger power flows, making them contribute more in the graph distance calculation. By using the power flow |f union | in the G l state to define the edge weight v pl , the method can pay more attention to the edges that have a greater impact on the power grid operation and accurately capture the power distribution reshaping caused by topological changes. On the other hand, some edges (such as low-flow or spare lines) have less impact on the power distribution, and even if topological changes occur, their impact on the entire network is not significant. By adding the design of v pl , the weights of these edges can be effectively reduced, avoiding their excessive participation in the graph distance calculation, thereby improving the sensitivity of the graph distance to key changes. This factor is normalized to ensure that the value range is in [0,1], and the weights of all edges are comparable.

[0046] The graph distance is obtained by accumulating the influence intensities of all topological change lines, and its calculation formula is:

[0047]

[0048] where E i ΔE j is equivalent to (E i \E j ) ∪ (E j \E i ), representing the symmetric difference of the edge sets of graphs G i and G j , that is, the set of all newly added and deleted lines. This definition not only considers the changes in the topological structure but also introduces the actual impact of each line on the power flow distribution, thus making the distance calculation have physical meaning and context awareness.

[0049] This method can quantify the physical impact of topological changes as a distance metric in the graph model, providing context information for anomaly detection. For example, when sensor data anomalies are detected at a certain moment, the graph distance can be combined to analyze whether the current topological state is significantly different from the historical state, so as to identify whether the anomaly is caused by topological changes. Secondly, by comprehensively considering the impact of global and local topological changes on the power flow, this method improves the robustness and accuracy of anomaly detection, especially in complex dynamic scenarios.

[0050] The graph distance calculation method based on branch outage distribution factors can effectively capture the impact of topological changes on the power distribution of smart grids, providing important context support for anomaly detection. It has clear physical meaning, high calculation efficiency, can naturally adapt to the dynamic characteristics of smart grids, and has broad application potential in the complex and changeable power grid operation environment.

[0051] In smart grids, anomaly detection is an important part of ensuring the safe and stable operation of the power grid. Its main task is to quickly and accurately identify potential abnormal events from the measurement data of sensors, such as sensor failures, line problems, or malicious attacks. Due to the dynamics and complexity of smart grids, traditional anomaly detection methods face many challenges, including how to cope with the impact of power grid topological changes on measurement data, how to integrate multi-source information, and how to improve real-time performance and calculation efficiency. To address these issues, this application achieves precise detection of complex anomalies in smart grids by combining time-weighted modeling, weighted distribution generation, and deviation determination. The following will elaborate on the algorithm's process and implementation details.

[0052] This part of the algorithm includes three main stages: time-weighted calculation based on graph distance, weighted distribution model generation, and anomaly determination. These three stages complement each other and sequentially complete the entire process from data preprocessing to anomaly detection.

[0053] Time-weighted calculation based on graph distance: In smart grids, the dynamic changes in the power grid topology may cause changes in the statistical characteristics of measurement data. Therefore, directly applying historical data to current anomaly detection may introduce biases. To solve this problem, the algorithm first calculates the time-weighted factor based on graph distance to screen out historical data that is more similar to the current topological state as a reference for distribution modeling.

[0054] Assume the current topological state is represented as G t , and the topological states at historical moments are G t-1 , G t-2 , …, G t-wd , where wd is the historical window size. Through the graph distance calculation method D(G t ,G t-k ) proposed in this application, the difference between the current topological state and historical topological states can be quantified. Then, weights are assigned to each historical moment according to the following formula:

[0055] w k =max(λ * -D(G t ,G t-k ),0) (3-6)

[0056] where, λ *is a threshold value used to control the weight of historical data far from the current state to zero, thus reducing the interference of irrelevant data on the model. The result of time weighting calculation is a set of weights W s ={w t-1 , w t-2 , …, w t-wd}, each weight corresponding to a historical moment, reflecting the relevance of the data at that moment to the current topological state. This weight will directly affect the generation process of the weighted distribution model.

[0057] Generation of weighted distribution model and calculation of anomaly scores: After obtaining the time weighting factor, the remaining steps are that the algorithm will use these weights to weight the historical data to generate a statistical model describing the current sensor data distribution. For historical data, we focus on three metrics that have been very effective in detecting anomalies in power grid sensor data in previous studies. These three metrics are the edge anomaly metric, the average anomaly metric, and the dispersion anomaly metric.

[0058] 1. Edge anomaly metric: Measures the change in the maximum power flow in the line connected to the current sensor.

[0059] 2. Average anomaly metric: Measures the average change in power flow on the lines connected to the sensor.

[0060] 3. Dispersion anomaly metric: Measures the standard deviation of the change in power flow on all lines connected to the sensor.

[0061] For sensor s, the set of its historical measurement values is denoted as X s ={x s,t-1 , x s,t-2 , …, x s,t-wd}. Using the time weighting factor, the algorithm calculates the central tendency and dispersion of the weighted distribution, which are represented by the weighted median M s and the weighted interquartile range R s respectively:

[0062] M s = WeightedMedian(X s , W s ) (3-7)

[0063] R s = WeightedIQR(X s , W s ) (3-8)

[0064] The weighted median M s represents the central position of the data, and the weighted interquartile range R sIt reflects the degree of dispersion of the distribution. Compared with the mean and variance, this weighting method is more robust and can effectively suppress the influence of outliers on the distribution model, thus improving the stability and accuracy of the model.

[0065] Anomaly determination and pseudo-label assignment: After generating the weighted distribution model, the algorithm detects whether the measured value at the current moment is an anomaly through a deviation determination method. For the observed value x at the current moment t s,t , its deviation z s is defined as:

[0066]

[0067] The deviation z s represents the ratio of the distance between the current value and the distribution center to the distribution range. By setting a threshold τ, the observed values with a large deviation are marked as anomalies. If z s > τ, then x s,t is determined to be an anomaly. The selection of the threshold τ needs to comprehensively consider the operating characteristics of the power grid and application requirements, and is determined through experimental parameter tuning or according to experience. At the same time, the algorithm assigns pseudo-labels to these data determined to be anomalies to record the specific information of the abnormal data (such as timestamp, location, and deviation). The pseudo-label mechanism aims to optimize the subsequent detection process. By marking the confirmed abnormal data, it avoids repeated detection and reduces the interference of irrelevant data on the model.

[0068] For the data that has been marked with pseudo-labels, its weight is set to zero in the subsequent detection.

[0069] By reducing the weight of the pseudo-labeled data in data statistics, the algorithm can significantly improve the stability and accuracy of the distribution model and reduce the negative impact of abnormal data on subsequent detection. In addition, the pseudo-label mechanism also supports continuous monitoring of the marked anomalies. If some pseudo-labeled data returns to normal at a subsequent moment (for example, the deviation is lower than the recovery threshold τ recovery), it can be re-incorporated into the detection model.

[0070] To comprehensively evaluate the anomaly detection algorithm proposed in this application, the experiment designed various types of abnormal data to simulate the complex scenarios that may be encountered in the operation of the smart grid. These data cover random anomalies, local area anomalies, and global topology anomalies, aiming to verify the applicability and robustness of the algorithm under different anomaly patterns.

[0071] The experimental dataset is mainly based on the case300 model of matpower. These models provide detailed power grid topology, node load distribution, and line parameters, and can truly reflect the basic characteristics of power grid operation. To generate the input time series of the load (i.e., the active power and reactive power of each node), we used the 20-day campus load data recorded by Carnegie Mellon University (CMU) from July 29 to August 17, 2016.

[0072] In the design of random anomalies, the measurements of individual sensors are mainly simulated, aiming to test the sensitivity of the algorithm to isolated anomaly points. Such anomalies are generated by superimposing random offsets on normal data, and the offset amplitude follows a normal distribution, with the standard deviation set to 10% to 50% of the normal measurement range. The injection time and location of random anomalies are randomly selected to simulate the actual situation of sporadic interference or short-term sensor failures.

[0073] The design of local area anomalies focuses on the fault scenarios of certain lines or equipment. By modifying the operating state of specific lines (such as simulating disconnection or short circuit), the impact on the power flow distribution of adjacent nodes is calculated, and these impact amounts are injected into the corresponding sensor data. Such anomalies usually manifest as the measurements of multiple sensors in the area deviating from the normal range simultaneously, and their intensity and range are adjusted according to the type and location of the line fault. For example, in the case300 system, a key line was selected to simulate a fault, resulting in the power flow offset amplitude of the surrounding nodes reaching 30% to 70% of the normal value.

[0074] The generation of global topology anomalies is achieved by dynamically modifying the power grid topology, such as randomly disconnecting or adding several lines, adjusting the load distribution, etc. Such anomalies aim to test the adaptability of the algorithm in complex dynamic scenarios. To simulate the real topology change process, the experiment calculates the impact of topology modification on the entire power grid according to the line outage distribution factor (LODF), and injects these impacts into the corresponding sensor data. In the case300 system experiment, the simulation of load fluctuations is further added to dynamically adjust the power demand of multiple nodes to enhance the complexity of the anomaly data.

[0075] In addition, to verify the performance of the algorithm in the scenario of superposition of multiple types of anomalies, a mixed anomaly scenario is also designed. This scenario combines random anomalies, local area anomalies, and global topology anomalies to simulate the possible situation of multiple anomalies occurring simultaneously in a complex operating environment. In the case300 system, the experiment randomly selects several sensors to add random anomalies, and simultaneously simulates a regional fault and a global topology adjustment, thus generating highly complex anomaly data.

[0076] Through the generation and design of the above abnormal data, the experiment can comprehensively cover the abnormal types and their characteristics that may be encountered in the smart grid. These data include both simple isolated abnormalities and multi-point, regional, and global abnormalities, providing a solid foundation for evaluating the detection capabilities of algorithms in diverse scenarios. Next, the experiment will comprehensively verify the performance of the algorithm based on these data and further reveal its advantages and potential through comparative analysis.

[0077] To comprehensively evaluate the performance of the anomaly detection algorithm proposed in this application in the smart grid environment, a series of rigorous evaluation metrics have been designed for the experiment, covering the performance of the algorithm in multiple dimensions such as accuracy, coverage, stability, and efficiency. These metrics can not only reflect the algorithm's ability to detect abnormal data but also quantify its reliability and applicability in actual operation.

[0078] Based on the method proposed in this application and the actual application scenarios of the smart grid, performance metrics centered on F1-score and AUC (area under the ROC curve) are introduced, combined with analysis of running time and scalability, to comprehensively reflect the detection capabilities and practical application value of the model.

[0079] F1-score is an important metric for measuring the performance of classification models, especially suitable for imbalanced datasets or anomaly detection scenarios. This metric combines Precision and Recall, and through weighted harmonic mean, it can not only reflect the model's ability to reduce False Positives but also its effectiveness in capturing True Positives

[56] .

[0080] Precision refers to the proportion of samples actually detected as abnormal among those detected as abnormal, and its formula is:

[0081]

[0082] Where TP (True Positive) represents the number of samples correctly detected as abnormal, and FP (False Positive) represents the number of normal samples misjudged as abnormal. Recall measures the proportion of all actual abnormal samples correctly detected, and its formula is:

[0083]

[0084] Where FN (False Negative) represents the number of abnormal samples not detected. In model evaluation, the advantages of these two metrics are different: Precision focuses on the ability to reduce false positives, while Recall focuses on improving the coverage of anomaly detection.

[0085] The core of the F1-score lies in combining precision and recall and calculating a comprehensive metric using the harmonic mean. Its formula is

[0086]

[0087] This formula is equally important for precision and recall and can achieve a good balance between the false positive rate and false negative rate of the model. When the F1-score is close to 1, it indicates that the model performs excellently in both precision and recall; while when the F1-score is close to 0, it means that there are significant problems in both key aspects of anomaly detection for the model.

[0088] The design of the F1-score takes into account the trade-off between precision and recall and is applicable to scenarios where a comprehensive evaluation of the model's performance is required. In the anomaly detection of smart grids, it can effectively measure the model's ability to detect abnormal data and its ability to protect normal data, providing a clear direction for model optimization. It can be considered that the F1-score is an important benchmark for evaluating the performance of smart grid anomaly detection algorithms.

[0089] AUC (Area Under Curve) is another key metric used to measure the comprehensive performance of the model at different thresholds. To comprehensively evaluate the performance of the anomaly detection model, AUC (Area Under Curve) is an important evaluation metric, especially in the task of smart grid anomaly detection, which is of great significance for the model's comprehensive performance at different thresholds. AUC is the area under the ROC curve (Receiver Operating Characteristic Curve) and measures the model's ability to distinguish between positive and negative samples

[57] . The ROC curve uses the False Positive Rate (FPR) as the horizontal axis and the True Positive Rate (TPR) as the vertical axis, showing the performance of the model at different decision thresholds. The formula for TPR is shown below, and the formula for FPR is shown below. TPR (recall) is defined as the proportion of actual positive samples that are correctly detected, while FPR represents the proportion of actual negative samples that are misjudged as positive samples.

[0090] By changing the threshold to draw the ROC curve, the trade-off between reducing false positives and capturing anomalies of the model can be observed.

[0091]

[0092] The physical meaning of AUC is the probability that the model correctly judges a positive sample to be better than a negative sample when randomly selecting a positive sample and a negative sample. Its value range is [0, 1]. When AUC = 1, it indicates that the model can perfectly distinguish positive and negative samples. When AUC = 0.5, it means that the performance of the model is no different from random guessing. When AUC < 0.5, it indicates that the model performance is poor and may even wrongly tend to negative samples. As a threshold-independent metric, AUC can comprehensively reflect the overall performance of the model under different classification criteria, which is particularly important for scenarios where it is difficult to fix the threshold in anomaly detection.

[0093] In the smart grid, the proportion of abnormal samples is often extremely low, while the normal samples account for the vast majority. This data imbalance problem makes it possible that simply relying on metrics such as accuracy may not truly reflect the model performance. AUC can not only effectively evaluate the comprehensive ability of the model in such scenarios, but also intuitively display the classification characteristics and trade-off relationships of the model through the ROC curve. Especially in a dynamic environment, the change of AUC can reveal the adaptability of the model to different abnormal patterns, thus providing guidance for subsequent optimization.

[0094] The calculation of AUC can be achieved through numerical integration or sorting methods. In the discrete case, AUC can be expressed as the normalization of the ranking of positive sample scores, and its formula is:

[0095]

[0096] where n r and n n are the numbers of positive and negative samples respectively, and rank(x i ) is the ranking value of the positive sample score. In this way, AUC can be efficiently calculated in actual data and combined with the ROC curve for more in-depth analysis.

[0097] In addition, to evaluate the actual deployment performance of the model, it is also necessary to pay attention to its running time and scalability. The running time reflects the efficiency of the model in processing a single detection task. Especially in real-time detection scenarios, the time performance is a key factor to ensure the rapid response of the system.

[0098] The experimental results show that the power grid anomaly detection method based on the graph model proposed in this application, hereinafter referred to as RGM-AD (Reinforced Graph Model for Anomaly Detection), demonstrates remarkable superiority in the anomaly detection task. Through experimental comparisons with different numbers of selected sensors (from 20 to 40 sensors), the RGM-AD algorithm has achieved excellent results in both F1-score and AUC metrics. Especially when the number of sensors increases, the detection performance shows a steady upward trend, reflecting the strong robustness of this method in dynamic power grid anomaly detection.

[0099] In terms of the Mean_F1-score metric, the performance of the RGM-AD algorithm remains stable under different numbers of sensors and reaches a maximum value of 0.78 when the number of sensors is 35 and 40. This represents an increase of approximately 0.08 compared to the RGM-AD algorithm and is far higher than other baseline methods such as VAR, LOF, and Parzen. Especially when the number of sensors is small (20 to 30), the RGM-AD algorithm can still stabilize at the level of 0.76, while VAR can only achieve relatively low performance between 0.18 and 0.28, indicating that the RGM-AD algorithm has better adaptability for anomaly detection with limited sensor resources.

[0100] In terms of the Mean_AUC metric, the RGM-AD algorithm also demonstrates its excellent detection ability. When the number of sensors gradually increases from 20 to 40, the AUC value increases from 0.9604 to 0.9772, always remaining at a relatively high level. Compared with DynWatch, RGM-AD is slightly inferior when the number of sensors is low (20 and 25), but when the number of sensors reaches 30 and above, RGM-AD shows better detection performance, exceeding the highest AUC value (0.9763) of the DynWatch algorithm. In addition, compared with baseline methods such as LOF and Parzen, the RGM-AD algorithm has obvious advantages in the AUC value under all experimental conditions, further verifying the efficiency and robustness of this method in dynamic topology scenarios.

[0101] It is worth noting that the performance of the VAR method does not improve significantly with the increase in the number of sensors and even shows a decline in performance when the number of sensors is large (35 and 40). This indicates that traditional time series modeling methods are difficult to effectively handle complex anomaly patterns in dynamic power grids. Although the LOF and Parzen methods have achieved certain results in the AUC metric (reaching maximum values of 0.8146 and 0.7913 respectively), their F1-score metric is relatively low, indicating that these methods have deficiencies in anomaly localization and false alarm reduction.

[0102] Based on the comprehensive experimental results, the RGM-AD algorithm, by making full use of the sensor grouping strategy and graph model, successfully combines local sensitivity and global topological information, thus achieving efficient detection of complex anomalies in the dynamic power grid under limited resource conditions. Compared with other baseline methods, RGM-AD has obvious advantages in detection accuracy and robustness, can better meet the actual power grid monitoring requirements, and provides an effective solution for anomaly detection in dynamic environments.

[0103] Table 3-1 Experimental results of F1-score index

[0104]

[0105]

[0106] Table 3-2 Experimental results of AUC index

[0107]

Claims

1. A method for detecting anomalies in a power grid based on a graph model, comprising the following steps: 1) Construct a power grid graph model: various devices in the power grid are taken as nodes, and lines and power flow connections are taken as edges, forming a power grid graph model with node sets and edge sets; 2) Graph distance calculation based on branch disconnection distribution factor: Compare the topological states of the power grid at different times, use the branch disconnection distribution factor to measure the impact of adding or deleting lines on the power flow distribution of other lines, and calculate the graph distance between the topological states at two times in combination with the edge weight; 3) Anomaly detection algorithm based on graph distance calculation: including time-weighted calculation based on graph distance, weighted distribution model generation, anomaly judgment and pseudo-label allocation.

2. The method for detecting anomalies in a power grid based on a graphical model according to claim 1, characterized in that: Time-weighted calculation based on graph distance: obtain the grid topology state at the current moment, compare it with the topology state at each moment in the historical window, and calculate the graph distance between each moment; the graph distance is determined by the branch disconnection distribution factor, which is used to quantify the impact of a certain line on the power flow of other lines when it is disconnected; according to the size of the graph distance, the weight of historical data that differs too much from the current topology is reduced to zero, and more representative historical data is screened out, and a distance-decreasing function is used to assign a time-weighted factor to it.

3. The method for detecting anomalies in a power grid based on a graphical model according to claim 2, characterized in that: Weighted distribution model generation: Weighted statistics are performed on the historical sensor measurements that are screened and retained, and the weighted median and weighted interquartile range are calculated to describe the center and dispersion of the distribution, respectively, to obtain a reference distribution model that is more robust to outliers.

4. The method for detecting anomalies in a power grid based on a graphical model according to claim 3, characterized in that: Anomaly determination and pseudo-label assignment: Compare the current measurement value with the weighted distribution model, and determine whether it is abnormal based on the deviation; when the deviation exceeds the preset threshold, mark the sensor measurement data as abnormal and assign a pseudo-label; for data marked as pseudo-labels, their weights for the distribution model will be set to zero in subsequent tests to prevent abnormal data from continuously interfering with the update of the model; if the abnormal value returns to normal levels at a subsequent moment, its pseudo-label can be removed and re-included in the distribution statistics.

5. The method for detecting anomalies in a power grid based on a graphical model according to claim 1, characterized in that: The graph distance calculation based on the branch disconnection distribution factor includes the following steps: a) Obtain the power grid topology at the current time and historical time, and identify the newly added or deleted lines; b) Calculate the branch disconnection distribution factor to measure the impact of a single line disconnection on the power flow distribution of other lines; c) The impact values ​​of the edge weights and the branch disconnection distribution factors are combined, and all newly added or deleted lines are accumulated to obtain the graph distance between the topological states at two moments. This distance can be used to quantify the physical impact of the topological change into a unified metric.

6. The method for detecting anomalies in a power grid based on a graphical model according to claim 5, characterized in that: The edge weight is obtained by normalizing the power flow of the line or its importance in the grid topology, and the contribution value of the high-weight line is amplified when calculating the graph distance.

7. The method for detecting anomalies in a power grid based on a graphical model according to claim 1, characterized in that: The time-weighted calculation strategy based on graph distance in step 3) includes: a) Calculate the graph distance between the current topological state and the topological state at the historical moment; b) if the graph distance is greater than a preset threshold, the historical data is discarded; if the graph distance does not exceed the threshold, a time weighting factor is assigned to the historical data according to a distance decreasing rule; c) Complete the screening and weighting of historical measurement data based on the above weighting factors.

8. The method for detecting anomalies in a power grid based on a graphical model according to claim 4, characterized in that: The calculation of the deviation of the sensor measurement value in the abnormality determination and pseudo-label assignment step includes at least the following abnormality indicators: a) Edge anomaly index: the maximum power flow change in the sensor connection line; b) Average abnormal index: the average value of the power flow change of the sensor connection line; c) Discreteness anomaly index: the standard deviation of the power flow change in the sensor connection line.

9. The method for detecting anomalies in a power grid based on a graphical model according to claim 2, characterized in that: When performing time weighting, historical data are also screened secondary according to the sensor's geographical location or its electrical distance from the current node to improve the positioning accuracy of local anomalies or regional faults.

10. The method for detecting anomalies in a power grid based on a graphical model according to claim 4, characterized in that: After completing the anomaly judgment and pseudo-label assignment, it also includes an iterative update step for the weighted distribution model. When the data with assigned pseudo-labels remains in an abnormal state for more than a preset time or the deviation continues to increase during a continuous detection cycle, the data of the sensor is removed from the current weighted statistical window to prevent long-term outliers from causing continuous interference to the weighted distribution model.

Citation Information

Patent Citations

  • Water environment monitoring method based on Internet of Things

    CN118275640A

  • Power distribution network panoramic dynamic topology anomaly detection method and system based on graph calculation

    CN118353011A

Cited By

  • Machine room fault positioning dynamic environment monitoring system

    CN120742162A