Power system anomaly identification method and device and electronic equipment
By combining a power grid anomaly identification model with a knowledge graph, and utilizing federated learning and path association analysis, the problem of accurate identification and cause analysis of abnormal events in the power system has been solved, realizing intelligent and real-time anomaly management and improving the stability and security of the power grid.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies have low accuracy in identifying and analyzing the causes of abnormal events in power systems. In particular, they are difficult to capture the inherent correlation and deep-seated causal relationships of abnormal events in complex power grids, and they cannot effectively utilize massive amounts of operational data to improve the accuracy of identification.
A power grid anomaly identification model is adopted to obtain current operational feature data through federated learning. The path association score is calculated by combining the target knowledge graph to identify the root cause entity of the abnormal event and its scope of influence. The combination of federated learning and knowledge graph realizes intelligent and real-time anomaly identification.
It improves the accuracy of identifying and analyzing the causes of abnormal events in the power system, enabling precise location of the root causes and the scope of their impact, thereby enhancing the stability and security of the power grid.
Smart Images

Figure CN121786677A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grids, and more specifically, to a method, apparatus, and electronic device for identifying power system anomalies. Background Technology
[0002] In the operation and maintenance of power systems, the identification and causal analysis of abnormal events are crucial for ensuring the stability and security of the power grid. However, related technologies have significant limitations in addressing this issue, especially when analyzing the causes of anomalies in complex power networks, where their accuracy often falls short of practical requirements. Methods in related technologies, such as rule-based expert systems or threshold judgments, while having some applications in specific scenarios, lack flexibility and adaptability in the face of the ever-increasing scale and complexity of power grids, making it difficult to capture the inherent correlations and deep-seated causal relationships of abnormal events. Furthermore, with the rapid development of smart grids, massive amounts of operational data are generated in power systems. This data contains potential signals of grid anomalies, but related technologies often fail to effectively utilize this data to improve the accuracy of anomaly identification. On the one hand, power data from a single region may be limited by local characteristics, resulting in poor generalization ability of the trained models; on the other hand, the complex causal relationships between entities within the power system are difficult to capture using simple statistical methods or machine learning models, especially when abnormal events involve the interaction of multiple power grid components, making it difficult for related technologies to accurately identify the root cause entities and their scope of influence.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a method, apparatus, and electronic device for identifying power system anomalies, in order to at least solve the technical problem of low accuracy in identifying abnormal events and analyzing the causes of anomalies in power systems in related technologies.
[0005] According to one aspect of the present invention, a power system anomaly identification method is provided, comprising: acquiring current operating characteristic data of a power system in a target area; based on the current operating characteristic data, employing a power grid anomaly identification model to obtain an initial anomaly identification result of the power system in the target area, wherein the power grid anomaly identification model is obtained through federated learning based on historical operating characteristic data of multiple areas and corresponding historical anomaly identification results, and the initial anomaly identification result includes a target anomaly event and an anomaly probability of the target anomaly event; based on the initial anomaly identification result, querying a target knowledge graph to obtain a path association score of the target anomaly event, wherein the target knowledge graph includes entities in the power system of multiple areas, entity vectors corresponding to the entities, and relationships between entities, the path association score is used to quantify the potential causal association strength between the target anomaly event and other entities within the power system, and the entity vector is used to indicate the semantic features and historical behavioral features of the entity; and determining a target anomaly identification result of the power system in the target area based on the initial anomaly identification result and the path association score, wherein the target anomaly identification result includes the root cause entity of the target anomaly event, and the comprehensive causal score and influence range measure of the root cause entity.
[0006] According to another aspect of the present invention, a power system anomaly identification device is also provided, comprising: a feature data acquisition module, configured to acquire current operating feature data of a power system in a target area; an initial anomaly identification module, configured to obtain an initial anomaly identification result of the power system in the target area based on the current operating feature data and using a power grid anomaly identification model, wherein the power grid anomaly identification model is obtained through federated learning based on historical operating feature data of multiple areas and corresponding historical anomaly identification results, and the initial anomaly identification result includes a target anomaly event and an anomaly probability of the target anomaly event; a path association score acquisition module, configured to query a target knowledge graph based on the initial anomaly identification result to obtain a path association score of the target anomaly event, wherein the target knowledge graph includes entities in the power system of multiple areas, entity vectors corresponding to the entities, and relationships between entities, the path association score is used to quantify the potential causal association strength between the target anomaly event and other entities within the power system, and the entity vector is used to indicate the semantic features and historical behavioral features of the entity; and a target anomaly identification module, configured to determine a target anomaly identification result of the power system in the target area based on the initial anomaly identification result and the path association score, wherein the target anomaly identification result includes the root cause entity of the target anomaly event, and the comprehensive causal score and influence range measure of the root cause entity.
[0007] According to another aspect of the present invention, a non-volatile storage medium is also provided, which stores a plurality of instructions adapted for a power system anomaly identification method to be loaded by a processor and executed at any one of them.
[0008] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the power system anomaly identification methods.
[0009] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of any one of the power system anomaly identification methods.
[0010] In this embodiment of the invention, current operational characteristic data of the power system in the target area is acquired. Based on the current operational characteristic data, a power grid anomaly identification model is used to obtain initial anomaly identification results of the power system in the target area. The power grid anomaly identification model is obtained through federated learning based on historical operational characteristic data of multiple areas and corresponding historical anomaly identification results. The initial anomaly identification results include the target anomaly event and its anomaly probability. Based on the initial anomaly identification results, a target knowledge graph is queried to obtain the path association score of the target anomaly event. The target knowledge graph includes entities in the power system of multiple areas, entity vectors corresponding to the entities, and relationships between entities. The path association score is used to quantify the strength of potential causal relationships between the target anomaly event and other entities within the power system. This method uses semantic features and historical behavioral features of entities to indicate their characteristics. Based on the initial anomaly identification results and path association scores, it determines the target anomaly identification results for the power system in the target area. The target anomaly identification results include the root cause entity of the target anomaly event, as well as the comprehensive causal score and impact range measurement of the root cause entity. This achieves intelligent and real-time identification of anomaly events by integrating the power grid anomaly identification model obtained through federated learning with the path association analysis method in the dynamic knowledge graph. By combining the semantic features and historical behavioral features of entities, it aims to accurately locate the root cause of anomaly events and their impact range, thereby improving the technical effect of identifying anomaly events and analyzing the causes of anomalies in the power system. This solves the technical problem of low accuracy in identifying anomaly events and analyzing the causes of anomalies in the power system in related technologies. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0012] Figure 1 This is a flowchart of a power system anomaly identification method according to an embodiment of the present invention;
[0013] Figure 2 This is a schematic diagram of the structure of an optional power system anomaly identification system according to an embodiment of the present invention;
[0014] Figure 3 This is a schematic diagram of a power system anomaly identification device according to an embodiment of the present invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] According to an embodiment of the present invention, a method for identifying power system anomalies is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0018] Figure 1 This is a flowchart of a power system anomaly identification method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0019] Step S102: Obtain the current operating characteristic data of the power system in the target area.
[0020] Optionally, real-time operating parameters can be collected from the power system in the target area, including, but not limited to, voltage, current, temperature, and power. This data can be provided by, but is not limited to, sensors, measuring devices, and automation systems within the power system, forming the foundational information for the current state of the power grid. For smart grids, the current operating characteristic data can also include the real-time status of advanced grid models, such as power flow calculation results and voltage distribution maps.
[0021] Step S104: Based on the current operating characteristic data, the power grid anomaly identification model is used to obtain the initial anomaly identification result of the power system in the target area. The power grid anomaly identification model is obtained through federated learning based on the historical operating characteristic data of multiple areas and the corresponding historical anomaly identification results. The initial anomaly identification result includes the target anomaly event and the anomaly probability of the target anomaly event.
[0022] Optionally, a power grid anomaly identification model trained through federated learning can be used to analyze real-time acquired operational characteristic data to identify whether abnormal events exist in the power system and their probabilities. Federated learning is a distributed machine learning technique that allows multiple participants (such as local servers corresponding to power grid operators in different regions) to jointly train a model without sharing raw data, thereby protecting data privacy while improving model performance. Based on a large amount of historical operational data and known anomaly cases, the power grid anomaly identification model can learn the differences between normal operation and abnormal conditions of the power system, providing a scientific basis for early warning of abnormal events.
[0023] In an optional embodiment, before obtaining the initial anomaly identification result of the power system in the target area based on the current operating feature data and using the power grid anomaly identification model, the method further includes: training the current local model parameters of the target area based on the training dataset to obtain the trained local model parameters, and sending the trained local model parameters to the central processing unit (CPU) for the CPU to update the global model parameters. The training dataset includes a first number of historical operating feature data of the power system in the target area, and corresponding historical anomaly identification results. The CPU receives the updated global model parameters returned by the CPU. Based on the updated global model parameters, the CPU updates the trained local model parameters to obtain the updated local model parameters. The updated local model parameters are used as the new current local model parameters, and the above operations are repeated until a preset termination condition is reached. The CPU is also used to obtain a power grid anomaly identification model based on the updated global model parameters obtained when the preset termination condition is reached, and the CPU is also used to send the power grid anomaly identification model to the local servers corresponding to each of the multiple areas.
[0024] Optionally, in each participating region, including the target region, a local model is first trained based on a training dataset unique to that region (including historical operational feature data and corresponding historical anomaly identification results). This training dataset is confidential, processed only locally, and not shared with other regions to ensure data privacy and security. Through this training process, the local model parameters are optimized to better adapt to the region's unique power system characteristics and anomaly patterns. The trained local model parameters are then transmitted to the central processing unit (central server), but this transmission does not include the original data; it only contains updated information about the model parameters. The central processing unit adjusts and optimizes the global model parameters by aggregating the updates from the local model parameters of all participating regions. This process is the core of federated learning, allowing for collaborative improvement of model parameters while ensuring the independence and privacy of data from each region. The optimized global model parameters are then sent back to the local servers in each region by the central processing unit. Upon receiving the updated global parameters, each region's local model updates its own model parameters accordingly, forming the updated local model parameters. This update process may be repeated multiple times until the changes in the model parameters reach a preset termination condition, such as a convergence criterion or the upper limit of the training epochs. In this way, the local model in each region gradually aligns with the globally optimized model, improving the model's generalization ability. Once the preset termination condition is met, the central processing unit (CPU) constructs a power grid anomaly identification model based on the final global model parameters and distributes it to the local servers in each region. These servers deploy the latest power grid anomaly identification model for real-time analysis and prediction of the power system's operating status within their respective regions, identifying potential anomalies early. Through the federated learning mechanism, the model benefits from the collective wisdom of multiple regions, demonstrating good anomaly identification performance even on data-constrained local servers. The application of federated learning in power system anomaly identification not only overcomes the problem of data silos but also effectively protects the privacy of power grid operation data. Through cross-regional model parameter collaboration, it significantly improves the model's accuracy and adaptability. This method is particularly suitable for the power industry because the operating characteristics of power systems differ significantly across regions. Federated learning can fully utilize this diversity to construct more robust and comprehensive anomaly identification models, helping power system operators respond to various anomalies promptly and accurately, reducing power outages and ensuring stable power grid operation.
[0025] Optionally, multi-source heterogeneous data from power grid operation, dispatch monitoring, and equipment can be uniformly collected to form a raw dataset. Let the collected data sequence be denoted as... ,in, Indicates time A data record acquired at any given time. This process ensures a complete and continuous input foundation for subsequent data quality governance. During the data cleaning phase, operations such as outlier removal, missing value imputation, and time alignment are primarily performed. To identify noisy data that significantly deviates from the normal range, threshold or statistical interval determination methods are used, defining the cleaning rules for numerical fields as follows: In the formula, This represents the data values after cleaning. and These are the minimum and maximum acceptable values for the field, respectively. For those marked as... Missing or outlier items can be filled in by interpolation or mean, and can be represented as: In the formula, This represents a strategy function that uses mean, linear interpolation, or neighbor-sample imputation to obtain the imputed value. Structured feature vectors are constructed from the cleaned data for subsequent model training and quality evaluation. The extracted feature vectors are denoted as: In the formula, For at any time The generated multidimensional feature vector, The feature mapping function can include operations such as normalization, time window statistics, and differential feature construction. For example, normalization can be represented as: In the formula, Represents the characteristic mean. The standard deviation of the features is used to unify data with different dimensions to a stable numerical range. After feature unification and structuring, the final feature sequence can be directly used for training and diagnosis. In the formula, That is, the structured feature set output by the preprocessing module is used to build the training dataset and the test dataset.
[0026] Optionally, the construction process of the power grid anomaly identification model is mainly divided into three stages: local training, model aggregation, and cross-domain evaluation. In the local training stage, each node (i.e., the local server) independently trains its model based on the preprocessed dataset. Let the... The dataset of each node contains Each sample (i.e., the corresponding second number of historical operational feature data, and the corresponding historical anomaly identification results) has local model parameters denoted as... By minimizing the training error within a node, the node obtains updated local model parameters, as shown below: In the formula, Represents a node Local training data, This indicates that one or more rounds of local model training will be performed on this data. This step ensures that the model can adapt to the characteristics of the node's own data while avoiding leakage of raw data. During the model aggregation phase, each node only uploads the trained model parameters. The data is then fed to the central server, where it is weighted and merged based on the amount of data from each node to obtain unified global model parameters. The polymerization process can be represented as: In the formula, Represents the total number of participating nodes. Represents a node The number of samples, For nodes Uploaded local model parameters. Weighted aggregation ensures that nodes with larger datasets contribute more weight in the fusion process, improving the robustness of the global model. The aggregated model or global model parameters are then synchronously returned to all nodes as initial parameters for the next round of local training, achieving cross-node knowledge sharing. During the cross-domain evaluation phase, each node uses its own test dataset (including the corresponding second-largest amount of historical running feature data and corresponding historical anomaly identification results) to verify the global model performance, checking the model's adaptability in different regions and business scenarios. Let the nodes... The test set is The prediction accuracy of the model at this node is expressed as: In the formula, This indicates that the model is at the node. The number of correctly predicted samples in the test set. This represents the total number of samples in the test set. Cross-domain performance evaluation can promptly identify whether the model exhibits biases in certain regions, further guiding iterative model optimization. This process is repeated until a preset termination condition is met, and based on the updated global model parameters obtained when the preset termination condition is met, the power grid anomaly identification model is obtained.
[0027] Step S106: Based on the initial anomaly identification results, query the target knowledge graph to obtain the path association score of the target anomaly event. The target knowledge graph includes entities in the power system of multiple regions, entity vectors corresponding to the entities, and relationships between entities. The path association score is used to quantify the strength of potential causal relationships between the target anomaly event and other entities within the power system. The entity vector is used to indicate the semantic features and historical behavioral features of the entity.
[0028] Optionally, after obtaining preliminary anomaly identification results, knowledge graph technology can be further utilized to explore the causal relationships between anomalous events and other power grid entities (such as generators, transmission lines, and substations). Knowledge graphs store the attributes, states, and complex interactions between entities within the power system, providing richer contextual information than simple data analysis. By calculating path association scores, the strength of potential causal connections between anomalous events and various entities is assessed, thereby aiding in the identification of key factors that may lead to anomalies. Entity vectors are numerical representations of each entity in the graph, incorporating its semantic characteristics and historical behavior, facilitating the understanding of entity roles and dynamic changes by knowledge graph analysis algorithms.
[0029] Optionally, the target knowledge graph in this embodiment can be obtained by dynamically updating the initial knowledge graph in real time based on the output of the power grid anomaly identification model over time. Before executing step S106, the initial knowledge graph can be constructed in the following manner, including two stages: constructing entities and constructing relationships. In the entity construction stage, the physical objects and abstract concepts in the power system are first mapped to entity sets. , represented as Each entity It can represent substations, switches, busbars, measurement points, alarm events, or business entities. To facilitate subsequent calculations and reasoning, a low-dimensional entity vector representation is assigned to each entity. ,in For the embedding dimension, vector Used to capture the semantic and historical behavioral features of entities. The entity construction process also retains a set of entity metadata. For example, device ID, the site to which it belongs. Measurement type Timestamp range, etc., are denoted as This metadata is used to provide contextual constraints during subsequent association and updates. In the association phase, the set of relationships between entities is defined. Each relation is represented as a triple. That is, the head entity Relationship types With tail entity To measure the reliability or association strength of triples, a scoring function is introduced. (Taking the negative distance form to make it more reliable as larger as possible), it can be represented as: In the formula, These are vector representations of the head entity, relation type, and tail entity, respectively. For embedded dimensions, Denotes the vector norm. When The smaller the value, the more closely the three factors match, and the higher the score will be. The larger the value (negative values close to 0), the more reliable the relationship. To facilitate unified management of the relationship structure within the system, all pairwise relationships between entities can be organized into an adjacency matrix. In the matrix Representing entities With entity There exists a certain type of relationship (such as "belonging", "connection", "triggering", etc.), among which This represents the total number of entities in the knowledge graph. Each relation edge is assigned a weight. This is used to represent the strength of the association or the confidence level of the relationship between entity pairs. Weight It can be derived from the scoring function. Historical statistics or model inference.
[0030] In an optional embodiment, before querying the target knowledge graph based on the initial anomaly identification result to obtain the path association score of the target anomaly event, the method further includes: if the target anomaly event is detected as a newly added anomaly event, determining an incremental triple set, a new entity set, and a newly added entity vector based on the initial anomaly identification result, wherein the incremental triple set includes newly added entities and entity relationships related to the target anomaly event; performing a union operation on the incremental triple set and the triple set in the current knowledge graph, performing a union operation on the new entity set and the entity set in the current knowledge graph, and updating the current entity vector in the current knowledge graph to obtain the target knowledge graph.
[0031] Optionally, if the knowledge graph update in the current round is the first update, the current knowledge graph is the initial knowledge graph; if the knowledge graph update in the current round is not the first update, the current knowledge graph is obtained by iteratively updating the initial knowledge graph through historical rounds.
[0032] Optionally, the incremental triple set includes entities related to the target anomalous event and their relationships, such as the device entity that triggered the anomalous event, related events, or business entities, as well as the types of relationships between these entities. The incremental triple set is then combined with the triple set in the current knowledge graph, and the new entity set is combined with the entity set in the current knowledge graph to obtain the updated knowledge graph (i.e., the target knowledge graph). The execution entity for the above-mentioned update process of the current knowledge graph can be the local server corresponding to the target region or the central server. When a new anomalous event is detected, the characteristics of the anomalous event and its relationships with other entities in the power system are analyzed based on the initial anomalous event identification results generated by the power grid anomaly identification model. This determines an incremental triple set, which contains new triples containing the new anomalous event and related entities and their relationships. Simultaneously, entities appearing for the first time in this event are identified, forming a new entity set. The establishment of the incremental triples and the new entity set provides the necessary basic information for updating the knowledge graph. Next, the incremental triplet set is joined with the existing triplet set in the knowledge graph, adding the new relationships to the knowledge graph to reflect the latest interactions between power grid entities. Simultaneously, the new entity set is joined with the existing entity set to ensure all relevant entities are included in the knowledge graph. For these new entities, their corresponding entity vectors need to be initialized or updated. An entity vector is a numerical representation that summarizes the entity's attributes and historical behavior, helping the knowledge graph algorithm understand and infer the entity's role and influence.
[0033] Through the above steps, the target knowledge graph is dynamically expanded and updated. The updated knowledge graph contains more information about anomalous events and a more comprehensive network of power grid entity relationships. Subsequently, this updated knowledge graph is used to calculate the path association score of the target anomalous event, more accurately assessing the strength of causal relationships between the event and other entities, thereby assisting in the analysis of anomalous events and root cause identification. This process of dynamically updating the knowledge graph ensures the accuracy and comprehensiveness of anomalous event analysis, enabling power grid anomaly identification methods to continuously adapt to changes in the power system and promptly capture newly emerging anomaly patterns and potential risks. This method is particularly important in the operation, management, and maintenance of smart grids, effectively improving the efficiency of anomalous event early warning and handling, avoiding false alarms or missed alarms caused by the lag of static knowledge bases, thereby enhancing the overall stability and security of the power grid.
[0034] In one optional embodiment, updating the current entity vector in the current knowledge graph includes: determining the weight values corresponding to the current entity vector and the newly added entity vector included in the current knowledge graph; performing a weighted operation based on the current entity vector, the newly added entity vector, and the weight values corresponding to the current entity vector and the newly added entity vector to update the current entity vector in the current knowledge graph, thereby obtaining the entity vector in the target knowledge graph.
[0035] Optionally, when updating entity vectors in the knowledge graph, it is first necessary to determine the weight values corresponding to the current entity vector and the newly added entity vector. The weight values can be determined based on factors such as the timeliness, reliability, and importance of the vector information. For example, if the newly added entity information is directly related to a recent anomaly, its weight value may be higher, and vice versa. Setting weight values helps to achieve a balance between old and new data, avoiding excessive impact on entity vectors due to abnormal fluctuations of a single data point. Once the weight values are determined, a weighted calculation will be performed to merge the current entity vector with the newly added entity vector. This calculation process can employ various mathematical methods, such as linear weighting and exponential weighting, depending on the design of the vector update strategy. Through weighted calculation, not only can historical information in the entity vectors be preserved, but the latest anomaly event information can also be introduced in a timely manner, ensuring the dynamic adaptability and accuracy of the entity vectors. The result of the weighted calculation is used to update the entity vectors in the knowledge graph, generating the latest entity vectors in the target knowledge graph. This update process can be performed at the backend of the knowledge graph maintenance, ensuring that the vector information of all entities reflects the latest operating status of the power system and the impact of anomalies in real time. The updated entity vectors will serve as the foundation for subsequent anomaly identification, analysis, and response, and are crucial for improving the intelligence level and response speed of power system operation and maintenance. The dynamic update mechanism of entity vectors is key to maintaining the vitality and adaptability of the knowledge graph. In power system operation and maintenance, this mechanism helps the system promptly capture changes in entity states, especially the occurrence of grid anomalies, enabling the knowledge graph to more accurately simulate the complex relationship network of the power system and improve the depth and breadth of anomaly analysis. By scientifically setting weight values and using weighted calculation methods, it can be ensured that the update of entity vectors reflects both the important information of current anomalies and the long-term trends of historical data, thereby achieving more accurate prediction and decision-making in grid anomaly management.
[0036] Optionally, during the dynamic update phase of the knowledge graph, the knowledge graph evolves incrementally over time and with the model output, recording time points. The knowledge graph, that is, the current knowledge graph is ,in, This represents the set of entities in the current knowledge graph; This represents the set of triples (i.e., entity relations) in the current knowledge graph. When the federated model (i.e., the power grid anomaly detection model) detects a new anomaly or when new measurement / equipment information arrives, the system generates an incremental set of triples. With possible new entity sets And the update is completed through a simple union operation: In the formula, and These can originate from two sources: first, data streams (such as new sensor deployments or changes in measurement points); and second, inferences from the federated modeling module (such as detected anomalies or potential causal chains). To maintain the timeliness of entity and relation vectors, an exponentially weighted update rule is applied to the affected entities: In the formula, For a moment entity vectors, For the instantaneous vector representation computed from new observations or new triples (e.g., by averaging the vectors of newly associated neighbors or using mini-batch training), the parameters Weighting coefficients for preserving historical memory (e.g.) (This indicates a primary focus on historical data, with minor adjustments based on new information). For newly added relationships, their initial weights... Can be determined by a scoring function The confidence value is obtained through normalization or directly assigned by the detection module. .
[0037] Based on the constructed and real-time updated semantic network described above, the dynamic knowledge graph module supports graph reasoning and path scoring to aid in understanding anomalies. For example, for suspected root cause entities... With the affected entity It can calculate the path association strength score between two entities. This is the product or weighted sum of the edge weights along the path: In the formula, Traverse all from arrive The path, This represents the weight of the edges along the path; a higher score indicates a stronger semantic association between the two entities. This score, combined with the anomaly confidence score of the federated model, can provide interpretable semantic evidence for anomaly root cause analysis.
[0038] Step S108: Based on the initial anomaly identification results and the path association score, determine the target anomaly identification results of the power system in the target area. The target anomaly identification results include the root cause entity of the target anomaly event, as well as the comprehensive causal score and impact range measure of the root cause entity.
[0039] Optionally, by combining the probability assessment provided by the anomaly identification model and the path association score calculated from the knowledge graph, the root entity of the anomaly event and the potential impact range of its abnormal state can be comprehensively determined. The root entity is the core trigger for the anomaly event, while the comprehensive causal score reflects the magnitude of its substantial influence. Measuring the impact range helps define the affected domain of the anomaly, guiding subsequent operation and maintenance strategies. This process aims to ensure that anomaly management in the power system is both accurate and efficient, helping to locate the source of problems in a timely manner, reduce system recovery time, and improve the overall stability and security of the power grid. The method in this embodiment can not only identify anomaly events in the power grid but also deeply analyze the causal relationships of events, locating specific root entities, providing a powerful tool for fault diagnosis and prevention in the power system.
[0040] In one optional embodiment, the target anomaly identification result of the power system in the target area is determined based on the initial anomaly identification result and the path association score. This includes: obtaining a target anomaly score for the target anomaly event based on the anomaly probability of the target anomaly event in the initial anomaly identification result and the path association score; if the target anomaly score is greater than a preset score threshold, determining a set of candidate root cause entities corresponding to the target anomaly event based on the target knowledge graph; determining the comprehensive causal score corresponding to each of the multiple candidate entities included in the set of candidate root cause entities; determining the root cause entity and the root cause path corresponding to the target anomaly event from the multiple candidate entities based on the comprehensive causal scores corresponding to each of the multiple candidate entities, wherein the root cause path is the connection path between the target anomaly event and the root cause entity; and determining the influence range measure of the root cause entity based on the weight of the edges between the entities included in the root cause path, wherein the weight of the edges between the entities is used to quantify the association strength between the entities.
[0041] Optionally, after obtaining the initial anomaly identification results (including the anomaly event and its probability) and path association scores, a comprehensive evaluation of the anomaly event is further performed to obtain a target anomaly score. This score not only considers the probability of the anomaly occurring but also combines the strength of the causal relationship between the anomaly event and other power grid entities, providing a more comprehensive indicator of the anomaly's importance. When the anomaly score exceeds a preset scoring threshold, it means that the anomaly identification result has high credibility, and further investigation into the underlying causes of the anomaly is needed. At this point, based on the target knowledge graph, entities with potential causal relationships with the target anomaly event are selected to form a candidate root cause entity set. These entities may be the direct cause of the anomaly or key points leading to the anomaly through a series of indirect influences. For each entity in the candidate root cause entity set, its comprehensive causal score needs to be determined. The comprehensive causal score is a key indicator for measuring the degree of influence of an entity on the anomaly event. It combines the semantic features and historical behavioral information contained in the entity vector with the strength of the causal relationship between the current event and the entity reflected in the path association score. Through calculation, it can be determined which entities play a more important role in the anomaly chain. Among all candidate entities, the entity with the highest comprehensive causal score is selected as the root cause entity. The direct and indirect causal paths between this root cause entity and the anomaly are determined through edge relationships in the knowledge graph. This path reveals the specific route of anomaly propagation, providing guidance for subsequent fault isolation and repair. Finally, by analyzing the edge weights (representing the strength of the association between entities) on the root cause path, the impact range of the root cause entity can be estimated. This impact range measurement helps operations and maintenance personnel understand the scope of the anomaly, develop effective emergency response plans, reduce unnecessary intervention, and improve processing efficiency. The entire process begins with the initial identification of the anomaly and gradually delves into the precise location of the root cause entity and the measurement of its impact range. Utilizing the power grid anomaly identification model obtained through federated learning and the deep analysis capabilities of the knowledge graph, the intelligence and targeting of power system anomaly management can be significantly improved. Compared to fault detection and troubleshooting processes in related technologies, this method can respond to power grid anomalies faster and more accurately, reduce power outage time, and improve the quality of power supply services.
[0042] In one optional embodiment, a target anomaly score is obtained based on the anomaly probability of the target anomaly event in the initial anomaly identification result and the path association score, including: determining a first weight corresponding to the anomaly probability of the target anomaly event and a second weight corresponding to the path association score; and performing a weighted calculation based on the anomaly probability of the target anomaly event, the path association score, the first weight, and the second weight to obtain the target anomaly score.
[0043] Optionally, when calculating the target anomaly score, weights are assigned to the anomaly probability and path association score of the anomaly event, namely, a first weight and a second weight. These two weights reflect the relative importance of two different types of information in assessing the anomaly event. For example, if the anomaly probability of an event is very high, the first weight may be set higher, and vice versa. Similarly, if the path association score shows a strong causal relationship between the anomaly event and other entities, the second weight will also be increased accordingly. The selection of weights should be based on the actual operating experience of the power system and the prediction of the consequences of the anomaly event to ensure the rationality and effectiveness of the scoring system. Next, the anomaly probability and path association score of the target anomaly event are weighted using the determined first and second weights. This calculation process combines the probability of the anomaly event with the intensity of its potential impact to form a target anomaly score that reflects the overall risk level of the anomaly event. The weighting calculation can employ various mathematical methods, such as linear weighting and exponential weighting, depending on the weight allocation strategy and the design of the calculation formula for the target anomaly score. The target anomaly score obtained through the above weighting calculation provides a quantitative risk assessment standard for anomalies in the power system. A higher score indicates a greater likelihood of an anomaly occurring and a more significant potential impact on the power grid. Maintenance personnel can use the score to determine whether emergency response measures are needed and which anomalies to prioritize, thereby optimizing resource allocation and improving the efficiency and accuracy of anomaly handling.
[0044] This embodiment introduces a weighting mechanism to scientifically balance the probability information of abnormal events with the strength of causal relationships, forming a scoring system that can reflect both the probability of an event occurring and the scale of its impact. This provides a more powerful and intelligent decision support tool for power system anomaly identification and management, and helps to quickly locate potential faults.
[0045] In one optional embodiment, determining the comprehensive causal score corresponding to each of the multiple candidate entities included in the candidate root cause entity set includes: obtaining the comprehensive causal score of any candidate entity among the multiple candidate entities by: determining the model anomaly confidence aggregate score of any candidate entity, wherein the model anomaly confidence aggregate score is obtained based on the historical anomaly probability corresponding to any candidate entity; determining the semantic association strength between any candidate entity and the target anomaly event based on the target knowledge graph; determining the weight values corresponding to the model anomaly confidence aggregate score and the semantic association strength respectively; performing a weighted operation based on the model anomaly confidence aggregate score, the semantic association strength, and the weight values corresponding to the model anomaly confidence aggregate score and the semantic association strength respectively to obtain the comprehensive causal score of any candidate entity; and obtaining the comprehensive causal score corresponding to each of the multiple candidate entities by using the method of obtaining the comprehensive causal score of any candidate entity.
[0046] Optionally, the model anomaly confidence aggregate score is based on a statistical summary of the anomaly confidence of the observations related to the candidate entity, such as using the mean or median of the federated learning model output. Semantic association strength can be obtained by finding the shortest or most reliable path from the candidate entity to the target anomaly event in the target knowledge graph, specifically using the path association score or based on the weighted product or weighted sum of the edges between entities on the path.
[0047] Optionally, firstly, a power grid anomaly identification model trained using federated learning analyzes real-time operational characteristic data from the power system in the target area to generate an initial anomaly identification result containing anomaly events and their probabilities. Next, by querying the target knowledge graph, path association scores between the anomaly event and its associated entities are calculated to quantify the strength of potential causal relationships. Based on this, a weighting mechanism is introduced, assigning different levels of importance to the anomaly probability and path association score. By weighted calculation, the information from both is combined to obtain the final score for the target anomaly event—the target anomaly score. A higher score indicates a greater severity and wider impact of the anomaly event, requiring priority processing. Once the target anomaly score exceeds a preset threshold, the initially identified anomaly event is considered to have high credibility and importance. At this point, the knowledge graph is further searched for all entities that may have a causal relationship with the anomaly event, forming a candidate root cause entity set. This step is similar to clue gathering in an investigation, aiming to narrow the scope of the investigation and focus on entities that may be directly or indirectly related to the anomaly event. Next, each candidate root cause entity is analyzed in depth, combining its historical operational data, semantic features, and other relevant factors to calculate a comprehensive causal score reflecting the entity's influence on the anomaly event. This score calculation considers the characteristics of the entity itself and the direct or indirect correlation between the anomaly and the event, helping to accurately determine the cause and propagation mechanism of the anomaly. Among the candidate entities, the entity with the highest comprehensive causal score is selected as the most likely root cause entity. Subsequently, the causal propagation path between the anomaly and the root cause entity is traced through the edge relationships in the knowledge graph. Finally, by analyzing all entities on the root cause path and their interrelationships, especially the edge weights, the impact range of the root cause entity on the entire power system can be assessed. This metric provides a quantitative indicator of the potential impact of the anomaly, helping maintenance personnel quickly determine the necessary countermeasures, effectively isolate faults, and prevent the expansion of systemic risks. This embodiment's method, by integrating big data analysis, federated learning, knowledge graphs, and graph algorithms, forms a highly intelligent framework for power system anomaly identification and processing. Compared to methods in related technologies, it can accurately locate the root cause of anomalies and predict the impact range in a shorter time, thereby accelerating fault recovery and enhancing the stability and security of the power system.
[0048] Optionally, the model output and feature input generated by the federated modeling module can be merged first. Anomaly scoring is performed on each observation using the current operational characteristics data of the power system in the target area and semantic information from the knowledge graph. Let the global model parameters be... (From aggregation), model for time... eigenvectors The anomaly confidence level (i.e., anomaly probability) is determined by the function Given, represented as: In the formula, This represents the probability (or confidence level) that the model classifies the sample as an anomaly. Meanwhile, based on knowledge graphs... Calculate the entity to which the sample belongs. Path association strength with neighboring entities This is used to quantify the anomaly propagation potential of the sample in the semantic network. The final anomaly score is obtained by linearly fusing the model confidence score with the graph path score. In the formula, To integrate weights, Representing entities The neighborhood group, The calculation method can refer to the aforementioned path product or weighted sum. When Exceeding the threshold Time (i.e.) This observation was marked as an anomaly to be traced; parameters It can be optimized based on cross-domain evaluation results.
[0049] Based on the structured relationships and causal clues of knowledge graphs, semantic localization and ranking of labeled abnormal samples are performed. First, a set of candidate root cause entities is determined. For example, devices / events connected to the abnormal entity within several hops or related to the abnormal type; for candidate entities Calculate the comprehensive causal score A weighted combination of model confidence and path / edge weights: In the formula, The entity to which the exception belongs. This is the aggregated anomaly confidence score based on the relevant observations of this entity (such as the mean of the output of the neighbor node model). For semantic association functions (e.g., maximum path score) (or the product of the edge weights along the path). These are weighting coefficients. To quantify the scope of influence, a measure of the scope of influence can be defined: In the formula, For entities in a knowledge graph With entity edge weights, thresholds Used to ignore low-confidence associations. Finally, by... and The entities are sorted, and the highest-scoring entities are selected as the most likely root causes. Root cause paths (i.e., the most credible paths from the root cause entity to the anomalous entity) are then generated as the basis for semantic interpretation.
[0050] The detected anomalies, ranked root cause candidates, impact range, and confidence level are output as structured records, with appended time windows, relevant measurement waveforms, or statistical summaries to facilitate rapid assessment by operations and maintenance personnel. Each diagnostic record can be represented as a triple: In the formula, The time when the anomaly occurred. The entity corresponding to the exception is listed in the collection. One root cause candidate and the corresponding causal scores and influence Finally, the report will be output to the operation and maintenance work order system or display interface in a preset format, and supports exporting comma-separated values / JavaScript object representation (CSV / JSON) for subsequent auditing and model retraining.
[0051] In summary, the anomaly diagnosis and source tracing module outputs the probability of the federated model. Semantic association with knowledge graph Combined, according to a unified scoring mechanism Identify anomalies and then base them on a comprehensive causal score. With influence The root causes are sorted and semantically interpreted, and the results are delivered to operations and maintenance in the form of a structured report, forming a traceable closed-loop governance process that is consistent with the naming of parameters in the front-end data processing and federated modeling modules.
[0052] Through steps S102 to S108, the goal is to achieve intelligent and real-time identification of abnormal events by integrating the power grid anomaly identification model obtained through federated learning with the path association analysis method in the dynamic knowledge graph. By combining the semantic features and historical behavioral features of entities, the root cause and scope of influence of abnormal events can be accurately located, thereby improving the technical effect of identifying abnormal events and analyzing the causes of abnormal events in the power system. This solves the technical problem of low accuracy in identifying abnormal events and analyzing the causes of abnormal events in the power system in related technologies.
[0053] Based on the above embodiments and optional embodiments, this invention proposes an optional implementation method for power system anomaly identification. It achieves collaborative modeling and privacy protection of multi-source distributed data through federated learning, and utilizes dynamic knowledge graphs to realize real-time semantic layer evolution and causal reasoning, thereby constructing an intelligent closed-loop mechanism of "perception-diagnosis-source tracing." This improves the intelligence and interpretability of power data quality governance, and enables continuous optimization of power data and autonomous evolution of the system. The method uses "perception-diagnosis-source tracing" as its core closed loop. Figure 2 This is a schematic diagram of an optional power system anomaly identification system according to an embodiment of the present invention. The method is applied to, for example... Figure 2 The power system anomaly identification system shown mainly includes four functional modules: data preprocessing module, federated modeling module, dynamic knowledge graph module, and anomaly diagnosis and tracing module. Among them:
[0054] The core objective of the data preprocessing module is to standardize multi-source heterogeneous data from power grid operation, dispatching, monitoring, and equipment sources, providing high-quality input for subsequent federated modeling. This module mainly comprises three key steps: data acquisition, data cleaning, and feature extraction. Specifically, it includes:
[0055] S11, Data Acquisition: This involves the unified acquisition of heterogeneous data from multiple sources, including power grid operation, dispatch monitoring, and equipment, to form a raw dataset. Let the acquired data sequence be denoted as... ,in, Indicates time A data record acquired at any given time. This process ensures that the total number of samples collected is complete and continuous, providing a solid foundation for subsequent data quality governance.
[0056] S12, Data Cleaning: The data cleaning stage mainly involves outlier removal, missing value imputation, and time alignment. To identify noisy data that significantly deviates from the normal range, threshold or statistical interval determination methods are used. The cleaning rules for numerical fields are the same as in the previous embodiments, and will not be repeated here.
[0057] S13, Feature Extraction: Construct a structured feature set from the cleaned data for subsequent model training and quality evaluation. The specific method for obtaining the structured feature set is the same as in the previous embodiments and will not be repeated here.
[0058] The federated modeling module aims to utilize the local data (including a primary set of historical operational feature data and corresponding historical anomaly identification results) held by each node (i.e., the local server in each region) for distributed training, achieving cross-node knowledge sharing and data quality assessment, while simultaneously performing cross-domain anomaly detection without exposing the original data. This module mainly includes three key steps: local training, model aggregation, and cross-domain evaluation. Specifically, it includes:
[0059] S21, Local Training: In the local training phase, each node (i.e., the local server) independently trains the model based on the preprocessed dataset. Let the... The dataset of each node contains The local model parameters of each sample are denoted as . By minimizing the training error within a node, the node obtains updated local model parameters. The specific formula is the same as in the previous embodiment and will not be repeated here. This step ensures that the model can adapt to the node's own data characteristics while avoiding the leakage of original data.
[0060] S22, Model Aggregation: In the model aggregation phase, each node only uploads the trained model parameters. The data is then fed to the central server, where it is weighted and merged based on the amount of data from each node to obtain unified global model parameters. The aggregation process is the same as in the previous embodiments and will not be repeated here. The weighted aggregation method ensures that nodes with larger amounts of data contribute higher weights in the fusion, improving the robustness of the global model. The aggregated model (i.e., the global model parameters) is then synchronously returned to all nodes as the initial parameters for the next round of local training, realizing cross-node knowledge sharing.
[0061] S23, Cross-domain evaluation: In the cross-domain evaluation phase, each node uses its own test dataset (including the corresponding second number of historical operational feature data and the corresponding historical anomaly identification results) to verify the global model performance in order to test the model's adaptability in different regions and different business scenarios. The specific verification method is the same as in the aforementioned embodiments, and will not be repeated here.
[0062] Cross-domain performance evaluation can promptly identify whether the model has biases for certain regions, further guiding model iteration and optimization.
[0063] The dynamic knowledge graph module focuses on equipment, events, measurement points, and their relationships within the power system. It constructs a dynamically evolving semantic network and updates the knowledge in real time by incorporating time series data and business context. This module primarily includes three key steps: entity construction, relationship building, and dynamic updating. Specifically, it includes:
[0064] S31, Constructing Entities: In the entity construction phase, the physical objects and abstract concepts in the power system are first mapped to entity sets. , represented as Each entity It can represent substations, switches, busbars, measurement points, alarm events, or business entities. To facilitate subsequent calculations and reasoning, a low-dimensional entity vector representation is assigned to each entity. ,in For the embedding dimension, vector Used to capture the semantic and historical behavioral features of entities. The entity construction process also retains a set of entity metadata. For example, device ID, the site to which it belongs. Measurement type Timestamp range, etc., are denoted as This metadata is used to provide contextual constraints during subsequent associations and updates.
[0065] S32, Association: In the association phase, define the set of relationships between entities. Each relation is represented as a triple. That is, the head entity Relationship types With tail entity To measure the reliability or association strength of triples, a scoring function is introduced. (The negative distance form is used so that a larger value is more reliable). The specific formula for this scoring function is the same as in the previous embodiment, and will not be repeated here. To facilitate unified management of the relationship structure in the system, the pairwise relationships between all entities can be organized into an adjacency matrix. In the matrix Representing entities With entity There exists a certain type of relationship (such as "belonging", "connection", "triggering", etc.), among which This represents the total number of entities in the knowledge graph. Each relation edge is assigned a weight. This is used to represent the strength of the association or the confidence level of the relationship between entity pairs. Weight It can be derived from the scoring function. Historical statistics or model inference.
[0066] S33, Dynamic Update: In the dynamic update phase, the knowledge graph evolves incrementally with time and model output, recording time points. The knowledge graph, that is, the current knowledge graph is When the federated model (i.e., the power grid anomaly detection model) detects a new anomaly or when new measurement / equipment information arrives, the system generates an incremental triplet set. With possible new entity sets The update is completed through a simple union operation. To maintain the timeliness of entity vectors and relation vectors, an exponentially weighted update rule can be applied to the affected entities. The specific update form is the same as in the previous embodiment and will not be repeated here. For newly added relations, their initial weights... Can be determined by a scoring function The confidence value is obtained through normalization or directly assigned by the detection module. .
[0067] Based on the constructed and real-time updated semantic network described above, the dynamic knowledge graph module supports graph reasoning and path scoring to aid in understanding anomalies. For example, for suspected root cause entities... With the affected entity It can calculate the path association strength score between two entities. The score is the product or weighted sum of the edge weights along the path; a higher score indicates a stronger semantic association between the two entities. This score, combined with the anomaly confidence score of the federated model, can provide interpretable semantic evidence for anomaly root cause analysis.
[0068] The anomaly diagnosis and tracing module takes the output of the federated model and the dynamic knowledge graph as input. It is responsible for confirming detected anomalies, performing semantic root cause analysis, and generating diagnostic reports that can be used for operational decision-making. This module mainly includes three key steps: anomaly detection, root cause analysis, and report generation.
[0069] S41, Anomaly Detection: First, the model output generated by the federated modeling module is combined with the feature input. Anomaly scoring is performed on each observation using the current operational characteristics data of the power system in the target area and semantic information from the knowledge graph. Let the global model parameters be... (From aggregation), model for time... eigenvectors The confidence level (i.e., the probability of anomaly) of the anomaly. Meanwhile, based on knowledge graphs Calculate the entity to which the sample belongs. Path association strength with neighboring entities This method quantifies the anomaly propagation potential of the sample in the semantic network. The final anomaly score is obtained by linearly fusing the model confidence score with the graph path score.
[0070] when Exceeding the threshold Time (i.e.) This observation was marked as an anomaly to be traced; parameters It can be optimized based on cross-domain evaluation results.
[0071] S42, Root Cause Analysis: Semantic localization and ranking of labeled anomalous samples based on structured relationships and causal clues from a knowledge graph. First, a set of candidate root cause entities is determined. For example, devices / events connected to the abnormal entity within several hops or related to the abnormal type; for candidate entities Calculate the comprehensive causal score This is a weighted combination of model confidence and path / edge weights. To quantify the scope of influence, an influence scope metric can be defined. .in, and The specific method for obtaining it is the same as in the aforementioned embodiments, and will not be repeated here. Finally, according to... and The entities are sorted, and the highest-scoring entities are selected as the most likely root causes. Root cause paths (i.e., the most credible paths from the root cause entity to the anomalous entity) are then generated as the basis for semantic interpretation.
[0072] S43, Report Generation: The detected anomalies, ranked root cause candidates, impact range, and confidence level are output as structured records, with appended time windows, relevant measurement waveforms, or statistical summaries for rapid operational assessment. Each diagnostic record can be represented as a triple: Finally, the report will be output to the operation and maintenance work order system or display interface in a preset format, and supports exporting to CSV / JSON for subsequent auditing and model retraining.
[0073] In summary, the anomaly diagnosis and source tracing module outputs the probability of the federated model. Semantic association with knowledge graph Combined, according to a unified scoring mechanism Identify anomalies and then base them on a comprehensive causal score. With influence The root causes are sorted and semantically interpreted, and the results are delivered to operations and maintenance in the form of a structured report, forming a traceable closed-loop governance process that is consistent with the naming of parameters in the front-end data processing and federated modeling modules.
[0074] It should be noted that related technologies for power data quality governance suffer from systemic problems such as isolated detection, static knowledge, and limited collaboration, making it difficult to achieve continuous high-quality data governance. This embodiment proposes a power data quality governance and anomaly identification method based on federated learning and dynamic knowledge graphs, aiming to achieve multi-domain collaboration and knowledge-driven intelligent governance while ensuring data privacy. This method integrates the distributed modeling capabilities of federated learning with the semantic evolution reasoning capabilities of dynamic knowledge graphs, constructing a closed-loop mechanism of "perception-diagnosis-source tracing," which enables intelligent discovery, cause diagnosis, and source tracing of abnormal data, thereby improving the intelligence, real-time performance, and autonomy of power data quality governance and power system anomaly diagnosis. The main features of this embodiment are: 1) A data quality governance mechanism based on federated learning: Cross-domain data quality modeling is achieved through local training and parameter aggregation, completing anomaly detection and quality assessment without sharing original data, solving the problem of power data not being centrally aggregated across regions and units. 2) Dynamic Knowledge Graph Construction and Update Method for Power Systems: Based on power entities such as equipment, measurement points, and events, and their relationships, a knowledge graph that continuously evolves with time, state, and business changes is constructed and used for anomaly correlation analysis and semantic interpretation. 3) Anomaly Diagnosis and Source Tracing Mechanism Integrating Federated Model Output and Knowledge Graph Inference: The anomaly results output by federated learning are combined with semantic paths and causal relationships in the knowledge graph to achieve interpretable anomaly localization, root cause identification, and correlation tracing. 4) End-to-End Process Framework for Cross-Domain Data Quality Assessment: Covering the complete governance chain of data preprocessing, federated modeling, knowledge graph updating, anomaly diagnosis, and report generation, achieving an intelligent and traceable closed loop for data quality governance.
[0075] As an optional embodiment, this embodiment uses a regional power grid control center as the application scenario, and substation operation data, fault alarm data, and dispatch monitoring data as inputs to illustrate a typical engineering implementation process of the method in this embodiment. The node system (i.e., local server) of this embodiment is deployed in three sub-regional control centers (A, B, and C) under the regional power grid. Each region only locally stores its own operation measurement data, equipment status data, and alarm records, and collaboratively models using federated learning. The central server is deployed in the regional central control center to perform model aggregation, knowledge graph merging, and anomaly diagnosis and display. Taking region A (containing approximately...) as an example... The processing flow is illustrated using a sample of measurement data.
[0076] The data preprocessing module performs the following operations:
[0077] Data Acquisition: The system acquires heterogeneous data from multiple sources, including voltage, current, load, frequency, switch status, and alarm events, from Supervisory Control and Data Acquisition (SCADA), Phasor Measurement Unit (PMU), and intelligent terminals. The raw data forms a sequence. The data is aligned to a 100ms interval using a unified timestamp.
[0078] Data cleaning: Set compliance ranges for each type of measurement value, such as the 10kV line voltage range [9.5kV, 10.5kV]. Example: A voltage measurement d_t=13.2kV is out of range → judged as an outlier and set to null. Missing and outlier values are filled using linear interpolation. Finally, the cleaned sequence samples were obtained. .
[0079] Feature extraction: Constructing feature vectors from measurements such as voltage and current. The method uses a 5-point sliding window with mean and difference features, and performs normalization: Region A ultimately generates a feature sequence: .
[0080] The federated modeling module performs the following operations: each of the three regions conducts local training, which is then aggregated by the provincial survey center. Specifically:
[0081] Local training: Each node utilizes Train a homogeneous deep model (e.g., a 3-layer LSTM with an attention layer). Update the local parameters for region A: Each of the three nodes underwent five rounds of local training.
[0082] Model aggregation: Central coordination server weights and merges data based on volume. The aggregated global model returns three regions, which serve as the initial parameters for the next training round. After 20 rounds of training, the global model achieves an average accuracy of 98.3% on the test set across all nodes.
[0083] Cross-domain evaluation: The accuracy of node anomaly detection was measured separately for Region A: =98.1%, Region B: =97.8%, Region C: =98.9%. The evaluation results were used to adjust the model learning rate and the number of local training epochs.
[0084] The dynamic knowledge graph module performs the following operations:
[0085] A cross-regional semantic knowledge graph (KG) is constructed using a unified provincial dispatch center. Specifically, this includes: Entity construction: The system loads 12 types of entities from the main data platform, including: 110kV / 220kV substation entities, line switches, busbars, protection devices, measurement points, over-limit alarms, and tripping events, forming an entity set E with a size of m≈8,000. An embedding vector is generated for each entity. .
[0086] Relation Construction: Constructing relation triples: ( Training relation vectors based on the TransE model. Calculate the triplet score: This forms a weighted adjacency matrix. (8k×8k).
[0087] Dynamic update example: When a line experiences an over-limit alarm "10kV-Nanhu Line-Current Over-Limit", the system automatically creates a new event entity. New triplet (circuit, trigger, And update the relevant line entity vectors using EWMA: After updating, immediately participate in subsequent root cause reasoning.
[0088] The anomaly diagnosis and tracing module executes the following process, using a "sudden increase in line current" event in region B as an example:
[0089] Anomaly detection: Global model for samples Output anomaly probability: Knowledge graph path association score: Fusion Score: The system determined this observation to be an "anomaly to be traced".
[0090] Root cause analysis: Set of candidate root cause entities Includes 3 related devices: Main transformer protection device Line #3 switch Line #3 terminal measurement point. System calculation of comprehensive causal score: ,in =0.7. Example score: , ; , ; , The system determined the root cause path to be... The line #3 switch malfunctioned, and the most reliable path was generated: line #3 switch → tripping event → current surge measurement point → target anomaly.
[0091] Report generation: The system automatically generates diagnostic triples: ReportItem=( , =“Sudden increase in current in line #3”, {( , 0.91, 1.27), ( , 0.76, 0.82), ( The report was pushed to the maintenance work order system. After verification, the maintenance personnel found that the #3 switch on the line was momentarily unable to operate due to mechanical wear.
[0092] This embodiment also provides a power system anomaly identification device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0093] According to an embodiment of the present invention, an apparatus embodiment for implementing the above-described power system anomaly identification method is also provided. Figure 3 This is a schematic diagram of the structure of a power system anomaly identification device according to an embodiment of the present invention, as shown below. Figure 3 As shown, the aforementioned power system anomaly identification device includes: a feature data acquisition module 300, an initial anomaly identification module 302, a path association score acquisition module 304, and a target anomaly identification module 306, wherein:
[0094] The feature data acquisition module 300 is used to acquire the current operating feature data of the power system in the target area.
[0095] The initial anomaly identification module 302 is connected to the feature data acquisition module 300. It is used to obtain the initial anomaly identification result of the power system in the target area based on the current operating feature data and the power grid anomaly identification model. The power grid anomaly identification model is obtained through federated learning based on the historical operating feature data of multiple areas and the corresponding historical anomaly identification results. The initial anomaly identification result includes the target anomaly event and the anomaly probability of the target anomaly event.
[0096] The path association score acquisition module 304 is connected to the initial anomaly identification module 302. It is used to query the target knowledge graph based on the initial anomaly identification result and obtain the path association score of the target anomaly event. The target knowledge graph includes entities in the power system of multiple regions, entity vectors corresponding to the entities, and relationships between entities. The path association score is used to quantify the strength of potential causal association between the target anomaly event and other entities within the power system. The entity vector is used to indicate the semantic features and historical behavioral features of the entity.
[0097] The target anomaly identification module 306 is connected to the path association score acquisition module 304. It is used to determine the target anomaly identification result of the power system in the target area based on the initial anomaly identification result and the path association score. The target anomaly identification result includes the root cause entity of the target anomaly event, as well as the comprehensive causal score and the impact range measure of the root cause entity.
[0098] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0099] It should be noted that the feature data acquisition module 300, the initial anomaly identification module 302, the path association score acquisition module 304, and the target anomaly identification module 306 mentioned above correspond to steps S102 to S108 in the embodiments. The instances and application scenarios implemented by the above modules and their corresponding steps are the same, but they are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run on a computer terminal.
[0100] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.
[0101] The aforementioned power system anomaly identification device may also include a processor and a memory. The aforementioned feature data acquisition module 300, initial anomaly identification module 302, path association score acquisition module 304, target anomaly identification module 306, etc., are all stored in the memory as program modules, and the processor executes the aforementioned program modules stored in the memory to realize the corresponding functions.
[0102] The processor contains a core that retrieves the corresponding program modules from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.
[0103] According to an embodiment of this application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein, when the program runs, it controls the device where the non-volatile storage medium is located to execute any of the aforementioned power system anomaly identification methods.
[0104] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals, and the non-volatile storage medium includes stored programs.
[0105] Optionally, a program that controls the device containing the non-volatile storage medium to execute any of the above-mentioned power system anomaly identification method steps during program execution.
[0106] According to an embodiment of this application, an embodiment of a processor is also provided. Optionally, in this embodiment, the processor is used to run a program, wherein the program executes any of the above-described power system anomaly identification methods.
[0107] According to an embodiment of this application, an embodiment of a computer program product is also provided, which, when executed on a data processing device, is adapted to execute a program that initializes the power system anomaly identification method steps described above.
[0108] This invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the above-described power system anomaly identification methods.
[0109] The order of the above embodiments of the present invention is merely for description and does not represent the superiority or inferiority of the embodiments.
[0110] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0111] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of modules described above can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between modules, and may be electrical or other forms.
[0112] The modules described above as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0113] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0114] If the aforementioned integrated modules are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0115] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for identifying power system anomalies, characterized in that, include: Obtain current operational characteristic data of the power system in the target area; Based on the current operating characteristic data, a power grid anomaly identification model is used to obtain the initial anomaly identification result of the power system in the target area. The power grid anomaly identification model is obtained through federated learning based on historical operating characteristic data of multiple areas and the corresponding historical anomaly identification results. The initial anomaly identification result includes the target anomaly event and the anomaly probability of the target anomaly event. Based on the initial anomaly identification results, the target knowledge graph is queried to obtain the path association score of the target anomaly event. The target knowledge graph includes entities in the power system of multiple regions, entity vectors corresponding to the entities, and relationships between entities. The path association score is used to quantify the strength of potential causal association between the target anomaly event and other entities within the power system. The entity vectors are used to indicate the semantic features and historical behavioral features of the entities. Based on the initial anomaly identification results and the path association scores, the target anomaly identification results of the power system in the target area are determined. The target anomaly identification results include the root cause entity of the target anomaly event, as well as the comprehensive causality score and impact range measure of the root cause entity.
2. The method according to claim 1, characterized in that, The step of determining the target anomaly identification result of the power system in the target area based on the initial anomaly identification result and the path association score includes: Based on the anomaly probability of the target anomaly event in the initial anomaly identification result and the path association score, the target anomaly score of the target anomaly event is obtained. If the target anomaly score is greater than a preset score threshold, a set of candidate root cause entities corresponding to the target anomaly event is determined based on the target knowledge graph. Determine the comprehensive causal score corresponding to each of the multiple candidate entities included in the candidate root cause entity set; Based on the comprehensive causal scores corresponding to each of the multiple candidate entities, the root cause entity and the root cause path corresponding to the target abnormal event are determined from the multiple candidate entities, wherein the root cause path is the connection path between the target abnormal event and the root cause entity. Based on the weights of the edges between entities included in the root cause path, a measure of the influence range of the root cause entity is determined, wherein the weights of the edges between the entities are used to quantify the association strength between the entities.
3. The method according to claim 2, characterized in that, The step of obtaining the target anomaly score of the target anomaly event based on the anomaly probability of the target anomaly event in the initial anomaly identification result and the path association score includes: Determine the first weight corresponding to the anomaly probability of the target anomaly event, and the second weight corresponding to the path association score; Based on the anomaly probability of the target anomaly event, the path association score, the first weight, and the second weight are weighted and calculated to obtain the target anomaly score.
4. The method according to claim 2, characterized in that, Determining the comprehensive causal score corresponding to each of the multiple candidate entities included in the candidate root cause entity set includes: The comprehensive causal score of any candidate entity among the multiple candidate entities is obtained in the following manner: Determine the model anomaly confidence aggregate score for any candidate entity, wherein the model anomaly confidence aggregate score is obtained based on the historical anomaly probability corresponding to any candidate entity; Based on the target knowledge graph, determine the semantic association strength between any candidate entity and the target anomalous event; Determine the weight values corresponding to the model anomaly confidence aggregate score and the semantic association strength, respectively. Based on the aggregated confidence score of the model anomaly, the semantic association strength, and the weight values corresponding to the aggregated confidence score of the model anomaly and the semantic association strength, a weighted operation is performed to obtain the comprehensive causal score of any candidate entity. The comprehensive causal score corresponding to each of the multiple candidate entities is obtained by using the method of obtaining the comprehensive causal score of any candidate entity.
5. The method according to claim 1, characterized in that, Before obtaining the initial anomaly identification result of the power system in the target area based on the current operating characteristic data and using the power grid anomaly identification model, the method further includes: Based on the training dataset, the current local model parameters of the target area are trained to obtain the trained local model parameters, and the trained local model parameters are sent to the central processing unit for the central processing unit to update the global model parameters. The training dataset includes a first number of historical operating feature data of the power system in the target area, and the corresponding historical anomaly identification results. Receive the updated global model parameters returned by the central processing unit; Based on the updated global model parameters, the trained local model parameters are updated to obtain the updated local model parameters; The updated local model parameters are used as the new current local model parameters, and the above operation is repeated until the preset termination condition is met. The central processing unit is further configured to obtain the power grid anomaly identification model based on the updated global model parameters obtained when the preset termination condition is met, and the central processing unit is further configured to send the power grid anomaly identification model to the local servers corresponding to the multiple regions.
6. The method according to any one of claims 1 to 5, characterized in that, Before querying the target knowledge graph based on the initial anomaly identification result to obtain the path association score of the target anomaly event, the method further includes: If the target abnormal event is detected as a newly added abnormal event, an incremental triple set, a new entity set, and a newly added entity vector are determined based on the initial abnormality identification result. The incremental triple set includes newly added entities and entity relationships related to the target abnormal event. The incremental triple set is joined with the triple set in the current knowledge graph, the new entity set is joined with the entity set in the current knowledge graph, and the current entity vector in the current knowledge graph is updated to obtain the target knowledge graph.
7. The method according to claim 6, characterized in that, The step of updating the current entity vector in the current knowledge graph includes: Determine the weight values corresponding to the current entity vector and the newly added entity vector included in the current knowledge graph; Based on the current entity vector, the newly added entity vector, and the weight values corresponding to the current entity vector and the newly added entity vector, a weighted operation is performed to update the current entity vector in the current knowledge graph, thereby obtaining the entity vector in the target knowledge graph.
8. A power system anomaly identification device, characterized in that, include: The feature data acquisition module is used to acquire the current operating feature data of the power system in the target area. The initial anomaly identification module is used to obtain the initial anomaly identification result of the power system in the target area based on the current operating characteristic data and the power grid anomaly identification model. The power grid anomaly identification model is obtained through federated learning based on the historical operating characteristic data of multiple areas and the corresponding historical anomaly identification results. The initial anomaly identification result includes the target anomaly event and the anomaly probability of the target anomaly event. The path association score acquisition module is used to query the target knowledge graph based on the initial anomaly identification result to obtain the path association score of the target anomaly event. The target knowledge graph includes entities in the power system of multiple regions, entity vectors corresponding to the entities, and relationships between entities. The path association score is used to quantify the potential causal association strength between the target anomaly event and other entities within the power system. The entity vectors are used to indicate the semantic features and historical behavioral features of the entities. The target anomaly identification module is used to determine the target anomaly identification result of the power system in the target area based on the initial anomaly identification result and the path association score, wherein the target anomaly identification result includes the root cause entity of the target anomaly event, as well as the comprehensive causal score and the impact range measure of the root cause entity.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, which are adapted to be loaded by a processor and executed by the power system anomaly identification method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the power system anomaly identification method according to any one of claims 1 to 7.