A method and system for locating fault root cause in a cross-regional power system

CN122818285APending Publication Date: 2026-09-25GUANGZHOU ELECTRIC POWER COMM NETWORK LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611256390.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,在跨区域异常场景中,异常信息分散在多个区域业务运维平台,各区域之间缺少有效的信息共享和协同分析机制,导致异常传播路径难以追溯,根因定位耗时较长

Benefits of technology

在本申请的实施例中,提供了一种跨区域电力系统故障根因定位方法及系统,通过预构建异常关联规则库,结合区域传播路径的自动匹配与反向追溯机制,可以实现对跨区域故障根因的精准定位,具有能够快速准确识别跨区域电力业务信息系统故障的起始区域和目标诱发节点,有效减少因人工排查造成的低效性和错误率高的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818285A_ABST
    Figure CN122818285A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power system fault diagnosis, and specifically provides a cross-regional power system fault root cause positioning method and system, comprising the following steps: real-time receiving of abnormal alarm information, extraction of abnormal feature data, and matching to obtain a regional propagation path; reverse tracing along the regional propagation path to identify a starting region and a target induced node; obtaining parameter variation characteristics of the target induced node, environmental influence characteristics of the starting region, and system configuration characteristics, and determining a fault root cause of the target induced node through a root cause determination model. The present application has the effect of quickly and accurately identifying the starting region and the target induced node of the cross-regional power business information system fault, and effectively reducing the inefficiency and high error rate caused by manual investigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system fault diagnosis technology, and in particular to a method and system for locating the root cause of faults in cross-regional power systems. Background Technology

[0002] As the scale of power business information systems continues to expand and the level of interconnectivity continues to improve, the coupling relationships between cross-power business information systems are becoming increasingly complex. When an anomaly occurs in one power business information system, the impact of the anomaly often spreads rapidly to other regions through data interfaces, synchronization component associations, and business scheduling associations, forming a chain reaction and potentially even triggering large-scale anomalies. Therefore, accurately locating the root cause of anomalies in cross-regional power business information systems is of great significance for ensuring the safe and stable operation of power grid information systems.

[0003] Existing anomaly diagnosis methods primarily focus on the internal structure of a single power business information system, relying on platform-specific alarm information and operational log data for analysis. However, in cross-regional anomaly scenarios, anomaly information is scattered across multiple regional business operation and maintenance platforms. The lack of effective information sharing and collaborative analysis mechanisms between regions makes it difficult to trace the anomaly propagation path and results in lengthy root cause localization. Furthermore, traditional methods typically employ expert experience combined with manual log analysis, meticulously checking alarm information one by one, which is inefficient and prone to missing crucial clues. Moreover, existing technologies easily overlook the inter-regional correlation characteristics and anomaly propagation patterns, making it difficult to accurately identify the initial triggering business node and the true root cause of the fault. This is especially problematic when anomalies occur simultaneously in multiple regions, as the propagation result of the anomaly can easily be misjudged as the source, thus affecting the formulation of subsequent anomaly handling and prevention measures. Summary of the Invention

[0004] In view of the aforementioned problems, this application is proposed to provide a method and system for locating the root cause of faults in a cross-regional power system, which overcomes or at least partially solves the aforementioned problems, comprising: The above-mentioned objective of this application is achieved through the following technical solution: A method for locating the root cause of faults in a cross-regional power system, involving multiple power business information systems, includes the following steps: [The method is described in the original text, but the provided excerpt ends here.] It receives abnormal alarm information from multiple power business information systems in real time, extracts abnormal feature data, and matches the abnormal feature data with an abnormal association rule base to obtain the regional propagation path; Tracing back along the regional propagation path, we can identify the starting area of ​​the regional propagation path and the earliest target triggering node in the region where anomalies occur. The region where the starting point is located is determined as the starting region, and the parameter change characteristics of the target inducing node at the time of the anomaly are obtained; The environmental impact characteristics and system configuration characteristics of the starting area are obtained, and combined with the parameter change characteristics, the root cause of the target inducing node is determined by a preset root cause determination model.

[0005] The second objective of this invention is achieved through the following technical solution: A cross-regional power system fault root cause localization system, comprising: The path acquisition module is used to receive abnormal alarm information from multiple power business information systems in real time, extract abnormal feature data, and match the abnormal feature data with an abnormal association rule base to obtain the regional propagation path. The node determination module is used to trace back along the regional propagation path and identify the starting area of ​​the regional propagation path and the target initiating node that caused the earliest anomaly in the area. The feature acquisition module is used to determine the region where the starting point is located as the starting region and to acquire the parameter change features of the target inducing node at the time of the anomaly occurrence. The root cause determination module is used to acquire the environmental impact characteristics and system configuration characteristics of the starting area, and combine them with parameter change characteristics to determine the root cause of the target inducing node's failure through a preset root cause determination model.

[0006] This application has the following advantages: In the embodiments of this application, a method and system for locating the root cause of cross-regional power system faults are provided. By pre-building an anomaly association rule base and combining the automatic matching and reverse tracing mechanism of regional propagation paths, the root cause of cross-regional faults can be accurately located. It has the ability to quickly and accurately identify the starting area and target inducing node of faults in cross-regional power business information systems, effectively reducing the inefficiency and high error rate caused by manual investigation. Attached Figure Description

[0007] Figure 1 A flowchart of a cross-regional power system fault root cause localization method provided in an embodiment of the present invention; Figure 2 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0008] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0009] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0010] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows: Power business information systems refer to various information systems used to support business operations and management in all aspects of power production, dispatching, trading, and marketing. These systems may include energy management systems, distribution management systems, dispatch automation systems, market trading systems, customer service systems, etc. There are usually complex data interactions and business linkages between these systems, which work together to ensure the safe and stable operation of the power system.

[0011] The anomaly association rule base is a pre-established knowledge set that stores the patterns, rules, and relationships of anomaly propagation between power business information systems. It describes how specific types of anomalies propagate between different systems or regions, and the anomalies they may trigger. Through this rule base, the system can match real-time received anomaly alarm information to infer the propagation path of the anomaly.

[0012] Abnormal alarm information refers to the alarm information generated and reported by the power business information system when it detects that the preset threshold is exceeded, the preset rules are violated, or unexpected behavior occurs during the operation of the system. It usually includes alarm time, alarm source, alarm type, alarm level, and related business identifiers.

[0013] Anomaly characteristic data is key information extracted from real-time received anomaly alarm information, used to describe the essential attributes of anomaly events. Typically, anomaly characteristic data includes business node identifiers, anomaly occurrence time, anomaly type, anomaly value, affected systems or components, etc., and is the basis for anomaly propagation path matching and root cause analysis.

[0014] Regional propagation path refers to the trajectory of an anomaly event transmitted between power business information systems in different geographical or logical regions. It describes the order and direction in which an anomaly spreads from one system in one region to another, and is a key clue for tracing the source of the anomaly.

[0015] The target triggering node refers to the business node in the anomaly propagation path that is the first to experience an anomaly and trigger subsequent chain reactions. Typically, the target triggering node is the direct cause or initial trigger point of the failure, and identifying it is the first step in determining the root cause of the failure.

[0016] Parameter change characteristics refer to the changes in key operating parameters of the target inducing node at the moment of an anomaly and for a period of time before and after it occurs, such as CPU utilization, memory utilization, network traffic, process status, service response time, etc., which can reflect the internal operating status and potential problems of the node.

[0017] Environmental impact characteristics refer to the external environmental conditions of the starting area at the time of the anomaly, which may include temperature, humidity, network bandwidth fluctuations, external electromagnetic interference, natural disasters, etc., which may affect the operation of the power business information system and are related to the occurrence of the fault.

[0018] System configuration characteristics refer to the configuration change records of relevant power business information systems within the starting area within a preset time window before the occurrence of an anomaly. These may include business service start / stop records, connection pool parameter adjustment records, permission configuration change records, software version upgrades, etc. Changes in system configuration are sometimes the direct cause of failures.

[0019] The root cause analysis model is a pre-defined analytical model used to comprehensively analyze the parameter variation characteristics of the target inducing node, the environmental impact characteristics of the starting region, and the system configuration characteristics, thereby determining the root cause of the target inducing node's failure. The root cause analysis model can be built based on machine learning algorithms, expert system rules, or causal reasoning mechanisms, aiming to improve the accuracy and automation of root cause analysis.

[0020] In one embodiment, such as Figure 1 As shown, this application discloses a method for locating the root cause of faults in a cross-regional power system, involving multiple power business information systems. The method has a pre-established anomaly association rule base and specifically includes the following steps: S10: Receive abnormal alarm information from multiple power business information systems in real time, extract abnormal feature data, and match the abnormal feature data through an abnormal association rule base to obtain the regional propagation path; In the embodiments of this application, the scheme pre-establishes an anomaly association rule base, which can be established in various ways. For example, known anomaly propagation patterns can be manually sorted and entered based on historical fault data and expert experience. When a component of system A experiences a specific type of anomaly, it may cause a service of system B to experience another type of anomaly; this association can be defined as a rule. Furthermore, rules in the rule base can be stored in the form of condition-action pairs, such as "If [system A, component X, anomaly type Y] occurs, it may lead to [system B, service Z, anomaly type W]", which is intuitive and easy to understand, enabling the rapid establishment of preliminary anomaly association knowledge. The system receives anomaly alarm information from multiple power business information systems in real time. These alarms may come from different regions and different business systems, and can be aggregated through a unified alarm collection platform or message queue service. After receiving the alarm information, it is necessary to extract anomaly feature data from these alarms. For example, the business node identifier (such as server IP, service name), the time of the anomaly (timestamp), and the anomaly type (such as CPU overrun, memory overflow, network connection interruption) can be extracted from the text content of the alarm message through keyword matching or regular expression parsing. The extracted anomaly feature data is then used to match against a pre-established anomaly association rule base. For example, the extracted anomaly feature data is compared with each rule in the rule base to determine if there is an anomaly propagation pattern that matches the rule definition. If a match is successful, one or more regional propagation paths can be obtained. For example, if the rule "System X anomaly in region A causes System Y anomaly in region B" is matched, the regional propagation path may be determined as "region A → region B".

[0021] Furthermore, when there are no abnormal association rules in the abnormal association rule base that match the time-series abnormal sequence with a degree exceeding a preset threshold, an emergency handling process including the following steps can be executed: (a) Based on the business nodes in the time-series anomaly sequence, extract all transmission chains containing these business nodes from the anomaly transmission knowledge graph to form a candidate path set; (b) Calculate the similarity between each path in the candidate path set and the temporal anomaly sequence in terms of node order and temporal relationship; (c) Select the path with the highest similarity as the temporary region propagation path; (d) Use this failure event and its corresponding temporary propagation path as new samples to update the anomaly propagation knowledge graph; The update of the anomaly propagation knowledge graph includes: adding weight attribute values ​​to existing directed propagation edges in the path; adding directed propagation edges and assigning initial weights to propagation relationships that do not exist in the path; and recording the temporal characteristics of this fault for updating the propagation delay attribute.

[0022] (e) When the frequency of occurrence of a newly added directed transmission edge reaches a preset threshold, a new abnormal association rule is generated based on the transmission pattern and added to the abnormal association rule library.

[0023] The preset threshold can be set between 0.6 and 0.8. When the matching degree of all rules is lower than the preset threshold, it is determined to be a new type of fault mode.

[0024] Through the above mechanism, the system can continuously learn new fault propagation patterns during operation, gradually improve the knowledge graph and rule base, thereby enhancing its ability to identify new cross-regional faults.

[0025] S20: Tracing back along the regional propagation path to identify the starting area of ​​the regional propagation path and the earliest target triggering node in the region where anomalies occur. In the embodiments of this application, after obtaining the regional propagation path, reverse tracing can be performed along that path. For example, if the regional propagation path is "Region A → Region B → Region C", the reverse tracing will start from Region C, proceed sequentially to Region B, and finally reach Region A. During the reverse tracing process, the system will identify the region where the starting point of the regional propagation path is located. For example, in the above path, Region A will be identified as the starting point region. Simultaneously, the system will also identify the target triggering node that first experienced an anomaly within the region. This can be achieved by analyzing the historical alarm records or operation logs of all business nodes within the starting point region to find the specific business node that first experienced an anomaly alarm within that region. For example, in Region A, a database service node might be identified as the first to experience an anomaly.

[0026] Furthermore, the steps for tracing back along the regional transmission path specifically include: The propagation direction of each entity node in the parsing region propagation path is determined by the direction attribute of the propagation edge in the anomaly association rule; Following the reverse direction of propagation, starting from the end node of the regional propagation path, each entity node is traversed step by step forward. During the reverse traversal, the timestamp information of each business node on the path is obtained within a preset range at the time of the exception. Compare the timestamp information of each business node and identify the business node with the earliest timestamp as the target triggering node; The region where the target-induced node is located is determined as the starting point region.

[0027] The timestamp information may include at least one of the following: the time of the fault alarm, the time of the parameter abnormality start, and the time of the protection action.

[0028] As a preferred implementation, when the timestamp of a certain service node is earlier than the timestamp of its upstream node in the regional propagation path, the service node can be used as the starting node of the propagation chain.

[0029] Furthermore, when comparing the timestamps of various business nodes, considering the impact of cross-regional clock synchronization errors, a tolerance threshold for timestamp comparison can be set, typically 200-500ms. When the difference between the timestamps of two business nodes is less than this tolerance threshold, they are considered to have occurred "quasi-simultaneously". The order is not determined by timing, but by another preset rule. This preset rule can be set by those skilled in the art based on actual scenarios and requirements.

[0030] S30: Determine the region where the starting point is located as the starting region, and obtain the parameter change characteristics of the target inducing node at the time of the anomaly occurrence; In the embodiments of this application, after identifying the originating region and the target triggering node, the originating region is determined as the starting region. Subsequently, the system acquires the parameter change characteristics of the target triggering node at the time of the anomaly. For example, for a database service identified as the target triggering node, the system queries historical data of its operating parameters such as CPU utilization, memory usage, disk I / O, and database connection count before and after the anomaly occurred, and analyzes the trends and abrupt changes of these parameters. These parameter change characteristics provide direct evidence for subsequent root cause analysis.

[0031] S40: Obtain the environmental impact characteristics and system configuration characteristics of the starting area, and combine them with the parameter change characteristics to determine the root cause of the target inducing node failure through a preset root cause determination model.

[0032] In the embodiments of this application, environmental impact characteristics may include temperature, humidity, network bandwidth fluctuations in the starting area at the time of the anomaly, and even information such as whether there is external construction or natural disaster. These environmental impact characteristics can be obtained through environmental sensors, network monitoring systems, or external data sources. System configuration characteristics may include configuration change records of relevant systems within the starting area within a preset time window before the anomaly occurred, such as service start / stop records, connection pool parameter adjustment records, and permission configuration change records, typically stored in a configuration management database or change management system. Finally, the obtained parameter change characteristics of the target inducing node, environmental impact characteristics of the starting area, and system configuration characteristics are input into a preset root cause determination model. The root cause determination model comprehensively analyzes the received multi-dimensional characteristics and determines the root cause of the target inducing node's failure through internal logical judgment or algorithm calculation. For example, if parameter change characteristics show a continuous increase in CPU utilization, system configuration characteristics show no recent changes, and environmental impact characteristics show network bandwidth fluctuations, the root cause determination model may infer that the service performance degradation is caused by an external network problem.

[0033] Furthermore, determining the root cause of the failure of the target inducing node based on the probability distribution can specifically include the following steps: From the probability distribution of each candidate root cause output by the root cause determination model, select the candidate root cause with the highest probability value, and determine whether the probability of the candidate root cause with the highest probability value exceeds the preset confidence threshold. If the confidence threshold is exceeded, the candidate root cause is determined as the root cause of the target inducing node's failure. If the preset confidence threshold is not exceeded, a multi-factor comprehensive judgment is performed; The multi-root cause comprehensive judgment includes: selecting the top N candidate root causes with the highest probability values, where N is a preset value, such as 2 or 3; and outputting multiple possible root causes and their corresponding probabilities for further manual confirmation by maintenance personnel.

[0034] Alternatively, the multi-root cause comprehensive judgment may also include: triggering a deep detection process for the target inducing node to obtain more diagnostic data.

[0035] The preset confidence threshold can be set between 0.7 and 0.85, which means that when the probability of a certain root cause exceeds 70%-85%, the root cause determination can be considered to have high credibility.

[0036] In another embodiment, when the probability difference between the two candidate root causes with the highest probability values ​​is less than a preset difference threshold, it can also be determined as an ambiguous situation that requires manual confirmation. In this case, multiple candidate root causes are output for comprehensive judgment.

[0037] For example, as a specific implementation, suppose a power company has power business information systems spanning three geographical regions, such as Region X, Region Y, and Region Z. Region X primarily operates a SCADA system, Region Y primarily operates an EMS system, and Region Z primarily operates a DMS system. On a certain workday, Region Z's DMS system suddenly reports a large number of alarms, subsequently triggering alarms in Region Y's EMS system, and ultimately affecting Region X's SCADA system as well.

[0038] First, the solution in this embodiment has a pre-established anomaly association rule base. This anomaly association rule base may contain rules based on historical fault data and expert experience, such as "an anomaly in the data acquisition module of the SCADA system may lead to anomalies in the data processing module of the EMS system", and "an anomaly in the scheduling instruction issuing module of the EMS system may lead to anomalies in the execution module of the DMS system". These rules are stored in a structured form, describing the potential propagation paths of anomalies between different systems and regions.

[0039] When an anomaly occurs, the system receives real-time alarm information from the DMS system in region Z, the EMS system in region Y, and the SCADA system in region X. For example, the DMS system reports "Scheduling instruction execution failure," the EMS system reports "Data processing delay," and the SCADA system reports "Data acquisition interruption." The system extracts anomaly characteristic data from these alarm messages, including the service node identifier of the alarm, the time of anomaly occurrence, and the anomaly type. For example, from the SCADA system alarm, the system extracts the "Data Acquisition Service" node, the time of anomaly occurrence T1, and the anomaly type "Data Acquisition Interruption." Subsequently, the system matches the extracted anomaly characteristic data using a pre-established anomaly association rule base. Through matching, the system can identify a regional propagation path, such as "Region X → Region Y → Region Z," indicating that the anomaly may have started in region X and propagated sequentially to region Y and region Z.

[0040] Based on this, the system traces backwards along the regional propagation path "Region X → Region Y → Region Z". The starting point of the backward tracing is Region Z, then Region Y, and finally Region X. During this process, Region X is identified as the starting region. Within Region X, the system further analyzes the alarms and operation logs of all business nodes in the SCADA system to identify the target triggering node where the earliest anomaly occurred. For example, the analysis reveals that the node named "Data Acquisition Service" in Region X was the first to generate an anomaly alarm at time T1; therefore, the "Data Acquisition Service" node is identified as the target triggering node.

[0041] Based on this, region X was determined as the starting region. The system obtains the parameter change characteristics of the "Data Acquisition Service" node at the time T1 when the anomaly occurred. For example, querying the historical data of this service before and after time T1, such as CPU utilization, memory usage, and network I / O, reveals that its CPU utilization continuously increased before time T1 and reached its peak at time T1, accompanied by a large number of error log outputs.

[0042] Simultaneously, the system acquires the environmental impact characteristics and system configuration characteristics of the starting area X. The environmental impact characteristics may include the weather conditions of area X at time T1 (no extreme weather), network bandwidth fluctuations (stable network), and whether there is external construction disturbance (no). The system configuration characteristics may include the configuration change records of the SCADA system in area X within a preset time window before time T1. For example, a query may find that the "Data Acquisition Service" has not undergone any recent version upgrades, parameter adjustments, or start / stop operations.

[0043] Finally, the parameter change characteristics of the "Data Acquisition Service" (sudden increase in CPU utilization, error logs), environmental impact characteristics (stable environment), and system configuration characteristics (no recent changes) are input into a preset root cause analysis model. The model comprehensively analyzes this input information. For example, if there is a sudden increase in CPU usage without any external environmental or configuration changes, the root cause analysis model might infer that the anomaly is caused by a "memory leak" or "algorithm defect" within the "Data Acquisition Service." Ultimately, the root cause analysis model might output "memory leak" as the most likely root cause of the failure.

[0044] As can be seen from the above examples, the solution in this embodiment forms a logically rigorous and complete process, from macroscopic location of cross-regional anomaly propagation paths to microscopic root cause analysis of triggering nodes. Specifically, the anomaly association rule base can provide initial propagation clues, real-time alarm reception and matching are the foundation of dynamic tracing, reverse tracing and triggering node identification are key to focusing on the root cause, and the acquisition of multi-dimensional features and root cause model analysis are the means to ultimately determine the root cause. Through the synergy of these various technical features, accurate location of the root cause of cross-regional power system faults is achieved.

[0045] The method provided in this embodiment, by pre-establishing an anomaly association rule base, can receive anomaly alarm information from multiple power business information systems in real time and quickly match regional propagation paths. Compared with the existing technology that relies on human experience and manual log analysis to trace anomaly propagation paths, it can significantly improve efficiency and accuracy, effectively solving the problem of difficulty in tracing cross-regional anomaly propagation paths.

[0046] Furthermore, the solution in this embodiment can trace back along the regional propagation path to identify the starting area of ​​the regional propagation path and the earliest target triggering node in that area. This avoids the problem in existing technologies where the result of anomaly propagation is easily misjudged as the source of the anomaly, ensuring that the starting point of root cause analysis is the true triggering point. For example, in the example above, by tracing back, the "data acquisition service" node in region X can be accurately identified as the earliest triggering node.

[0047] Furthermore, the solution in this embodiment, by acquiring the environmental impact characteristics and system configuration characteristics of the starting area, and combining them with the parameter change characteristics of the target inducing node, performs a comprehensive analysis through a preset root cause determination model. This overcomes the limitations of existing technologies, which rely on a single analytical dimension and struggle to accurately identify the true root cause of the fault. Through the fusion analysis of multi-dimensional features, the solution in this embodiment can more comprehensively and deeply reveal the essential cause of the fault, thereby providing a reliable basis for the formulation of subsequent fault handling and prevention measures.

[0048] Overall, the solution in this embodiment, by pre-building an anomaly association rule base and combining it with the automatic matching and reverse tracing mechanism of regional propagation paths, can achieve accurate location of the root cause of cross-regional faults. It has the ability to quickly and accurately identify the starting area and target inducing node of cross-regional power business information system faults, effectively reducing the inefficiency and high error rate caused by manual investigation.

[0049] The following will further explain a method for locating the root cause of faults in a cross-regional power system in this exemplary embodiment.

[0050] In one embodiment, the construction of the anomaly association rule base specifically includes: Collect operation log data of the power business information system and obtain the associated characteristic data of the power business information system, including topological connection relationship and business service linkage logic; In the embodiments of this application, collecting operational log data of the power business information system refers to acquiring various events, statuses, errors, and warnings automatically recorded by the power business information system during its operation. This collection can be achieved by deploying a log collection agent to extract data from the log files of each system in real time; alternatively, it can be achieved through a unified log management platform to centrally receive and store the log streams reported by each system. The associated characteristic data describes the interconnections and dependencies between components and services within and between the power business information system. Among these, the associated characteristic data includes topological connection relationships and business service linkage logic. Topological connection relationships refer to the physical or logical connection methods between system components, such as servers, network devices, databases, and application service instances. Specifically, this can be obtained by parsing the network topology diagram, reading asset information from the configuration management database, or automatically scanning and constructing the network through network discovery tools. Business service linkage logic refers to the call order, data flow, and functional dependencies between different business services. Specifically, this can be obtained by analyzing business process documents, service interface definitions, or by monitoring API call chains and message queue interactions between services.

[0051] Based on the topological connection relationship, identify cross-regional data interfaces and key nodes, and based on the business service linkage logic, identify cross-regional synchronization component associations and business scheduling associations; In the embodiments of this application, a data interface refers to a boundary point or channel for data exchange between different regional systems, while a critical node refers to a component in the topology with high connectivity, high traffic, or carrying core functions, whose anomalies may have a significant impact on the entire or multiple regions. Identifying data interfaces and critical nodes can be done by analyzing the topology connection diagram to find routers, gateways, message middleware, etc., that connect different regional subnets, or by calculating the centrality indicators of nodes in the graph, such as degree centrality and betweenness centrality. Synchronization component relationships refer to the dependencies between business components in different regions that require maintaining real-time data consistency or state synchronization, such as the synchronous replication of a distributed database. Business scheduling relationships refer to the business operations in one region that trigger or depend on the completion of business operations in another region, such as the issuance and execution of cross-regional power dispatch instructions. Both synchronization component relationships and business scheduling relationships can be identified and obtained by analyzing the configuration of distributed transactions, message queue patterns, or workflow engines in the business service linkage logic.

[0052] Establish the association mapping relationship of the power business information system based on the relationship between data interfaces, key nodes, synchronization components and business scheduling; In the embodiments of this application, the association mapping relationship refers to a unified and structured representation of all cross-regional connection points and dependencies identified above. The association mapping relationship can be a graph model, where nodes represent system components, services, or regions, and edges represent data interfaces, key nodes, synchronization component associations, or business scheduling associations, and may include attributes describing their type, direction, and regional affiliation.

[0053] An anomaly propagation knowledge graph is constructed based on runtime log data, associated characteristic data, and associated mapping relationships, and an anomaly association rule base is established based on the anomaly propagation knowledge graph.

[0054] In the embodiments of this application, the anomaly propagation knowledge graph is a knowledge base represented in the form of a graph, where entities are nodes, such as system components, business services, and anomaly types, and the anomaly propagation paths and timing between entities are directed edges. The construction process of the anomaly propagation knowledge graph involves extracting historical anomaly events from runtime log data, combining them with correlation characteristic data and correlation mapping relationships, mapping these anomaly events to entities and relationships in the graph, thereby forming anomaly propagation chains, and recording the propagation time delay. The anomaly association rule base refers to a series of rules generated based on repetitive anomaly propagation patterns and rules discovered in the anomaly propagation knowledge graph. These anomaly association rules are typically stored in the form of "if A occurs, it may lead to B," and include information such as triggering conditions, timing constraints, and regional propagation priorities. The anomaly association rules can be extracted from the knowledge graph using graph mining algorithms, association rule learning algorithms, or pattern recognition methods based on expert experience.

[0055] As an example, and as one specific implementation method, the construction of an anomaly association rule base may include: First, operational log data is collected in real-time or periodically from multiple power business information systems, including power dispatch automation systems, energy management systems, distribution management systems, and enterprise-level monitoring platforms. This includes recording events such as equipment temperature exceeding limits at a substation, abnormal packet loss rates in a network link in a certain area, and service response timeouts. Simultaneously, correlational characteristic data of these systems is obtained. For example, by reviewing system architecture documents and network topology diagrams, it is determined that there is a direct fiber optic connection between the SCADA server in area A and the DMS server in area B. Analysis of business process definitions reveals that the power generation plan adjustment service in area A triggers a recalculation of the load forecasting service in area B, demonstrating the business service linkage logic.

[0056] Furthermore, based on the aforementioned topological connections, the data exchange gateway between Region A and Region B is identified as the cross-regional data interface, and the central database server in Region B is marked as a critical node because its failure would affect multiple business systems. Based on the business service linkage logic, a data synchronization mechanism is identified between the distributed file systems of Region A and Region B, forming a synchronization component association relationship; simultaneously, a command-feedback business scheduling association relationship exists between the scheduling instruction issuing platform in Region A and the execution unit in Region B.

[0057] Based on this, and according to the identified relationships between data interfaces, key nodes, synchronization components, and business scheduling, a mapping relationship is established for the power business information system. For example, a graph database can be constructed, where nodes represent specific system components or services, edges represent the connections or dependencies between them, and the type and direction of the edges are labeled.

[0058] Building upon this foundation, an anomaly propagation knowledge graph is constructed using operational log data, related characteristic data, and established relational mappings. For example, by analyzing historical logs, it is discovered that when a network device in region A experiences a high packet loss rate alarm, a service in region B that relies on that network typically reports a connection timeout error within 3-5 minutes. In the anomaly propagation knowledge graph, "high packet loss rate of network device in region A" and "connection timeout of service in region B" are designated as entity nodes, and a directed edge is established from the former to the latter. This directed edge is assigned timestamp and propagation delay attributes.

[0059] Finally, from the constructed anomaly propagation knowledge graph, recurring anomaly feature patterns and regional propagation rules are extracted through graph pattern matching and statistical analysis, and anomaly association rules are generated. For example, a rule can be generated: "If network devices in region A experience a high packet loss rate, then services in region B that rely on that network may experience connection timeouts within 3-5 minutes, and this propagation path has a high priority." These generated rules, along with their triggering conditions, timing constraints, and regional propagation priorities, will be stored in the anomaly association rule base for matching during real-time fault localization.

[0060] The above technical solution systematically collects operational log data and related characteristic data. Based on this, it identifies cross-regional data interfaces, key nodes, synchronization components, and business scheduling relationships, thereby establishing a comprehensive correlation mapping. Furthermore, based on this data and correlation mapping, an anomaly propagation knowledge graph is constructed to intuitively and accurately reflect the propagation path and sequence of historical anomalies. Finally, an anomaly correlation rule base is extracted and established from the anomaly propagation knowledge graph, ensuring the comprehensiveness, accuracy, and timeliness of the rule base. Therefore, when locating faults in real time, this anomaly correlation rule base can be used to achieve more accurate real-time anomaly feature data matching, thereby accurately identifying regional propagation paths. This improves the efficiency and reliability of fault root cause location and reduces misjudgments or omissions caused by incomplete or inaccurate rules.

[0061] In one embodiment, step S20, which involves "constructing an anomaly propagation knowledge graph based on runtime log data, associated characteristic data, and associated mapping relationships, and establishing an anomaly association rule base based on the anomaly propagation knowledge graph," specifically includes: The business services, data interfaces, and nodes in each power business information system are treated as entity nodes, and the association mapping relationship is treated as the association edge between entity nodes. In the embodiments of this application, entity nodes are the basic units in a knowledge graph that represent concrete or abstract concepts in the real world. In this embodiment, business services, data interfaces, and nodes are abstracted as entity nodes. For example, a business service can refer to a specific application function module in a power system, a data interface is a channel for data exchange between different systems or modules, and a node can be a physical device, logical component, or regional boundary. Defining these elements as entity nodes helps to unify the graph representation of complex power system structures and functions. This can be achieved by using unique identifiers of these elements as node IDs and storing their type, attributes, and other metadata, or by mapping these elements to concepts in an ontology through a predefined ontology model, thereby achieving a higher level of semantic representation. Association edges represent the relationships between entity nodes. In this embodiment, association mapping relationships are defined as association edges between entity nodes, where association mapping relationships can include topological connection relationships, business service linkage logic, data flow relationships, control dependency relationships, etc. Representing association mapping relationships as edges enables knowledge graphs to intuitively display the interactions and dependencies between elements within a power system. This can be achieved by creating edges with types and attributes in a graph database. For example, an edge can represent relationships such as "call", "connection", and "dependence", and can be accompanied by attributes such as connection protocol and data type. Alternatively, semantic web technology can be used to represent entity nodes and association edges using RDF triples.

[0062] Based on historical anomaly records in the operation log data, determine the anomaly propagation path and anomaly propagation sequence between each entity node; In the embodiments of this application, the anomaly propagation path refers to the propagation trajectory of an anomaly event in the power business information system, that is, the process starting from one entity node, passing through a series of interconnected entity nodes, and finally reaching one or more other entity nodes. The anomaly propagation path is usually inferred based on the observed sequence and correlation of anomaly events in historical anomaly records, and its function is to reveal the diffusion pattern of anomaly events in the system. The anomaly propagation path can be determined by analyzing the source and propagation chain of anomaly alarms in the logs, combining the correlation edges between entity nodes to construct the propagation sequence of anomaly events, or by using sequence pattern mining algorithms to automatically discover frequently occurring anomaly propagation sequences from a large number of historical anomaly logs. The anomaly propagation timing refers to the temporal order and time interval of anomaly event propagation between different entity nodes. It quantifies the time required for an anomaly to propagate from one node to the next, and its function is to provide temporal dimension information of anomaly propagation. The anomaly propagation timing can be determined by parsing the timestamps of anomaly alarms from each entity node in historical anomaly records and calculating the time difference between adjacent anomaly nodes, or by combining information such as communication delays and processing delays within the system to calibrate and refine the anomaly occurrence times recorded in the logs.

[0063] Construct an anomaly propagation knowledge graph based on anomaly propagation paths and anomaly propagation sequences; In the embodiments of this application, constructing an anomaly propagation knowledge graph refers to integrating the previously identified entity nodes, associated edges, anomaly propagation paths, and anomaly propagation sequence to form a structured, queryable knowledge representation. This anomaly propagation knowledge graph not only includes the system's static topology and business relationships but also incorporates dynamic anomaly propagation behavior information. Its role is to provide a unified data foundation for subsequent anomaly feature pattern extraction and association rule establishment. The construction of the anomaly propagation knowledge graph can be achieved by using a graph database, treating entity nodes as vertices, association mapping relationships as basic edges, representing anomaly propagation paths as a series of directed edges, and adding anomaly propagation sequence as an attribute to these edges; or by using semantic web standards such as RDF / OWL to represent all information as triples.

[0064] Extract abnormal feature patterns and regional transmission rules from the abnormal transmission knowledge graph, generate abnormal association rules, and establish an abnormal association rule library.

[0065] In the embodiments of this application, anomaly feature patterns refer to recurring anomaly propagation sequences or subgraphs with specific structures and temporal characteristics in the anomaly propagation knowledge graph. These represent typical behaviors of the power system under specific fault or anomaly conditions. Their function is to abstract complex anomaly propagation phenomena into identifiable and reusable patterns, providing a basis for rapid localization and diagnosis. The extraction of anomaly feature patterns can be achieved through graph pattern mining algorithms to find frequent subgraphs or path patterns in the anomaly propagation knowledge graph, or by combining machine learning methods to cluster anomaly propagation paths in the graph. Regional propagation patterns refer to specific patterns and priorities of anomaly events propagating between different regions. These typically involve cross-regional boundary nodes, data interfaces, and inter-regional business dependencies, revealing the propagation characteristics of cross-regional faults. Regional propagation patterns can be identified by analyzing anomaly propagation paths involving cross-regional boundary nodes in the anomaly propagation knowledge graph, identifying typical paths and temporal characteristics of anomalies propagating from one region to another, or by using community detection algorithms to identify different regional communities in the knowledge graph and analyze the patterns of anomaly propagation between communities. Anomaly association rules are a set of rules that formally represent anomaly feature patterns and regional propagation laws. Each rule typically includes a trigger condition and a corresponding result. Their function is to make implicit knowledge in the knowledge graph explicit, facilitating subsequent real-time matching and reasoning. Anomaly association rules can be generated by using extracted anomaly feature patterns as the first part of the rule and the corresponding regional propagation path or potential affected area as the second part. Statistical indicators such as confidence and support can be added, or association rule mining algorithms can be used to mine association rules with high support and confidence from anomaly propagation paths. The anomaly association rule base is a knowledge base that stores all generated anomaly association rules. Its function is to provide a queryable and maintainable set of rules for matching and reasoning during real-time fault localization. The anomaly association rule base can be established by storing rules in a relational database, with each rule as a record containing fields such as rule ID, trigger condition, result, and confidence, or by using a dedicated rule engine or knowledge base system to manage these rules.

[0066] For example, as a specific implementation method, when constructing the anomaly propagation knowledge graph, specific components in the power business information system such as "SCADA service," "EMS data interface," and "Region A gateway node" can be defined as entity nodes. Simultaneously, the associated mapping relationships such as "SCADA service calling EMS data interface" and "EMS data interface connecting to Region A gateway node" can be defined as association edges between entity nodes. When the operation log data records historical anomaly records such as "SCADA service experiences high CPU utilization anomaly at time T1," "EMS data interface experiences data transmission delay anomaly at T1+5 seconds," and "Region A gateway node experiences network packet loss anomaly at T1+10 seconds," the system can determine an anomaly propagation path from "SCADA service" to "EMS data interface" and then to "Region A gateway node."

[0067] Simultaneously, the propagation time from the SCADA service to the EMS data interface can be calculated to be 5 seconds, and the propagation time from the EMS data interface to the regional A gateway node can be calculated to be 5 seconds. Based on this information, directed edges can be created in the knowledge graph, such as an edge from "SCADA service" to "EMS data interface", and assigned the attribute "propagation delay = 5 seconds".

[0068] Furthermore, when extracting abnormal feature patterns from the constructed anomaly propagation knowledge graph, if the sequence "SCADA service CPU high → EMS data interface latency → Region A gateway packet loss" is found to occur frequently in historical records, it can be identified as an abnormal feature pattern. Simultaneously, if "Region A gateway node anomaly" frequently leads to "Region B database service anomaly," then the regional propagation pattern "from Region A to Region B" can be identified.

[0069] Finally, based on the above information, an anomaly association rule can be generated, for example: "If the SCADA service CPU is high and the EMS data interface is delayed, it may cause packet loss at the gateway in region A and further propagate to region B, with a higher priority." At this time, this rule, along with its corresponding anomaly characteristic pattern and regional propagation characteristics, will be stored in the anomaly association rule base for subsequent real-time matching.

[0070] Through the above technical solution, this application can efficiently and accurately transform heterogeneous operation log data and related characteristic data in the power business information system into a structured anomaly propagation knowledge graph. This anomaly propagation knowledge graph not only clearly shows the static relationships between various components within the system, but also dynamically captures the actual propagation path and timing information of abnormal events within the system. Based on this, by extracting representative anomaly feature patterns and regional propagation rules from the knowledge graph and formalizing them into anomaly association rules, the solution in this embodiment can establish a high-quality anomaly association rule base. This improves the construction efficiency and rule effectiveness of the anomaly association rule base, thereby enhancing the accuracy and reliability of cross-regional power system fault root cause localization.

[0071] In one embodiment, the step of "constructing an anomaly propagation knowledge graph based on the anomaly propagation path and anomaly propagation sequence" specifically includes: Based on the anomaly propagation path, directed transmission edges are established between entity nodes, and timestamp and propagation delay attributes are assigned to the directed transmission edges according to the anomaly propagation sequence. In the embodiments of this application, a directed propagation edge represents a directional connection from one entity node to another. Furthermore, based on the propagation sequence of the anomaly, the directed propagation edge can be assigned a timestamp attribute and a propagation delay attribute. The timestamp attribute records the specific time point at which the anomaly event propagates along the directed propagation edge; for example, it can record the time the anomaly arrives at the source node and the target node. The propagation delay attribute quantifies the time required for the anomaly to propagate along the directed propagation edge. For example, it can be obtained by calculating the time difference between the time the anomaly occurs at the source node and the time it occurs at the target node, or by statistically analyzing the average propagation time of this path in historical data.

[0072] Calculate the frequency of occurrence of each directed transmission edge in the historical anomaly records, and use the frequency of occurrence as the weight attribute of the corresponding directed transmission edge; In the embodiments of this application, the weight attribute can reflect the activity level or importance of a specific anomaly propagation path in historical data. For example, the frequency of occurrence can be determined by counting the number of times a directed propagation edge is activated by anomaly events over a past period, or different initial weights can be assigned to different types of anomaly propagation paths by combining expert experience. The weight attribute can provide a quantitative basis for subsequent anomaly pattern recognition and rule generation, helping to distinguish between common and occasional propagation paths.

[0073] Identify cross-regional transmission edges, and label the source and target regions connected by the cross-regional transmission edges to generate region crossing identifiers; In the embodiments of this application, a cross-regional transmission edge refers to a directed transmission edge connecting entity nodes in different power business information system regions. The region crossing identifier can be a Boolean value indicating whether the corresponding edge is a cross-regional transmission edge, or it can contain more detailed information, such as unique identifiers for the source and target regions. For example, whether it is a cross-regional transmission edge can be determined by analyzing the deployment location or business domain of the entity node and annotating it according to predefined regional division rules. The region crossing identifier helps to clearly distinguish between intra-regional and cross-regional propagation in the knowledge graph, providing crucial information for locating cross-regional faults.

[0074] An anomaly propagation knowledge graph is constructed based on directed propagation edges and entity nodes that have timestamp attributes, propagation delay attributes, weight attributes, and region crossing identifiers. In the embodiments of this application, the anomaly propagation knowledge graph not only includes the topological structure of anomaly propagation, but also incorporates key information such as time, frequency and region, thereby reflecting the propagation law of power system anomalies more comprehensively and deeply.

[0075] Store the anomaly propagation knowledge graph as a graph data structure and create node indexes and edge indexes.

[0076] In the embodiments of this application, the graph data structure is a database model specifically designed for storing and querying graph data, such as an attribute graph model, which can efficiently represent nodes, edges, and their attributes. Establishing node and edge indexes is to optimize the query performance of the knowledge graph; for example, a specific node can be quickly found by its node ID, or edges that meet specific conditions can be quickly retrieved by their attributes. As an example, graph databases such as Neo4j and JanusGraph can be used for storage, and their built-in indexing mechanisms can be utilized to improve data access efficiency.

[0077] For example, as a specific implementation method, suppose there are two entity nodes in a certain power business information system: a "power grid dispatch service" located in North China and a "data synchronization interface" located in East China.

[0078] When the "Power Grid Dispatch Service" experiences an anomaly, the "Data Synchronization Interface" will also experience an anomaly after a certain period. Based on the anomaly propagation path, a directed propagation edge can be established between the "Power Grid Dispatch Service" and the "Data Synchronization Interface". If the "Power Grid Dispatch Service" experiences an anomaly at 10:00:00, and the "Data Synchronization Interface" experiences an anomaly at 10:00:15, then this directed propagation edge can be assigned a timestamp attribute to record the time of the anomaly, and a propagation delay attribute of 15 seconds can be calculated.

[0079] Furthermore, if analysis of historical anomaly records reveals that this abnormal propagation path from "Power Grid Dispatch Service" to "Data Synchronization Interface" has occurred 50 times in the past month, then this directed propagation edge can be assigned a weight attribute with a value of 50. Since "Power Grid Dispatch Service" and "Data Synchronization Interface" belong to two different regions, North China and East China, respectively, this directed propagation edge will be identified as a cross-regional propagation edge, and a regional crossing identifier will be generated, clearly marking the source region as "North China" and the target region as "East China".

[0080] Ultimately, all these entity nodes with detailed attributes and directed transmission edges can be integrated and stored in a graph database to form an anomaly transmission knowledge graph. To improve query efficiency, node indexes can be created for entity nodes such as "Power Grid Dispatch Service" and "Data Synchronization Interface," and edge indexes can be created for directed transmission edges with specific transmission delay or weight attributes.

[0081] Through the above technical solution, in constructing the anomaly propagation knowledge graph, not only is the topological structure and temporal relationship of anomaly propagation clarified, but the temporal, frequency, and regional characteristics of propagation are further quantified. This enables the anomaly propagation knowledge graph to more comprehensively and accurately reflect the propagation patterns of anomalies in complex power systems. Specifically, the propagation delay attribute helps to accurately determine the speed and temporal dependence of anomaly propagation; the weight attribute can highlight high-frequency, high-risk propagation paths, providing priority references for fault early warning and root cause localization; and the regional crossing identifier clearly defines the propagation path and impact range of cross-regional faults. Based on this, the solution in this embodiment can improve the expressive power and analytical depth of the anomaly propagation knowledge graph, providing a data foundation for subsequently extracting anomaly feature patterns and regional propagation patterns from the knowledge graph, thereby improving the accuracy and efficiency of root cause localization of cross-regional power system faults.

[0082] In one embodiment, the step of "extracting abnormal feature patterns and regional transmission rules from the abnormal transmission knowledge graph, generating abnormal association rules, and establishing an abnormal association rule base" specifically includes: Path tracing is performed on the directed transmission edges in the anomaly transmission knowledge graph to extract the transmission chains from a single source node to multiple target nodes, and the transmission chains are sorted in descending order of weight attributes to obtain a set of transmission chains. In the embodiments of this application, path tracing of directed transmission edges in the anomaly propagation knowledge graph refers to the process of starting from an entity node and traversing along directed transmission edges until reaching other entity nodes. Specifically, this can be implemented using graph traversal algorithms such as depth-first search or breadth-first search to discover all possible anomaly propagation paths. A propagation chain refers to a path connected by a series of continuous directed transmission edges, representing the complete process of an anomaly propagating from a source to multiple downstream nodes. The weight attribute represents the frequency of occurrence of each directed transmission edge in historical anomaly records; the higher the weight, the more common or important the propagation path. Arranging the paths in descending order of weight attributes indicates that anomaly propagation paths that occur more frequently and are more representative in history will be given priority, thus focusing on more important propagation chains in subsequent pattern recognition.

[0083] Identify conduction chains with the same topological structure features from the set of conduction chains, cluster the identified conduction chains into conduction chain clusters, and generate abnormal feature patterns based on the conduction chain clusters; In the embodiments of this application, having the same topological structure features means that the structural characteristics of entity nodes in the transmission chain are similar, such as the connection method, order, and type. For example, multiple transmission chains may exhibit the pattern of "database anomaly → application service anomaly → user interface anomaly," even if the specific database instances or application service names are different, their topological structure features are similar. Clustering is an unsupervised learning method used to group data points with similar features. Here, clustering algorithms, such as K-means, DBSCAN, or hierarchical clustering, can group transmission chains according to their topological similarity, forming transmission chain clusters. A transmission chain cluster is a set of transmission chains with similar topological structure features, and each cluster represents a typical anomaly propagation pattern. An anomaly feature pattern is a representative anomaly propagation structure and behavioral feature abstracted from the transmission chain cluster, describing how an anomaly propagates from one or a group of source nodes to other nodes under specific conditions, as well as the entity types and connection relationships involved in the propagation process.

[0084] Identify transmission chains associated with region crossing markers and obtain the region transmission characteristics between the source region and the target region corresponding to the region crossing markers; In embodiments of this application, the region crossing identifier is used to mark the source and target regions connected by the cross-regional propagation edge. The region propagation characteristics describe the features of anomaly propagation between different regions. For example, the propagation from region A to region B typically involves which key interfaces or nodes, what the average propagation delay is, and the transformation or amplification effects that may occur at specific region boundaries, which helps to understand the macroscopic laws of cross-regional anomaly propagation.

[0085] Triggering conditions and timing constraints are set based on abnormal feature patterns, and regional propagation priorities are set based on regional propagation characteristics to generate abnormal association rules; In the embodiments of this application, the triggering condition refers to the condition under which the anomaly association rule is activated when the real-time received anomaly alarm information matches a certain anomaly characteristic pattern. The timing constraint refers to the required time sequence and time interval of events occurring during the anomaly propagation process. Based on regional propagation characteristics, priorities are set for cross-regional anomaly propagation paths; for example, propagation from the core area to the edge area may have a higher priority, or propagation between certain key areas may be considered more important. The anomaly association rule is a rule that integrates information such as anomaly characteristic patterns, triggering conditions, timing constraints, and regional propagation priorities to guide the matching of real-time anomaly alarm information and the determination of regional propagation paths.

[0086] The anomaly association rules are associated with the corresponding anomaly feature patterns and regional propagation features and stored in the anomaly association rule library.

[0087] In the embodiments of this application, the associated storage means that in the anomaly association rule base, each anomaly association rule not only includes its own definition, but also establishes a clear link relationship with the anomaly feature pattern and regional transmission feature on which it is based. This facilitates the maintenance, updating, and interpretation of the rules, and also makes it convenient to quickly retrieve and understand the background information of the rules during the matching process. The anomaly association rule base is a database or knowledge base that stores all generated anomaly association rules for use when matching real-time anomaly alarm information.

[0088] For example, as a specific implementation method, in a power business information system, the path of anomaly propagation knowledge graph can be traced first. For instance, starting from an entity node of "database server CPU overload", the path can be traced to the "application service response timeout" node, and then to the "user interface lag" node, thus forming a propagation chain.

[0089] If multiple similar propagation chains exist, such as "Database server A CPU overload → Application service X response timeout → User interface Y lag" and "Database server B CPU overload → Application service Z response timeout → User interface W lag," these two propagation chains share similar topological characteristics of "database CPU overload → application service response timeout → user interface lag." In this case, these propagation chains can be clustered into a propagation chain cluster, from which the abnormal characteristic pattern of "database CPU overload causing application service response timeout, thus affecting user experience" can be abstracted. Furthermore, if a propagation chain is traced from "gateway device malfunction in region 1" to "business system interruption in region 2," the region crossing identifier associated with this propagation chain can be identified, and the region propagation characteristics from "region 1 to region 2" can be obtained. For example, this cross-region propagation typically involves specific data interfaces and average latency.

[0090] Based on these abnormal characteristic patterns, trigger conditions can be set, such as triggering when the database CPU utilization exceeds 90% for 5 consecutive minutes and the application service response time exceeds 3 seconds. Simultaneously, timing constraints can be set, requiring the database CPU overload event to occur before the application service response timeout event, with a time interval of less than 1 minute. For regional propagation characteristics, regional propagation priorities can be set; for example, propagation from the core data center region to the regional substation has a higher priority than propagation from the regional substation to the user side.

[0091] Finally, by integrating the set triggering conditions, timing constraints, and regional propagation priorities, a corresponding anomaly association rule can be generated, and it, along with the corresponding anomaly feature pattern and regional propagation feature, can be stored in the anomaly association rule library.

[0092] Through the above technical solution, this application can extract representative and operable anomaly feature patterns and regional propagation rules from massive historical anomaly data. This pattern recognition and rule generation mechanism enables the rules in the anomaly association rule base to be more accurate and generalizable. Based on this, in real-time fault location, it can more efficiently and accurately match anomaly alarm information, identify the regional propagation path of anomalies, improve the accuracy and efficiency of fault root cause location, and reduce misjudgments or omissions caused by coarse or incomplete rules.

[0093] In one embodiment, step S30, which involves "receiving abnormal alarm information from multiple power business information systems in real time, extracting abnormal feature data, and matching the abnormal feature data using an abnormal association rule base to obtain the regional propagation path," specifically includes: The system receives abnormal alarm information reported by various power business information systems in real time, and extracts abnormal feature data based on the abnormal alarm information. The abnormal feature data includes business node identifier, time of abnormal occurrence, and abnormal type. In the embodiments of this application, receiving abnormal alarm information can be achieved by deploying a message queue system, such as Apache Kafka or RabbitMQ, enabling each power business information system to publish abnormal alarm information to designated topics, and the location method to subscribe to these topics to achieve real-time data acquisition; alternatively, a RESTful API interface can be provided for each system to proactively push alarm information. When extracting abnormal feature data, predefined parsing rules can be used, such as matching key fields in alarm logs using regular expressions, or using natural language processing technology to identify core information such as business node identifiers, the time of abnormal occurrence, and the type of abnormality from unstructured alarm text. The business node identifier is used to uniquely identify the system component or service where the abnormality occurred, the time of abnormal occurrence records the point in time when the event occurred, and the type of abnormality describes the specific nature of the abnormality.

[0094] Based on the time of occurrence of the anomaly, the anomaly feature data are arranged in chronological order to generate a time-series anomaly sequence; In the embodiments of this application, the temporal arrangement can be implemented in various ways. For example, the received abnormal feature data can be stored in an ordered set in memory, such as a skip list or a red-black tree, or the timestamp index of the database can be used for sorting and querying when the data is entered into the database. The generated temporal anomaly sequence is a list of abnormal events arranged in chronological order, which can intuitively show the dynamic process of an anomaly from its occurrence to its propagation.

[0095] Business nodes and their regional distribution are extracted based on time-series anomaly sequences, and corresponding anomaly association rules are retrieved from the anomaly association rule base based on the business nodes. In the embodiments of this application, business nodes are extracted from the time-series anomaly sequence, that is, all occurrences of business node identifiers are collected. The extraction of regional distribution can be achieved by associating business node identifiers with a preset regional mapping table, for example, determining the region based on server IP address ranges, device ID prefixes, or service names. Subsequently, the extracted business nodes are used as query conditions to search in an anomaly association rule base. The search method can include exact matching based on business node identifiers or fuzzy matching based on attributes such as node type and region, to obtain all anomaly association rules that may be related to the current anomaly sequence.

[0096] Calculate the matching degree between the time-series abnormal sequence and the abnormal association rule, select the abnormal association rule with the highest matching degree, and determine the regional propagation path based on the selected abnormal association rule.

[0097] In the embodiments of this application, the calculation of the matching degree can comprehensively consider multiple dimensions. For example, it can compare the consistency between the occurrence order of business nodes in the time-series anomaly sequence and the propagation order defined in the rule, evaluate the degree of consistency between the time interval between anomaly occurrences and the propagation delay in the rule, and verify whether the order of region crossing matches the region crossing identifier in the rule. By assigning different weights to these dimensions and performing a weighted summation, a comprehensive matching degree score can be obtained. After calculating the matching degree of all candidate anomaly association rules, the system will select the rule with the highest matching degree. Once the best matching rule is determined, the propagation path of the current cross-regional anomaly event can be determined based on the predefined propagation path information in this anomaly association rule.

[0098] For example, as a specific implementation method, suppose that in a certain power system, the monitoring system reports three abnormal alarm messages in real time: the first alarm comes from the "protection device of substation A", the time of the abnormality is "10:00:05", and the type of abnormality is "protection action tripping"; the second alarm comes from the "data acquisition server of the dispatch center", the time of the abnormality is "10:00:10", and the type of abnormality is "data link interruption"; the third alarm comes from the "application service of the regional control center", the time of the abnormality is "10:00:18", and the type of abnormality is "business processing timeout".

[0099] First, the system receives these three alarm messages and extracts the abnormal feature data from them: 1. Business node identifier: "Protection device", time of anomaly occurrence: "10:00:05", anomaly type: "Protection action trip". 2. Business node identifier: "Data acquisition server", time of anomaly occurrence: "10:00:10", anomaly type: "Data link interruption". 3. Business node identifier: "Application service", time of anomaly occurrence: "10:00:18", anomaly type: "Business processing timeout".

[0100] Furthermore, based on the time of occurrence of the anomaly, these anomaly characteristic data are arranged in chronological order to generate a chronological anomaly sequence: [(protection device, 10:00:05, protection action trip), (data acquisition server, 10:00:10, data link interruption), (application service, 10:00:18, business processing timeout)].

[0101] Based on this, the system extracts the business nodes "protection device", "data acquisition server" and "application service" according to the time-series anomaly sequence, and identifies that these business nodes belong to "substation A area", "dispatch center area" and "regional control center area" respectively.

[0102] Based on this, the system uses this business node information to retrieve several potentially related anomaly association rules from the anomaly association rule base. For example, it might retrieve a rule R. 001 The rule is described as "protection device tripping → data link interruption → service processing timeout", and the node sequence, time delay range, and region crossing order associated with the rule are highly consistent with the current sequence.

[0103] Finally, the system calculates the matching degree between the current time-series anomaly sequence and each retrieved anomaly association rule. For rule R... 001 The system will evaluate its business node matching rate, such as all three nodes matching, timing deviations (e.g., the 5-second time interval from 10:00:05 to 10:00:10 matches the delay range defined in the rule, and the 8-second time interval from 10:00:10 to 10:00:18 also matches), and the consistency of the region crossing order with the region crossing identifier defined in the rule. Assume R... 001 The system can select R if it has the highest matching degree. 001 And based on R 001 The determined regional propagation path is: Substation A area → Dispatch Center area → Regional Control Center area.

[0104] Through the above technical solution, this application can identify the propagation path of anomalies from massive amounts of abnormal alarms in power business information systems in real time and accurately. Specifically, by extracting features and arranging the abnormal alarm information in a temporal sequence, an anomaly sequence with a time dimension can be formed, providing high-quality input for subsequent pattern matching. Furthermore, by retrieving and calculating the matching degree from the anomaly association rule base based on business nodes, the rules that best match the current anomaly propagation pattern can be effectively filtered out, thereby reducing false positives and false negatives caused by simple matching. Based on this, the solution of this embodiment has the effect of improving the accuracy and reliability of anomaly propagation path determination and improving the accuracy and efficiency of fault diagnosis.

[0105] In one embodiment, the step of "calculating the matching degree between the time-series anomaly sequence and the anomaly association rule, selecting the anomaly association rule with the highest matching degree, and determining the regional propagation path based on the selected anomaly association rule" specifically includes: The system counts the number of business nodes in a time-series anomaly sequence that match the business services associated with the anomaly association rules, and calculates the business matching rate based on the number of matches. In the embodiments of this application, a business node is the smallest unit of anomaly occurrence in the power business information system, while the anomaly association rule predefines the business service sequence through which the anomaly may propagate. By counting the number of matches between the two, the similarity between the anomaly sequence and the rule at the business level can be quantified. This number of matches can be obtained by performing an intersection operation on the business node identifiers in the time-series anomaly sequence and the predefined business service identifiers in the anomaly association rule; the number of intersection elements is the number of matches. The business matching rate can be calculated by dividing the number of matches by the total number of business services in the anomaly association rule or the total number of business nodes in the time-series anomaly sequence. Alternatively, a unique hash value can be assigned to each business node and business service, and then the set of hash values ​​in the time-series anomaly sequence can be compared with the set of hash values ​​in the anomaly association rule; the number of identical hash values ​​is then counted as the number of matches.

[0106] Calculate the time interval between the occurrence times of adjacent anomalous feature data in the time-series anomaly sequence, and compare the time interval with the propagation delay attribute of the corresponding directed propagation edge in the anomaly association rule to obtain the time-series deviation; In the embodiments of this application, anomaly propagation typically has a specific time delay. By comparing the actual observed time interval between anomaly occurrences with the preset propagation delay in the rules, it can be determined whether the temporal characteristics of anomaly propagation are consistent with the rules. The smaller the temporal deviation, the higher the temporal fit. For each pair of adjacent anomaly feature data in the temporal anomaly sequence, the difference between their anomaly occurrence times can be calculated as the actual time interval. Then, the directed propagation edge corresponding to the propagation between these two pairs is found in the anomaly association rules, and its propagation delay attribute is obtained. The temporal deviation can be defined as the absolute value of the difference between the actual time interval and the propagation delay attribute, or its relative error. Furthermore, more complex statistical methods can be used, such as fitting the actual time interval and the propagation delay attribute to a Gaussian or exponential distribution, and calculating the similarity or distance between the two distributions as the temporal deviation.

[0107] Determine whether the region crossing order involved in the time-series anomaly sequence is consistent with the region crossing identifier associated with the anomaly association rule to obtain a consistency determination result; In the embodiments of this application, cross-regional propagation is one of the complex characteristics of power system faults, and accurately determining the order of region crossings is crucial for locating the source of the fault. Specifically, the sequence of regions involved when the anomaly occurred can be extracted from the time-series anomaly sequence, for example, by identifying the region information associated with the service node. Simultaneously, the associated region crossing identifier is obtained from the anomaly association rule. This identifier typically contains the order information of the source and target regions, and then the order of these two region sequences is directly compared to see if they are completely consistent. Alternatively, fuzzy matching or partial matching can be used. For example, if the anomaly association rule describes the region propagation as A→B→C, while the time-series anomaly sequence shows A→X→B→C, where X is an intermediate region between A and B, but the overall crossing order remains A to B then to C, then it can be considered to have a certain degree of consistency.

[0108] The matching degree between the time-series abnormal sequence and the abnormal association rule is calculated based on the business matching rate, time-series deviation, and consistency judgment results. In the embodiments of this application, a weighted summation method can be used to determine the matching degree between the time-series abnormal sequence and the abnormal association rule. Alternatively, a machine learning model, such as a support vector machine or a neural network, can be used to train the model and output the matching degree, taking the business matching rate, time-series deviation, and consistency judgment result as input features.

[0109] Furthermore, based on the business matching rate, time series deviation, and consistency judgment results, the matching degree between the time series anomaly sequence and the anomaly association rule is calculated, specifically including: The matching degree is calculated using a weighted fusion method, and the formula is as follows: ,in, This represents the business matching rate, with a value ranging from 0 to 1. To normalize the timing deviation, the calculation method is to limit the ratio of the actual timing deviation to a preset threshold within the range of 0-1; This is the regional consistency coefficient. It can be 1 when the consistency determination result is consistent, and 0 when the result is inconsistent. , , All are weighting coefficients, and their sum is 1 and all are greater than 0.

[0110] As a preferred implementation, the weighting coefficient can be adjusted according to the operating characteristics of different power grid information systems. For example, for systems with a relatively fixed topology, it can be set as follows: For fast protection systems with strict timing constraints, the following settings can be adopted: In addition, the preset threshold can be determined based on the statistical distribution of conduction delay in historical fault data, for example, by taking the mean conduction delay plus twice the standard deviation as the threshold.

[0111] Select the anomaly association rule with the highest matching degree, and determine the regional propagation path based on the selected anomaly association rule.

[0112] In the embodiments of this application, after calculating the matching degree of all candidate anomaly association rules, these matching degree values ​​can be directly compared, and the anomaly association rule with the highest value can be selected. If multiple rules have the same highest matching degree, further filtering can be performed according to a preset priority, or all rules with the highest matching degree can be used as possible propagation paths. In this case, the selected anomaly association rule itself contains information about the regional propagation path, such as its associated entity node sequence and directed propagation edge sequence, which can be directly used as the regional propagation path.

[0113] It should be noted that if multiple abnormal association rules calculate the same highest matching degree, the regional propagation priority built into each candidate rule can be extracted, and the abnormal association rule with the higher regional propagation priority level can be selected as the final matching rule.

[0114] Among them, the regional propagation priority built into the anomaly association rule is generated based on the regional propagation characteristics. It is used to distinguish the risk weights of cross-regional chain failures and local failures in the same scenario, and to prioritize the selection of high-risk cross-regional propagation paths.

[0115] For example, as a specific implementation, suppose that in a cross-regional power system, a series of abnormal alarm messages are received in real time, and after processing, a time-series abnormality sequence is generated. This time-series abnormality sequence shows that the abnormality occurs sequentially in service S1 of region A, service S2 of region A, and service S3 of region B. Meanwhile, there are multiple rules in the abnormality association rule base. One rule, R1, describes the propagation of the abnormality from S1 in region A to S2 in region A, and then to S3 in region B. The propagation delay from S1 to S2 is 5 seconds, and the propagation delay from S2 to S3 is 10 seconds. The region crossing identifier is A→B.

[0116] First, when calculating the business matching rate, the business nodes in the time-series anomaly sequence are compared with the business services associated with rule R1. All three are found to match, resulting in a matching count of 3. If rule R1 has a total of 3 business services, the business matching rate is 100%. Second, when calculating the time-series deviation, assuming the actual time interval between S1 and S2 in the time-series anomaly sequence is 6 seconds, and the actual time interval between S2 and S3 is 9 seconds, and the propagation delay attribute between S1 and S2 in rule R1 is 5 seconds, and the propagation delay attribute between S2 and S3 is 10 seconds, then the time-series deviation between S1 and S2 is |6-5|=1 second, and the time-series deviation between S2 and S3 is |9-10|=1 second. The average time-series deviation or the weighted average time-series deviation can then be calculated.

[0117] Furthermore, when judging the consistency determination result, the region crossing order involved in the time-series abnormal sequence is A→B, which is completely consistent with the region crossing identifier A→B associated with rule R1. Therefore, the consistency determination result is "consistent".

[0118] Finally, the calculated business matching rate, timing deviation, and consistency judgment result are input into a preset matching degree calculation model. This matching degree calculation model can be a simple weighted summation function, for example, matching degree = 0.5 * business matching rate + 0.3 * (1 - normalized timing deviation) + 0.2 * consistency judgment result. Through calculation, the matching degree of rule R1 can be obtained. This process is repeated for all candidate rules, and finally, the rule with the highest matching degree, such as R1, is selected, and the propagation path described by it is determined as the current regional propagation path.

[0119] Through the above technical solution, this application overcomes the problem that in complex cross-regional power systems, relying solely on service node matching or single-dimensional matching is insufficient to accurately identify abnormal propagation paths. Specifically, by comprehensively considering service matching rate, timing deviation, and consistency of regional crossing order, the solution of this embodiment provides a multi-dimensional matching degree calculation method, enabling the system to more comprehensively and accurately evaluate the matching degree between real-time abnormal sequences and preset abnormal association rules, thereby improving the accuracy of regional propagation path location. Especially when facing multiple potential fault modes with similar service nodes but different propagation timing or regional crossing order, the solution of this embodiment can effectively distinguish and identify the propagation path that best matches the actual situation, providing a more reliable basis for subsequent fault root cause tracing, thereby improving the efficiency and accuracy of fault location.

[0120] In one embodiment, step S60, which involves "obtaining the environmental impact characteristics and system configuration characteristics of the starting area, and combining them with parameter change characteristics to determine the root cause of the target inducing node's failure using a preset root cause determination model," specifically includes: The environmental impact characteristics of the starting area at the time of the anomaly are obtained, including temperature conditions, humidity conditions, network bandwidth fluctuations, and external disturbances. In the embodiments of this application, environmental impact characteristics refer to external or internal non-configurable factors that may affect the operation of the power system. Temperature and humidity conditions can be obtained, for example, by real-time collection of environmental monitoring data through an IoT sensor network deployed in the initial area, or by obtaining regional meteorological data through an integrated meteorological service interface. Network bandwidth fluctuations can be obtained, for example, by real-time monitoring of network device traffic, latency, packet loss rate, and other indicators in the initial area using network performance monitoring tools, or by identifying abnormal traffic patterns in network device logs. External disturbances can be obtained, for example, by integrating external event source identification, such as news APIs, social media analysis, geological disaster early warning systems, etc., or by manual input, expert system judgment, etc.

[0121] The system configuration features of the starting region within a preset time window before the occurrence of the anomaly are obtained. The system configuration features include business service start / stop records, connection pool parameter adjustment records, and permission configuration change records. In the embodiments of this application, system configuration features refer to the configuration or operation records within the power business information system, reflecting the system's operating status and management activities. The preset time window can be configured based on system characteristics and empirical values; for example, it could be several minutes, hours, or days before an anomaly occurs. Business service start / stop records can be obtained by collecting and parsing operating system or application server logs through, for example, system log management tools, or through the audit logs of the service management platform. Connection pool parameter adjustment records can be obtained through, for example, the audit logs of a database management system, the configuration management logs of an application server, or configuration file change records stored in a version control system. Permission configuration change records can be obtained through, for example, an identity and access management system, operating system security logs, or the audit logs of a directory service.

[0122] The parameter change characteristics, environmental impact characteristics, and system configuration characteristics of the target-induced node are input into the root cause determination model to obtain the probability distribution of each candidate root cause. In the embodiments of this application, the root cause determination model is a model used to analyze multi-source heterogeneous data and infer the most likely root cause of a failure. The data input to the root cause determination model undergoes preprocessing steps such as data cleaning, standardization, and feature engineering. The probability distribution of candidate root causes output by the root cause determination model means that the model does not directly provide a single definitive root cause, but rather a set of possible root causes and their corresponding probabilities of occurrence, which helps decision-makers assess risks and take appropriate measures. The root cause determination model can be a machine learning-based model, such as support vector machines, random forests, or neural networks, trained using historical failure data; alternatively, it can be a model based on expert systems or rule engines, making judgments through predefined causal rules and reasoning logic; or it can be a model combining Bayesian networks, graph neural networks, and other techniques capable of handling complex causal relationships.

[0123] The root cause of the target-induced node failure is determined based on the probability distribution.

[0124] In the embodiments of this application, based on the probability distribution output by the root cause determination model, the candidate root cause with the highest probability can be selected as the final root cause of the failure, or all candidate root causes with probabilities higher than a certain threshold can be selected as potential root causes for further analysis. For example, the candidate root cause with the highest probability value can be directly selected; or, a confidence threshold can be set, and all candidate root causes with probabilities exceeding the confidence threshold can be included as a list of possible root causes; or, human experience or expert knowledge can be combined to perform a secondary judgment and confirmation on the probability distribution.

[0125] For example, as a specific implementation method, suppose that in a certain cross-regional power system, a certain substation has been located as the starting area through the aforementioned steps, and one of the relays is the target triggering node that first causes an anomaly, and its parameter change is characterized by a sudden drop in output voltage.

[0126] To further determine the root cause of the fault, the system first obtained the environmental impact characteristics of the substation: through sensors deployed in the substation, the real-time temperature at the time of the anomaly was 45℃, which is higher than the normal value, and the humidity was 90%; through the network monitoring system, it was found that the network bandwidth of the substation fluctuated significantly before the anomaly occurred; at the same time, through external information sources, it was identified that the area had been struck by lightning before the anomaly occurred, i.e., external disturbance.

[0127] Furthermore, the system obtained the system configuration characteristics of the substation: by querying the system log, it was found that the control software of the relay had been updated once within 1 hour before the anomaly occurred; by querying the database audit log, it was found that the SCADA system connection pool parameters associated with the relay were adjusted 30 minutes before the anomaly occurred; by querying the access control system log, no recent change records for the control permissions of the relay were found.

[0128] Based on this, the voltage drop parameter change characteristics of the relay, the aforementioned environmental impact characteristics, and the system configuration characteristics are input into a preset root cause determination model. After analysis, the root cause determination model outputs the probability distribution of each candidate root cause. For example, the probability of "bug introduced by relay software update" is 60%, the probability of "lightning strike causing relay hardware damage" is 25%, the probability of "improper adjustment of SCADA system connection pool parameters" is 10%, and the probability of "sensor false alarm" is 5%.

[0129] Ultimately, based on this probability distribution, the system can determine that "the bug introduced by the relay software update" is the root cause of the failure in the target node.

[0130] Through the above technical solution, this application can accurately obtain key environmental and configuration information affecting the determination of the root cause of power system faults, and combine it with the parameter change characteristics of the target inducing node, inputting it into the root cause determination model. Based on this, the root cause determination model comprehensively analyzes multi-source heterogeneous data and provides candidate root causes in the form of probability distributions, thereby improving the accuracy and reliability of fault root cause location. This effectively solves the problem that in complex cross-regional power systems, it is difficult to accurately identify the root cause of faults based solely on local parameter changes.

[0131] In one embodiment, the root cause determination model includes a feature decoupling layer, a spatiotemporal correlation layer, a causal reasoning layer, and a probability calculation layer. The step of "inputting the parameter change features, environmental impact features, and system configuration features of the target inducing node into the root cause determination model to obtain the probability distribution of each candidate root cause" specifically includes: The feature decoupling layer decomposes the parameter change features into trend components and abrupt change components, and extracts the peak time of the abrupt change components. In the embodiments of this application, the root cause determination model can be an ensemble model based on machine learning, such as a hybrid model combining decision trees, support vector machines, and neural networks, which learns the complex mapping relationship between features and root causes by training on historical fault data; or it can be an inference engine based on expert systems and knowledge graphs, which performs logical deduction through predefined causal rules and domain knowledge to infer the root cause of the fault. The feature decoupling layer decomposes the original parameter change features, separating components of different properties so that subsequent analysis can focus more on key information. It can employ signal processing techniques, such as wavelet transform, Fourier transform, or empirical mode decomposition, to decompose time-series data into components of different frequencies or scales, thereby separating trends and abrupt changes; or it can use statistical filtering methods, such as Kalman filtering or moving average filtering, to smooth the data to extract trend components, and identify abrupt change components by the difference between the original data and the smoothed data. Trend components can reflect long-term, slow changes in parameters; abrupt change components can reflect short-term, drastic fluctuations or abnormal jumps in parameters. Separating these two components helps to distinguish between slow system degradation and sudden failures. Peak moments refer to the time when a mutation component reaches its maximum value or the most significant change. Accurately identifying peak moments helps to align with system configuration changes or external events in time, thereby discovering potential causal relationships.

[0132] The temporal correlation between the peak time of the mutation component and the time of operation in the system configuration features, as well as the spatial distribution coupling strength between environmental influence features and mutation components, are calculated through the spatiotemporal correlation layer. In the embodiments of this application, the spatiotemporal correlation layer is used to analyze the temporal and spatial correlations between different features. Temporal correlation calculation can be implemented using cross-correlation functions, dynamic time warping, or sequence pattern mining methods to quantify the synchronicity or lag of different events. Spatial distribution coupling strength calculation can utilize geographic information system data and network topology information, combined with spatial statistical methods or graph neural networks, to assess the degree of mutual influence of feature changes between different regions or nodes. Temporal correlation measures the degree of temporal sequence, synchronicity, or lag between two or more events, used to determine whether there is a temporal correlation between system configuration operations and parameter mutations. Spatial distribution coupling strength measures the degree of mutual influence between feature changes in different geographical locations or logical regions, used to determine whether environmental impact features and parameter mutations are spatially consistent or interact.

[0133] The causal inference layer determines the activation conditions of each candidate root cause based on a preset causal rule base, and calculates and generates a set of causal contribution based on temporal correlation, spatial distribution coupling strength, the degree of deterioration of trend components and the suddenness of mutation components. In the embodiments of this application, the causal inference layer is used to infer the causal relationship between each candidate root cause and the fault based on existing knowledge and extracted correlation information. Specifically, the causal inference layer can be based on a Bayesian network or a causal graph model, constructing a directed acyclic graph between variables and using conditional probability distributions for inference to determine causal relationships and activation conditions; or it can employ a rule-based expert system, using a pre-defined rule set combined with fuzzy logic or deterministic logic to match and infer input features to identify activation conditions and calculate contributions. The causal rule base is a predefined set of rules describing how specific conditions lead to specific results, providing foundational knowledge for causal inference and guiding the model to determine the activation conditions of each candidate root cause. The causal contribution set is a numerical set that quantifies the contribution of each candidate root cause to the observed anomaly, providing input for probability calculations and reflecting the importance of each factor in the causal chain.

[0134] The probability distribution of each candidate root cause is obtained by calculating the probability distribution of each root cause based on the activation conditions and the set of causal contributions through the probability calculation layer.

[0135] In the embodiments of this application, the probability calculation layer is used to quantify the results of causal inference into the probability distribution of each candidate root cause, providing a quantitative basis for the final root cause determination. The probability calculation layer can employ classification algorithms such as logistic regression, Softmax regression, or support vector machines, taking the causal contribution set as input and outputting the probability value of each candidate root cause; or it can be based on a Bayesian inference framework, combining prior probabilities and likelihood functions, and calculating the probability distribution of each candidate root cause by updating the posterior probability. Activation conditions refer to a series of preconditions that must be met for a candidate root cause to be triggered or occur; they are the key output of causal inference, used to determine which potential root causes are active.

[0136] For example, as a specific implementation, the feature decoupling layer is used to decompose the parameter change features into trend components and abrupt change components. This includes: collecting the time series of operating parameters of the target inducing node within a preset observation window before and after the occurrence of the anomaly; decomposing the parameter time series using a time series decomposition method, such as STL decomposition or wavelet decomposition, to obtain trend components and residual components; performing abrupt change detection on the residual components, such as using sliding window analysis of variance or the CUSUM method, to identify abrupt change points that significantly deviate from the normal fluctuation range; extracting the abrupt change components, i.e., the magnitude and duration of parameter changes near each abrupt change point; and identifying the peak time of the abrupt change components, i.e., the time when the parameter change magnitude is the largest.

[0137] The spatiotemporal correlation layer is used to calculate the correlation strength between abrupt change features and external factors, including calculating temporal correlation and spatial distribution coupling strength. An example is provided below: The calculation involves determining the temporal correlation between the peak of the mutation and the timing of the scheduling operation; and calculating the spatial coupling strength between the target-induced node location and the environmental influencing factors. Specifically, the temporal correlation can be calculated based on the deviation between the time difference and the typical response delay, while the spatial coupling strength can be calculated based on the attenuation model of geographical distance and the intensity of environmental factors. Furthermore, those skilled in the art can reproduce the calculation process based on common knowledge and available information, taking into account the needs of actual scenarios, and therefore will not be elaborated upon here.

[0138] The causal inference layer determines the activation conditions of each candidate root cause based on a pre-defined causal rule base and calculates the causal contribution, including: Construct a causal rule library, which stores multiple causal rules, each of which includes the root cause type, activation condition, and feature combination pattern; For each candidate root cause, check whether its activation conditions are met. This can be achieved by calculating the degree of degradation of the trend component and the burst strength of the mutation component. The degree of degradation can be calculated, for example, using slope analysis or deviation from the normal baseline, while the burst strength can be calculated, for example, using the rate of change or peak amplitude.

[0139] If all activation conditions are met, the candidate root cause will be activated. For activated candidate root causes, the causal inference layer will further calculate their causal contribution and generate a set of causal contributions.

[0140] Finally, the probability calculation layer calculates the probability distribution of each candidate root cause based on the activation conditions and causal contribution, including: calculating the initial probability based on the causal contribution, correcting the prior probability by combining the actual scenario and historical experience, and then normalizing the adjusted probabilities of all candidate root causes to obtain the final probability distribution.

[0141] Through the synergy of the above four-layer architecture, the root cause determination model can comprehensively consider the temporal characteristics of parameter changes, environmental factors, system configuration factors, and historical experience knowledge, and infer the root cause of the failure from multiple dimensions, thereby improving the accuracy of the determination.

[0142] Furthermore, the causal rule base can be constructed by combining historical fault data mining with expert knowledge labeling, specifically including: Historical fault sample sets were collected, and the characteristic trend components and mutation component sequences corresponding to the occurrence of each type of root cause were extracted. The Apriori association rule algorithm was used to mine frequent itemsets between root cause types and feature combination patterns. Support thresholds and confidence thresholds were set, and feature combinations that met the threshold conditions were selected as activation conditions for the corresponding root causes. The determination thresholds of activation conditions were modified based on the experience of domain experts, and each causal rule was labeled with root cause type, activation condition threshold, and feature combination pattern, and stored in the causal rule library.

[0143] Additionally, as an example, the causal contribution can be calculated using a weighted fusion of the degree of degradation and the intensity of suddenness, with the following formula: ,in, For the first The causal contribution of each candidate root cause; For the first The degree of degradation of the trend component corresponding to each candidate root cause, with a value range of 0 to 1, is obtained by normalizing the ratio of the current slope value to the degradation threshold. For the first The burst intensity of the corresponding mutation component of each candidate root cause, with a value ranging from 0 to 1, is obtained by normalizing the ratio of the current peak amplitude to the burst threshold. , All are contribution weight coefficients, satisfying And all are greater than 0.

[0144] Through the above technical solution, this application can decompose the parameter change characteristics of the target inducing node, effectively distinguish between slow system degradation and sudden anomalies, and deeply explore the spatiotemporal correlation between parameter mutations and system configuration operations and environmental influences. Based on this, combined with a pre-set causal rule base, multi-dimensional deep causal reasoning can be performed to quantify the contribution of each candidate root cause to the fault, ultimately providing the root cause of the fault in the form of a probability distribution. This improves the accuracy and reliability of fault root cause localization, effectively solving the technical problems of low efficiency and poor accuracy in root cause localization caused by the difficulty in effectively integrating multi-source heterogeneous data and accurately identifying causal relationships in complex power business information systems.

[0145] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0146] In one embodiment, a cross-regional power system fault root cause localization system is provided, which corresponds one-to-one with the cross-regional power system fault root cause localization method described in the above embodiments. The cross-regional power system fault root cause localization system includes: The path acquisition module is used to receive abnormal alarm information from multiple power business information systems in real time, extract abnormal feature data, and match the abnormal feature data with an abnormal association rule base to obtain the regional propagation path. The node determination module is used to trace back along the regional propagation path and identify the starting area of ​​the regional propagation path and the target initiating node that caused the earliest anomaly in the area. The feature acquisition module is used to determine the region where the starting point is located as the starting region and to acquire the parameter change features of the target inducing node at the time of the anomaly occurrence. The root cause determination module is used to acquire the environmental impact characteristics and system configuration characteristics of the starting area, and combine them with parameter change characteristics to determine the root cause of the target inducing node's failure through a preset root cause determination model.

[0147] For specific limitations regarding a cross-regional power system fault root cause localization system, please refer to the limitations of a cross-regional power system fault root cause localization method described above, which will not be repeated here. Each module in the aforementioned cross-regional power system fault root cause localization system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0148] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 2 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database is used for data storage, data processing, and data analysis. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a cross-regional power system fault root cause localization method.

[0149] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0151] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for locating the root cause of faults in a cross-regional power system, involving multiple power business information systems, characterized in that, The method has a pre-established abnormal association rule base and includes the following steps: It receives abnormal alarm information from multiple power business information systems in real time, extracts abnormal feature data, and matches the abnormal feature data with an abnormal association rule base to obtain the regional propagation path; Tracing back along the regional propagation path, we can identify the starting area of ​​the regional propagation path and the earliest target triggering node in the region where anomalies occur. The region where the starting point is located is determined as the starting region, and the parameter change characteristics of the target inducing node at the time of the anomaly are obtained; The environmental impact characteristics and system configuration characteristics of the starting area are obtained, and combined with the parameter change characteristics, the root cause of the target inducing node is determined by a preset root cause determination model.

2. The cross-regional power system fault root cause localization method according to claim 1, characterized in that, The construction of the anomaly association rule base includes: Collect operation log data of the power business information system and obtain the associated characteristic data of the power business information system, including topological connection relationship and business service linkage logic; Based on the topological connection relationship, identify cross-regional data interfaces and key nodes, and based on the business service linkage logic, identify cross-regional synchronization component associations and business scheduling associations; Establish the association mapping relationship of the power business information system based on the relationship between data interfaces, key nodes, synchronization components and business scheduling; An anomaly propagation knowledge graph is constructed based on runtime log data, associated characteristic data, and associated mapping relationships, and an anomaly association rule base is established based on the anomaly propagation knowledge graph.

3. The cross-regional power system fault root cause localization method according to claim 2, characterized in that, The steps of constructing an anomaly propagation knowledge graph based on runtime log data, related characteristic data, and related mapping relationships, and establishing an anomaly association rule base based on the anomaly propagation knowledge graph, specifically include: The business services, data interfaces, and nodes in each power business information system are treated as entity nodes, and the association mapping relationship is treated as the association edge between entity nodes. Based on historical anomaly records in the operation log data, determine the anomaly propagation path and anomaly propagation sequence between each entity node; Construct an anomaly propagation knowledge graph based on anomaly propagation paths and anomaly propagation sequences; Extract abnormal feature patterns and regional transmission rules from the abnormal transmission knowledge graph, generate abnormal association rules, and establish an abnormal association rule library.

4. The cross-regional power system fault root cause localization method according to claim 3, characterized in that, The steps for constructing an anomaly propagation knowledge graph based on the anomaly propagation path and anomaly propagation sequence specifically include: Based on the anomaly propagation path, directed transmission edges are established between entity nodes, and timestamp and propagation delay attributes are assigned to the directed transmission edges according to the anomaly propagation sequence. Calculate the frequency of occurrence of each directed transmission edge in the historical anomaly records, and use the frequency of occurrence as the weight attribute of the corresponding directed transmission edge; Identify cross-regional transmission edges, and label the source and target regions connected by the cross-regional transmission edges to generate region crossing identifiers; An anomaly propagation knowledge graph is constructed based on directed propagation edges and entity nodes that have timestamp attributes, propagation delay attributes, weight attributes, and region crossing identifiers. Store the anomaly propagation knowledge graph as a graph data structure and create node indexes and edge indexes.

5. The cross-regional power system fault root cause localization method according to claim 4, characterized in that, The steps of extracting abnormal feature patterns and regional transmission rules from the abnormal transmission knowledge graph, generating abnormal association rules, and establishing an abnormal association rule base specifically include: Path tracing is performed on the directed transmission edges in the anomaly transmission knowledge graph to extract the transmission chains from a single source node to multiple target nodes, and the transmission chains are sorted in descending order of weight attributes to obtain a set of transmission chains. Identify conduction chains with the same topological structure from the set of conduction chains, cluster the identified conduction chains into conduction chain clusters, and generate abnormal feature patterns based on the conduction chain clusters; Identify transmission chains associated with region crossing markers and obtain the region transmission characteristics between the source region and the target region corresponding to the region crossing markers; Triggering conditions and timing constraints are set based on abnormal feature patterns, and regional propagation priorities are set based on regional propagation characteristics to generate abnormal association rules; The anomaly association rules are associated with the corresponding anomaly feature patterns and regional propagation features and stored in the anomaly association rule library.

6. The cross-regional power system fault root cause localization method according to claim 1, characterized in that, The steps of receiving abnormal alarm information from multiple power business information systems in real time, extracting abnormal feature data, and matching the abnormal feature data using an abnormal association rule base to obtain the regional propagation path specifically include: The system receives abnormal alarm information reported by various power business information systems in real time, and extracts abnormal feature data based on the abnormal alarm information. The abnormal feature data includes business node identifier, time of abnormal occurrence, and abnormal type. Based on the time of occurrence of the anomaly, the anomaly feature data are arranged in chronological order to generate a time-series anomaly sequence; Based on the time-series anomaly sequence, business nodes and regional distribution are extracted, and corresponding anomaly association rules are retrieved from the anomaly association rule base based on the business nodes; Calculate the matching degree between the time-series abnormal sequence and the abnormal association rule, select the abnormal association rule with the highest matching degree, and determine the regional propagation path based on the selected abnormal association rule.

7. The cross-regional power system fault root cause localization method according to claim 6, characterized in that, The steps of calculating the matching degree between the time-series anomaly sequence and the anomaly association rule, selecting the anomaly association rule with the highest matching degree, and determining the region propagation path based on the selected anomaly association rule specifically include: The system counts the number of business nodes in a time-series anomaly sequence that match the business services associated with the anomaly association rules, and calculates the business matching rate based on the number of matches. Calculate the time interval between the occurrence times of adjacent anomalous feature data in the time-series anomaly sequence, and compare the time interval with the propagation delay attribute of the corresponding directed propagation edge in the anomaly association rule to obtain the time-series deviation; Determine whether the region crossing order involved in the time-series anomaly sequence is consistent with the region crossing identifier associated with the anomaly association rule to obtain a consistency determination result; The matching degree between the time-series abnormal sequence and the abnormal association rule is calculated based on the business matching rate, time-series deviation, and consistency judgment results. Select the anomaly association rule with the highest matching degree, and determine the regional propagation path based on the selected anomaly association rule.

8. The cross-regional power system fault root cause localization method according to claim 1, characterized in that, The step of obtaining the environmental impact characteristics and system configuration characteristics of the starting area, and combining them with parameter change characteristics, to determine the root cause of the target inducing node's failure through a preset root cause determination model, specifically includes: The environmental impact characteristics of the starting area at the time of the anomaly are obtained, including temperature conditions, humidity conditions, network bandwidth fluctuations, and external disturbances. The system configuration features of the starting region within a preset time window before the occurrence of the anomaly are obtained. The system configuration features include business service start / stop records, connection pool parameter adjustment records, and permission configuration change records. The parameter change characteristics, environmental impact characteristics, and system configuration characteristics of the target-induced node are input into the root cause determination model to obtain the probability distribution of each candidate root cause. The root cause of the target-induced node failure is determined based on the probability distribution.

9. The cross-regional power system fault root cause localization method according to claim 8, characterized in that, The root cause determination model includes a feature decoupling layer, a spatiotemporal correlation layer, a causal inference layer, and a probability calculation layer. The step of inputting the parameter change features, environmental impact features, and system configuration features of the target inducing node into the root cause determination model to obtain the probability distribution of each candidate root cause specifically includes: The feature decoupling layer decomposes the parameter change features into trend components and abrupt change components, and extracts the peak time of the abrupt change components. The temporal correlation between the peak time of the mutation component and the time of operation in the system configuration features, as well as the spatial distribution coupling strength between environmental influence features and mutation components, are calculated through the spatiotemporal correlation layer. The causal inference layer determines the activation conditions of each candidate root cause based on a preset causal rule base, and calculates and generates a set of causal contribution based on temporal correlation, spatial distribution coupling strength, the degree of deterioration of trend components and the suddenness of mutation components. The probability distribution of each candidate root cause is obtained by calculating the probability distribution of each root cause based on the activation conditions and the set of causal contributions through the probability calculation layer.

10. A cross-regional power system fault root cause localization system, characterized in that, include: The path acquisition module is used to receive abnormal alarm information from multiple power business information systems in real time, extract abnormal feature data, and match the abnormal feature data with an abnormal association rule base to obtain the regional propagation path. The node determination module is used to trace back along the regional propagation path and identify the starting area of ​​the regional propagation path and the target initiating node that caused the earliest anomaly in the area. The feature acquisition module is used to determine the region where the starting point is located as the starting region and to acquire the parameter change features of the target inducing node at the time of the anomaly occurrence. The root cause determination module is used to acquire the environmental impact characteristics and system configuration characteristics of the starting area, and combine them with parameter change characteristics to determine the root cause of the target inducing node's failure through a preset root cause determination model.