A power alarm tracing method based on a knowledge graph and a large model
Patent Information
- Application Number
- CN202611047963.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-29
AI Technical Summary
在告警数量较多的情况下,容易因视觉信息过载而遗漏关键告警或误判故障传播关系,从而降低故障处理效率,甚至可能导致事故进一步扩大
[0023]本发明实施例的具有以下有益效果:。
Smart Images

Figure CN122840070A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power fault diagnosis technology, and in particular to a power alarm tracing method based on knowledge graphs and large models. Background Technology
[0002] With the continuous improvement of automation levels in large hydropower stations, substations, and integrated energy stations, power monitoring systems can monitor the real-time operating status of the power grid and equipment, and generate corresponding alarm information when equipment abnormalities, protection actions, or power grid faults occur. These alarm messages typically originate from multiple monitoring systems such as SCADA systems, integrated automation systems, and protection information management systems, reflecting changes in equipment operating status, protection actions, and fault development processes, providing a basis for on-duty personnel to conduct fault analysis and accident handling.
[0003] However, in actual operation, when abnormal conditions such as power grid fluctuations, busbar faults, equipment short circuits, or protection activation occur, a large number of devices will generate related alarm messages in a very short period of time. Different monitoring systems will also simultaneously send a large number of original alarm messages, resulting in a massive alarm storm that occurs in a short period of time. These alarms include both core alarms that directly reflect the cause of the fault and numerous derivative alarms generated by equipment linkage, protection actions, status changes, and signal transmission. On-duty personnel need to quickly identify the true source of the fault from a large number of interconnected alarm messages, posing a significant challenge to the safe and stable operation of the power plant.
[0004] Existing technologies typically employ expert systems or rule-based alarm processing methods, which involve manually pre-establishing alarm filtering rules, alarm correlation rules, and fault diagnosis rules to filter, converge, and analyze massive amounts of alarms. However, these solutions heavily rely on manually writing and maintaining rules. When new equipment is added to the power plant, equipment is upgraded, operating modes are adjusted, or protection configurations change, a large number of rules need to be modified simultaneously. This results in a large workload and high maintenance costs for the rule base, and conflicts or omissions can easily occur between different rules, affecting the accuracy of alarm analysis results.
[0005] On the other hand, when faced with complex faults, multi-device linkage faults, or multiple protection actions occurring simultaneously, traditional rule-based matching methods often struggle to accurately understand the semantic relationships between different alarms. Mechanical matching based solely on preset rules can easily lead to alarm convergence failures, misjudgments of fault causes, or inaccurate determination of the root cause of the fault. Furthermore, when actual fault scenarios exceed the coverage of existing rules, current technologies struggle to adapt to new fault types and operating conditions.
[0006] Furthermore, most existing technologies can only display or simply categorize alarm information, lacking the ability to analyze fault propagation by combining it with the topological relationships of power plant equipment. On-duty personnel typically need to analyze and compare each alarm individually, combining numerous indicator light alarms, primary electrical wiring diagrams, secondary protection logic diagrams, and their own operational experience, to deduce the possible location and cause of the fault. With a large number of alarms, visual information overload can easily lead to overlooking critical alarms or misjudging fault propagation relationships, thereby reducing fault handling efficiency and potentially causing the accident to escalate further.
[0007] Meanwhile, existing monitoring systems typically only passively display alarm information, lacking the ability to comprehensively analyze faults by combining equipment topology, fault propagation paths, and operating procedures. They are unable to proactively output targeted fault analysis results and structured handling suggestions. Therefore, on-duty personnel still need to rely on their personal experience to consult relevant operating procedures and formulate handling plans, which not only increases the workload of fault handling but also easily affects the efficiency of accident handling due to differences in experience in complex fault scenarios.
[0008] Therefore, how to intelligently semantically analyze alarm information, accurately locate the root cause of faults, and automatically generate fault analysis results and handling suggestions based on the massive amounts of correlated alarms generated by power monitoring systems has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0009] The present invention aims to at least partially solve one of the technical problems in the related art.
[0010] The main objective of this invention is to provide a power alarm tracing method based on knowledge graphs and large models.
[0011] The second objective of this invention is to provide an electronic device.
[0012] A third objective of this invention is to provide a non-transitory computer-readable storage medium.
[0013] To achieve the above objectives, a first aspect of the present invention proposes a power alarm tracing method based on knowledge graphs and large models, comprising:
[0014] S1. Obtain the original alarm messages generated by the power monitoring system, and perform time alignment and format standardization processing on the original alarm messages to form standardized alarm data. S2. Input the standardized alarm data into a pre-deployed large language model, perform semantic parsing on the standardized alarm data, extract the device entity, action entity and status entity corresponding to the alarm, and form an alarm entity set; S3. Based on the alarm entity set, map each device entity to the corresponding device node in the pre-constructed power knowledge graph; S4. Using the device node as the starting node, perform a reverse graph traversal along the physical connection relationship and protection logic relationship of the device in the power knowledge graph to determine the common ancestor node corresponding to multiple alarms as candidate root cause nodes. S5. Determine the root cause node of the fault based on the candidate root cause node and its corresponding graph traversal path. S6. Input the root cause node of the fault and the corresponding graph traversal path into the large language model to generate fault analysis results and handling suggestions.
[0015] In one embodiment of the present invention, the standardization process in step S1 includes a unified alarm time format, a unified device identifier format, and a unified alarm field format.
[0016] In one embodiment of the present invention, in step S2, the large language model performs semantic parsing on the standardized alarm data, including identifying the device name, protection action, fault status and device operating status in the alarm, and removing derivative alarms that are not related to fault propagation.
[0017] In one embodiment of the present invention, in step S3, the power knowledge graph includes multiple device nodes and relational edges connecting each device node, wherein the relational edges include at least physical connection relationships and protection logic relationships between devices.
[0018] In one embodiment of the present invention, step S4, the reverse graph traversal includes searching from each device node to its upstream node, and determining the common ancestor node of each search path as a candidate root cause node.
[0019] In one embodiment of the present invention, in step S5, the root cause node of the fault is determined according to the length of the graph traversal path corresponding to the candidate root cause node and the number of associated alarms.
[0020] In one embodiment of the present invention, in step S6, the large language model generates a fault analysis result containing the fault cause, fault propagation process and handling suggestions based on the fault root cause node, graph traversal path and pre-stored operation procedure.
[0021] To achieve the above objectives, a second aspect of this application provides an electronic device, including a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, for implementing the method described in the first aspect embodiment.
[0022] To achieve the above objectives, a third aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect.
[0023] The embodiments of the present invention have the following beneficial effects:
[0024] 1. This invention applies the semantic parsing capabilities of a large language model to the raw alarm messages generated by a power monitoring system. It automatically extracts equipment entities, action entities, and status entities from standardized alarm data and identifies core alarms and derived alarms based on alarm semantics. Compared to processing methods that rely on manually pre-written alarm filtering and association rules, this invention can perform unified semantic parsing on alarm messages from different manufacturers, with different expressions, and under different fault scenarios. This reduces missed and false alarms caused by an increase in the number of rules, rule conflicts, or untimely rule updates, lowers the workload of manual configuration and maintenance of alarm processing rules, and improves adaptability to new alarm types and different text expressions.
[0025] 2. This invention performs time alignment and format standardization on original alarm messages from different sources, and utilizes a locally deployed large language model to perform semantic noise reduction on massive alarm messages. This transforms a large number of duplicate alarms, derivative alarms, and auxiliary status alarms generated at the time of a fault into a smaller set of alarm entities with clear semantics. Based on this, alarm entities are mapped to device nodes in the power knowledge graph, and a reverse graph traversal is performed. This enables the analysis of the correlation between multiple alarms in a shorter time, reducing the need for on-duty personnel to check alarms one by one, consult wiring diagrams and protection logic, thereby improving the efficiency of fault analysis and the timeliness of emergency response in alarm storm scenarios.
[0026] 3. This invention utilizes a power knowledge graph constructed based on the primary electrical wiring diagram and secondary protection logic of a power plant to perform joint analysis of the physical topology and protection logic relationships of the device nodes corresponding to alarms. By traversing the reverse graph upstream from multiple alarm-corresponding device nodes, the common ancestor node of multiple alarms is determined. The root cause node is then identified based on the number of alarms associated with the candidate root cause node and the length of the corresponding graph traversal path. This allows multiple discrete alarms to be converged into a root cause device with clear topological basis, avoiding judgment of fault causes based solely on a single alarm or manual experience, and improving the consistency and accuracy of fault root cause localization.
[0027] 4. This invention, while identifying the root cause node of a fault, retains the graph traversal path from the root cause node to each alarm-corresponding device node, and forms a fault propagation topology based on the physical connection relationships of the devices and the protection logic relationships. The root cause node and its corresponding fault propagation path are differentiated and highlighted on a large visualization screen in the central control center, enabling on-duty personnel to intuitively obtain the location of the fault, the range of affected equipment, and the propagation process of protection actions. This reduces the cognitive processing required for manual comparison of numerous alarm messages, electrical wiring diagrams, and protection logic diagrams, improving the understandability of the fault analysis results.
[0028] 5. This invention inputs the root cause nodes of the fault, the graph traversal path, and related alarm information into a large language model, and combines this with pre-stored power plant operation procedures to generate a root cause analysis report including the fault cause, the fault propagation process, and handling suggestions. This allows the node relationships and graph traversal results in the graph database to be converted into natural language content that is easy for on-duty personnel to understand, and ensures that the handling suggestions correspond to the actual faulty equipment, fault status, and power plant operation procedures, providing auxiliary basis for centralized control personnel to conduct fault verification, equipment isolation, and operation mode adjustments.
[0029] 6. This invention adopts a large language model deployed locally and privately, and completes alarm semantic parsing and root cause analysis report generation through an edge server equipped with a GPU or NPU. This enables alarm data, power plant equipment topology and operating procedures to be processed in the internal network or controlled network environment of the power plant. While meeting the real-time processing requirements, it reduces the data security risks caused by the transmission of power operation data to external networks.
[0030] 7. The power knowledge graph used in this invention can update equipment nodes, node attributes, and relationship edges according to changes in power plant equipment commissioning, maintenance, renovation, and wiring methods, exhibiting good scalability. Furthermore, the power knowledge graph and fault root cause analysis results can be correlated with a digital twin platform, enabling fault alarms, equipment topology, and equipment operating status to be displayed in a unified equipment model, providing a structured data foundation for subsequent equipment status analysis, fault trend analysis, and predictive maintenance. Attached Figure Description
[0031] The above-described and additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which: Figure 1 The flowchart illustrates a power alarm tracing method based on knowledge graphs and large models, provided in this embodiment of the invention. Detailed Implementation
[0032] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0034] The following describes a power alarm tracing method based on knowledge graphs and large models according to an embodiment of the present invention, with reference to the accompanying drawings.
[0035] Example 1 This embodiment provides a power alarm tracing method based on knowledge graphs and large models. For example... Figure 1 As shown, the method includes the following steps: S1. Obtain the original alarm messages generated by the power monitoring system, and perform time alignment and format standardization processing on the original alarm messages to form standardized alarm data.
[0036] In this embodiment, step S1 is used to obtain the original alarm messages generated by the power monitoring system, and complete the time alignment and format standardization processing to form standardized alarm data that can be used for semantic parsing by the subsequent large language model.
[0037] Specifically, the data source in this embodiment is real-time alarm information generated by the power plant monitoring system. This monitoring system includes, but is not limited to, SCADA systems, integrated automation systems, protection information management systems, and other monitoring systems capable of generating equipment operation alarm information. The raw alarm messages generated by these systems are uniformly received through an alarm data access unit, which includes a data cleaning module and a protocol gateway. The protocol gateway is used to adapt to the data communication protocols of different monitoring systems, enabling unified access to alarm messages from different sources. The data cleaning module is used to preprocess the received raw alarm messages, providing a unified data format for subsequent analysis.
[0038] Furthermore, since alarm messages generated by different monitoring systems differ in data structure, time format, device naming conventions, and field content, it is necessary to first perform time alignment processing on the original alarm messages. In this embodiment, time alignment processing refers to converting and correcting the timestamps of alarm messages from different sources based on a unified time benchmark, so that all alarm information is represented using a unified time format. This ensures that alarm information generated by different devices during the same fault process can be arranged in the actual order of occurrence, avoiding deviations in subsequent fault propagation relationship analysis due to clock skew between systems.
[0039] After time alignment, the original alarm messages are standardized in format. Specifically, alarm fields from different manufacturers and systems are uniformly mapped, converting data fields with different names but the same meaning into unified field names, and standardizing device identifier formats, alarm type formats, and status description formats. Specifically, device identifiers are uniformly represented by predefined device codes or device names, ensuring that different naming methods for the same device in different systems can correspond to a unique device identifier; alarm types are uniformly converted to preset standard types; and status descriptions are uniformly represented using standardized expressions, thereby eliminating data differences caused by different naming rules between different monitoring systems.
[0040] Furthermore, the data cleaning module performs integrity checks on the original alarm messages, processing alarm data that is missing key fields, has abnormal field formats, or is uploaded repeatedly. Specifically, data missing non-key fields can be supplemented according to preset rules, alarm messages with duplicate content are deduplicated, and data with obvious errors or that cannot be identified is filtered out to ensure that the data entering subsequent processing flows has integrity and consistency.
[0041] After the above processing, each alarm message forms a unified data structure. Preferably, each standardized alarm data message includes at least a unified timestamp, device identifier, alarm type, alarm content, and device operating status. Those skilled in the art will understand that, without affecting the purpose of this invention, the standardized data fields can be supplemented with other information reflecting the device operating status according to actual application needs.
[0042] Through the aforementioned time alignment and format standardization processes, a large number of heterogeneous original alarm messages from different monitoring systems are converted into standardized alarm data in a unified format. On the one hand, this ensures that the subsequent large language model can perform semantic parsing based on a unified data structure, improving the accuracy of semantic recognition. On the other hand, it also ensures that device entities in the subsequent knowledge graph can establish a unique correspondence with device identifiers in the standardized alarm data, providing unified data input for the semantic parsing, extraction of device entities, action entities, and status entities from the standardized alarm data using the large language model in step S2. Simultaneously, the completion of time sequence and format standardization also provides a reliable data foundation for subsequent fault propagation path analysis and root cause tracing based on the knowledge graph.
[0043] S2. Input the standardized alarm data into a pre-deployed large language model, perform semantic parsing on the standardized alarm data, and extract the device entity, action entity and status entity corresponding to the alarm to form an alarm entity set.
[0044] In this embodiment, step S2 is based on the standardized alarm data obtained in step S1. The standardized alarm data is input into a pre-deployed large language model to perform semantic parsing on the standardized alarm data, extract the device entity, action entity and status entity corresponding to the alarm, form an alarm entity set, and provide structured input data for subsequent knowledge graph entity mapping.
[0045] Specifically, the large language model inference unit in this embodiment includes a lightweight power domain large language model deployed on an edge server. The edge server is equipped with an AI computing card, which can employ AI computing acceleration devices such as GPUs or NPUs to improve the inference efficiency of the large language model, thereby meeting the real-time processing needs of massive alarm data when power plant faults occur. Preferably, the large language model adopts a local private deployment method to avoid transmitting alarm data to external networks during operation, thus ensuring real-time inference while improving the security of power plant operation data.
[0046] Furthermore, the large language model is a lightweight model fine-tuned from a power industry corpus. It has pre-learned professional knowledge such as power equipment names, protection device names, power grid operation terminology, typical fault types, and equipment operation procedures, thereby enabling it to accurately understand the content of professional alarm messages generated by the power monitoring system. Those skilled in the art will understand that the specific model type of the large language model does not constitute a limitation of the present invention, as long as it can achieve semantic understanding and entity extraction of power alarm messages.
[0047] In this embodiment, the standardized alarm data generated in step S1 is sent to the large language model according to a preset input format. The large language model first performs natural language understanding on the alarm text, analyzes the semantic information in the alarm text, and identifies the device object, device action, and device operating status expressed by the alarm content. Among them, the device entity is used to represent the specific device that triggered the alarm; the action entity is used to represent the action event or protection action that occurred on the device; and the status entity is used to represent the current operating status or fault status of the device.
[0048] Furthermore, since the same fault may trigger a large number of related alarms, including both core alarms that directly reflect the cause of the fault and derivative alarms generated by protection actions, signal linkages, or changes in equipment status, the large language model performs semantic noise reduction on standardized alarm data while completing entity recognition. Specifically, the large language model combines power industry expertise to determine the semantic relationships between alarms, identify the core actions in the alarm content that truly reflect the essence of the fault, and filter out derivative alarm information that is not directly related to fault propagation analysis. For example, for alarms containing auxiliary action descriptions such as "channel abnormality" or "plate retraction," the large language model can identify them as derivative actions and not use them as the main basis for subsequent root cause analysis; while for alarms that directly reflect the equipment fault status, such as "PT fuse blown" or "zero-sequence overcurrent protection action," they are identified as core action entities and retained for subsequent knowledge graph reasoning.
[0049] Preferably, for alarm messages that have different descriptions but represent the same equipment or the same action, the large language model further performs semantic normalization processing. For example, alarm messages generated by equipment from different manufacturers may use different names to describe the same equipment or the same protection action. Based on its knowledge of the power industry, the large language model maps different text expressions to consistent equipment entities, action entities, and status entities, improving the consistency and accuracy of subsequent knowledge graph entity mapping.
[0050] After completing the semantic parsing of all standardized alarm data, the large language model outputs the corresponding structured parsing results. Each parsing result includes at least a device entity, an action entity, and a status entity, and retains the unified timestamp and device identification information of the corresponding alarm. Subsequently, all the parsed structured entities are unified into an alarm entity set. Each alarm entity maintains a correspondence with the original alarm message to accurately locate the corresponding device node during subsequent knowledge graph mapping.
[0051] Through the above steps, this embodiment converts the standardized alarm data output in step S1 into a set of alarm entities with clear semantic meaning, realizing semantic understanding and noise reduction of massive alarm messages. On the one hand, the large language model can automatically eliminate a large number of redundant alarms generated by equipment linkage, improving the efficiency of subsequent analysis; on the other hand, the extracted equipment entities, action entities, and status entities all adopt a unified expression method, providing an accurate and consistent data foundation for mapping each equipment entity to the corresponding equipment node in the pre-constructed power knowledge graph in step S3, thereby ensuring that the subsequent graph traversal and tracing process can be carried out around the analysis of the real faulty equipment.
[0052] S3. Based on the alarm entity set, map each device entity to the corresponding device node in the pre-constructed power knowledge graph.
[0053] In this embodiment, step S3 is based on the alarm entity set formed in step S2. According to the device entity information contained in each alarm entity in the alarm entity set, each device entity is mapped to the corresponding device node in the pre-built power knowledge graph, so as to establish the association between real-time alarm data and power plant equipment topology and provide a starting node for subsequent reverse graph traversal.
[0054] Specifically, the power knowledge graph is pre-stored in a knowledge graph storage engine. The knowledge graph storage engine includes a graph database server, which stores various equipment nodes within the power plant, their attribute information, and the edges between them, and supports subsequent queries and graph traversal operations on the equipment nodes and their edges.
[0055] In this embodiment, the power knowledge graph is pre-constructed based on the primary electrical wiring diagram and secondary protection logic of the power plant. The primary electrical wiring diagram determines the physical connections between circuit breakers, disconnectors, busbars, transformers, voltage transformers, current transformers, and other primary equipment. The secondary protection logic determines the logical relationships between protection devices, measurement circuits, control circuits, and the protected equipment. These devices and their connections are converted into triples in the form of "entity-relationship-entity," and each triple is written into the graph database server, thereby forming a power knowledge graph that reflects the physical topology of the power plant equipment and the topology of the protection logic.
[0056] For example, the relationship "10kV circuit breaker - belongs to - 10kV busbar" can be stored as a ternary in the power knowledge graph. Here, "10kV circuit breaker" and "10kV busbar" correspond to different equipment nodes, and "belongs to" corresponds to the relationship edge connecting the two equipment nodes. Similarly, based on the actual equipment configuration and protection setup of the power plant, the connection relationships between voltage transformers and busbars, the protection relationships between protection devices and protected equipment, the attribution relationships between circuit breakers and corresponding bays, and the power supply relationships between upstream and downstream equipment can be established.
[0057] Furthermore, the power knowledge graph includes multiple device nodes and relational edges connecting these device nodes. The device nodes represent equipment objects within the power plant capable of participating in alarm generation, fault propagation, or protection actions, and at least include bus nodes, circuit breaker nodes, disconnector nodes, transformer nodes, voltage transformer nodes, current transformer nodes, and protection device nodes. Depending on the actual equipment configuration of the power plant, the device nodes may also include feeder nodes, switchgear nodes, protection and control unit nodes, and other device nodes related to fault tracing.
[0058] Each device node is assigned a unique node identifier and associated with corresponding device attributes. These attributes may include one or more of the following: device code, standard device name, device type, voltage level, assigned bay, assigned bus, installation location, and device alias. The node identifier corresponds to the standardized device identifier from step S1, and the device alias records different names used by different monitoring systems or equipment manufacturers for the same device, thus improving the adaptability of device entity mapping.
[0059] The relationship edges include at least physical device connection relationships and protection logic relationships. Physical device connection relationships represent the electrical connections, power supply affiliations, or upstream / downstream connections between devices as determined by the primary electrical wiring diagram. Protection logic relationships represent the protection associations between protection devices and protected equipment as determined by secondary protection logic, or the logical associations between a protection action and a corresponding equipment state change. By simultaneously storing physical device connection relationships and protection logic relationships, the power knowledge graph can reflect both the physical propagation path of faults between primary devices and the logical propagation path of fault-induced protection device actions and related signal changes.
[0060] When mapping device entities, the alarm entities are first read one by one from the alarm entity set formed in step S2, and the device entity, unified timestamp, device identifier, action entity, and status entity contained in each alarm entity are obtained. Then, using the device identifier corresponding to the device entity as the primary query condition, device nodes with the same node identifier are searched in the power knowledge graph. When a unique matching device node is found, a mapping relationship is established between the device entity and the device node.
[0061] In some implementations, the alarm entity does not contain a directly usable unique device identifier, or the device name in the alarm message does not completely match the standard device name in the power knowledge graph. In such cases, the corresponding device node can be determined from the power knowledge graph by matching the normalized device name obtained in step S2 with attributes such as device type, voltage level, bay to which it belongs, installation location, and device alias.
[0062] Furthermore, when multiple candidate device nodes are obtained based on the device name, the matching range can be narrowed down using other device attributes contained in the alarm entity. Specifically, the consistency of device type, voltage level, bay, and bus can be compared sequentially, and candidate device nodes with consistent attributes can be identified as target device nodes. For example, when the device name in the alarm entity is "bus PT" and there are multiple bus voltage transformers in the substation, the voltage level, bus number, or bay corresponding to the alarm entity can be combined to determine the voltage transformer node that actually corresponds to the alarm entity from among the multiple candidate device nodes.
[0063] Preferably, when an alarm entity cannot form a unique correspondence with any device node in the power knowledge graph, the alarm entity is marked as an entity to be verified, and its original alarm identifier and semantic parsing result are retained to avoid erroneous mapping interfering with subsequent root cause tracing. For alarm entities that can uniquely identify the corresponding device node, a corresponding entity mapping record is generated. The entity mapping record includes at least the alarm entity identifier, device entity, target device node identifier, action entity, status entity, and unified timestamp.
[0064] In this embodiment, the same device node may correspond to multiple alarm entities. For example, during the same fault process, the same protection device or the same circuit breaker may continuously generate multiple alarms such as protection start-up, protection action, switch tripping, and status change. For multiple alarm entities mapped to the same device node, the action entity, status entity, and unified timestamp of each alarm entity are retained and associated with the same device node. This allows the subsequent graph traversal process to utilize both the topological relationship between device nodes and the alarm action and occurrence time corresponding to that device node to determine the fault propagation process.
[0065] For the core alarms retained after semantic denoising in step S2, their corresponding device entities are preferentially mapped to device nodes in the power knowledge graph. For alarms identified as derived alarms but still having an auxiliary role in fault propagation, their entity mapping results can be retained, and core alarms and derived alarms can be distinguished by alarm type identifiers to assist in verifying the fault propagation path during subsequent graph traversal. Thus, while reducing interference from irrelevant alarms, it is possible to retain device status information related to the fault propagation process.
[0066] After mapping each device entity in the alarm entity set, a set of device nodes corresponding to this fault alarm is obtained. Each device node in the set is associated with at least one alarm entity and retains the action entity, status entity, and unified timestamp of the corresponding alarm entity. This set of device nodes serves as the starting node set for the reverse graph traversal in step S4.
[0067] Through the above steps, the text semantic entities obtained in step S2 are converted into device nodes with clear topological locations in the power knowledge graph, realizing the association between real-time alarm information and the primary electrical wiring structure and secondary protection logic of the power plant. Therefore, subsequent steps can start from the device node that has triggered the alarm, and traverse backwards along the physical connection relationships and protection logic relationships of the devices in the power knowledge graph to trace the common upstream node of multiple alarms and determine the candidate root cause node.
[0068] S4. Using the device node as the starting node, perform a reverse graph traversal along the physical connection relationship and protection logic relationship of the device in the power knowledge graph to determine the common ancestor node corresponding to multiple alarms as candidate root cause nodes.
[0069] In this embodiment, step S4 is to obtain the set of device nodes corresponding to the current fault alarm after obtaining the set of device nodes in step S3, and then to take each device node in the set of device nodes as the starting node, and to perform a reverse graph traversal along the physical connection relationship and protection logic relationship of the device in the pre-constructed power knowledge graph to obtain the upstream tracing path corresponding to each starting node, and to determine the common ancestor node corresponding to multiple alarms according to the intersection relationship between different tracing paths, and to take the common ancestor node as the candidate root cause node.
[0070] Specifically, the device node set obtained in step S3 includes device nodes mapped from core alarms, and device nodes corresponding to derivative alarms retained for assisting in verifying fault propagation paths. Each device node is associated with the action entity, status entity, and unified timestamp of the corresponding alarm entity. The graph reasoning algorithm module in the tracing and interaction unit reads the device node set and its associated information, and sets each device node as the starting node for the reverse graph traversal.
[0071] In this embodiment, the reverse graph traversal refers to searching layer by layer from the device node that generated the alarm to the upstream device node that may have triggered the alarm, following the direction opposite to the fault propagation direction. The relationship edges in the power knowledge graph are pre-set with relationship types and connection directions based on the primary electrical wiring diagram and secondary protection logic. The graph reasoning algorithm module determines the upstream neighboring node corresponding to the current device node based on the relationship edge type and connection direction, and adds the upstream neighboring node to the next layer of the node set to be traversed.
[0072] For the physical connection relationships of equipment, the reverse graph traversal is performed by tracing back in the direction of power transmission, the upstream and downstream connection direction of equipment, or the equipment affiliation. For example, when the starting node is a feeder circuit breaker node, the tracing can proceed along its physical connection relationship to the corresponding bus node; when the starting node is a bus node, the tracing can continue to the incoming circuit breaker node, transformer node, or voltage transformer node connected to that bus. In this way, the physical topology path extending from the equipment that generated the alarm to the upstream primary equipment that may have caused the equipment malfunction can be obtained.
[0073] For protection logic relationships, the reverse graph traversal follows the triggering logic of the protection action in reverse order. For example, when the starting node is the equipment node corresponding to the circuit breaker trip, the protection logic relationship can be traced back to the protection device node that triggered the circuit breaker trip; when the starting node is a protection device node, the tracing can continue to the primary equipment node protected by the protection device or the current transformer node or voltage transformer node that provides current and voltage measurement signals to the protection device. Thus, the protection logic path extending from equipment state changes or protection actions to their triggering source can be obtained.
[0074] Furthermore, during the reverse graph traversal, the graph reasoning algorithm module establishes a corresponding set of visited nodes, a set of nodes to be traversed, and a trace path record for each device node. The set of visited nodes records device nodes that have already been traversed, preventing the same node from being searched repeatedly or causing loop traversal due to circular connections in the graph. The set of nodes to be traversed stores upstream device nodes that need to be visited after the current search level. The trace path record stores the device nodes and relational edges traversed from the starting node to each upstream device node.
[0075] Specifically, for any device node that serves as the starting node, the device node is first added to the set of visited nodes, and the physical relationship edges and protection logic relationship edges connected to the device node are queried. Then, based on the connection direction of these relationship edges, relationship edges pointing to the upstream source of the fault are filtered to obtain the upstream neighboring nodes corresponding to the current device node. For upstream neighboring nodes that have not yet been visited, they are added to the set of nodes to be traversed, and the connection relationships between the current device node, its corresponding relationship edge, and the upstream neighboring nodes are recorded, forming a tracing path.
[0076] After completing the node search at the current level, each device node in the set of nodes to be traversed is taken as the new current device node, and its upstream neighboring nodes are queried in the same way. This process is repeated layer by layer until the preset traversal termination condition is reached, thereby obtaining the set of upstream nodes and the corresponding tracing path for each alarm-related device node.
[0077] In this embodiment, the traversal termination condition may include at least one of the following: the current device node does not have an upstream neighbor node that matches the reverse tracing direction; the current tracing path has reached a preset power boundary node or topology boundary node; the current traversal level has reached a preset maximum traversal depth; or the current device node has been recorded in the corresponding visited node set. Those skilled in the art will understand that the maximum traversal depth can be preset according to the power plant topology scale and device level to limit the expansion of irrelevant paths and improve graph traversal efficiency.
[0078] Furthermore, during the reverse graph traversal, the unified timestamp associated with each alarm entity is used to perform time-series verification on the tracing path. Specifically, for adjacent device nodes in the same tracing path, the occurrence time of the alarms associated with the adjacent device nodes is used to determine whether they conform to the time sequence of fault propagation from upstream to downstream devices or the time sequence of protection device activation to action. When the alarm occurrence sequence corresponding to adjacent device nodes is significantly inconsistent with the preset fault propagation logic, the further expansion of that branch can be terminated, or the branch can be marked as a low-association path. By performing time-series verification, the interference of device nodes that are only connected in the topology but have no direct association with the current fault process on the tracing results can be reduced.
[0079] For the device nodes corresponding to the core alarms identified in step S2, a reverse graph traversal is performed first, and their tracing paths are used as the primary basis for determining the common ancestor node. For the device nodes corresponding to derived alarms, their tracing paths are used to verify whether the candidate fault propagation paths are consistent with the equipment state changes and protection action processes. Thus, while avoiding a large number of derived alarms expanding the search scope, auxiliary information reflecting the fault propagation process is retained.
[0080] After completing the reverse graph traversal of each starting node, multiple upstream node sets are obtained. Each upstream node set includes the corresponding starting node itself and one or more upstream device nodes obtained by tracing back from that starting node along the physical connection relationships and protection logic relationships of the devices. The graph inference algorithm module performs intersection analysis on the upstream node sets corresponding to different starting nodes to determine the device nodes that appear simultaneously in multiple upstream node sets.
[0081] In this embodiment, device nodes that are simultaneously located in the upstream tracing paths of two or more starting nodes and can reach the starting nodes through physical device connections or protection logic relationships are identified as common ancestor nodes. The common ancestor node indicates that multiple alarms share a common upstream origin in the power knowledge graph, and therefore may be the source of faults triggering multiple device alarms or protection actions.
[0082] Specifically, when a device node can be connected to multiple alarm-generating device nodes along different downstream paths, it indicates that an anomaly in that device node may simultaneously trigger the multiple alarms through the device's physical topology or protection logic. For example, when multiple feeder devices, protection devices, or circuit breakers generate alarms, the reverse tracing paths may converge to the same bus node; when bus undervoltage, protection action, and related equipment status changes occur simultaneously, the corresponding tracing paths may also converge to the voltage transformer node connected to that bus. The bus node or voltage transformer node can be identified as a common ancestor node.
[0083] Furthermore, when multiple common ancestor nodes exist in the upstream node set of multiple starting nodes, all device nodes that satisfy the common ancestor condition are retained, and the graph traversal path between each common ancestor node and each starting node is recorded. The paths from the common ancestor node to each starting node together constitute the corresponding candidate fault propagation topology, which is used to characterize the propagation process that may cause multiple alarms after the common ancestor node becomes abnormal.
[0084] For upstream nodes that can only be associated with a single alarm device node and cannot explain other alarm device nodes, they will not be considered as the common ancestor node for multiple alarms. For device nodes that can be associated with some alarms but not all alarms, they can be retained based on the number of core alarms associated with them, in order to accommodate situations where some alarms are not included in the device node set during the same fault process due to communication delays, data loss, or semantic noise reduction processing.
[0085] Preferably, when the alarm data is complete and all core alarms have been accurately mapped, the upstream device nodes that can be jointly connected to all core alarm corresponding device nodes are determined as common ancestor nodes. When alarm data is missing, the upstream device nodes that can be jointly connected to a preset number or preset proportion of core alarm corresponding device nodes can be determined as common ancestor nodes. The preset number or preset proportion can be set according to the power plant equipment scale and the completeness of the alarm data.
[0086] Furthermore, for multiple common ancestor nodes located on the same tracing path, their node identifiers, device attributes, associated alarm counts, and path lengths to each starting node are retained. This information collectively constitutes the association information of candidate root cause nodes and serves as the data basis for further determining the root cause node in step S5.
[0087] In a specific application scenario, when a power grid fault occurs, the alarm data access unit receives hundreds of original alarm messages within seconds. After time alignment and format standardization in step S1, and semantic noise reduction and entity extraction in step S2, a set of alarm entities is obtained, including "PT fuse blown," "zero-sequence overcurrent protection activated," and related equipment status changes. In step S3, each equipment entity is mapped to voltage transformer nodes, protection device nodes, circuit breaker nodes, and busbar nodes in the power knowledge graph. In step S4, a reverse graph traversal is performed from these equipment nodes. Each tracing path extends upstream layer by layer along the physical connection relationships and protection logic relationships of the equipment, intersecting at the corresponding busbar node or voltage transformer node. The equipment node corresponding to the intersection of these paths is then identified as a candidate root cause node.
[0088] Through the aforementioned reverse graph traversal process, the multiple device nodes corresponding to hundreds of discrete alarms are transformed into a fault propagation topology with clear upstream and downstream relationships, and a common ancestor node is determined from the tracing paths corresponding to multiple alarms. The common ancestor node and its corresponding graph traversal path together form a set of candidate root cause nodes, which are then used by step S5 to further determine the final root cause node based on the length of the graph traversal path corresponding to the candidate root cause node and the number of associated alarms.
[0089] S5. Determine the root cause node of the fault based on the candidate root cause node and its corresponding graph traversal path.
[0090] In this embodiment, step S5 is to compare each candidate root cause node with the candidate root cause node based on the number of alarms associated with the candidate root cause node and the length of the graph traversal path from the candidate root cause node to each alarm corresponding device node after obtaining the candidate root cause node and the graph traversal path from the candidate root cause node to each alarm corresponding device node, and determine the fault root cause node from the candidate root cause nodes.
[0091] Specifically, each candidate root cause node output in step S4 is associated with a corresponding node identifier, device attributes, associated alarm information, and one or more graph traversal paths. The graph traversal path represents the fault propagation association link between the candidate root cause node and the device node that generated the alarm, based on the physical connection relationship and protection logic relationship between the devices. The associated alarm information includes at least the alarm entity identifier connected to the candidate root cause node, the corresponding device node, the action entity, the status entity, and a unified timestamp.
[0092] In this embodiment, the graph inference algorithm module first counts the number of alarms that can be associated with each candidate root cause node through the corresponding graph traversal path. The number of associated alarms refers to the number of alarm-corresponding device nodes that can be reached from the candidate root cause node along the device physical connection relationship or protection logic relationship recorded in step S4.
[0093] When multiple alarm entities are associated with the same alarm corresponding to a device node, the count can be performed based on the number of alarm entities or the number of device nodes, depending on the actual processing needs. Preferably, to avoid duplicate counting caused by the same device generating multiple alarms consecutively in a short period of time, the number of device nodes corresponding to the alarm is used as the main statistical basis for the number of associated alarms, and all action entities, status entities, and unified timestamps associated with that device node are retained as auxiliary judgment information.
[0094] Furthermore, when counting the number of associated alarms, core alarms and derived alarms can be distinguished. For core alarms identified and retained in step S2, their corresponding device nodes are counted in the number of main associated alarms; for derived alarms used to assist in verifying the fault propagation path, their corresponding device nodes can be used as auxiliary association information. This avoids the excessive influence of a large number of derived alarms with low correlation to the root cause of the fault on the determination of the root cause node.
[0095] In one implementation, when a candidate root cause node can be associated with more core alarm-related device nodes, it indicates that the candidate root cause node can explain more alarm phenomena, and therefore it is more likely to be the root cause node of the fault. Conversely, when a candidate root cause node can only be associated with a small number of alarm-related device nodes, it indicates that its coverage of the current alarm set is low, and it is usually not preferentially identified as the root cause node of the fault.
[0096] After obtaining the number of associated alarms corresponding to each candidate root cause node, the graph traversal path length corresponding to each candidate root cause node is further determined. The graph traversal path length refers to the number of relation edges traversed from the candidate root cause node to the corresponding alarm device node, or the number of device node levels traversed. In this embodiment, a unified path length calculation method is used to compare all candidate root cause nodes to ensure consistency in the judgment results between different candidate root cause nodes.
[0097] Specifically, when there are multiple reachable paths between a candidate root cause node and a single alarm-corresponding device node, the shortest effective path that conforms to the physical connection direction of the device, the protection logic direction, and the alarm occurrence time order is preferably selected, and the length of the shortest effective path is taken as the path length between the candidate root cause node and the corresponding alarm-corresponding device node.
[0098] A valid path refers to a graph traversal path in which all device nodes and relational edges can form a continuous physical connection link or protection logic link, and the alarm occurrence sequence on the path matches the fault propagation process. Paths that do not conform to electrical connection relationships, protection action logic, or time sequence will not be used as the basis for determining the root cause node of the fault.
[0099] When a candidate root cause node is associated with multiple alarm-corresponding device nodes, the path length from the candidate root cause node to each alarm-corresponding device node is determined, and the path length information corresponding to the candidate root cause node is obtained based on each path length. The path length information can be represented as the sum of the path lengths, the average path length, or the longest path length.
[0100] Preferably, in this embodiment, the average path length is used to represent the overall topological distance between the candidate root cause node and the device node corresponding to its associated alarm. The smaller the average path length, the closer the candidate root cause node is to multiple alarm-corresponding device nodes in the power knowledge graph, and its anomaly is more likely to trigger the multiple alarms through shorter physical device connection links or protection logic links.
[0101] Those skilled in the art will understand that, without changing the technical concept of determining the root cause node of a fault based on the path length of graph traversal, the sum of path lengths or other parameters that can characterize the topological distance between the candidate root cause node and the device node corresponding to the alarm can also be used.
[0102] After calculating the number of associated alarms and the length of the graph traversal path, the graph inference algorithm module compares each candidate root cause node. Specifically, it prioritizes candidate root cause nodes that can be associated with a large number of alarms and have shorter corresponding graph traversal paths as the root cause nodes. This ensures that the determined root cause nodes not only cover the numerous alarms generated during the current fault but also maintain a relatively direct topological association with the corresponding device nodes in the power knowledge graph.
[0103] In one specific implementation, candidate root cause nodes are first sorted according to the number of associated alarms, from most to least, and the candidate root cause node with the most associated alarms is determined as the priority candidate node. When multiple candidate root cause nodes have the same number of associated alarms, the path length information corresponding to the multiple candidate root cause nodes is compared, and the candidate root cause node with the smallest average path length is determined as the fault root cause node.
[0104] In another implementation, nodes with a preset number or proportion of associated alarms can be selected from the candidate root cause nodes first, and then the node with the shortest graph traversal path length can be selected as the root cause node. The preset number or proportion can be preset according to the scale of the power plant equipment, the completeness of alarm data, and actual operating requirements.
[0105] Furthermore, when multiple candidate root cause nodes meet the associated alarm quantity requirements and have the same or similar path lengths, the graph traversal path recorded in step S4 can be used to perform fault propagation consistency verification on the candidate root cause nodes. Specifically, it checks whether the path from the candidate root cause node to each alarm's corresponding device node can continuously explain the corresponding device anomalies, protection actions, and switch status changes, and checks whether the unified timestamp of each alarm in the path conforms to the time sequence of fault propagation from the root cause node to downstream devices.
[0106] When the graph traversal path corresponding to a candidate root cause node can simultaneously explain multiple core action entities and status entities, and the occurrence time sequence of each alarm conforms to the preset device action logic, the candidate root cause node is retained. For candidate root cause nodes in the path that have connection interruptions, inconsistent protection logic, or obvious conflicting alarm time sequences, their selection priority is reduced or they are excluded from the candidate root cause node set.
[0107] For example, in the same fault process, if a candidate root cause node can form a continuous graph traversal path through the physical connection relationship between the voltage transformer and the bus, and the protection logic relationship between the protection device and the protected equipment, and the occurrence time of the relevant alarms conforms to the sequence of protection device startup, protection action and equipment status change after the fault occurs, then the candidate root cause node can provide a complete explanation of the alarm process.
[0108] Conversely, if a candidate root cause node is connected to some alarm device nodes in the topology, but its corresponding path cannot explain the main protection action, or the alarm occurrence time on the path is inconsistent with the fault propagation sequence, then it will not be given priority to be identified as the root cause node.
[0109] Furthermore, during the comparison of candidate root cause nodes, the graph inference algorithm module retains all valid graph traversal paths corresponding to the finally selected node. The valid graph traversal paths include the physical propagation path and the protection logic propagation path from the fault root cause node to the corresponding device node of each core alarm, and retain the node identifier, device attributes, relation edge type, action entity, status entity and unified timestamp of each device node on the path.
[0110] When a unique candidate root cause node is determined based on the above processing, that candidate root cause node is designated as the root cause node of the current fault. When a unique node cannot be determined due to missing alarm data, branching of the device topology, or simultaneous anomalies of multiple devices, multiple top-ranked candidate root cause nodes can be retained. The candidate root cause node with a large number of associated alarms and a short path length is designated as the primary root cause node, while the remaining candidate root cause nodes serve as auxiliary analysis nodes.
[0111] Preferably, when a fault is caused by a single device malfunction, a candidate root cause node is determined as the final root cause node, thereby converging hundreds of discrete alarms into a single root cause node. For cases with multiple independent fault sources, steps S4 and S5 can be executed separately for different sets of alarm device nodes to determine the root cause node corresponding to each set of alarm device nodes.
[0112] In a specific application scenario, when the reverse graph traversal paths of multiple alarm-corresponding device nodes converge to the 10kV bus node and the bus voltage transformer node, respectively, the graph inference algorithm module counts the number of alarms associated with the two candidate root cause nodes and calculates the path length to each alarm-corresponding device node. When the bus voltage transformer node can be associated with multiple core alarms such as "PT fuse blown", "bus undervoltage", and related protection actions, and its graph traversal path to each alarm device node is shorter than the path corresponding to other candidate root cause nodes, the bus voltage transformer node is determined as the fault root cause node.
[0113] After the root cause node of the fault is determined, root cause analysis data is generated. This root cause analysis data includes at least the node identifier, device name, device type, device attributes, number of associated alarms, and graph traversal path from the root cause node to the corresponding device node of each alarm. This root cause analysis data serves as the input data for step S6, used by the large language model in conjunction with the power plant operation procedures to generate fault analysis results and handling suggestions.
[0114] By utilizing the above steps, the coverage of multiple alarms by candidate root cause nodes and their topological distance to the corresponding device nodes of each alarm can be used to determine the root cause node of the fault that can explain the alarm process more completely and directly from multiple common ancestor nodes. This avoids relying solely on a single alarm or human experience to make root cause judgments and provides a data foundation for generating subsequent fault analysis results with clear fault causes and propagation paths.
[0115] S6. Input the root cause node of the fault and the corresponding graph traversal path into the large language model to generate fault analysis results and handling suggestions.
[0116] In this embodiment, step S6 is to input the fault root cause node, the graph traversal path corresponding to the fault root cause node, and the relevant alarm information into a pre-deployed large language model after determining the fault root cause node and forming root cause analysis data in step S5. The large language model combines the pre-stored power plant operation procedures to generate fault analysis results and handling suggestions, and outputs the fault analysis results and handling suggestions to the visualization interactive interface of the central control center.
[0117] Specifically, the root cause analysis data output in step S5 includes at least the node identifier, device name, device type, device attributes, number of associated alarms, and graph traversal path from the root cause node to the corresponding device node of each alarm. Each graph traversal path includes device nodes arranged in the direction of fault propagation, relational edges connecting adjacent device nodes, action entities and status entities associated with each device node, and corresponding unified timestamps.
[0118] After receiving the root cause analysis data, the large language model inference unit first organizes the fault root cause nodes and their corresponding graph traversal paths according to a preset data organization method, converting the node data and relation data output from the graph database into structured input content that can be parsed by the large language model. The structured input content includes at least root cause device information, fault propagation paths, key alarm information, and the occurrence time of each key alarm.
[0119] Furthermore, the root cause device information includes one or more of the following: standard device name, device type, voltage level, associated bay, associated bus, and installation location corresponding to the fault root cause node; the fault propagation path includes the device nodes and relation edge types sequentially traversed from the fault root cause node to each alarm-corresponding device node; the key alarm information includes the action entity, status entity extracted in step S2, and the unified timestamp formed in step S1.
[0120] For example, when the root cause node of the fault determined in step S5 is the 10kV bus voltage transformer node, the structured content of the input large language model may include the equipment name and bus information of the voltage transformer node, the physical connection relationship between the node and the 10kV bus node, the protection logic relationship between the bus node and the protection device node, and the action entity, status entity and unified timestamp corresponding to key alarms such as "PT fuse blown", "bus undervoltage" and "protection device action".
[0121] In this embodiment, the large language model also receives the power plant operation procedures corresponding to the root cause node of the fault. The power plant operation procedures are pre-stored in a local data storage device, and their content includes one or more of the following: equipment anomaly handling procedures, fault isolation requirements, operation mode adjustment requirements, safety operation requirements, and equipment inspection requirements. Based on the equipment type, equipment location, and fault status corresponding to the root cause node, the procedure content related to the current fault is obtained from the power plant operation procedures, and the procedure content, along with the root cause analysis data, is input into the large language model.
[0122] Preferably, the power plant operation procedures are stored and invoked locally, and the large language model is deployed locally in a private manner. Therefore, the equipment topology information, alarm data, operation procedures, and handling suggestions involved in fault analysis are all processed within the power plant's internal network or a controlled network environment, without needing to be sent to an external server.
[0123] Furthermore, to ensure that the large language model generates fault analysis results in a uniform format, a prompt template for constraining the output content can be pre-set. This prompt template instructs the large language model to sequentially generate the fault root cause, fault propagation process, and handling suggestions based on the input fault root cause node, graph traversal path, key alarm information, and operating procedures, and it must not generate handling content unrelated to the input data and operating procedures.
[0124] Specifically, the large language model first reads the equipment information of the root cause node of the fault, and determines the fault object and fault type based on the action entity and state entity associated with the root cause node of the fault. Then, according to the arrangement order of each equipment node and relation edge in the graph traversal path, the fault propagation link represented by the graph structure is converted into the fault propagation process in natural language form. Finally, based on the fault object, fault type and corresponding power plant operation procedures, a handling suggestion matching the current fault is generated.
[0125] The fault analysis results include at least fault root cause information and fault propagation process. The fault root cause information describes the root cause device and its abnormal state corresponding to the fault; the fault propagation process describes the physical propagation relationship and protection logic propagation relationship between the fault root cause node and each alarm-corresponding device node, and reflects the chronological order of major alarms or device actions according to a unified timestamp.
[0126] For example, when the graph traversal path represents the bus voltage transformer fuse blowing, causing abnormal bus voltage measurement, and further triggering bus undervoltage alarm and related protection actions, the large language model can generate fault analysis results including root cause equipment, direct fault status, affected equipment, and protection action process based on the above path.
[0127] The proposed solutions are generated based on the equipment type, fault status, and operating procedures corresponding to the root cause of the fault, and include at least the equipment to be inspected and fault isolation measures. In different implementations, the proposed solutions may also include adjustments to the operating mode, deployment of backup equipment, on-site verification items, and confirmations required before power restoration.
[0128] Specifically, when the root cause node of the fault is a voltage transformer node and the fault status is a blown fuse, the large language model extracts content related to the abnormal handling of the voltage transformer from the operating procedures, and generates handling suggestions such as inspecting the corresponding voltage transformer cabinet, isolating the faulty equipment, and verifying the bus voltage status. When the root cause node of the fault is a bus node, circuit breaker node, or protection device node, the large language model calls the corresponding operating procedure content that matches the type of equipment and fault status.
[0129] Furthermore, the proposed handling suggestions are auxiliary handling information generated based on the root cause node of the fault, the fault propagation path, and the operating procedures. Before actually performing equipment operations, the on-duty personnel can confirm them in conjunction with the real-time operating status of the power plant and on-site safety requirements to avoid directly executing incompatible operations when the equipment operating status changes.
[0130] After generating the fault analysis results and handling suggestions, the large language model forms a root cause analysis report according to a preset output structure. The root cause analysis report includes at least the fault occurrence time, the root cause device, the fault status, key alarms, the fault propagation process, and handling suggestions. Preferably, the root cause analysis report also includes the device location of the root cause node and the range of affected devices, so that control room personnel can quickly identify the objects to be inspected on-site.
[0131] In a specific application scenario, when step S5 determines that the root cause node of the fault is the voltage transformer node of the 10kV II section bus, and the graph traversal path indicates that the blown voltage transformer fuse caused abnormal bus voltage measurement, bus undervoltage alarm, and related protection actions, the large language model, combined with the corresponding operating procedures, can generate the following fault analysis results: Through root cause tracing, the 10kV II section bus undervoltage alarm was caused by the blown voltage transformer fuse of this bus, and the fault propagated along the voltage measurement circuit and protection logic, triggering related alarms; the corresponding handling suggestions include checking the voltage transformer cabinet and fuse status, isolating the faulty equipment according to the operating procedures, and confirming the actual operating status of the bus.
[0132] Those skilled in the art will understand that the above natural language descriptions are only used to illustrate the generation method of fault analysis results and handling suggestions, and the specific output content should be determined according to the actual fault root cause nodes, graph traversal paths and corresponding operating procedures.
[0133] Furthermore, the tracing and interaction unit sends the root cause analysis report to the central control center's visualization screen. This visualization screen displays the root cause device, fault analysis results, and handling suggestions. Based on the root cause analysis data, it calls the corresponding nodes and relationship edges in the power knowledge graph to form a fault propagation topology diagram corresponding to this fault.
[0134] Specifically, in the fault propagation topology diagram, the root cause nodes are displayed separately, and the valid graph traversal paths from the root cause nodes to the corresponding alarm device nodes are highlighted. Thus, while viewing the natural language root cause analysis report, control personnel can intuitively obtain the location of the root cause device in the power plant topology and the fault propagation path.
[0135] Preferably, the visualization screen synchronously displays the original number of alarms, the number of core alarms retained after semantic noise reduction, the root cause node of the fault, and the main fault propagation path, thereby converging the hundreds of discrete alarms generated when the fault occurs into a single root cause node of the fault and its corresponding fault propagation topology.
[0136] In this embodiment, each processing step from obtaining the original alarm message in step S1 to generating the fault analysis results and handling suggestions in step S6 can be completed between the edge server, graph database server, and central control center equipment within the power plant. The large language model inference unit performs semantic parsing and report generation through an edge server equipped with a GPU or NPU, the knowledge graph storage engine performs device node queries and graph traversal through a graph database server, and the tracing and interaction unit displays the root cause determination results through the graph inference algorithm module and a large visualization screen.
[0137] Through the above steps, the root cause nodes and their corresponding graph traversal paths obtained in step S5 are transformed into structured fault analysis results that are easy for on-duty personnel to understand. Based on the power plant operation procedures, handling suggestions matching the faulty equipment and fault status are generated. Simultaneously, the root cause nodes and fault propagation topology are highlighted on a large visual screen, enabling control personnel to simultaneously obtain both text-based root cause analysis reports and graphical fault propagation chains, thereby achieving rapid convergence of alarm information and assisted handling of fault root causes.
[0138] Example 2 To implement the methods of the above embodiments, the present invention also provides an electronic device, which includes a memory and a processor; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.
[0139] Example 3 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0140] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0141] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0142] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A power alarm tracing method based on knowledge graphs and large models, characterized in that, Includes the following steps: S1. Obtain the original alarm messages generated by the power monitoring system, and perform time alignment and format standardization processing on the original alarm messages to form standardized alarm data. S2. Input the standardized alarm data into a pre-deployed large language model, perform semantic parsing on the standardized alarm data, extract the device entity, action entity and status entity corresponding to the alarm, and form an alarm entity set; S3. Based on the alarm entity set, map each device entity to the corresponding device node in the pre-constructed power knowledge graph; S4. Using the device node as the starting node, perform a reverse graph traversal along the physical connection relationship and protection logic relationship of the device in the power knowledge graph to determine the common ancestor node corresponding to multiple alarms as candidate root cause nodes. S5. Determine the root cause node of the fault based on the candidate root cause node and its corresponding graph traversal path. S6. Input the root cause node of the fault and the corresponding graph traversal path into the large language model to generate fault analysis results and handling suggestions.
2. The power alarm tracing method according to claim 1, characterized in that, In step S1, the standardization process includes a unified alarm time format, a unified device identifier format, and a unified alarm field format.
3. The power alarm tracing method according to claim 1, characterized in that, In step S2, the large language model performs semantic parsing on the standardized alarm data, including identifying the device name, protection action, fault status and device operating status in the alarm, and removing derivative alarms that are not related to fault propagation.
4. The power alarm tracing method according to claim 1, characterized in that, In step S3, the power knowledge graph includes multiple device nodes and relational edges connecting each device node. The relational edges include at least the physical connection relationship of the devices and the protection logic relationship.
5. The power alarm tracing method according to claim 4, characterized in that, In step S4, the reverse graph traversal includes searching from each device node to its upstream node to determine the common ancestor node of each search path as a candidate root cause node.
6. The power alarm tracing method according to claim 5, characterized in that, In step S5, the root cause node of the fault is determined based on the length of the graph traversal path corresponding to the candidate root cause node and the number of associated alarms.
7. The power alarm tracing method according to claim 1, characterized in that, In step S6, the large language model generates a fault analysis result containing the fault cause, fault propagation process, and handling suggestions based on the fault root cause node, graph traversal path, and pre-stored operating procedures.
8. An electronic device, characterized in that, Including processor and memory; The processor runs a program corresponding to the executable program code stored in the memory to implement the method as described in any one of claims 1-7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.