A power system causal verification attack tracing method and system
By constructing a semantic attack attribution knowledge graph in the power sector and introducing causal verification methods based on business rules and physical feasibility constraints, the problems of high false alarm rate and low physical credibility in power system attack attribution are solved, achieving high-precision attack attribution and defense strategy formulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NARI INFORMATION & COMM TECH
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-10
AI Technical Summary
Existing power system attack tracing technologies suffer from high false alarm rates, low physical reliability, and poor practicality of results. They cannot effectively distinguish between causal relationships and false correlations, lack power business logic constraints and physical feasibility verification, leading to broken attack chains and misjudgments.
A causal verification attack tracing method for power systems is adopted. This method constructs an attack tracing knowledge graph that deeply integrates the semantics of the power field, combines it with the Do-Calculus algorithm for causal reasoning, and introduces power business rule constraints and physical feasibility constraints to screen candidate events and causal edges, thereby constructing an attack causal chain.
It achieves high-precision attack attribution tracing, reduces false alarm rate, improves the ability to identify complex power-specific attacks, ensures the professional credibility and practicality of attribution results, provides clear attack propagation paths and causal logic, and supports precise defense measures.
Smart Images

Figure CN122372343A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system network security technology, and more specifically, relates to an attack tracing method and system that integrates power control closed-loop semantic ontology, business rule embedding causal verification, and fault mechanism constraint attack chain reconstruction. Background Technology
[0002] With the digital transformation of the power grid, the number of connected devices has surged, including terminals in new energy power plants, smart converged terminals, and smart meters. This has significantly expanded the attack surface. Typical attacks include forged control commands targeting protocols such as IEC 61850, device hijacking attacks by implanting malicious code into PLCs, and lateral propagation attacks exploiting vulnerabilities. Attack attribution is a crucial aspect of traditional power system cybersecurity. Accurately locating the source of an attack and reconstructing its propagation path provides a basis for attack response and defense reinforcement. Current attribution technologies suffer from several drawbacks. Data is scattered across different subsystems, leading to insufficient data fusion, broken attack chains, high false positive rates, difficulty in distinguishing attacks from normal operation and maintenance, poor scenario adaptability, inability to identify power-specific attacks, low practicality of results that fail to guide effective defense, and a lack of correlation analysis mechanisms adapted to new power scenarios, making efficient attribution difficult.
[0003] Existing technical document 1 (CN120658465A) discloses a lightweight dynamic causal reasoning method for tracing attacks on power IoT terminals. Its shortcomings are that the causal reasoning relies on the attention weights between nodes in the graph neural network, which is a statistical correlation modeling method and cannot distinguish between causal associations and false correlations. Furthermore, its knowledge graph does not deeply integrate power business semantics, resulting in the attack chain being broken at the business level. The generation path also lacks power physical constraints and has low credibility.
[0004] Existing technical document 2 (CN121189798A) discloses a knowledge graph-based power information operation violation risk monitoring system. Its shortcomings are that causal reasoning only uses Do-Calculus as a general statistical tool to identify key risk nodes, without embedding business rules into causal verification. It cannot determine whether the correlation is reasonable in the power business logic, cannot distinguish the causal difference between compliant operation and malicious tampering, resulting in misjudging normal operation and maintenance as an attack and thus false alarms. It lacks the ability to reconstruct the complete attack propagation path and cannot locate the source of the attack and block its propagation.
[0005] Existing technical document 3 (CN121151119A) discloses an APT attack tracing and path reconstruction method and system. Its shortcomings are that the general network security entity lacks dedicated modeling of the characteristics and protocol semantics of power equipment, and the path exploration does not embed the physical mechanism constraints of power system fault propagation. It may generate false paths that are physically infeasible, such as RTU anomalies leading to the modification of upper-level PLC settings, resulting in low credibility of the tracing results. Summary of the Invention
[0006] To address the technical problems of high false alarm rates, low physical credibility, and poor practicality of results in existing power system attack tracing technologies, this invention proposes a causal verification attack tracing method and system for power systems. Based on a causal verification mechanism constrained by both business rules and physical feasibility, this invention constructs an attack tracing knowledge graph deeply integrated with power domain semantics and utilizes the Do-Calculus algorithm for causal reasoning. During the causal edge filtering process, this invention collaboratively introduces and applies power business rule constraints and physical feasibility constraints. Specifically, business rule constraints verify whether event sequences conform to power operation procedures and safety logic, fundamentally distinguishing between malicious attacks and normal operation and maintenance, thereby accurately resolving the problem of accidental association misjudgment and reducing the related false alarm rate. Physical feasibility constraints verify the feasibility of attack propagation paths within the power system through multiple physical laws such as time delay, direction, and pattern, thus completely resolving the problem of unreliable attack chains and ensuring that the generated attack chains conform to the power system fault propagation laws. The synergistic effect of these dual constraints ensures extremely high accuracy and professional credibility of the final tracing results, enhancing the ability to identify concealed and complex power-specific attacks.
[0007] The present invention adopts the following technical solution.
[0008] The first aspect of the present invention provides a method for tracing the source of causal verification attacks in a power system, comprising the following steps: Collect multi-source data from the power system, extract triples of entity-basic relation-basic attribute and entity-basic relation-entity, and generate triple datasets; Based on the triple dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing the source of power attacks is constructed. Based on equipment anomaly events, candidate events and candidate causal edges are selected from the power attack tracing knowledge graph to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between equipment anomaly events and each candidate event. Directed candidate causal edges that simultaneously satisfy power business rule constraints and physical feasibility constraints constitute the attack causal chain. Based on equipment anomaly events and the attack causal chain, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
[0009] Preferably, generating the triplet dataset includes: Based on the standardized fields of device entities, attack entities, business entities, and vulnerability entities in the preset entity field dictionary, device entities, attack entities, business entities, and vulnerability entities are extracted from multi-source data respectively. Based on the preset entity-basic relationship-basic attribute triplet mapping table, the protocol fields and log fields parsed from the multi-source data are mapped to the basic attributes of the corresponding entities, and the entity-basic relationship-basic attribute triplet is generated with the entity, the basic relationship between the entity and its basic attributes, and the basic attributes of the entity. Based on the preset entity-basic relation-entity triplet mapping table, the data representing the relationship between entities in the multi-source data are mapped to the basic relations between entities, and entity-basic relation-entity triplets are generated based on entities and the basic relations between entities; Construct an intermediate dataset of triples based on entity-base relation-base attribute triples and entity-base relation-entity triples; Clean the intermediate dataset of triples and output the cleaned dataset of triples.
[0010] Preferably, based on the triplet dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing power attack origins is constructed, including: Based on each entity, its business attributes, and the business relationships between entities, an ontology model for a knowledge graph of power attack tracing is constructed. Extract a set of new knowledge triples that conform to the ontology model from the triple dataset; The new knowledge triple set is fused and semantically disambiguated to generate a fused knowledge set. The fused knowledge set is then used to complete the relationships and generate a knowledge graph for tracing the source of power attacks.
[0011] Preferably, the set of new knowledge triples that conform to the ontology model is extracted from the triple dataset, including: The triplet dataset is divided into structured data and semi-structured data; According to the preset mapping rules, the entities, basic relationships and basic attributes in the structured data are mapped to the entities, business relationships and business attributes corresponding to the ontology model, and the candidate knowledge triple set A is determined. Based on regular expressions specific to the power scenario, entities, basic relationships, and basic attributes in semi-structured data are mapped to entities, business relationships, and business attributes corresponding to the ontology model, thus determining the candidate knowledge triple set B. Generate a new set of knowledge triples based on the entities in the candidate knowledge triple set A and the candidate knowledge triple set B.
[0012] Preferably, dividing the triplet dataset into structured data and semi-structured data includes: Determine whether the entities, basic relations, and basic attributes of the triples in the triple dataset directly correspond semantically to the entities, business relations, and business attributes in the ontology model. Specifically, for entity-basic relation-basic attribute triples, determine whether their entities, basic relations, and basic attributes can all be found to have direct corresponding entities, business relations, and business attributes in the ontology model; for entity-basic relation-entity triples, determine whether their head entity, basic relation, and tail entity can all be found to have direct corresponding entities and business relations in the ontology model. If the entity, basic relation, and basic attribute of a triple do not directly correspond to all the entities, business relations, and business attributes in the ontology model, then the triple is determined to be semi-structured data. If the entity, basic relation, and basic attribute of a triple directly correspond to the entity, business relation, and business attribute in the ontology model, then it is determined whether the target data item pointed to by the basic relation in the triple is atomic data. If the target data item pointed to by the basic relation is atomic data, the triple is determined to be structured data. If the target data item pointed to by the basic relation is not atomic data, the triple is determined to be semi-structured data. Specifically, for an entity-basic relation-basic attribute triple, the target data item pointed to by the basic relation refers to the value of the basic attribute; for an entity-basic relation-entity type triple, the target data item pointed to by the basic relation refers to the identifier of the tail entity.
[0013] Preferably, the physical feasibility constraints include time delay compatibility verification, causal directionality verification, and cascading failure mode matching verification; Among them, the delay compatibility check is used to determine that the directed candidate causal edge passes the delay compatibility check if the actual time interval between the candidate events associated with each directed candidate causal edge in the attack causal graph is greater than or equal to the minimum propagation delay, based on the predefined minimum propagation delay of the power system attack action. Causality directionality verification is used to determine if the direction of each directed candidate causal edge in the attack causal graph is consistent with the allowed direction defined by the network topology and the direction of business data flow, based on a predefined network topology and business data flow. Cascaded fault mode matching verification is used to determine that a directed candidate causal edge passes the cascaded fault mode matching verification if it matches any known propagation pattern in the pre-set library of typical power system faults and attack propagation patterns based on a pre-set library of typical power system faults and attack propagation patterns. If a directed candidate causal edge passes the time delay compatibility check, causal directionality check, and cascading failure mode matching check in sequence, it is determined to meet the physical feasibility constraint.
[0014] Preferably, locating the attack source and reconstructing the attack chain in the power attack attribution knowledge graph includes: Within the knowledge graph for tracing power attacks, candidate events that are logically related to abnormal device events within a preset time window are selected to form a candidate event set; Using candidate events in the candidate event set as nodes, candidate causal edges are constructed based on time proximity and logical association rules to generate an attack causal graph; For candidate causal edges in the attack causal graph, the Do-Calculus method is used to conduct intervention tests to determine the average causal effect between the device malfunction event and the candidate event. Based on the average causal effect, directed candidate causal edges are extracted to form a set of directed causal edges. The power business rules are used to verify the consistency of the power business rules for candidate events associated with directed candidate causal edges. Candidate events associated with directed candidate causal edges that pass the verification are marked as suspicious attack paths. Suspicious attack paths are physically feasible through physical feasibility constraints. Suspicious attack paths that pass the physical feasibility verification are used as the verified attack causal chains. Starting from the abnormal equipment event, trace the attack causal chain backwards along the verified path, query the power attack tracing knowledge graph to obtain the business relationship corresponding to the abnormal equipment event, thereby determining the source of the attack and reconstructing the attack chain from the source of the attack to the device corresponding to the abnormal equipment event.
[0015] Preferably, the verified attack causal chain includes: The Do-Calculus method is used to intervene in the candidate events corresponding to the candidate causal edges in the causal graph. Based on historical data, the average causal effect of the device malfunction event before and after the intervention in the candidate events is calculated. Based on the relationship between the average causal effect and the preset significance threshold, the direction of the candidate causal edge is determined, and directed candidate causal edges with a clear unidirectional causal direction are selected. Multiple attack causal chains are constructed based on directed candidate causal edges. The candidate events corresponding to each attack causal chain are checked for consistency with the power business rules. When the consistency check is passed, the rationality of the attack causal chain is determined. When the consistency check is not passed, the corresponding attack causal chain is filtered out. A reasonable attack causal chain is physically feasible. An attack causal chain that passes the physical feasibility verification is considered a valid attack causal chain.
[0016] A second aspect of the present invention provides a power system causality verification attack tracing system, wherein running the power system causality verification attack tracing system described in the first aspect of the present invention includes: The data acquisition module is used to collect multi-source data from the power system; The data preprocessing module is used to extract triples of entity-basic relation-basic attribute and entity-basic relation-entity, clean the data of the two types of triples, and generate triple datasets. The knowledge graph construction module is used to build a knowledge graph for tracing the source of power attacks based on the triple dataset, combined with the business attributes of entities and the business relationships between entities; The attack localization module is used to filter candidate events and candidate causal edges from the power attack tracing knowledge graph based on abnormal equipment events to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between abnormal equipment events and each candidate event. Directed candidate causal edges that simultaneously satisfy power business rule constraints and physical feasibility constraints form an attack causal chain. Based on abnormal equipment events and attack causal chains, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
[0017] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements a power system causal verification attack tracing method according to the first aspect.
[0018] The fourth aspect of the present invention is a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements a power system causal verification attack tracing method according to the first aspect.
[0019] Compared with the prior art, the beneficial effects of the present invention include at least the following: By organizing the triplet data according to the preset power entities (such as attacks, equipment, and services) and relationships (such as initiation and impact), a structured semantic network (knowledge graph) is formed. This enables deep semantic fusion and association of multi-source heterogeneous data, so that power system data such as "network attack events", "abnormal equipment status", and "service scheduling instructions" are no longer isolated points, but connected in the same queryable and reasonable graph. This overcomes the defect that simple data splicing cannot reflect deep business logic, improves the event association capability across devices, protocols, and business layers, and makes it possible to restore the complete potential path from the attack entry point to the business impact, thus alleviating the problem of attack chain breakage at the logical level. Traditional tracing techniques rely on time series or statistical correlations to address technical issues. For example, if an attacking IP and a compromised device appear simultaneously, it cannot distinguish between causal and coincidental relationships. This can lead to misjudging normal remote maintenance IP communications as attack propagation, resulting in a high false positive rate and difficulty in locating the true source of the attack. This invention, after causal calculation, immediately uses business rules (such as scheduling instruction verification process, five-defense logic, and operation ticket sequence) to conduct compliance reviews on candidate causal edges. This effectively separates normal business operations, allowing the system to focus on identifying abnormal event sequences that violate established business logic. By using business rule constraints, the false positive rate caused by normal operation and maintenance is reduced, and the accuracy of attack event determination is improved. Strict physical feasibility constraints are imposed on suspicious paths identified through business rules, including time delay compatibility (such as protection action time limits and communication transmission delays), directionality (such as power flow and control command flow), and cascading pattern matching. This ensures that the final attack causal chain strictly conforms to the physical operating laws and fault propagation mechanisms of the power system. Physical feasibility constraints enhance the physical basis and professional credibility of the tracing conclusions.
[0020] The synergy of business rule constraints and physical feasibility constraints enables this invention to identify targeted attacks hidden within normal processes or exploiting system physical characteristics. For example, it can detect and reconstruct "supply chain attacks using legitimate accounts to perform low-frequency, minor parameter tampering" or "targeted attacks that forge specific protocol messages to trigger protection malfunctions." This ability to identify and reconstruct advanced persistent threats (APTs) specific to the power industry has been fundamentally enhanced, improving the ability to detect complex and covert attacks and reconstruct complete attack chains.
[0021] By reducing the false alarm rate by an order of magnitude, the time and effort spent by security operations personnel reviewing false alarms is significantly reduced, allowing them to focus on real threats and drastically reducing interference and decision-making costs in security operations. Defense strategies based on highly reliable, physically consistent attack chains (such as hardening specific devices and adjusting access policies) are more accurate and effective, avoiding system operational risks and resource waste caused by ineffective or excessive defense measures based on erroneous or unreliable attribution conclusions, and effectively reducing secondary risks caused by misjudgments or mistrust.
[0022] The final output of this method is a causally verified complete attack chain that traces back from the target anomaly to the source of the attack, rather than just an isolated list of IPs or devices. It provides a clear attack propagation path, key nodes at each stage, and causal logic, enabling operations and maintenance personnel to clearly understand the attack's spread logic and scope of impact. This provides a direct and reliable basis for decision-making in developing targeted defense measures (such as hardening critical jump server devices and blocking specific propagation paths), greatly enhancing the practical value of the tracing results in guiding actual defense hardening. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the power system causal verification attack tracing process provided in accordance with the embodiments of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0025] like Figure 1 As shown, Embodiment 1 of the present invention provides a method for tracing the source of causal verification attacks in a power system, comprising the following steps: Step 1: Collect multi-source data from the power system, including equipment log data, network data, and business and security data.
[0026] In a preferred but non-limiting embodiment of the present invention, step 1 includes: Step 1.1: Adopt a distributed acquisition architecture and deploy distributed acquisition agents in substations and dispatch centers.
[0027] Distributed data acquisition architecture refers to deploying data acquisition functions, i.e., acquisition agents, at multiple different physical locations and logical levels within the power system (such as, but not limited to, various substations and dispatch centers), rather than centralizing them on a single server. Its core idea is to collect data near the source of the data (such as protection devices and RTUs within a substation), reducing the pressure and latency of long-distance transmission of raw data over the network. Each acquisition agent works independently, responsible for the acquisition, caching, and preliminary processing of data within its designated area or type. Through network collaboration, they collectively form a complete acquisition system. When new acquisition points need to be added (such as building a new substation), only new agents need to be deployed; there is no need to modify the central architecture. A failure of a single agent does not affect other nodes, improving the overall reliability of the system.
[0028] Step 1.2: Based on the distributed data acquisition agent deployed in Step 1.1, collect device log data, network data, and business and security data to generate multi-source data.
[0029] More preferably, step 1.2 includes: Step 1.2.1, collecting equipment log data includes: reading PLC control instruction execution logs (including instruction content, execution results, and parameter changes) via the IEC61850MMS protocol; collecting RTU status monitoring data (including voltage, current, and equipment operating status codes) via the Modbus protocol; and obtaining server process logs and system logs via the SSH protocol. The collection frequency is set to 1 time / second (based on the minimum execution cycle of the power PLC control instruction being 1 second, this frequency can completely capture each control instruction and parameter change).
[0030] Step 1.2.2, collecting network data includes: deploying a dedicated network packet capture module to parse IEC61850 GOOSE messages (extracting dataset identifier, sender ID, and data value), 376.1 protocol, 698 protocol, 104 protocol, Modbus protocol traffic (extracting function code, register address, and data content), and TCP / IP communication data (extracting source IP, destination IP, port number, and communication frequency). A circular buffer is used to store the original messages, with a buffer capacity of 1GB to avoid data loss.
[0031] Step 1.2.3, collecting business and security data includes: obtaining load scheduling records (including scheduling instruction ID, adjustment threshold, issuance time, and associated devices) through the scheduling system API interface; collecting alarm data (including attack type, trigger time, suspected IP, and exploit vulnerability ID) through the security device SDK interface; and ensuring security through SSL / TLS 1.3 encryption for data transmission.
[0032] Multi-source data acquisition serves as the foundational data support. Distributed acquisition agents are deployed in substations and dispatch centers to collect four major categories of power system data across all dimensions: equipment logs, network data, business data, and security data. Three acquisition methods cover all dimensions of power system data, ensuring comprehensiveness and real-time performance of data collection. Equipment logs, network traffic, and business data are collected in parallel at different locations by different agent instances, rather than by a central program remotely logging into all devices to retrieve them.
[0033] It is worth noting that existing methods for tracing attacks on new power systems can only correlate single-type log data, such as network device logs, and cannot integrate heterogeneous power-specific data such as smart fusion terminals, smart meters, PLC control command logs, RTU status monitoring data, and load dispatching business data. This results in the inability to reconstruct the complete attack chain from attack initiation, intermediate propagation, to device impact, leading to fragmented tracing results. This invention actively collects and standardizes multi-source data, transforming it into a unified "entity-relationship-attribute" triple structure. This achieves preliminary integration and unified representation of multi-dimensional power data, including equipment, network, business, and security data. It solves the problem of fragmented attack chain analysis caused by the inability of existing methods to integrate heterogeneous power-specific data such as smart terminal logs, PLC logs, RTU data, and load dispatching records. It improves the correlation between data from different sources and in different formats, and reduces the risk of broken correlations of attack events due to isolated data sources and inconsistent formats.
[0034] Step 2: Extract the triples of entity-basic relation-basic attribute and entity-basic relation-entity to generate a triple dataset.
[0035] In a preferred but non-limiting embodiment of the present invention, step 2 includes: Step 2.1: Based on the standardized fields of device entity, attack entity, business entity and vulnerability entity in the preset entity field dictionary, extract device entity, attack entity, business entity and vulnerability entity from the multi-source data respectively.
[0036] More preferably, the entities include four categories: device entities, attack entities, business entities, and vulnerability entities.
[0037] The entity field dictionary includes standardized fields for device entities, attack entities, business entities, and vulnerability entities.
[0038] The basic attributes of a device entity include: device ID, device type, model, deployment location, and operating status. For example, but not limited to, the device ID is resolved from the device identifier field as Sub220-PLC-001, the device type is resolved from the asset library or protocol identifier as PLC, RTU, or Server, the model is resolved from the asset library or log header as S7-1200, the deployment location is obtained from the topology configuration as the No. 1 main transformer bay of the 220kV substation, and the operating status includes the trip signal, normal, and alarm signals resolved from the GOOSE message data value 0x01.
[0039] The basic attributes of the attacking entity include: attacking IP, attack type, attack time, and exploit ID. For example, but not limited to, the attacking IP is resolved from security alerts or network traffic as 202.120.36.18, the attack type is resolved from the alert log event type as port scanning, forged MMS messages, and brute-force attack, the attack time is resolved from the alert log trigger time as 2024-06-10 09:15:30, and the exploit ID is resolved from the alert log associated vulnerability field as CVE-2023-1234.
[0040] The basic attributes of a business entity include: business ID, business type, control instruction, execution result, and associated device ID. For example, but not limited to, the business ID is parsed from the scheduling instruction ID as CMD-20231025001; the business type is parsed from the scheduling instruction type as current protection setting adjustment and remote tripping; the control instruction is parsed from the MMSWrite service or scheduling instruction content as writing a single register 0x0023 and closing; the execution result is parsed from the PLC execution log result field as success, failure, and timeout; the issuance time is parsed from the scheduling record issuance time as 2024-06-10 09:20:00; and the associated device ID is parsed from the scheduling instruction target device or log association as Sub220-PLC-001.
[0041] The basic attributes of a vulnerable entity include: vulnerability ID, vulnerability type, vulnerability description, remediation plan, and discovery time. For example, but not limited to, the vulnerability ID is resolved from the vulnerability database or alarm-related vulnerability ID to CVE-2023-1234; the vulnerability type is resolved from the vulnerability database type to authentication bypass and buffer overflow; the vulnerability description is resolved from the vulnerability database description to an authentication bypass vulnerability in the MMS protocol processing of a certain PLC model; the remediation plan is resolved from the vulnerability database remediation plan to upgrade the firmware to version V4.5.0; and the discovery time is resolved from the vulnerability database release time to 2023-11-05.
[0042] Step 2.2: Based on the pre-defined triplet mapping table, establish the basic relation mappings between entities and their basic attributes, and between entities themselves, generating an intermediate triplet dataset. The triplet mapping table includes an entity-basic relation-basic attribute triplet mapping table and an entity-basic relation-entity triplet mapping table. More preferably, step 2.2 includes: Step 2.2.1: Based on the preset entity-basic relationship-basic attribute triplet mapping table, the protocol fields and log fields parsed from the multi-source data are mapped to the basic attributes of the corresponding entities, and the entity-basic relationship-basic attribute triplet is generated with the entity, the basic relationship between the entity and its basic attributes, and the basic attributes of the entity.
[0043] The entity-basic relation-basic attribute triple structure is <entity, basic relation between entity and basic attribute, basic attribute value of entity>.
[0044] Establish the mapping between entities and their basic attributes: Map the fields in the protocol and log data of the power system multi-source data in step 1 to the basic attributes of the entities, as shown in Table 1.
[0045] Table 1. Entity-Basic Relationship-Basic Attribute Triple Mapping Table
[0046] The basic relationships between an entity and its underlying attributes, including but not limited to being identified as, containing, having, operating in, discovered in, and described as.
[0047] Entity-Based Relationship-Based Attribute triples, such as but not limited to: <Device: Sub220-PLC-001, identified as LD01 / LLN0>, <Device: Sub220-PLC-001, includes status, trip signal>, <Device: Sub220-PLC-001, has IP address, 192.168.1.10>, <Attack: 202.120.36.18, attack type is port scan>, <Attack: 202.1 Vulnerability 20.36.18, occurred on 2024-06-10 09:15:30>, <Business: CMD-20231025001, contains instruction to write a single register 0x0023>, <Business: CMD-20231025001, execution result: successful>, <Vulnerability: CVE-2023-1234, description: ...>, <Vulnerability: CVE-2023-1234, remediation solution: upgrade firmware to V4.5.0>, etc.
[0048] Step 2.2.2: Based on the preset entity-basic relation-entity triplet mapping table, the data representing the relationship between entities in the multi-source data are mapped to the basic relations between entities, and entity-basic relation-entity triplets are generated based on entities and the basic relations between entities.
[0049] The structure of the entity-base relation-entity triple is <head entity, base relation between head entity and tail entity, tail entity>.
[0050] Establish entity-to-entity mapping: Based on the interaction features in the multi-source power system data from step 1, establish the associations between entities, as shown in Table 2.
[0051] Table 2 Entity-Relationship-Entity Triple Mapping Table
[0052] In multi-source data, the data representing the relationship between entities is called interaction feature. Interaction features include data from the power system multi-source data in step 1 that indicates the relationship between entities. For example, but not limited to, mapping two communicating device entities based on the source / destination MAC addresses of GOOSE packets in network communication traffic to generate a triple <Device: Sub220-SW, sending GOOSE to, Device: Sub220-PLC-001>; mapping attacking entities to vulnerable entities based on the "attack IP" and "associated vulnerability ID" in security alarm data to generate a triple <attacking entity, exploit, vulnerable entity>; mapping attacking entities to accessed device entities based on the source and destination IP communication features in network data to generate a triple <attacking entity, initiating access, device entity>; and mapping business entities issuing instructions to device entities executing instructions based on service scheduling records to generate a triple <business entity, acting on, device entity>. Furthermore, after extracting standardized fields from unstructured logs using regular expressions, they are also uniformly converted and saved in JSON format according to the above mapping table.
[0053] The basic relationships between entities include, but are not limited to: send to, receive from, access, act on, utilize, associate, and occur in.
[0054] The data source is the entity-basic relationship-entity triplet of network communication, such as but not limited to <Device: Sub220-SW, sending GOOSE to, Device: Sub220-PLC-001>, <Attack: 202.120.36.18, accessing, Device: Sub220-PLC-001>.
[0055] The data source is the entity-basic relationship-entity triplet of security alerts, such as but not limited to: <Attack: 202.120.36.18, Exploitation, Vulnerability: CVE-2023-1234>.
[0056] The data source is the entity-basic relationship-entity triplet of business scheduling, such as but not limited to: <Business: CMD-20231025001, acting on, equipment: Sub220-PLC-001>.
[0057] The entity-basic relationship-basic attribute triple is responsible for characterizing the intrinsic, static features of a single entity, reducing scattered fields in the original data to standardized attribute descriptions centered on the entity. By uniformly extracting, naming, and assigning values to basic relationships and attributes from logs, messages, and databases, the chaotic multi-source data is transformed into standardized feature description units centered on the entity, providing unified features for entity identification, classification, and subsequent analysis.
[0058] Example: <Device: PLC, identified as, Sub220-PLC-001> describes the entity "PLC" with the basic attribute "Sub220-PLC-001", and the basic relationship between the entity and the basic attribute is "identified as". Extract basic attributes such as device type, attack IP, instruction content, etc. and relationships such as identified as from logs and messages to generate triples like <Device: PLC, identified as, Sub220-PLC-001>. The core of this triple is to declare the existence of an entity ("PLC") and assign it a unique and standardized identifier ("Sub220-PLC-001"), and at the same time, it may be accompanied by other attribute triples (such as <Sub220-PLC-001, model, S7-1200>) to enrich its characteristics.
[0059] Finally, a set of entity collections with clear identifiers and standardized characteristics is output, solving the problem of data fragmentation and ensuring that the same entity is uniformly referred to and described regardless of where the data comes from.
[0060] Based on the standardized entity collection established by the entity-basic relationship-basic attribute triples, the entity-basic relationship-entity triples define how entities are connected, responsible for depicting the dynamic interactions and associations between different entities, discovering and declaring the specific interaction behaviors that occur between these entities, and thus connecting isolated "entity nodes" into a network with the associated "edges" between entity nodes.
[0061] Based on the interaction characteristics in the data (such as network communication, instruction issuance), establish preliminary connections between entities. Example: <Attack IP: 192.168.5.201, access, Device: Sub220-PLC-001> describes the basic relationship "access" that exists between the entity "192.168.5.201" and the entity "Sub220-PLC-001". The core is to record a specific association event that occurs between two identified entities, directly solving the problems of attack chain breakage and association loss. By establishing basic relationships across devices and networks, isolated security events, network traffic, and device states are initially associated at the data level, providing the original relationship clues to prevent the attack chain from breaking at the technical level, outputting a preliminary dynamic association network with entities as nodes and interaction events as edges, directly solving the "association loss" problem, and providing the original and factual association clues to prevent the attack chain from breaking at the technical level.
[0062] The entity-basic relation-basic attribute triple defines the entity nodes and their detailed characteristics in the graph, while the entity-basic relation-entity triple defines the edges between entity nodes in the graph. Using the entity-basic relation-basic attribute triple, "an entity with IP address 192.168.5.201" and "a device with ID Sub220-PLC-001" are identified from the logs. Then, through interaction records in network traffic, a relation triple is generated, establishing an "access" edge between the two. A preliminary, machine-readable "event-relationship" network containing rich attributes and associated edges is thus constructed, which is an indispensable basic data structure for causal reasoning.
[0063] Traditional methods can only concatenate fields, while this invention achieves data unification at the entity level through attribute triples and data association at the interaction level through relation triples. The combination of entity-basic relation-basic attribute triples and entity-basic relation-entity not only unifies the data format, but also establishes a complete expression at the semantic level in which entities containing entity attributes are associated with other entities through basic relations.
[0064] The accurate entity characteristics provided by the entity-basic relationship-basic attribute triple are the factual basis for business rule validation. For example, the validation of the rule "value modification requires an instruction" relies on the <business entity: CMD-XXX, modified to..., instruction content> provided by the entity-basic relationship-basic attribute triple. Without accurate attribute extraction, rule validation is impossible.
[0065] The entity-basic relationship-entity triplet provides the connection relationship between entities, which is the path basis for physical feasibility verification. Delay verification requires the time difference between events (the event sequence associated with the dependency relationship edge), and directionality verification requires judging the direction of the relationship (such as whether "server → PLC" conforms to the control flow direction). Without accurate relationship extraction, physical verification loses its evaluation object.
[0066] In the early stages of data modeling, semantic labels are applied to all entities, basic attributes, and basic relationships using two types of triples. This allows subsequent dual constraints to act directly and efficiently on these structured knowledge elements, thereby enabling the deep embedding of domain knowledge (rules, physical properties) into the data model and reasoning process, which is not possible with traditional methods.
[0067] The triples defined in this invention have a set of "relationships" and "attributes" that are specific to the power industry scenario. For example, "basic relationships" include "send GOOSE to" and "modify register," which are interactions specific to power industrial control protocols; "basic attributes" include "set value area code" and "telemetry point table," which are parameters specific to power equipment.
[0068] General tracing methods use common entities and relationships such as "processes," "files," and "network connections," which cannot accurately describe attacks like "GOOSE message replay." Addressing the technical problems of existing methods not being adapted to the characteristics of new power systems and failing to handle the correlation analysis between power-specific protocol data such as IEC61850 and Modbus and security data—for example, failing to identify the causal relationship between forged GOOSE message attacks and PLC parameter anomalies, and lacking sufficient tracing capabilities for power-specific attacks—this invention designs power-specific triples, enabling the system to describe events using the "language" of the power field during the data extraction stage. This means that attack behavior is characterized at the data level as a series of power-specific "abnormal attribute changes" and "abnormal relationship establishments," making it possible to identify specialized attacks such as "forged settings" and "replay tripping," fundamentally solving the problem of poor scenario adaptability of traditional methods.
[0069] The entity-basic relation-attribute triples and entity-basic relation-entity triples, where attribute triples describe the specifications of each entity and relation triples describe the connection specifications between entities, allow disorganized data from different sources to be assembled into a "model" that accurately depicts the operation and attack process of the power system. It is based on this precise, structured, and domain-semantic-rich underlying data model that the upper-level knowledge graph fusion and, more importantly, the "dual constraint verification of business rules and physical feasibility" can be executed efficiently and reliably, ultimately achieving a leap in the accuracy, credibility, and practicality of attack attribution tracing.
[0070] Step 2.2.3: Construct intermediate datasets of triples based on entity-basic relation-basic attribute triples and entity-basic relation-entity triples.
[0071] Step 2.3 involves cleaning the intermediate triplet dataset and outputting the cleaned triplet dataset. Data cleaning includes deduplication based on data fingerprints, filtering based on invalid rules, and association verification based on power business logic. This three-level cleaning mechanism ensures high quality and reliability of the output data, providing a reliable data foundation for subsequent knowledge graph construction.
[0072] More preferably, based on data fingerprint deduplication, the MD5 hash value is calculated for each data in the triplet intermediate dataset, and the hash value is queried through the Redis cache hash table. If the hash value already exists, it is determined to be duplicate data, and the duplicate data is deleted.
[0073] Set invalid data rules to remove null fields, malformed data, and invalid data that exceeds a reasonable range from the intermediate dataset of triples. Monitor the proportion of invalid data, and trigger a data acquisition agent failure alarm if it exceeds 5%.
[0074] The correlation between data in the intermediate triplet dataset is verified based on the power business rules. If the data does not match the power business rules, it is marked as data to be verified and pushed to the operation and maintenance personnel for manual review. After the review is approved, it is re-added to the triplet dataset.
[0075] Step 3: Based on the triplet dataset, and combining the business attributes of entities and the business relationships between entities, construct a knowledge graph for tracing the source of power attacks.
[0076] Step 3.1: Construct an ontology model for the knowledge graph of power attack tracing based on each entity, the business attributes of each entity, and the business relationships between entities.
[0077] Preferably, entities include attack entities, device entities, business entities, and vulnerability entities.
[0078] The business attributes of an attacking entity include attack ID (format: Attack-Date-Serial Number), attack tools, attack intent, location, and threat level, as well as basic attributes such as attack IP, attack type, attack time, and exploit ID; The business attributes of the device entity include firmware version, asset importance level, security partition, associated business list, and known vulnerability list. For PLCs, the business attributes also include set of setpoint area numbers, current setpoint area, and setpoint modification permission status. For RTUs, the business attributes also include telemetry point table, teleindication point table, dead zone value, and sampling period. The basic attributes also include device ID, device type, model, deployment location, and operating status. The business attributes of a business entity include business status, dispatcher ID, operation ticket number, and five-prevention verification result. It also includes the business ID (format: Biz-date-serial number), business type, control instruction content, instruction execution result, execution time, and associated device ID in the basic attributes. The business attributes of the vulnerable entity include severity level, affected device types, typical attack scenarios in the power industry, vulnerability ID (CVE number / power industry vulnerability number), vulnerability type, affected device types, vulnerability description, remediation plan, and discovery time. (Vulnerability ID, vulnerability type, vulnerability description, remediation plan, and discovery time are also included.) Preferred business relationships between entities include: initiation, utilization, dissemination, and influence.
[0079] The attacking entity initiates the attack on the device entity, and the attributes of the initiation relationship include communication protocol and communication frequency; The attacking entity exploits the vulnerable entity, and the attributes of the exploitation relationship include the exploitation method; The propagation of device entities involves the propagation of other device entities, and the attributes of the propagation relationship include the propagation medium and the propagation time. The attacking entity affects the business entity, and the attributes of the affected relationship include the duration of business interruption and the type of data anomaly.
[0080] This invention constructs a three-dimensional semantic ontology for power control closed loops: Equipment entity: Defines the functional characteristics of power control equipment entities, including PLC entities (attributes include "set value zone number set", "current set value zone", "set value modification permission status", and "return calibration timeout"), RTU entities (attributes include "telemetry point table", "telecommunication signaling point table", "dead zone value", and "sampling period"), and smart meter entities (attributes include "load curve sampling interval", "demand period", and "communication protocol type"); defines the relationships "collection-collected" (RTU collects PLC data), "control-controlled" (PLC controls circuit breakers), and "reporting-receiving" (RTU reports to the dispatcher); Protocol Dimension Ontology: Defines semantic field entities for protocols such as IEC61850 and Modbus, including GOOSE message entities (attributes include "dataset reference (LD / LN)", "control bits (Test / ConfRev)", and "time to live (TAL)") and MMS message entities (attributes include "service object (Write / Read)", "variable access specification", and "data type"); defines the relationship "mapping-mapped" (protocol fields are mapped to device functions); Business Dimension Ontology: Defines atomic operation entities for power dispatching services, including control instruction entities (attributes include "instruction type (remote dispatch / remote control)", "target equipment", "expected parameter value", and "permission verification status") and dispatching process entities (attributes include "operation ticket number", "five-prevention verification result", and "dispatcher password verification result"); defines the relationship "trigger-execution" (instruction triggers equipment operation) and "verification-permission" (process verification result determines instruction execution); this invention constructs a three-dimensional semantic ontology of "equipment-protocol-service", realizing semantic-level deep fusion of multi-source power data and solving the problems of attack chain breakage and fragmented tracing.
[0081] It is worth noting that this invention, targeting the closed-loop characteristics of power system control (scheduling-execution-feedback), designs a three-dimensional semantic ontology encompassing equipment dimension (PLC / RTU / smart meter functional characteristics), protocol dimension (IEC61850 GOOSE / MMS, Modbus semantic fields), and business dimension (control commands, scheduling processes). It achieves cross-layer semantic fusion through an automatic alignment algorithm of "protocol fields - equipment functions - business operations." Compared to existing technologies that only perform "pseudo-fusion" at the field level, this invention can completely restore the vertical semantic association of "forged MMS message fields → PLC setting modification function → current protection setting adjustment business," solving the problem of attack chain breakage at the business stage and achieving full-link tracing of attack initiation, equipment impact, and business disturbance. Addressing the insufficient multi-source data fusion in traditional methods, this invention achieves vertical semantic fusion of "protocol parsing - equipment control - business operations" through a three-dimensional semantic ontology, rather than field-level splicing. For example, when parsing the MMS message "Write service + setting area 1 + address 0x0023", it automatically associates it with the "PLC-001 setting modification" function and the "current protection setting adjustment" business operation, completely reconstructing the actual impact path of the attack on the power business and effectively avoiding fragmented tracing. Addressing the poor adaptability of traditional methods to power scenarios, this invention designs a dedicated entity for power control equipment (defining attributes such as PLC setting area number and RTU telemetry point table) and fault mechanism constraints (protection action time limit and power flow direction constraints), which can accurately identify power-specific attacks. For example, by modeling dedicated fields such as "dataset reference (LD / LN)" and "control bit (Test / ConfRev)" in the GOOSE message, it associates the causal relationship between "replay message" and "circuit breaker maloperation", filling the gaps in power scenarios that general methods cannot cover.
[0082] To address the shortcomings of traditional methods in integrating multi-source data, this invention constructs a power-specific knowledge graph that links equipment logs, network traffic, business data, and other multi-dimensional data. This allows for the complete reconstruction of the entire attack chain, from initiation and propagation to business impact. For example, it captures the complete process from forged MMS messages and PLC parameter tampering to RTU anomalies and load scheduling deviations, effectively avoiding fragmented tracing. This is because the entity-relationship structure of the knowledge graph is naturally suited for multi-source data association, and the power-specific ontology model ensures the effectiveness of data fusion. Traditional methods, on the other hand, can only associate single-type logs and cannot cover the full dimensions of power system data, making it difficult to reconstruct the complete attack chain.
[0083] Step 3.2: Extract a set of new knowledge triples that conform to the ontology model from the triple dataset.
[0084] More preferably, step 3.2 includes: Step 3.2.1: Divide the triplet dataset into structured data and semi-structured data.
[0085] More preferably, step 3.2.1 includes: Determine whether the entities, basic relations, and basic attributes of the triples in the triple dataset directly correspond semantically to the entities, business attributes, and business relations in the ontology model. Semantics include, for example, `device_id` and `attack_time`. Specifically, for entity-basic relation-basic attribute triples, determine whether their entities, basic relations, and basic attributes can all be found to have direct corresponding entities, business relations, and business attributes in the ontology model; for entity-basic relation-entity triples, determine whether their head entity, basic relation, and tail entity can all be found to have direct corresponding entities and business relations in the ontology model. If the entity, basic relation, and basic attribute of a triple do not directly correspond to all entities, business relations, and business attributes in the ontology model, that is, at least one type (entity, basic relation, or basic attribute) in the triple cannot be directly corresponded to the ontology model, then the triple is determined to be semi-structured data. This means that its core semantic information cannot be directly mapped through the type names of entities, basic relations, and basic attributes, and needs to be routed to a semi-structured data parser for deep pattern matching parsing.
[0086] If the entity, basic relation, and basic attribute of the triple directly correspond to the entity, business relation, and business attribute in the ontology model, that is, all components of the triple can directly correspond to the ontology model, then it is further determined whether the target data item pointed to by the basic relation of the triple is atomic data that does not require secondary splitting.
[0087] If the target data item pointed to by the base relation is atomic data, the triple is determined to be structured data. If the target data item pointed to by the base relation is not atomic data, the triple is determined to be semi-structured data. Specifically, for the entity-base relation-base attribute triple, the target data item pointed to by the base relation refers to the value of the base attribute; for the entity-base relation-entity type triple, the target data item pointed to by the base relation refers to the identifier of the tail entity.
[0088] The standard for atomic data is that its value itself is a complete and indivisible semantic unit, such as, but not limited to, a specific device ID "Sub220-PLC-001" or a standard ISO timestamp "2024-06-10 09:20:00", without the need for internal parsing of its string content to obtain multiple information fragments.
[0089] If for an entity-basic relation-basic attribute triple, the target data item pointed to by the basic relation has an atomic value for the basic attribute, then the semantics of the basic attribute of the triple can be directly and completely represented by the key-value pair of "basic attribute field name - basic attribute field value", without the need to parse the value content. It can be routed to the structured data rule extractor to perform direct key-value mapping.
[0090] If the target data item pointed to by the base relation of a triple is not atomic, the triple is determined to be semi-structured data. If, for an entity-base relation-base attribute triple, the target data item pointed to by the base relation refers to a non-atomic value of the base attribute, such as, but not limited to, a free text sentence containing multiple information fragments such as device, action, and source IP, then it means that pattern parsing is still required to extract the complete semantics of the triple, and the triple is routed to the semi-structured data parser extractor.
[0091] Step 3.2.2: According to the preset mapping rules, the entities, basic relationships and basic attributes in the structured data are mapped to the entities, business relationships and business attributes corresponding to the ontology model, and the candidate knowledge triple set A is determined.
[0092] Based on the mapping rules, the entity type (such as attack entity, device entity) of the triple is identified in the ontology model, and its unique identifier is determined. For example, in an alarm record, the mapping rules map the value of the attack IP field to a unique identifier of the "attack entity" type.
[0093] Other fields in the record are mapped to the entity's business attributes or business relationships with other entities, according to the rules. For example, the attack type field is mapped to one business attribute of the attacking entity, and the time field is mapped to another business attribute.
[0094] For example, according to a structured alert record: {Attack IP: 192.168.1.100, Attack type: Port scan, Time: 2023-10-25 10:00:00}.
[0095] The mapping process is as follows: The rule identifies the attacking IP field and creates an attacking entity identified as "192.168.1.100". The rule maps the attack type field to the business attribute "attack type" of the attacking entity, with a value of "port scan", generating a triple: <attacking entity: 192.168.1.100, has attribute, attack type: port scan>. The rule maps the time field to the business attribute "2023-10-25 10:00:00" of the attacking entity and maps the "occurrence time" to the business relationship between the attacking entity and its business attribute, generating a triple: <attacking entity: 192.168.1.100, occurrence time, 2023-10-25 10:00:00>. The two triples generated above will be added to the candidate knowledge triple set A.
[0096] Structured data refers to data with a highly predefined data model, containing key-value pairs with strictly preset fields (such as load scheduling records obtained through scheduling system APIs, structured alarm data obtained through security device SDKs, exported vulnerability scan reports, and device status tables). Because its field names and locations are fixed, specific device entities, vulnerability entities, and other nodes can be extracted directly using preset "key-value mapping rules" or "header field mapping rules," and the basic relationship triples between entities can be accurately extracted incidentally. Mapping rules include "key-value mapping rules" or "header field mapping rules."
[0097] Step 3.2.3: Based on the regular expressions specific to the power scenario, the entities, basic relationships, and basic attributes in the semi-structured data are mapped to the entities, business relationships, and business attributes corresponding to the ontology model, thus determining the candidate knowledge triple set B.
[0098] Semi-structured data refers to text or byte streams that lack a globally unified data dictionary but contain specific delimiters or header structures to separate semantic elements. Examples include, but are not limited to, server process logs, raw IEC61850 GOOSE packets captured by network packet capture, and Modbus traffic. Because field lengths are variable, they cannot be read directly column-by-column; therefore, dedicated regular expressions or specific protocol parsing scripts must be designed for pattern matching.
[0099] For each specific semi-structured data format, a regular expression specific to the power sector is applied. This expression defines patterns for accurately matching and extracting target fields (such as timestamps, device IDs, IP addresses, and specific protocol fields) from complex text. The extracted field values are then mapped to corresponding entities, business relationships, and business attributes in the ontology model according to predefined rules.
[0100] For example, consider a typical power terminal operation alarm log: [2023-10-25 10:15:30][Sub220-PLC-001][ERROR] Unauthorized Access from IP: 192.168.1.50.
[0101] Set the dedicated regular expression as: \[?P <time>\d{4}-\d{2}-\d{2}\d{2}:\d{2}:\d{2})\]\[(?P<device_id> [A-Za-z0-9\-]+)\]\[(?P <level>[A-Z]+)\]. IP:(?P <ip>\d{1,3}(\.\d{1,3}){3}).
[0102] Among them, \[(?P <time>\d{4}-\d{2}-\d{2}\d{2}:\d{2}:\d{2})\] represents matching and extracting the timestamp, used to find [2023-10-25 10:15:30], and storing the time string 2023-10-25 10:15:30 into a variable named time;\[(?P<device_id> [A-Za-z0-9\-]+)\] indicates matching and extracting the device ID, used to find [Sub220-PLC-001], and storing the device identifier Sub220-PLC-001 into a variable named device_id;\[(?P <level>[AZ]+)\] indicates matching and extracting log levels, used to find [ERROR], and storing the log level ERROR in a variable named level. This indicates matching any text within the string. In this example, it's used to match "UnauthorizedAccessfrom," ensuring the flexibility of regular expressions and adapting even if the message wording changes. IP: (?P <ip>\d{1,3}(\.\d{1,3}){3}) means matching and extracting the IP address, used to find the IP: 192.168.1.50, and storing the IP address 192.168.1.50 into a variable named ip.
[0103] The application successfully matched the alarm log using a special regular expression, extracting the key-value pair: time: "2023-10-25 10:15:30", device_id: "Sub220-PLC-001", level: "ERROR", ip: "192.168.1.50".
[0104] According to the mapping rules, the value of device_id "Sub220-PLC-001" is identified as a device entity, the value of time is mapped to the business attribute of the device entity, and the "time of occurrence" is mapped to the business relationship between the device entity and its business attribute, generating an entity-business relationship-business attribute triple: <Device entity: Sub220-PLC-001, time of occurrence, 2023-10-25 10:15:30>.
[0105] The IP address "192.168.1.50" is identified as the attacking entity. Based on the semantics of the alarm log, the business relationship "abnormal access" between it and the above-mentioned device entity is established, and the entity-business relationship-entity triplet is generated: <Attacking entity: 192.168.1.50, abnormal access, device entity: Sub220-PLC-001>.
[0106] The input triplet dataset, which extracts a new set of knowledge triples conforming to the ontology model from the triplet dataset, means that the process does not begin with raw, messy multi-source data (such as log text or network packets), but rather with a highly structured intermediate data product that has already undergone entity identification, preliminary relation association, and attribute standardization. Using the power domain ontology model, the semantics of the previously generated, standardized entity-basic relation-basic attribute triples and entity-basic relation-entity triples are semantically enhanced to generate new knowledge triples containing business knowledge. The feasibility and accuracy of this mapping depend entirely on the unambiguous, structured facts already provided by the input entity-basic relation-basic attribute triples and entity-basic relation-entity triples.
[0107] The entity-basic relationship-basic attribute triple provides a characteristic description of the entity, such as, but not limited to: <Device: Sub220-PLC-001, includes status, trip signal> (describes device status), <Attack: 202.120.36.18, attack type is, port scan> (describes attack characteristics). The entity-basic relationship-entity triple provides a record of specific interactions that have occurred between entities, such as, but not limited to: <Attack IP: 192.168.5.201, Access, Device: Sub220-PLC-001> (a network access was recorded). The entity-basic relationship-basic attribute triples and the entity-basic relationship-entity triples use standard IDs like Sub220-PLC-001, standardized descriptions like "access" and "identification," and standardized values like "trip signal." First, through the shared entity "Sub220-PLC-001," the basic attribute describing its state (trip signal) is automatically associated with the basic relationship recording that it was accessed by IP 192.168.5.201, forming a preliminary event context: "At some time, IP 192.168.5.201 accessed the PLC, and subsequently, the PLC's state became tripped." This provides standardized operation objects for subsequent mapping rules and regular expressions, greatly reducing the complexity and error rate of directly extracting high-level semantics from unstructured data.
[0108] The entity-basic relationship-entity triple has already established the association between entities (such as IP accessing PLC). This means that when generating "business relationships" in subsequent mappings, there is no need to rediscover the associations. Instead, semantic "annotation" is performed on the basis of the existing factual associations. For example, basic relationship access, combined with the malicious label of the IP and business rules, can be mapped to malicious scanning or attack attempts in the business relationship.
[0109] Subsequently, the business rule base and threat intelligence are used for inference: The input entity-basic relationship-basic attribute triple: <Device: Sub220-PLC-001, including status, trip signal>; entity-basic relationship-entity triple: <Attack IP: 192.168.5.201, access, Device: Sub220-PLC-001>. The entity-basic relationship-basic attribute triple and the entity-basic relationship-entity triple are associated (based on the shared entity Sub220-PLC-001), and mapping rules (such as threat intelligence, vulnerability databases) are invoked for contextual reasoning. A mapping rule might specify that "access from an untrusted IP, if the subsequently associated device is in an abnormal state, should be considered a suspicious attack." A threat intelligence database query reveals that IP 192.168.5.201 is marked as a malicious scanning source. Based on this, semantic mapping is performed to generate new, business-level knowledge triples, including entity-business relationship-business attribute triples and entity-business relationship-entity triples. Example of an entity-business relationship-business attribute triple: <Device abnormal event: Sub220-PLC-001 trip, associated threat intelligence, malicious scanning source IP>. This triple adds a business attribute "associated threat intelligence" to the entity "trip event", with the value "malicious scanning source IP".
[0110] The entity-business relationship-entity triple establishes a semantic relationship with business implications between two entities. Example: <Attacking entity: 192.168.5.201, causing (business relationship), device malfunction event: Sub220-PLC-001 trip>. This triple establishes a business relationship of "causing" between the attacking entity and the device malfunction event, implying a causal judgment, replacing the basic "access" relationship. Here, "causing" is no longer a basic "access," but rather carries the business semantics of causal judgment.
[0111] The mapping process from basic facts to business semantics is not an independent text transformation step, but an indispensable core component of the hierarchical and interpretable knowledge extraction pipeline of this invention. It is closely coupled with the definitions of the two types of basic triples mentioned above, together forming the data and knowledge foundation that supports the creativity of the entire method.
[0112] This invention enables the computationalization of semantic mapping, as traditional methods struggle to reliably extract complex business semantics such as cause and utilization from unstructured logs. This invention decomposes tasks and uses stable, explicit parsing rules to generate standardized basic triples to determine what happened. Then, based on structured facts, domain knowledge mapping rules are applied for semantic association and annotation. This two-stage design, prioritizing syntax over semantics, transforms the problem of ambiguous semantic understanding into precise rule-based reasoning on structured data, significantly improving the accuracy and reliability of semantic information extraction. This is a prerequisite for supporting subsequent high-precision causal verification.
[0113] The newly generated knowledge triples are the direct objects of operation in the subsequent causal verification stage for "business rule consistency verification" and "physical feasibility verification".
[0114] Business rule validation relies on the explicit business relationships (such as "caused") and business attributes (such as "associated with vulnerabilities") in the new triplet as the basis for judgment. For example, to validate the rule "whether the setting modification was authorized by the scheduler", it is necessary to query the business relationship triplet <schedule instruction: X, applied to, device: Y> and the business attribute triplet <device: Y, current setting, Z>.
[0115] Physical feasibility verification relies on the business relationship chain between entities in the new triplet to define the propagation path to be verified. The verification of physical constraints such as latency and direction is carried out on the relationship chain such as <Entity A, Business Relationship R1, Entity B>, <Entity B, Business Relationship R2, Entity C>, etc.
[0116] This forms a complete and traceable knowledge evolution chain. The entire data processing flow is as follows: raw data → basic triples (standardized facts) → new knowledge triples (business semantics) → causal verification and attack chain. This chain is clear and traceable. The output of each layer is a high-quality, structured, and semantically rich input for the next layer. This design ensures that the final attack attribution conclusion does not originate from a "black box" model, but is built on progressive and interpretable knowledge reasoning, significantly enhancing the credibility and persuasiveness of the attribution results.
[0117] Semantic mapping and knowledge generation are crucial bridges connecting data standardization and intelligent causal reasoning. Using two defined basic triples as precise and unambiguous input, it generates structured business knowledge tailored for subsequent dual-constraint verification. This hierarchical knowledge construction system is the core architectural innovation enabling high-precision, high-reliability attack tracing, fundamentally different from traditional end-to-end methods that attempt to solve both grammatical and semantic problems directly from raw data. This constitutes a prominent and substantial feature of this application.
[0118] Step 3.2.4: Generate a new set of knowledge triples based on the entities in candidate knowledge triple set A and candidate knowledge triple set B. If the database does not have a knowledge graph, store the entities in candidate knowledge triple set A and candidate knowledge triple set B as a knowledge graph in the database; if the database has a knowledge graph, align the entities in candidate knowledge triple set A and candidate knowledge triple set B with the entities in the existing knowledge graph to generate a new set of knowledge triples.
[0119] By using the graph database index, query the existing knowledge graph to see if there is an entity node with the same identifier as the extracted entity; If an entity node with the same identifier as the extracted entity exists in the existing knowledge graph, the attributes and business relationships in the two candidate knowledge triple sets are appended to the existing entity node in the existing knowledge graph.
[0120] If there is no entity with the same identifier as the extracted entity in the existing knowledge graph, a new entity node is created in the existing knowledge graph, and the attributes and business relationships of the extracted entity are added to the new entity node.
[0121] Knowledge fusion and graph alignment involve aligning the entity nodes extracted in the above steps with existing knowledge graph entities. Before storing the extracted entities and triples into the graph, a rapid matching index is performed using unique device identifiers (such as MAC addresses, asset codes, and CVE vulnerability numbers). If a newly extracted entity identifier already exists in the graph, the newly extracted attributes or relationships are appended to the existing entity nodes for merging, avoiding duplicate entity storage and ensuring a graph link accuracy rate of ≥98%.
[0122] Step 3.3 involves knowledge fusion and semantic disambiguation of the new knowledge triple set to generate a fused knowledge set. Based on the power topology control rule base and the standard vulnerability threat intelligence base, the fused knowledge set is used to complete relationships and generate a power attack tracing knowledge graph. This power attack tracing knowledge graph is stored in the Neo4j graph database and indexed, then output in segments by substation.
[0123] More preferably, step 3.3 includes: Step 3.3.1: Perform semantic disambiguation and alias unification on the set of new knowledge triples, and output the unified set of new knowledge triples.
[0124] Newly extracted entities from the initial knowledge set are compared and merged with existing entities in the knowledge graph using unique identifiers such as device ID and CVE vulnerability number; different representations of the same entity in different data sources are unified into a standard name.
[0125] Unlike the standardization of "data format (field structure)" in Step 2, this step aims to eliminate naming ambiguity (data value standardization) for the same entity in different business systems. For example, it maps "forged MMS message" in security alerts and "forged IEC61850MMS message" in threat intelligence to a standard attack type name; or it standardizes the mapping of vulnerability descriptions and device aliases due to differences in capitalization and abbreviations, ensuring that the same physical entity or concept uniquely corresponds to one node in the graph, avoiding broken links in graph queries.
[0126] Step 3.3.2, based on the preset power topology control rule library and the standard vulnerability threat intelligence library, supplement the missing relationships in the unified new knowledge triple set.
[0127] More preferably, Step 3.3.2 includes: Based on the pre-constructed power topology control rule library, automatically complete the missing physical control relationships or logical control relationships between device entities.
[0128] Preferably, formalize the physical wiring common sense and control logic common sense in the power field into graph completion rules. For example, during knowledge extraction, if only two isolated entities, "overcurrent protection device operates" and "circuit breaker trips", are extracted and their direct association is not found in the log; based on the control logic rule "the protection device in a specific interval of the substation directly controls the corresponding circuit breaker to operate through hard wiring" in the preset rule library, the system automatically completes the missing relationship of <overcurrent protection device - control trigger - circuit breaker>; or based on the hierarchical control rule, automatically complete the communication control relationship of <dispatching master station - sending instructions to - substation RTU>.
[0129] Based on the pre-constructed standard vulnerability threat intelligence library, automatically complete the missing vulnerability exploitation relationships between device entities and vulnerability entities, and between attack entities and vulnerability entities.
[0130] Preferably, formalize historical attack data into structured national / industry vulnerability databases (such as CNNVD, State Grid vulnerability library) and threat intelligence indicators (IOC). For example, during knowledge extraction, the device entity "Sub220-PLC-001 (model: S7-1200, firmware version: V4.1)" and the alarm entity "attempt to exploit CVE-2023-XXXX vulnerability" in network traffic are extracted, but their association is lacking in the log; by querying the standard vulnerability threat intelligence library, the system determines that "the affected range of CVE-2023-XXXX vulnerability includes S7-1200 series versions V4.0 - V4.2", and thus based on the comparison of device models and versions, automatically completes the missing association relationships of <Sub220-PLC-001 - has physical vulnerability - CVE-2023-XXXX> and <alarm entity - intends to exploit - Sub~220-PLC-001>.
[0131] Step 3.3.3, use the Neo4j graph database to store the knowledge graph, create indexes for device ID, attack IP, and vulnerability ID to improve query efficiency, adopt a data sharding strategy to divide data shards by substation, and control the data volume of a single shard within 1 million nodes to ensure storage stability. The data sharding strategy is based on the performance test of the Neo4j graph database. When the number of nodes in a single shard ≤ 1 million, the query response time for attack event association ≤ 1 second, meeting the real-time traceability requirements of the new power system.
[0132] Indexes are created for the device ID attribute of device entities, the attack IP attribute of attack entities, and the vulnerability ID attribute of vulnerability entities in the knowledge graph to improve query efficiency; a data sharding strategy based on substations is adopted to store and manage the knowledge graph data, and the number of nodes in a single shard is controlled within a preset threshold to ensure the storage and query performance of the graph database.
[0133] This system is responsible for generating and storing structured knowledge. First, it designs an ontology model for the power industry, clarifying core entities such as attack entities, equipment entities, business entities, and vulnerability entities, as well as core relationships such as initiation, exploitation, propagation, and impact. Then, it extracts triples that conform to the ontology model from the preprocessed data, unifies the data format and naming conventions through conflict resolution, and finally stores the fused knowledge graph in the Neo4j graph database, forming a standardized power attack tracing knowledge graph, which directly provides structured graph data support for subsequent attack causal graph construction and reasoning.
[0134] Unlike step 2, which only unifies heterogeneous formats and parses basic fields at the physical data level, step 3 aims to map preprocessed surface data (such as IP, MAC, and protocol codes) to the power control closed-loop ontology model at the semantic level. Through logical reasoning and semantic disambiguation, a knowledge graph with power business awareness capabilities is constructed. This layered architecture, which first cleans the underlying data and then integrates the upper-level knowledge, effectively solves the computational overload problem caused by directly performing complex ontology reasoning on massive concurrent power data.
[0135] It is worth noting that this invention employs a single-machine deployment of the Neo4j graph database with sharded index optimization, integrating semantic fusion, causal verification, and causal graph generation modules into edge computing nodes, thus avoiding the cross-node communication overhead of federated learning. Attack tracing is achieved on substation edge nodes (ARM architecture, 4GB memory), with inference latency ≤1 second and model size ≤500MB, meeting the deployment requirements of new power systems for real-time handling (≤1 second response) and resource-constrained deployment (≤8GB memory).
[0136] Step 4: Based on the abnormal equipment events, candidate events and candidate causal edges are selected from the power attack tracing knowledge graph to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between the abnormal equipment events and each candidate event. Directed candidate causal edges that simultaneously satisfy the constraints of power business rules and physical feasibility constraints form an attack causal chain. Based on the abnormal equipment events and the attack causal chain, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
[0137] In a preferred but non-limiting embodiment of the present invention, step 4 includes: Step 4.1: Within the power attack tracing knowledge graph, candidate events that are logically related to device anomaly events within a preset time window are selected to form a candidate event set.
[0138] More preferably, step 4.1 includes: Step 4.1.1: Receive equipment abnormal events input by maintenance personnel. The equipment abnormal event information includes equipment ID, abnormality type, and abnormality time. Step 4.1.2: Call Neo4j's Cypher query statement to query the knowledge graph. Use the knowledge graph to filter candidate events within the device anomaly event information that have a logical connection within 30 minutes before and after the anomaly time. The preset time window is set to 30 minutes before and after the anomaly time of the device anomaly event.
[0139] Step 4.1.3: Based on the principle of "time proximity + data correlation", candidate events that occurred within 30 minutes before the anomaly and have a logical connection with the equipment anomaly are selected to form a candidate event set.
[0140] Step 4.2: Using candidate events in the candidate event set as nodes, construct candidate causal edges based on time proximity and logical association rules to generate an attack causal graph.
[0141] More preferably, step 4.2 includes: Each event in the candidate event set is set as an event node. The event node attributes include event ID (format: Event-Date-Sequence Number), event type, occurrence time, associated entity, and event details. Based on the rules of temporal proximity and logical association, candidate causal edges are established between event nodes to form a preliminary attack causal graph.
[0142] Edge construction rules: If event a occurs earlier than event b and the time interval is ≤5 minutes, establish a preliminary related edge; if event a and event b have a logical relationship or the related entities are the same, strengthen the related edge and mark the edge attribute "data association basis", and finally construct the attack causal graph.
[0143] Step 4.3: For candidate causal edges in the attack causal graph, the Do-Calculus method is used for intervention testing to determine the average causal effect between the device malfunction event and the candidate event. Based on the average causal effect, directed candidate causal edges are extracted to form a set of directed causal edges. The candidate events associated with the directed candidate causal edges are checked for consistency with the power business rules through power business rule constraints. The candidate events associated with the directed candidate causal edges that pass the check are marked as suspicious attack paths. The suspicious attack paths are verified for physical feasibility through physical feasibility constraints. The suspicious attack paths that pass the physical feasibility verification are taken as the verified attack causal chains.
[0144] More preferably, step 4.3 includes: Step 4.3.1: Use the Do-Calculus method to intervene in the candidate events corresponding to the candidate causal edges in the attack causal graph (e.g., event X → event Y). Based on historical data, calculate the average causal effect of the device malfunction event Y before and after intervening in the candidate event X.
[0145] The core concept is to calculate the average causal effect ACE = P(Y|do(X)) - P(Y|do(X')). Here, P(Y|do(X)) represents the probability of the target event Y occurring when event X is artificially forced to occur, and P(Y|do(X')) represents the probability of the target event Y occurring when event X is artificially prevented from occurring.
[0146] Step 4.3.2: Based on the relationship between the average causal effect and the preset significance threshold, determine the direction of the candidate causal edges, and select directed candidate causal edges with a clear unidirectional causal direction to convert the undirected edges in the initial causal graph into directed edges with a clear causal direction.
[0147] If the average causal effect ACE value exceeds the preset significance threshold, then the statistical causal effect of event X on event Y is determined to be significant.
[0148] For example, if tests verify that the causal effect of X→Y is significantly higher than that of Y→X, then the causal direction is determined to be X leading to Y. This step transforms the potentially ambiguous bidirectional associations established based on time sequence in the initial causal graph into unidirectional edges with clear causal directions.
[0149] It is worth noting that this invention introduces a Do-Calculus causal verification mechanism embedded in business rules, achieving dual-constraint judgment of "statistical causality + business logic," significantly reducing the false positive rate in tracing the source. This invention formalizes power operation specifications (such as IEC61850 control service sequence, dispatching procedures, and five-prevention logic) into prior constraints for causal verification, simultaneously verifying the statistical causal effect and compliance with business rules during Do-Calculus intervention testing. Through a three-step process of "intervention testing - result inversion - business rule consistency verification," it can distinguish between "attack tampering" and "normal operation and maintenance" (such as identifying whether "set value modification has been authorized by dispatch"), reducing the false alarm rate of misjudging normal operation and maintenance as an attack from 23% in traditional methods to 3%, accurately locating the true source of the attack. Addressing the high false positive rate of traditional methods, this invention embeds power business rules into the Do-Calculus causal verification process, adding business rule consistency verification on top of statistical causal significance testing. For example, both normal modification of settings by maintenance personnel and tampering with settings by attackers are statistically manifested as "change in settings → equipment abnormality". However, the former conforms to the business process of "scheduling instruction - five-prevention verification - operation ticket execution". This invention can distinguish between the two through business rule constraints, which can greatly reduce the probability of misjudging normal maintenance as an attack.
[0150] To address the high false positive rate of traditional methods, this invention introduces Do-Calculus causal reasoning. Through intervention testing, it verifies the necessary correlation between events. For example, after blocking the attacking IP communication, PLC parameters no longer show abnormalities, thus ruling out accidental correlations such as normal maintenance. Traditional methods rely on time and statistical correlation, only reflecting the correlation between events and failing to distinguish between accidental and necessary correlations, easily misjudging normal maintenance operations as attacks. Causal reasoning, on the other hand, can deeply explore the essential causal relationships between events, significantly reducing the false positive rate in tracing the source. After filtering the candidate event set related to abnormal events from the knowledge graph, this invention introduces Do-Calculus to construct and verify the attack causal graph. Through intervention testing and other causal reasoning methods, it distinguishes between necessary causal relationships and accidental temporal correlations between events, reducing the false positive rate of misjudging normal maintenance operations (such as planned remote maintenance) as malicious attacks and improving the accuracy of locating the true source of the attack (external IP, internal compromised equipment).
[0151] Step 4.3.3: Based on the directed candidate causal edges, construct multiple attack causal chains. Perform consistency verification between the candidate events corresponding to each attack causal chain and the power business rules. If the consistency verification passes, determine the rationality of the attack causal chain. If the consistency verification fails, filter out the corresponding attack causal chain.
[0152] More preferably, step 4.3.3 includes: If a path fully complies with all business rules (e.g., there is a valid scheduling instruction before "PLC setting modification"), it is marked as a "normal and compliant operation path".
[0153] If a path has a significant statistical causal effect (high ACE value) but violates business rules (e.g., no valid scheduling instruction before "PLC setting modification"), then the path is marked as a "suspicious attack path" because its behavior pattern deviates from the normal operating procedure.
[0154] Power business rule constraints are used to determine whether candidate events associated with directed candidate causal edges conform to predefined power business logic rules. Each rule includes preconditions, time constraints, and postconditions. For example, rule r_1 is "If a PLC setting modification event occurs, there must be a record of dispatch instruction verification passing within 5 seconds", and rule r_2 is "If a GOOSE trip command is sent, a circuit breaker change confirmation frame must be received within 100 milliseconds". If the average causal effect meets the statistically significant causal effect but violates the business rules, it is judged as statistical causality but business abnormal and marked as a suspicious attack path.
[0155] Step 4.3.4 involves performing physical feasibility verification on reasonable attack causal chains. Attack causal chains that pass the physical feasibility verification are considered valid attack causal chains. Physical feasibility constraints include delay compatibility verification, causal directionality verification, and cascading failure mode matching verification.
[0156] More preferably, step 4.3.4 includes: The delay compatibility check is used to determine if the actual time interval between the candidate events associated with each directed candidate causal edge in the attack causal graph is greater than or equal to the minimum propagation delay, based on the predefined minimum propagation delay of the power system attack action.
[0157] The time delay compatibility verification utilizes known time delay parameters of the power system to establish a time delay compatibility test, including protection action time limits (line protection ≤ 100ms, main transformer protection ≤ 200ms), communication transmission delays (GOOSE ≤ 4ms, MMS ≤ 100ms, Modbus ≤ 10ms), and equipment response delays (PLC setting modification ≤ 1s, circuit breaker tripping ≤ 60ms). If the time difference Δt between two events exceeds the physically possible propagation delay range, even if statistically correlated, it is determined to be non-causal. Causality directionality verification is used to determine if the direction of each directed candidate causal edge in the attack causal graph is consistent with the allowed direction defined by the network topology and the direction of business data flow, based on a predefined network topology and business data flow.
[0158] Causality directionality verification, based on the power flow direction (power source → load) and control command flow direction (dispatch → station control → bay → process) of the power system, verifies whether the direction of suspicious attack paths is consistent with the predefined allowed directions and prohibits the generation of reverse causal edges.
[0159] The cascaded fault mode matching verification is used to determine if a directed candidate causal edge passes the cascaded fault mode matching verification if it matches any known propagation mode in the pre-set library of typical power system faults and attack propagation modes.
[0160] By using cascading fault mode matching constraints, the historical fault case library is called to compare whether the suspicious attack path conforms to the propagation pattern of known non-offensive events such as equipment failure and natural disturbance. If it matches the propagation pattern of non-offensive events, the suspicious attack path is eliminated.
[0161] The cascading failure mode matching constraint includes constructing typical failure propagation modes based on the N-1 security analysis of power systems and a historical failure case library, and performing pattern matching verification when generating the causal graph—only retaining attack propagation paths that conform to known failure modes.
[0162] If a directed candidate causal edge passes the time delay compatibility check, causal directionality check, and cascading failure mode matching check in sequence, it is determined to meet the physical feasibility constraint.
[0163] It is worth noting that this invention generates an attack causal graph based on power fault mechanism constraints, ensuring the physical credibility of the attack chain and improving the ability to identify power-specific attacks. This invention predefines directional constraints on causal edges (power flow direction, control command flow direction), time delay compatibility constraints (protection action time limit, communication transmission delay), and cascading fault mode matching (N-1 security analysis mode), transforming attack path generation from free exploration to physically guided constraint optimization. The generated attack chain conforms 100% to the physical laws of power system fault propagation, without anomalies such as reverse causal timeouts or delays. It can accurately identify power-specific attacks such as GOOSE message replay leading to circuit breaker malfunction and MMS forgery and setting tampering, avoiding false paths caused by the lack of physical constraints in general methods.
[0164] Step 4.4: Starting from the abnormal device event, trace the attack causal chain backwards along the verified path, query the power attack tracing knowledge graph to obtain the business relationship corresponding to the abnormal device event, thereby determining the source of the attack and reconstructing the attack chain from the source of the attack to the device corresponding to the abnormal device event.
[0165] More preferably, step 4.4 includes: Step 4.4.1: Starting from the node corresponding to the abnormal device event, traverse backward along the causal edges in the verified attack causal chain. During the traversal, query the power attack tracing knowledge graph simultaneously. Through the business relationships of initiation, propagation, exploitation, and impact in the graph, obtain the causal information of the attack source IP, the initial compromised device, the exploited vulnerability, the attack tools, and the scope of the affected business, and generate an enhanced attack chain.
[0166] Querying the knowledge graph for tracing the origins of power attacks includes: querying the detailed path of the attack's lateral movement within the system, the exploited vulnerabilities (CVE numbers), and the attack tools through propagation relationships; locating the initial attack source and obtaining information such as the attacker's IP, port, and location through initiation relationships; determining the scope and consequences of the affected business through impact relationships; and clarifying the details of the exploited vulnerabilities and the specific affected equipment assets through exploitation relationships.
[0167] Step 4.4.2 integrates the causal information in the enhanced attack chain according to the preset attack lifecycle stages, determines the attack source, and forms a complete attack chain from the attack source to the device corresponding to the abnormal device event, forming a complete attack chain description of the attack entry point, attack action, involved assets, and exploited vulnerabilities at each stage.
[0168] It is worth noting that existing source tracing results lack causal explanations that guide defense decision-making. While they output attack nodes, they only provide attack node sequences or risk decision points such as attacking IPs and compromised devices, lacking a visual representation of the attack propagation path and explanations of causal relationships. They fail to provide causal explanations such as "why this node is a critical propagation point" or "which edge to block to most effectively curb the attack." This makes it difficult for operations and maintenance personnel to directly deduce defense measures from source tracing results, understand the attack propagation logic, and optimize defense strategies accordingly, resulting in insufficient practicality. This invention addresses these technical problems by providing causal explanations that guide defense decision-making, achieving closed-loop decision support from attack source tracing to defense hardening. This invention not only outputs attack node sequences but also provides causal explanations of why this node is a critical propagation point and which edge to block to most effectively curb the attack. Through weak link causal localization, causal effect prediction of defense intervention (Do-Calculus calculation of the causal effect of defense measures), and minimal cut set defense strategy generation, it directly outputs executable defense suggestions such as blocking weak passwords on a server to cut off the attack path. This supports operations and maintenance personnel in optimizing defense strategies and solves the problem of insufficient practicality in existing source tracing methods that cannot guide defense decisions. To address the issue of low practicality of traditional methods, this invention achieves a closed loop from source tracing to defense through a defense decision-oriented causal explanation. Traditional methods only output isolated conclusions related to the attacking IP and the PLC device, while this invention outputs a complete causal chain: the attacking IP exploits a weak server password vulnerability → laterally propagates to the PLC → tampering with setpoints leads to business anomalies, significantly improving practicality.
[0169] Step 4.4.3: Based on the complete attack chain narrative, execute defense decisions and generate an attack attribution report.
[0170] Defense decision support analysis includes: identifying key defense failure points in the complete attack chain and marking them as high-priority hardening points; using the Do-Calculus model to predict the causal effect of implementing corresponding defense measures on each high-priority hardening point on blocking the attack chain; when there are multiple propagation paths of the attack, calculating the minimum set of nodes required to block all attack paths based on the minimum cut set algorithm, and generating a defense strategy recommendation with optimal cost; the attack attribution report should at least include a complete attack chain description, attack source location information, causal verification basis, and defense strategy recommendations.
[0171] Embodiment 2 of the present invention provides a power system causality verification attack tracing system, which runs a power system causality verification attack tracing method of Embodiment 1, including: The data acquisition module is used to collect multi-source data from the power system; The data preprocessing module is used to extract triples of entity-basic relation-basic attribute and entity-basic relation-entity, clean the data of the two types of triples, and generate triple datasets. The knowledge graph construction module is used to build a knowledge graph for tracing the source of power attacks based on the triple dataset, combined with the business attributes of entities and the business relationships between entities; The attack localization module is used to filter candidate events and candidate causal edges from the power attack tracing knowledge graph based on abnormal equipment events to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between abnormal equipment events and each candidate event. Directed candidate causal edges that simultaneously satisfy power business rule constraints and physical feasibility constraints form an attack causal chain. Based on abnormal equipment events and attack causal chains, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
[0172] The generated triplet dataset includes: Based on the standardized fields of device entities, attack entities, business entities, and vulnerability entities in the preset entity field dictionary, device entities, attack entities, business entities, and vulnerability entities are extracted from multi-source data respectively. Based on the preset entity-basic relationship-basic attribute triplet mapping table, the protocol fields and log fields parsed from the multi-source data are mapped to the basic attributes of the corresponding entities, and the entity-basic relationship-basic attribute triplet is generated with the entity, the basic relationship between the entity and its basic attributes, and the basic attributes of the entity. Based on the preset entity-basic relation-entity triplet mapping table, the data representing the relationship between entities in the multi-source data are mapped to the basic relations between entities, and entity-basic relation-entity triplets are generated based on entities and the basic relations between entities; Construct an intermediate dataset of triples based on entity-base relation-base attribute triples and entity-base relation-entity triples; Clean the intermediate dataset of triples and output the cleaned dataset of triples.
[0173] Based on the triplet dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing the origins of power attacks is constructed, including: Based on each entity, its business attributes, and the business relationships between entities, an ontology model for a knowledge graph of power attack tracing is constructed. Extract a set of new knowledge triples that conform to the ontology model from the triple dataset; The new knowledge triple set is fused and semantically disambiguated to generate a fused knowledge set. The fused knowledge set is then used to complete the relationships and generate a knowledge graph for tracing the source of power attacks.
[0174] The set of new knowledge triplets that conform to the ontology model extracted from the triplet dataset includes: The triplet dataset is divided into structured data and semi-structured data; According to the preset mapping rules, the entities, basic relationships and basic attributes in the structured data are mapped to the entities, business relationships and business attributes corresponding to the ontology model, and the candidate knowledge triple set A is determined. Based on regular expressions specific to the power scenario, entities, basic relationships, and basic attributes in semi-structured data are mapped to entities, business relationships, and business attributes corresponding to the ontology model, thus determining the candidate knowledge triple set B. Generate a new set of knowledge triples based on the entities in the candidate knowledge triple set A and the candidate knowledge triple set B.
[0175] The triplet dataset is divided into structured data and semi-structured data, including: Determine whether the entities, basic relations, and basic attributes of the triples in the triple dataset directly correspond semantically to the entities, business relations, and business attributes in the ontology model. Specifically, for entity-basic relation-basic attribute triples, determine whether their entities, basic relations, and basic attributes can all be found to have direct corresponding entities, business relations, and business attributes in the ontology model; for entity-basic relation-entity triples, determine whether their head entity, basic relation, and tail entity can all be found to have direct corresponding entities and business relations in the ontology model. If the entity, basic relation, and basic attribute of a triple do not directly correspond to all the entities, business relations, and business attributes in the ontology model, then the triple is determined to be semi-structured data. If the entity, basic relation, and basic attribute of a triple directly correspond to the entity, business relation, and business attribute in the ontology model, then it is determined whether the target data item pointed to by the basic relation in the triple is atomic data. If the target data item pointed to by the basic relation is atomic data, the triple is determined to be structured data. If the target data item pointed to by the basic relation is not atomic data, the triple is determined to be semi-structured data. Specifically, for an entity-basic relation-basic attribute triple, the target data item pointed to by the basic relation refers to the value of the basic attribute; for an entity-basic relation-entity type triple, the target data item pointed to by the basic relation refers to the identifier of the tail entity.
[0176] Physical feasibility constraints include time delay compatibility verification, causal directionality verification, and cascading failure mode matching verification. Among them, the delay compatibility check is used to determine that the directed candidate causal edge passes the delay compatibility check if the actual time interval between the candidate events associated with each directed candidate causal edge in the attack causal graph is greater than or equal to the minimum propagation delay, based on the predefined minimum propagation delay of the power system attack action. Causality directionality verification is used to determine if the direction of each directed candidate causal edge in the attack causal graph is consistent with the allowed direction defined by the network topology and the direction of business data flow, based on a predefined network topology and business data flow. Cascaded fault mode matching verification is used to determine that a directed candidate causal edge passes the cascaded fault mode matching verification if it matches any known propagation pattern in the pre-set library of typical power system faults and attack propagation patterns based on a pre-set library of typical power system faults and attack propagation patterns. If a directed candidate causal edge passes the time delay compatibility check, causal directionality check, and cascading failure mode matching check in sequence, it is determined to meet the physical feasibility constraint.
[0177] Locating the source of an attack and reconstructing the attack chain within a knowledge graph for tracing power system attacks includes: Within the knowledge graph for tracing power attacks, candidate events that are logically related to abnormal device events within a preset time window are selected to form a candidate event set; Using candidate events in the candidate event set as nodes, candidate causal edges are constructed based on time proximity and logical association rules to generate an attack causal graph; For candidate causal edges in the attack causal graph, the Do-Calculus method is used to conduct intervention tests to determine the average causal effect between the device malfunction event and the candidate event. Based on the average causal effect, directed candidate causal edges are extracted to form a set of directed causal edges. The power business rules are used to verify the consistency of the power business rules for candidate events associated with directed candidate causal edges. Candidate events associated with directed candidate causal edges that pass the verification are marked as suspicious attack paths. Suspicious attack paths are physically feasible through physical feasibility constraints. Suspicious attack paths that pass the physical feasibility verification are used as the verified attack causal chains. Starting from the abnormal equipment event, trace the attack causal chain backwards along the verified path, query the power attack tracing knowledge graph to obtain the business relationship corresponding to the abnormal equipment event, thereby determining the source of the attack and reconstructing the attack chain from the source of the attack to the device corresponding to the abnormal equipment event.
[0178] The verified attack causal chains include: The Do-Calculus method is used to intervene in the candidate events corresponding to the candidate causal edges in the causal graph. Based on historical data, the average causal effect of the device malfunction event before and after the intervention in the candidate events is calculated. Based on the relationship between the average causal effect and the preset significance threshold, the direction of the candidate causal edge is determined, and directed candidate causal edges with a clear unidirectional causal direction are selected. Multiple attack causal chains are constructed based on directed candidate causal edges. The candidate events corresponding to each attack causal chain are checked for consistency with the power business rules. When the consistency check is passed, the rationality of the attack causal chain is determined. When the consistency check is not passed, the corresponding attack causal chain is filtered out. A reasonable attack causal chain is physically feasible. An attack causal chain that passes the physical feasibility verification is considered a valid attack causal chain.
[0179] Embodiment 3 of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements a power system causal verification attack tracing method according to Embodiment 1.
[0180] Embodiment 4 of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a power system causal verification attack tracing method according to Embodiment 1.
[0181] Embodiment 5 of this invention, combined with an attack tracing scenario at a 220kV substation, provides a detailed explanation of the power system causal verification attack tracing method of Embodiment 1. In this embodiment, the attack involves an external attacker altering the parameters of the PLC equipment within the substation by forging IEC61850MMS messages, leading to abnormal RTU status and load dispatch deviations. This invention achieves complete reconstruction of the attack chain and source location through multi-module collaboration.
[0182] Three distributed data acquisition agents were deployed at the 220kV substation and the upper-level dispatch center (deployed in the No. 1 main transformer bay, the dispatch data network gateway, and the safety operation and maintenance platform, respectively), employing three types of acquisition methods to achieve full-dimensional data collection: (1) Equipment log acquisition: The PLC equipment control instruction log numbered "Sub220-PLC-001" is read through the IEC61850MMS protocol (port 102), and is collected once every 1 second to obtain key data such as "Instruction ID: CMD-20240610-008, Instruction content: Modify current protection setting to 1800A, Execution result: Success, Parameter change: Original 1200A→1800A"; the log numbered "Sub220-PLC-001" is collected through the Modbus RTU protocol (baud rate 9600bps). The status data of the RTU device "220-RTU-003" was obtained as follows: "Voltage: 221kV, Current: 1500A, Running status code: 0x02 (abnormal)". The operation and maintenance server with the number "Sub220-SVR-002" was logged in via SSH protocol (port 22) and the process log was collected. It was found that "Process name: mms-simulator.exe, Start time: 2024-06-10 09:15:32, Network connection: Established TCP connection with 192.168.5.201".
[0183] (2) Network data acquisition: Deploy a network packet capture module on the mirror port of the core switch of the substation, enable the ring buffer (1GB capacity) to store the original message, and parse the IEC61850 GOOSE message (data set identifier: LD01 / LLN0.StaPow, sender ID: 00-1A-2B-3C-4D-5E, data value: 0x01 (trip signal)), Modbus protocol traffic (slave address: 1, function code: 0x06 (write a single register), register address: 0x0023, data content: 0x0708 (corresponding to current setting 1800A)) and TCP / IP communication data (source IP: 192.168.5.201, destination IP: 10.20.30.40 (PLC device address), port number: 102, communication frequency: 5 times / minute).
[0184] (3) Business and security data collection: Load scheduling records are obtained through the scheduling system API interface, and the following information is obtained: "Business ID: Biz-20240610-012, Business type: Current protection setting adjustment, Control instruction: Maintain 1200A, Execution result: Abnormal (actual 1800A), Issuance time: 2024-06-10 09:20:00, Associated device ID: Sub220-PLC-001"; Alarm data is collected through the Intrusion Detection System (IDS) SDK interface, and the following information is obtained: "Attack type: Forged IEC61850MMS message, Trigger time: 2024-06-10 09:15:40, Suspected IP: 192.168.5.201, Associated vulnerability ID: CVE-2023-XXXX (PLC device MMS protocol authentication bypass vulnerability)". Data transmission is encrypted with SSL / TLS1.3 to avoid data leakage during transmission.
[0185] Based on standardized rules specific to the power sector, "entity-relationship-attribute" triples are extracted from the collected heterogeneous data: (1) Field uniformity: In accordance with the core field dictionary specification data format, the uniform field of the device entity "Sub220-PLC-001" is "Device ID: Sub220-PLC-001, Device type: PLC, Model: S7-1200, Deployment location: 220kV No.1 main transformer bay, Operating status: Abnormal"; the uniform field of the attack entity is "Attack ID: Attack-20240610-001, Attack type: Forged MMS message, Attack IP: 192.168.5.201, Attack port The attack time was 2024-06-10 09:15:32, and the vulnerability ID was CVE-2023-XXXX. The unified field of the business entity "Biz-20240610-012" is "Business ID: Biz-20240610-012, Business type: Current protection setting adjustment, Control instruction content: Maintain 1200A, Instruction execution result: Abnormal, Execution time: 2024-06-10 09:20:00, Associated device ID: Sub220-PLC-001".
[0186] (2) Protocol parsing: Parse the IEC61850 GOOSE message "Application identifier: 0x81, dataset reference: LD01 / LLN0.StaPow, data value: 0x01", generating a triple <GOOSE message - contains - dataset identifier: LD01 / LLN0.StaPow>; Parse the Modbus protocol "Slave address: 1, function code: 0x06, register address: 0x0023", generating a triple <Modbus traffic - modify - register address: 0x0023>; For unstructured server process logs, extract fields through the regular expression "Process name: (.?), startup time: (.?), network connection: established with (. ?)", and convert them to JSON format for storage. An example is as follows: { "Entity type": "Process", "Process name": "mms - simulator.exe", "Startup time": "2024 - 06 - 10 09:15:32", "Associated IP": "192.168.5.201", "Associated device": "Sub220 - SVR - 002" } In view of the problem that the traditional method has poor adaptability to the power scenario, the present invention designs parsing rules for power - specific protocols, and at the same time adapts to device characteristics such as PLC parameter over - range thresholds and RTU status codes, and can accurately identify power - specific attacks, such as attacks like GOOSE message replay causing circuit breaker misoperation. The rules of traditional general tracing methods do not cover the details of power protocols and device characteristics, and cannot effectively identify attack types unique to power systems; while the scenario - based design of the present invention exactly fills this gap and improves the tracing ability in the power scenario.
[0187] Ensure data quality through a three - level cleaning mechanism: (1) Duplicate data deletion: Calculate the MD5 hash value for each piece of collected data. For example, the hash value of the PLC control instruction log "CMD - 20240610 - 008" is "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6". Query through the Redis cache hash table, and no duplicate records are found, so the data is retained; For 3 duplicate Modbus messages (with the same hash value) in network packet captures, they are determined as duplicate data and deleted.
[0188] (2) Invalid data filtering: Set invalid data rules to remove null values or abnormal data such as "the voltage value in the RTU status data is 0kV" and "there is no attack type field in the attack alarm data"; monitor the proportion of invalid data in real time. In this case, the proportion of invalid data is 1.2%, which does not exceed 5%, so there is no need to trigger the acquisition agent fault alarm.
[0189] (3) Data fusion verification: Verify the correlation of data based on power business rules. For example, if there is a logical contradiction between "load dispatching instruction requires maintaining a current setpoint of 1200A" and "PLC actual parameter is modified to 1800A", mark it as "data to be verified" and push it to the operation and maintenance personnel for manual review. After review, if it is confirmed that the contradiction is an anomaly caused by an attack, add the triple <attack-cause-business anomaly: Biz-20240610-012> to the dataset.
[0190] Construction and Implementation of Knowledge Graph for Tracing Power Source Attacks: Ontology model instantiation: Based on the ontology model of four core entities—"attack, device, business, and vulnerability"—entities and relationships are instantiated in this case study: (1) Core entity instance: Attacker: Attack-20240610-001 (Attack type: Forged MMS message, Attacker IP: 192.168.5.201, Attack time: 2024-06-10 09:15:32, Vulnerability ID exploited: CVE-2023-XXXX) Equipment entities: Sub220-PLC-001 (Equipment type: PLC, Model: S7-1200, Deployment location: 220kV main transformer bay No.1, Operating status: Abnormal), Sub220-SVR-002 (Equipment type: Server, Model: RH2288H, Deployment location: Dispatch room, Operating status: compromised), Sub220-RTU-003 (Equipment type: RTU, Model: DTU3300, Deployment location: Main transformer control cabinet No.1, Operating status: Abnormal) Business Entity: Biz-20240610-012 (Business Type: Current Protection Setpoint Adjustment, Control Command: Maintain 1200A, Execution Result: Abnormal, Associated Device ID: Sub220-PLC-001) Vulnerable entity: CVE-2023-XXXX (Vulnerability type: authentication bypass, Affected device type: S7-1200 series PLC, Remediation: upgrade firmware to V4.5.0, Discovery time: 2023-11-05) (2) Examples of core relationships: Initiation: <Attack-20240610-001-Initiation (Communication Protocol: IEC61850MMS, Communication Frequency: 5 times / minute)-Sub220-PLC-001> Exploitation: <Attack-20240610-001-Exploitation (Vulnerability Exploitation Method: Forged Authentication Message)-CVE-2023-XXXX> Propagation: <Sub220-SVR-002-Propagation (Propagation Medium: TCP Connection, Propagation Time: 2024-06-10 09:15:30)-Sub220-PLC-001> Impact: <Attack-20240610-001-Impact (Business Interruption Duration: 15 minutes, Data Abnormality Type: Parameter Tampering)-Biz-20240610-012> Knowledge Extraction and Fusion: (1) Knowledge Extraction: For structured vulnerability scanning reports (including details of CVE-2023-XXXX vulnerabilities), use rule-based extraction to generate triples such as <Vulnerability-CVE-2023-XXXX-Attribute-Vulnerability Type: Authentication Bypass> for fields like "Vulnerability ID: CVE-2023-XXXX, Vulnerability Type: Authentication Bypass"; for semi-structured PLC logs, design a regular expression "Instruction ID: (.?), Instruction Content: (.?), Execution Result: (. ?)” to extract fields and generate triples <PLC-Sub220-PLC-001-Execution-Instruction: CMD-20240610-008>; Align the extracted entity "Sub220-PLC-001" with existing entities in the knowledge graph through device ID indexing, with a link accuracy rate of 99.2% to avoid duplicate entity storage.
[0191] (2) Knowledge Fusion: In the conflict resolution stage, unify the expression of "attack type" (unify "Forged MMS message attack" to "Forged IEC61850MMS message attack") and the device ID format (ensure that all device IDs follow the rule of "substation number-device type-sequence number"); in the knowledge completion stage, based on historical attack data (the same model of PLC was attacked due to CVE-2023-XXXX vulnerability in 2023), supplement the triple <Sub220-PLC-001-Exists-Vulnerability: CVE-2023-XXXX>; in the storage stage, use Neo4j graph database to create indexes for device ID, attack IP, and vulnerability ID, and divide data shards by substation (the data in this case is classified into the "220kV substation shard" with a data volume of 860,000 nodes) to ensure that the query response time ≤ 1 second.
[0192] Implementation of Causal Reasoning Attack Tracing: Location of suspected attack: The maintenance personnel discovered an "Abnormal Status of Sub220-RTU-003" (Abnormal type: current monitoring value exceeds threshold, abnormal time: 2024-06-10 09:20:15), which served as the trigger condition: (1) Call Neo4j's Cypher query statement: MATCH(d: Device{device_id: "Sub220-RTU-003"}) <- [r: associated] - (e: Event) WHERE e.event_time BETWEEN "2024-06-10 09:10:15" AND "2024-06-10 09:30:15" RETURN (2) Filter out the candidate event set, including "PLC device Sub220-PLC-001 parameter modification event (09:15:32)" "Server Sub220-SVR-002 starts mms-simulator.exe process event (09:15:32)" "Attack IP192.168.5.201 communicates with PLC event (09:15:30-09:15:40)" "Load dispatching business abnormal event (09:20:00)".
[0193] Attack causality graph construction: (1) Node definition: The four events in the candidate event set are used as nodes. For example, node 1: "PLC parameter modification event (Event-20240610-001, event type: parameter tampering, occurrence time: 09:15:32, associated entity: Sub220-PLC-001, event details: current set value 1200A→1800A)", node 2: "RTU status abnormal event (Event-20240610-002, event type: equipment abnormal, occurrence time: 09:20:15, associated entity: Sub220-RTU-003, event details: current monitoring value 1500A exceeds threshold 1200A)".
[0194] (2) Edge construction: The occurrence time of node 1 (09:15:32) is earlier than that of node 2 (09:20:15), and there is a logical connection between the abnormal RTU status and the PLC parameter tampering (RTU monitoring data comes from PLC acquisition). Establish the connection edge and mark "Data connection basis: RTU current monitoring value is calculated based on PLC acquisition data"; Similarly, establish the connection edge of "server process start event → PLC parameter modification event" "attack IP communication event → server process start event" "PLC parameter modification event → load scheduling business abnormal event" to form an attack causal graph.
[0195] Do-Calculus causality verification: (1) Intervention test: Intervene in the "PLC parameter modification event" (simulate restoring the PLC parameter to 1200A) and observe whether the "RTU status abnormal event" changes. The result shows that the RTU current monitoring value has recovered to 1500A (the original threshold is 1200A. Since the actual load has not changed, it needs to be corrected in conjunction with the business logic: the RTU current monitoring value is calculated by the PLC acquisition value + load compensation coefficient. After intervening in the PLC parameter, the monitoring value has not recovered because the actual load has not exceeded the compensation range. After blocking the attack IP, the PLC parameter has not been tampered with and the load compensation is normal. The monitoring value returns to the threshold, verifying the causal relationship of "attack IP communication → PLC parameter tampering → RTU abnormality".
[0196] (2) Result inversion: Based on the results of the intervention test, the causal direction is inverted to determine the causal chain direction of "attack IP 192.168.5.201 initiates communication → server starts malicious process → forges MMS message to tamper with PLC parameters → RTU status abnormality and business abnormality".
[0197] (3) Consistency verification: Combined with the deep consistency verification of business rules, the verification object is to conduct compliance review on each key state transition event in the causal chain, especially the "PLC parameter modification event (Event-20240610-001)".
[0198] The system calls the formalized business rule base, where a key rule r1 is instantiated: "If a 'PLC protection setting modification' event occurs, there must be a scheduling instruction with a status of 'in execution' or 'issued' and 'verification passed' before it is triggered. The target parameter of this instruction is consistent with the modified parameter." Search the knowledge graph for scheduling instructions that target Sub220-PLC-001 before Event-20240610-001 (09:15:32).
[0199] The most recent scheduling instruction Biz-20240610-012 was found, but its instruction content was "maintain 1200A", which is inconsistent with the modified result "1800A".
[0200] No valid scheduling instructions containing "modified to 1800A" were found.
[0201] The "parameter modification" event in this causal chain was determined to be a serious violation of business rule r1. Therefore, although the causal relationship was statistically significant, the chain was marked as "business anomaly," i.e., a highly suspicious attack path, standing out from a massive number of events. This step directly solved the "false positive" problem, accurately identifying abnormal operations without legitimate instructions as attack clues.
[0202] For suspicious attack paths that pass the business rules verification, perform a physical feasibility deep verification and a final physical feasibility review.
[0203] i. Latency compatibility constraints: MMS write service execution latency ≤ 1 second; RTU status collection and dead zone judgment period = 5 seconds.
[0204] Verification: Forged MMS message (09:15:32:00) → PLC parameter tampering (09:15:32:50), time difference 0.5 seconds, consistent. PLC parameter tampering (09:15:32) → RTU anomaly (09:20:15), time difference approximately 4 minutes and 43 seconds. Analysis revealed that the calculated values (based on the tampered PLC parameters) collected by the RTU at multiple cycle points such as 09:16:00, 09:16:05, etc., exceeded the threshold, but due to the "dead zone judgment" logic and communication delay, the alarm was not reported until 09:20:15. This timeline is consistent with the physical working principle and communication mechanism of the RTU, pass.
[0205] ii. Directional constraints: Attacks must propagate along reachable paths. From the external IP to the PLC, the attack must pass through multiple nodes such as firewalls and maintenance servers.
[0206] Verification: The path is external IP → server → PLC. After checking the network topology and access control rules in the knowledge graph, it is confirmed that server Sub220-SVR-002 does indeed have MMS communication permissions with PLC Sub220-PLC-001, and this server has an externally accessible RDP service. The path matches the actual network access relationships and attack surface exposure; therefore, it passes the verification.
[0207] iii. Cascaded Failure Mode Matching Constraint: The attack chain pattern should match the known threat pattern.
[0208] Verification: The path was compared with the threat intelligence database and matched the known attack pattern (TTP) of "using the operation and maintenance server as a springboard to forge industrial protocol messages for parameter injection". The verification was successful.
[0209] The dual-constraint collaborative result: Only the path that passes both "business rule consistency verification" (confirming that it is a malicious act) and "physical feasibility verification" (confirming that the act is physically feasible) will be ultimately determined as a "verified attack causal chain".
[0210] In this embodiment, the causal chain attack IP → server → forged message → PLC tampering → RTU anomaly is identified as malicious due to violation of business rules (no instruction modification), and is confirmed as credible because it conforms to physical laws (time delay, direction, and pattern are all reasonable). This ensures that the final output attack chain is not a statistical coincidence or theoretical deduction, but a high-confidence attack fact that has both business violation characteristics and physical implementation conditions, providing irrefutable evidence for subsequent accurate response.
[0211] Attack chain reconstruction and source identification: (1) Reverse tracing: Starting from the "RTU status abnormal event", trace back along the causal chain, and combine the "propagation" relationship (server → PLC) and "initiation" relationship (attack IP → server) in the knowledge graph to supplement the details of the attack source: the attack IP 192.168.5.201 is an external access IP (location: an unknown IDC data center), logs into the server Sub220-SVR-002 (weak password vulnerability, account: admin, password: 123456) through the remote desktop protocol (RDP), and then starts the mms-simulator.exe tool to forge MMS messages.
[0212] (2) Attack chain presentation: A complete attack chain is formed: "External attack IP 192.168.5.201 (09:10:00) → Login to server Sub220-SVR-002 using weak password (09:15:20) → Start mms-simulator.exe process (09:15:32) → Forge IEC61850MMS message (09:15:32) → Tamper with PLC Sub220-PLC-001 current setting (09:15:32) → RTU Sub220-RTU-003 status abnormal (09:20:15) + Load scheduling service abnormal (09:20:00)".
[0213] (3) Source output: Output key information of the attack source: attack IP 192.168.5.201 (location: a certain IDC data center, IP type: dynamic public IP, attack tool: mms-simulator.exe); initial compromised device Sub220-SVR-002 (model: RH2288H, firmware version: V2.0, vulnerability repair status: weak password vulnerability not repaired).
[0214] Source tracing report generation: Automatically generate a traceability report, which includes the following core contents: (1) Details of the attack source: Attacking IP 192.168.5.201 (location: a certain IDC data center, IP type: dynamic public IP, historical attack record: participated in 3 power equipment attacks in May 2024), the initial compromised device Sub220-SVR-002 (vulnerability: weak password, firmware version: V2.0, no security patch installed).
[0215] (2) Details of the propagation path: The attack went through three stages: Intrusion stage (09:10:00-09:15:20): The attacking IP logged into the server through a weak RDP password; Control stage (09:15:20-09:15:32): The malicious process was started to forge MMS messages; Damage stage (09:15:32-09:20:15): The PLC parameters were tampered with, causing equipment and business abnormalities.
[0216] (3) Causal verification basis: Through Do-Calculus intervention test, it is verified that "attack IP communication → PLC parameter tampering" and "PLC parameter tampering → RTU abnormality" are necessary causal relationships, and accidental related factors such as "operation and maintenance" and "equipment failure" are excluded.
[0217] (4) Defense recommendations: ① Investigate weak passwords for all servers and PLC devices and enable two-factor authentication; ② Upgrade the Sub220-PLC-001 firmware to V4.5.0 and fix the CVE-2023-XXXX vulnerability; ③ Deploy IEC61850 protocol deep detection equipment in the dispatch data network to intercept forged MMS messages; ④ Regularly verify the parameter consistency of RTU and PLC devices and promptly detect abnormal tampering.
[0218] This implementation case demonstrates that the present invention can effectively integrate multi-source power data, completely reconstruct the attack chain, and accurately locate the source of the attack. The tracing results have a clear causal explanation, providing maintenance personnel with a clear direction for defense optimization and solving the problems of fragmented tracing, high false positive rate, and poor adaptability of existing technologies.
[0219] Glossary of relevant technical terms (1) PLC (Programmable Logic Controller): A digital computing and operating electronic system used for industrial control. In power monitoring systems, it is mainly used for the execution of control instructions and parameter adjustment of substation equipment (such as switches and transformers). Its log contains key information such as instruction content, execution results, and parameter changes.
[0220] (2) RTU (Remote Terminal Unit): A terminal device in the power system used to remotely collect field data (voltage, current, equipment operating status code) and transmit it to the dispatch center. The data usually comes from PLC or field sensors and is an important data source reflecting the operating status of equipment.
[0221] (3) IEC61850 protocol: The power system communication standard developed by the International Electrotechnical Commission (IEC) is the core communication protocol of the power monitoring system. It includes two key subsets: MMS (Manufacturing Message Specification, used for equipment control commands and data transmission, default port 102) and GOOSE (Generic Object Oriented Substation Event, used for fast transmission of real-time signals such as tripping and alarms).
[0222] (4) Modbus protocol: an industrial communication protocol that is often used in power monitoring systems for data transmission between devices such as RTUs and smart meters and host computers (such as dispatch servers). Its main functions include reading device status and writing control instructions. The core fields include slave address (device identifier), function code (operation type, such as 0x06 for writing a single register), and register address (data storage location).
[0223] (5) Knowledge Graph: A semantic network based on graph structure, consisting of "entities" (such as attack IP, PLC device), "relationships" (such as "initiating" "affecting"), and "attributes" (such as the location of the attack IP, the model of the PLC). It can intuitively present the relationship between multiple sources of data and support complex relational queries and analysis.
[0224] (6) Do-Calculus algorithm: a mathematical tool for causal reasoning. It verifies the causal relationship between events by "intervention operations" (such as simulating the blocking of an event) rather than relying solely on statistical correlation. It can effectively eliminate accidental associations and is often used for causal relationship identification in complex systems (such as power monitoring systems).
[0225] (7) Circular Buffer: A data storage structure that uses a circular storage method. When the buffer is full, new data will overwrite the oldest stored data. It is suitable for temporary storage of real-time data such as network packets and logs, and can avoid data overflow. In this invention, it is used to store the original packets of network packet capture. The capacity is set to 1GB to balance storage requirements and data integrity.
[0226] (8) SSL / TLS1.3 protocol: a security protocol for network communication encryption, wherein TLS1.3 is the latest version. Compared with the old version (such as TLS1.2), it reduces the number of handshakes, shortens the connection time, and enhances the encryption strength. In this invention, it is used for encryption of business and security data transmission to ensure that the data is not stolen or tampered with during transmission.
[0227] This invention collects data through distributed acquisition agents deployed in substations and dispatch centers. It uses protocols such as IEC61850MMS / GOOSE, Modbus, and SSH to collect equipment logs, network traffic, service scheduling records, and security alarm data. It performs data standardization and cleaning, outputting an "entity-relationship-attribute" triple dataset. It configures a power control closed-loop semantic ontology library, performs 3D ontology construction and cross-dimensional semantic fusion, outputting a fused knowledge graph. It also configures a business rule library, a Do-Calculus inference engine, and a fault propagation mechanism library, performing causal verification of embedded business rules, causal graph generation of fault mechanism constraints, and attack chain reconstruction, outputting a source tracing report containing defense decision-oriented causal explanations.
[0228] The system configures a fault propagation mechanism library including directional constraints, latency parameters, and cascading modes. It then performs physical constraint attack causal graph generation and attack chain reconstruction; performs vulnerability location, defense intervention effect prediction, and minimum cut set defense strategy generation; and outputs a source tracing report. The semantic fusion and ontology construction, causal verification of business rule embedding, and causal graph generation of fault mechanism constraints are implemented on edge computing nodes, deployed on a single machine using the Neo4j graph database to meet real-time and resource-constrained requirements.
[0229] Design cross-dimensional semantic fusion rules: Design an automatic alignment algorithm for "protocol field - device function - business operation" - when parsing the MMS message "Write service + setting area 1 + address 0x0023", it automatically associates it with the "PLC-001 setting modification" function of the device dimension based on the ontology mapping rules, and further associates it with the "current protection setting adjustment" operation of the business dimension, so as to realize the unified semantic identification of data across the three layers of network-device-business.
[0230] This invention employs a Do-Calculus framework for a three-step verification process: intervention testing, result inversion, and business rule consistency verification. It uses power operation specifications (such as IEC 61850 control service timing and dispatching procedures) as intrinsic constraints for causal verification, rather than post-event verification. It constructs a three-dimensional ontology model of "equipment-protocol-service" to achieve semantic-level fusion of "GOOSE / MMS message fields → atomic operations of control commands → dispatching business processes," rather than field-level concatenation. It is based on the directionality of power system fault propagation, protection action delay, and cascading fault modes. Based on physical mechanisms, a predefined set of allowed edges and time-delay compatibility constraints for the causal graph structure are used to eliminate spurious associations that are physically impossible. In a field test at a 220kV substation, the false alarm rate of misjudging normal operation and maintenance as an attack was reduced from 23% to 3%, improving the accuracy of causal verification. The generated attack paths conform to the physical laws of power system fault propagation 100% and have no anomalies such as "reverse causality" or "timeout delay", improving the physical credibility of the attack chain. The source tracing report directly outputs executable defense suggestions such as "blocking a weak password on a certain server can cut off the attack path", making defense decisions directly guideable.
[0231] The power operation specifications are formalized into prior constraints for causal verification. In the Do-Calculus intervention test, the compliance of statistical causal effects with business rules is verified simultaneously, realizing the dual constraint judgment of "statistical causality + business logic". An ontology structure for attack process tracing is designed, defining four types of entities: "attack-device-business-vulnerability" and attack-specific relationships such as "initiation-exploitation-propagation-impact", supporting temporal-topological joint reasoning of the attack chain. Based on the causal verification results, the attack propagation path is dynamically reconstructed, and the attack source is traced back from the victim device. The complete attack chain is output, including attack stage division, propagation medium identification, and entry point location.
[0232] The response time from anomaly detection to attack chain reconstruction is ≤1 second, meeting the requirements for real-time attack tracing capabilities in power systems. In a field test at a 220kV substation, the accuracy rate of attack IP location was 100%, and the accuracy rate of initial compromised device identification was 98%, improving the accuracy of attack source location. The tracing results directly explain "why the server became a propagation node" (e.g., "weak password vulnerability allows RDP login"), supporting targeted hardening and improving the causal explainability of defense measures.
[0233] The dedicated ontology for power control equipment defines the control function attributes (such as setpoint area number, remote signaling point number, and sampling period) and protocol semantic fields (such as GOOSE's LD / LN references and MMS's object references) of devices such as PLCs, RTUs, and smart meters, realizing vertical semantic fusion of "protocol parsing - device control - business operation". Based on the hierarchical and partitioned control structure of the power system and the physical laws of fault propagation, it predefines the directional constraints, time delay compatibility constraints, and cascading pattern matching of causal edges to generate a causal graph with fault mechanism constraints, transforming attack path generation from "free exploration" to "physically guided constraint optimization". It adopts a graph database (Neo4j) single-machine deployment + sharded index optimization to avoid the cross-node communication overhead of federated learning, meet the real-time (≤1 second) and resource-constrained (memory ≤8GB) requirements of power edge devices, and generate a lightweight real-time inference architecture.
[0234] It can accurately identify power-specific attacks such as "GOOSE message replay causing circuit breaker malfunction" and "MMS forgery and setting tampering," which cannot be covered by general methods, thus improving the ability to identify power-specific attacks. The generated attack chain is 100% consistent with the power system control logic and fault propagation law, with no physically impossible false paths, thus improving the physical credibility of the attack chain. Attack tracing is achieved at the substation edge node (ARM architecture, 4GB memory), with inference latency ≤1 second and model size ≤500MB, improving the real-time performance of edge deployment.
[0235] This invention employs a knowledge graph construction technology that integrates multi-source heterogeneous power data. It designs an ontology model tailored to the characteristics of power systems, comprising four core entities: attacks, equipment, services, and vulnerabilities, along with specific relationships such as initiation, exploitation, propagation, and impact. This model is adaptable to power-specific data such as PLC control instruction logs, IEC61850 GOOSE messages, and load dispatch data. Through a hybrid extraction method combining rules and regular expressions, along with conflict resolution, it achieves structured integration of multi-source data, solving the fragmented tracing problem caused by traditional methods that only associate single-type logs. This is the core construction logic of a knowledge graph specific to power scenarios.
[0236] This invention presents an attack causal verification mechanism based on the Do-Calculus algorithm. It constructs an attack causal graph based on candidate events in a knowledge graph, and through a three-step process of intervention testing, result inversion, and consistency verification, eliminates accidental events that are temporally close but have no necessary connection, accurately identifying the causal relationship between attack behavior and equipment anomalies, thus reducing the false positive rate in tracing the source. This is an innovative application of causal reasoning in power attack tracing.
[0237] This invention is based on power-specific attack tracing and adaptation technology: it designs exclusive data parsing rules for power-specific protocols such as IEC61850 and Modbus, including GOOSE message dataset identification, Modbus function code and register address extraction, etc.; at the same time, it adapts to device characteristics such as PLC parameter out-of-bounds thresholds and RTU status codes, and can accurately associate the event chain of power-specific attacks such as forged GOOSE message replay and PLC malicious code injection, thus solving the problem of poor adaptability of traditional general tracing methods.
[0238] This invention is based on a source tracing output technology that includes causal explanations: the source tracing report integrates attack source details, causal verification basis, and defense recommendations. Attack source details include the firmware version of the initially compromised device and the historical behavior of the attacking IP. Causal verification basis includes intervention test results and power business logic support. Defense recommendations include vulnerability remediation solutions and protocol protection measures. It can clearly present the necessary causal relationship between the attack propagation logic and the event, rather than just outputting a list of isolated attack nodes, providing operation and maintenance personnel with actionable defense optimization directions.
[0239] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.< / ip> < / level> < / time> < / ip> < / level> < / time>
Claims
1. A method for tracing the source of causal verification attacks in power systems, characterized in that, include: Collect multi-source data from the power system, extract triples of entity-basic relation-basic attribute and entity-basic relation-entity, and generate triple datasets; Based on the triple dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing the source of power attacks is constructed. Based on equipment anomaly events, candidate events and candidate causal edges are selected from the power attack tracing knowledge graph to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between equipment anomaly events and each candidate event. Directed candidate causal edges that simultaneously satisfy power business rule constraints and physical feasibility constraints constitute the attack causal chain. Based on equipment anomaly events and the attack causal chain, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
2. The method for tracing the source of causal verification attacks in a power system according to claim 1, characterized in that: The generated triplet dataset includes: Based on the standardized fields of device entities, attack entities, business entities, and vulnerability entities in the preset entity field dictionary, device entities, attack entities, business entities, and vulnerability entities are extracted from multi-source data respectively. Based on the preset entity-basic relationship-basic attribute triplet mapping table, the protocol fields and log fields parsed from the multi-source data are mapped to the basic attributes of the corresponding entities, and the entity-basic relationship-basic attribute triplet is generated with the entity, the basic relationship between the entity and its basic attributes, and the basic attributes of the entity. Based on the preset entity-basic relation-entity triplet mapping table, the data representing the relationship between entities in the multi-source data are mapped to the basic relations between entities, and entity-basic relation-entity triplets are generated based on entities and the basic relations between entities; Construct an intermediate dataset of triples based on entity-base relation-base attribute triples and entity-base relation-entity triples; Clean the intermediate dataset of triples and output the cleaned dataset of triples.
3. A method for tracing the source of causal verification attacks in a power system according to claim 1, characterized in that: Based on the triplet dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing the origins of power attacks is constructed, including: Based on each entity, its business attributes, and the business relationships between entities, an ontology model for a knowledge graph of power attack tracing is constructed. Extract a set of new knowledge triples that conform to the ontology model from the triple dataset; The new knowledge triple set is fused and semantically disambiguated to generate a fused knowledge set. The fused knowledge set is then used to complete the relationships and generate a knowledge graph for tracing the source of power attacks.
4. A method for tracing the source of causal verification attacks in a power system according to claim 3, characterized in that: The set of new knowledge triplets that conform to the ontology model extracted from the triplet dataset includes: The triplet dataset is divided into structured data and semi-structured data; According to the preset mapping rules, the entities, basic relationships and basic attributes in the structured data are mapped to the entities, business relationships and business attributes corresponding to the ontology model, and the candidate knowledge triple set A is determined. Based on regular expressions specific to the power scenario, entities, basic relationships, and basic attributes in semi-structured data are mapped to entities, business relationships, and business attributes corresponding to the ontology model, thus determining the candidate knowledge triple set B. Generate a new set of knowledge triples based on the entities in the candidate knowledge triple set A and the candidate knowledge triple set B.
5. A method for tracing the source of causal verification attacks in a power system according to claim 4, characterized in that: The triplet dataset is divided into structured data and semi-structured data, including: Determine whether the entities, basic relations, and basic attributes of the triples in the triple dataset directly correspond semantically to the entities, business relations, and business attributes in the ontology model. Specifically, for entity-basic relation-basic attribute triples, determine whether their entities, basic relations, and basic attributes can all be found to have direct corresponding entities, business relations, and business attributes in the ontology model; for entity-basic relation-entity triples, determine whether their head entity, basic relation, and tail entity can all be found to have direct corresponding entities and business relations in the ontology model. If the entity, basic relation, and basic attribute of a triple do not directly correspond to all the entities, business relations, and business attributes in the ontology model, then the triple is determined to be semi-structured data. If the entity, basic relation, and basic attribute of a triple directly correspond to the entity, business relation, and business attribute in the ontology model, then it is determined whether the target data item pointed to by the basic relation in the triple is atomic data. If the target data item pointed to by the basic relation is atomic data, the triple is determined to be structured data. If the target data item pointed to by the basic relation is not atomic data, the triple is determined to be semi-structured data. Specifically, for an entity-basic relation-basic attribute triple, the target data item pointed to by the basic relation refers to the value of the basic attribute; for an entity-basic relation-entity type triple, the target data item pointed to by the basic relation refers to the identifier of the tail entity.
6. A method for tracing the source of causal verification attacks in a power system according to claim 1, characterized in that: Physical feasibility constraints include time delay compatibility verification, causal directionality verification, and cascading failure mode matching verification. Among them, the delay compatibility check is used to determine that the directed candidate causal edge passes the delay compatibility check if the actual time interval between the candidate events associated with each directed candidate causal edge in the attack causal graph is greater than or equal to the minimum propagation delay, based on the predefined minimum propagation delay of the power system attack action. Causality directionality verification is used to determine if the direction of each directed candidate causal edge in the attack causal graph is consistent with the allowed direction defined by the network topology and the direction of business data flow, based on a predefined network topology and business data flow. Cascaded fault mode matching verification is used to determine that a directed candidate causal edge passes the cascaded fault mode matching verification if it matches any known propagation pattern in the pre-set library of typical power system faults and attack propagation patterns based on a pre-set library of typical power system faults and attack propagation patterns. If a directed candidate causal edge passes the time delay compatibility check, causal directionality check, and cascading failure mode matching check in sequence, it is determined to meet the physical feasibility constraint.
7. A method for tracing the source of causal verification attacks in a power system according to claim 1, characterized in that: Locating the source of an attack and reconstructing the attack chain within a knowledge graph for tracing power system attacks includes: Within the knowledge graph for tracing power attacks, candidate events that are logically related to abnormal device events within a preset time window are selected to form a candidate event set; Using candidate events in the candidate event set as nodes, candidate causal edges are constructed based on time proximity and logical association rules to generate an attack causal graph; For candidate causal edges in the attack causal graph, the Do-Calculus method is used to conduct intervention tests to determine the average causal effect between the device malfunction event and the candidate event. Based on the average causal effect, directed candidate causal edges are extracted to form a set of directed causal edges. The power business rules are used to verify the consistency of the power business rules for candidate events associated with directed candidate causal edges. Candidate events associated with directed candidate causal edges that pass the verification are marked as suspicious attack paths. Suspicious attack paths are physically feasible through physical feasibility constraints. Suspicious attack paths that pass the physical feasibility verification are used as the verified attack causal chains. Starting from the abnormal equipment event, trace the attack causal chain backwards along the verified path, query the power attack tracing knowledge graph to obtain the business relationship corresponding to the abnormal equipment event, thereby determining the source of the attack and reconstructing the attack chain from the source of the attack to the device corresponding to the abnormal equipment event.
8. A method for tracing the source of causal verification attacks in a power system according to claim 7, characterized in that: The verified attack causal chains include: The Do-Calculus method is used to intervene in the candidate events corresponding to the candidate causal edges in the causal graph. Based on historical data, the average causal effect of the device malfunction event before and after the intervention in the candidate events is calculated. Based on the relationship between the average causal effect and the preset significance threshold, the direction of the candidate causal edge is determined, and directed candidate causal edges with a clear unidirectional causal direction are selected. Multiple attack causal chains are constructed based on directed candidate causal edges. The candidate events corresponding to each attack causal chain are checked for consistency with the power business rules. When the consistency check is passed, the rationality of the attack causal chain is determined. When the consistency check is not passed, the corresponding attack causal chain is filtered out. A reasonable attack causal chain is physically feasible. An attack causal chain that passes the physical feasibility verification is considered a valid attack causal chain.
9. A power system causal verification attack tracing system, characterized in that: The data acquisition module is used to collect multi-source data from the power system; The data preprocessing module is used to extract triples of entity-basic relation-basic attribute and entity-basic relation-entity, clean the data of the two types of triples, and generate triple datasets. The knowledge graph construction module is used to build a knowledge graph for tracing the source of power attacks based on the triple dataset, combined with the business attributes of entities and the business relationships between entities; The attack localization module is used to filter candidate events and candidate causal edges from the power attack tracing knowledge graph based on abnormal equipment events to construct an attack causal graph. In the attack causal graph, directed candidate causal edges are extracted based on the causal effect between abnormal equipment events and each candidate event. Directed candidate causal edges that simultaneously satisfy power business rule constraints and physical feasibility constraints form an attack causal chain. Based on abnormal equipment events and attack causal chains, the attack source is located and the attack chain is reconstructed in the power attack tracing knowledge graph.
10. A power system causal verification attack tracing system according to claim 9, characterized in that: The generated triplet dataset includes: Based on the standardized fields of device entities, attack entities, business entities, and vulnerability entities in the preset entity field dictionary, device entities, attack entities, business entities, and vulnerability entities are extracted from multi-source data respectively. Based on the preset entity-basic relationship-basic attribute triplet mapping table, the protocol fields and log fields parsed from the multi-source data are mapped to the basic attributes of the corresponding entities, and the entity-basic relationship-basic attribute triplet is generated with the entity, the basic relationship between the entity and its basic attributes, and the basic attributes of the entity. Based on the preset entity-basic relation-entity triplet mapping table, the data representing the relationship between entities in the multi-source data are mapped to the basic relations between entities, and entity-basic relation-entity triplets are generated based on entities and the basic relations between entities; Construct an intermediate dataset of triples based on entity-base relation-base attribute triples and entity-base relation-entity triples; Clean the intermediate dataset of triples and output the cleaned dataset of triples.
11. A power system causal verification attack tracing system according to claim 9, characterized in that: Based on the triplet dataset, and combining the business attributes of entities and the business relationships between entities, a knowledge graph for tracing the origins of power attacks is constructed, including: Based on each entity, its business attributes, and the business relationships between entities, an ontology model for a knowledge graph of power attack tracing is constructed. Extract a set of new knowledge triples that conform to the ontology model from the triple dataset; The new knowledge triple set is fused and semantically disambiguated to generate a fused knowledge set. The fused knowledge set is then used to complete the relationships and generate a knowledge graph for tracing the source of power attacks.
12. A power system causal verification attack tracing system according to claim 11, characterized in that: The set of new knowledge triplets that conform to the ontology model extracted from the triplet dataset includes: The triplet dataset is divided into structured data and semi-structured data; According to the preset mapping rules, the entities, basic relationships and basic attributes in the structured data are mapped to the entities, business relationships and business attributes corresponding to the ontology model, and the candidate knowledge triple set A is determined. Based on regular expressions specific to the power scenario, entities, basic relationships, and basic attributes in semi-structured data are mapped to entities, business relationships, and business attributes corresponding to the ontology model, thus determining the candidate knowledge triple set B. Generate a new set of knowledge triples based on the entities in the candidate knowledge triple set A and the candidate knowledge triple set B.
13. A power system causal verification attack tracing system according to claim 12, characterized in that: The triplet dataset is divided into structured data and semi-structured data, including: Determine whether the entities, basic relations, and basic attributes of the triples in the triple dataset directly correspond semantically to the entities, business relations, and business attributes in the ontology model. Specifically, for entity-basic relation-basic attribute triples, determine whether their entities, basic relations, and basic attributes can all be found to have direct corresponding entities, business relations, and business attributes in the ontology model; for entity-basic relation-entity triples, determine whether their head entity, basic relation, and tail entity can all be found to have direct corresponding entities and business relations in the ontology model. If the entity, basic relation, and basic attribute of a triple do not directly correspond to all the entities, business relations, and business attributes in the ontology model, then the triple is determined to be semi-structured data. If the entity, basic relation, and basic attribute of a triple directly correspond to the entity, business relation, and business attribute in the ontology model, then it is determined whether the target data item pointed to by the basic relation in the triple is atomic data. If the target data item pointed to by the basic relation is atomic data, the triple is determined to be structured data. If the target data item pointed to by the basic relation is not atomic data, the triple is determined to be semi-structured data. Specifically, for an entity-basic relation-basic attribute triple, the target data item pointed to by the basic relation refers to the value of the basic attribute; for an entity-basic relation-entity type triple, the target data item pointed to by the basic relation refers to the identifier of the tail entity.
14. A power system causal verification attack tracing system according to claim 9, characterized in that: Physical feasibility constraints include time delay compatibility verification, causal directionality verification, and cascading failure mode matching verification. Among them, the delay compatibility check is used to determine that the directed candidate causal edge passes the delay compatibility check if the actual time interval between the candidate events associated with each directed candidate causal edge in the attack causal graph is greater than or equal to the minimum propagation delay, based on the predefined minimum propagation delay of the power system attack action. Causality directionality verification is used to determine if the direction of each directed candidate causal edge in the attack causal graph is consistent with the allowed direction defined by the network topology and the direction of business data flow, based on a predefined network topology and business data flow. Cascaded fault mode matching verification is used to determine that a directed candidate causal edge passes the cascaded fault mode matching verification if it matches any known propagation pattern in the pre-set library of typical power system faults and attack propagation patterns based on a pre-set library of typical power system faults and attack propagation patterns. If a directed candidate causal edge passes the time delay compatibility check, causal directionality check, and cascading failure mode matching check in sequence, it is determined to meet the physical feasibility constraint.
15. A power system causal verification attack tracing system according to claim 9, characterized in that: Locating the source of an attack and reconstructing the attack chain within a knowledge graph for tracing power system attacks includes: Within the knowledge graph for tracing power attacks, candidate events that are logically related to abnormal device events within a preset time window are selected to form a candidate event set; Using candidate events in the candidate event set as nodes, candidate causal edges are constructed based on time proximity and logical association rules to generate an attack causal graph; For candidate causal edges in the attack causal graph, the Do-Calculus method is used to conduct intervention tests to determine the average causal effect between the device malfunction event and the candidate event. Based on the average causal effect, directed candidate causal edges are extracted to form a set of directed causal edges. The power business rules are used to verify the consistency of the power business rules for candidate events associated with directed candidate causal edges. Candidate events associated with directed candidate causal edges that pass the verification are marked as suspicious attack paths. Suspicious attack paths are physically feasible through physical feasibility constraints. Suspicious attack paths that pass the physical feasibility verification are used as the verified attack causal chains. Starting from the abnormal equipment event, trace the attack causal chain backwards along the verified path, query the power attack tracing knowledge graph to obtain the business relationship corresponding to the abnormal equipment event, thereby determining the source of the attack and reconstructing the attack chain from the source of the attack to the device corresponding to the abnormal equipment event.
16. A power system causal verification attack tracing system according to claim 15, characterized in that: The verified attack causal chains include: The Do-Calculus method is used to intervene in the candidate events corresponding to the candidate causal edges in the causal graph. Based on historical data, the average causal effect of the device malfunction event before and after the intervention in the candidate events is calculated. Based on the relationship between the average causal effect and the preset significance threshold, the direction of the candidate causal edge is determined, and directed candidate causal edges with a clear unidirectional causal direction are selected. Multiple attack causal chains are constructed based on directed candidate causal edges. The candidate events corresponding to each attack causal chain are checked for consistency with the power business rules. When the consistency check is passed, the rationality of the attack causal chain is determined. When the consistency check is not passed, the corresponding attack causal chain is filtered out. A reasonable attack causal chain is physically feasible. An attack causal chain that passes the physical feasibility verification is considered a valid attack causal chain.
17. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded onto the processor, it implements a power system causal verification attack tracing method according to any one of claims 1-8.
18. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, implements a method for tracing causal verification attacks on a power system according to any one of claims 1-8.
Citation Information
Patent Citations
Lightweight dynamic causal reasoning-based power Internet of Things terminal attack tracing method
CN120658465A
APT attack traceability and path restoration method and system
CN121151119A
Electric power information operation violation risk supervision system based on knowledge graph
CN121189798A