Threat detection method and system based on multi-stage attack process matching

By employing a threat detection method that matches the multi-stage attack process in new power systems, the shortcomings of traditional IDS in detecting multi-stage attacks in new power systems are addressed. This enables accurate identification and situational awareness of multi-stage attacks, and improves multi-source information processing capabilities and rapid verification of security solutions.

CN121966935APending Publication Date: 2026-05-01STATE GRID HENAN ELECTRIC POWER ELECTRIC POWER SCI RES INST +4
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID HENAN ELECTRIC POWER ELECTRIC POWER SCI RES INST
Filing Date
2025-12-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies lack efficient detection capabilities against multi-stage attacks in new power systems. Traditional IDS detection measures can only respond when an attack occurs and lack comprehensive situational awareness and accurate threat assessment of multi-source network environments.

Method used

A threat detection method based on multi-stage attack process matching is adopted, which achieves accurate identification and reconstruction of multi-stage attacks through heterogeneous traffic detection, preprocessing, multi-source event correlation, uncertainty handling and attack sequence graph generation.

Benefits of technology

It enables accurate identification and reconstruction of complex, multi-stage network attacks, enhances the ability to process multi-source heterogeneous information, and provides complete attack situation awareness and rapid security solution verification capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121966935A_ABST
    Figure CN121966935A_ABST
Patent Text Reader

Abstract

The invention discloses a threat detection method and system based on multi-stage attack process matching. The method comprises the following steps: preprocessing heterogeneous traffic in a novel power system, obtaining security alarm information, carrying out association analysis, constructing a coherent attack behavior sequence, and distributing a credibility index for each attack behavior. Uncertainties and conflict evidence present in the alarm data are processed and analyzed. And dynamically generating an attack sequence diagram according to the attack behavior sequence, traversing and reconstructing all possible attack scheme paths according to a time sequence, and selecting a potential attack path with the highest belief value. Through attack stage matching, the most credible attack sequence diagram and the most reasonable attack reduction path are searched, the optimal result pair is identified, and the current killing chain stage of an attacker is determined. According to the method, security analysis is carried out through multi-source information sources, heterogeneous information can be effectively integrated, uncertainty can be processed, an attack scheme can be accurately identified and restored, and complete attack situation awareness is realized.
Need to check novelty before this filing date? Find Prior Art

Description

A threat detection method and system based on multi-stage attack process matching Technical Field

[0001] This invention belongs to the field of simulation technology of distributed resource networks in new power systems for network security, and specifically relates to a threat detection method and system based on multi-stage attack process matching. Background Technology

[0002] With the continuous development of digital convergence, new power systems interconnect various distributed devices on networks. This convergence increases the attack surface and complexity of security threats. Attackers utilize multi-stage attacks with multi-step and cross-domain characteristics, posing a severe challenge to the stable operation of the power grid. Current research shows that the demand for large amounts of cross-domain information makes the requirements for automated response mechanisms more complex. Traditional IDS detection measures can only detect and alert when an attack occurs abnormally in the system, providing a basis for determining appropriate responses and remedial measures. To obtain complete situational awareness of the system, it is necessary to examine the rationality of communication and process data transmitted through industrial control networks and signs of possible attack traces in order to detect latent multi-stage attacks in a timely manner. Due to the lack of high-quality data samples and the uneven distribution of sample label categories, intrusion detection models suffer from low accuracy and high false alarm rates. Existing knowledge acquisition is limited by events from the same source, ignoring alerts from other security systems or logs from other communication network components, resulting in limited perception and assessment of potential wide-ranging network events.

[0003] To address security challenges, new power systems require accurate and descriptive threat activity detection capabilities. A holistic understanding of the attack development process across multiple domains is needed, rather than relying solely on a single source of information for network security analysis. The MITREATT&CK for ICS Attack Matrix provides an advanced information repository for preparing, selecting, and implementing appropriate mitigation, countermeasures, and recovery plans. Its framework, established from the attacker's perspective of the kill chain, maps to TTPs (Tactics, Techniques, and Procedures), enabling the detection of multi-stage attacks to contextually correlate attack indicators from different components and process critical data streams captured in multi-source network environments. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a threat detection method and system based on multi-stage attack process matching. By combining the correlation between security events and attack schemes, it can map domain-specific attack signs to an attack correlation graph, and map the attack behaviors corresponding to the multi-stage attack stages against new power systems to the kill chain stages, ultimately enabling the threat detection method to identify the current attack scenario.

[0005] To address the challenge of detecting multi-layered network attacks from a holistic perspective, this invention proposes an attack detection based on multi-stage attack process matching (AD-MSAPM) method. This method utilizes early warning context and multi-source information processing to achieve threat detection of security events in novel power systems. By combining the correlation between security events and attack schemes, these domain-specific attack signatures can be mapped to an attack correlation graph. This maps attack behaviors corresponding to the multi-stage attack phases targeting novel power systems to the kill chain phases, ultimately enabling the threat detection method to identify the current attack scenario.

[0006] The present invention adopts the following technical solution.

[0007] A threat detection method based on multi-stage attack process matching includes the following steps: Step 1, heterogeneous traffic detection is performed on a novel power system, and the heterogeneous traffic is preprocessed to generate security alarm information, obtaining information transmission routes; Step 2, the preprocessed security alarm information is correlated with events through a multi-source event correlator to construct a sequence of potential attack behaviors, and a trust index is assigned to each attack behavior; Step 3, uncertainty processing and combination of behavior quality functions from multiple evidence sources are performed to obtain a comprehensive behavior quality function; Step 4, based on the potential attack behavior sequence obtained in Step 2, corresponding behavior nodes are generated for each attack behavior, and metadata or semantic connections are added according to the attack scheme and kill chain stage to which the behavior node belongs. Based on the comprehensive behavior quality function, edge transition quality functions are assigned to each semantic connection to obtain an attack sequence graph; Step 5, based on the attack sequence graph and attack methods, potential attack paths are reconstructed through attack path restoration, and then... The quality distribution of all behaviors in the path is adjusted, and the distribution of the path quality function is evaluated. The potential attack path with the highest belief value and exceeding the preset threshold and the corresponding attack sequence graph are selected. Step 6: The potential attack path selected in Step 5 and the corresponding attack sequence graph are structurally consistent to ensure that the path logic and the topological semantics of the graph are completely matched, thereby identifying the combination of behaviors that best represents the attack activity and defining it as the optimal result pair. Then, the path quality function value in the optimal result pair is extracted and compared with the preset constraints to control the minimum confidence level of the result. The current state is determined as no attack or unknown attack based on the matching completeness of the optimal result pair. If the confidence level meets the standard and the match is successful, the kill chain stage is determined by mapping the behavior attributes of the end node of the path in the optimal result pair. Based on the topological and temporal information covered by the optimal result pair, an association result report containing a list of infected devices, specific attack schemes and attack time ranges is output.

[0008] Preferably, step 1 specifically includes: step 1.1, performing anomaly detection on heterogeneous traffic in the new power system, identifying and extracting contextual features strongly related to power business; step 1.2, converting the heterogeneous traffic with anomalies into a standardized format containing power business semantics, generating security alarm information containing standard fields; step 1.3, constructing a globally unique event ID through sensor identifiers, classifying same-source communication paths, and forming a traversable information transmission route.

[0009] Preferably, the security alarm information of the standard field includes the following fields: alarm ID, load, load type, timestamp, protocol information, device endpoint, and sensor ID.

[0010] Preferably, step 2 specifically includes: Step 2.1, Behavior analysis and event differentiation: Analyzing the behavior executed on the device that issued the alarm signal, considering the communication behavior of the infected device, and integrating system logs and security alarm information to identify and extract recurring suspicious events; Step 2.2, Kill chain mapping: Obtaining harmful behaviors based on security alarm information and suspicious events, mapping the harmful behaviors to the network kill chain model, and identifying and defining a single attacker's operation pattern; Step 2.3, Obtaining an access attempt indicator dataset based on suspicious events and security alarm information, and systematically grouping it according to time series and network host targets to obtain a behavior dataset arranged according to attack intent and time sequence; Step 2.4, Generating intrusion indicator key-value pairs based on the detected intrusion indicators. The format of the intrusion indicator key-value pairs is {source address-target address}. Local and remote access attempts are distinguished by the source address and target address; Step 2.5, Assigning a behavior quality function to the access attempt behavior based on the intrusion indicator key-value pairs, thereby obtaining a trust index corresponding to each access attempt behavior. The trust index serves as the probability that the alarm corresponding to each access attempt behavior represents a real attack behavior.

[0011] Preferably, step 2.5 specifically includes: setting a set of hypotheses for the access attempt behavior state. ,in, This indicates that the current access attempt is a genuine attack. This indicates that the current access attempt is a non-attack scenario; This indicates that the current access attempt is in a state of uncertainty, meaning it's impossible to distinguish between a real attack and a non-attack; Behavior Quality Function The domain is set All subsets of; among which, This indicates the degree to which the current evidence supports the conclusion that it was a real attack; This indicates the degree to which the current evidence supports the conclusion that there was no attack. This indicates the degree of support for uncertain states when the current access attempt is in a state of uncertainty. , and The values ​​of all values ​​belong to the interval [0,1] and satisfy: + + =1 Behavioral quality function Based on the security alert attributes and their associated features extracted from the system logs and real-time alerts for this access attempt behavior, the associated features include contextual information related to the behavior, alert severity, number of recurrences within a preset time window, degree of matching with intrusion indicators or attack technique templates, and degree of deviation of the communication pattern from the historical baseline.

[0012] Preferably, step 3 specifically includes: assuming the first The behavioral quality function corresponding to each source of evidence is: For any state Comprehensive behavioral quality distribution satisfy:

[0013] in, This represents each possible behavioral state in the result obtained after merging the first k evidence sources; This represents the various possible behavioral states provided by the (k+1)th source of evidence introduced; Describing the degree of conflict, set This indicates that the state is considered after combination. All state pairs ;gather The meaning is defined as: when for hour, At least include ;when for hour, At least include The remaining combinations that do not have clear consistency between the two sources of evidence are considered by the system to be in an uncertain state.

[0014] Preferably, the degree of conflict The formula for calculation is:

[0015] in, This represents a set of contradictory pairs of states between two sources of evidence, including at least... and .

[0016] Preferably, step 4 specifically includes: step 4.1, generating an attack sequence graph, wherein the nodes of the attack sequence graph contain metadata, overall scheme, attack behavior, and semantic connections of the kill chain stages; step 4.2, defining the edge connection types of the attack sequence graph, wherein the edge connection types include same source, same target, extended, non-overlapping, and arbitrary overlap; step 4.3, setting an edge transition quality function for each edge of the attack sequence graph, wherein the edge transition quality function is used to define the semantic and probabilistic characteristics of the connection.

[0017] Preferably, step 4.1 specifically includes: reading the sequence of potential attack behaviors arranged in chronological order, generating a behavior node for each attack behavior, wherein the behavior node records at least the behavior identifier, timestamp, source host and target host, corresponding {source address – target address} key-value pair, and kill chain stage label determined in step 2.2; for behavior sets that share the same attacker identifier, the same {source address – target address} key-value pair, or are in the same attack stage group, constructing the corresponding overall scheme node, and establishing semantic connections between the overall scheme node, behavior nodes, and kill chain stage nodes to obtain an initial attack sequence graph; preferably, step 5 specifically includes: step 5.1, traversing nodes and their reachable edges in the attack sequence graph in chronological order to generate candidate attack paths; step 5.2, using the edge quality function of each edge on the path as input, constructing the path quality function through a multiplicative chain combination method; step 5.3, calculating the comprehensive credibility of each potential attack path based on the path quality function, and selecting the path with the highest comprehensive credibility. If the credibility of the selected path exceeds a preset threshold, then the path is a potential attack path, and the attack behavior graph is reconstructed using the potential attack paths.

[0018] Preferably, step 6 specifically includes: step 6.1, identifying the optimal result pair based on the potential attack path and reconstructed attack behavior graph obtained in step 5.3; step 6.2, obtaining potential attack paths that meet the preset threshold limit by comparing the path quality function value corresponding to the potential attack path with a preset threshold; step 6.3, taking the stage of the last attack node in the attack reconstruction path as the current kill chain stage of the attacker; step 6.4, outputting the correlation process, including the detected attack behavior and scheme, the kill chain stage last perceived by the attacker, and the list of detected infected devices.

[0019] This invention also proposes a threat detection system based on multi-stage attack process matching to implement the aforementioned threat detection method, comprising: an alarm preprocessing module for formatting multi-source security alarm information and generating a normalized alarm containing alarm ID, timestamp, and sensor ID, while constructing a globally unique event ID; an event association module, acting as a multi-source event correlator, for mapping alarm data to a network kill chain model, constructing an attack behavior sequence, and assigning a quality function to each potential attack behavior; and a trust balancing module, employing an evidence reasoning framework for handling incomplete knowledge, using robust combination rules and logarithmic robust combination. The system integrates rule-based dependencies and conflict evidence, and introduces an environmental impact quality function to achieve a dynamic balance of alarm trust. A scheme association module dynamically generates an attack sequence graph with semantically connected nodes and typified edges based on the attack behavior sequence, and assigns initial quality function values ​​to the edges. An attack path reconstruction module traverses the attack sequence graph in chronological order, reconstructs all possible attack scheme paths after adjusting the quality distribution, and selects the path with the highest belief value. An attack phase matching module identifies the optimal result pair formed by the most trustworthy attack sequence graph and the most reasonable attack reconstruction path, thereby determining the current stage of the kill chain the attacker is in.

[0020] The present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the threat detection method when loaded onto the processor.

[0021] The present invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the threat detection method described above.

[0022] The beneficial effects of this invention are as follows: Compared with existing technologies, this invention proposes a threat detection method based on multi-stage attack process matching. Through multi-source heterogeneous information processing, uncertainty management, and attack process matching, it achieves accurate identification and reconstruction of complex, multi-stage network attacks. This invention employs an event correlation mechanism based on Dempster-Shafer theory, which can effectively integrate heterogeneous information from different security systems and network components, and handle uncertainties and conflicts, significantly improving the multi-source heterogeneous information processing capability. This effectively overcomes the problem of limited perception and assessment caused by the single information source in traditional methods. Simultaneously, by mapping the correlated attack behavior sequence to dynamically generated attack sequence graphs and kill chains, this method can not only accurately match attack patterns and reconstruct the attack evolution process, but also provide security personnel with complete attack situation awareness, solving the problem that existing technologies only provide scattered alert information and lack the ability to reconstruct the full attack picture. Furthermore, by integrating various network intelligence sources and realizing the systematic processing and correlation of this information, this invention effectively assesses threats, providing an advanced information base for preparing, selecting, and implementing appropriate mitigation, countermeasures, and recovery plans. This provides a technical foundation for achieving comprehensive situation awareness and helps promote the rapid verification and iteration of security solutions. Attached Figure Description

[0023] Figure 1 is a flowchart illustrating the threat detection method based on multi-stage attack process matching in this invention; Figure 2 is a structural diagram illustrating the threat detection system based on multi-stage attack process matching in this invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, other embodiments obtained by those skilled in the art without creative effort are all within the protection scope of this invention.

[0025] As shown in Figure 1, this invention proposes a threat detection method based on multi-stage attack process matching. The method includes the following steps: Step 1, heterogeneous traffic detection is performed on the new power system, and the heterogeneous traffic is preprocessed to obtain information transmission routes; Preprocessing of alarm information: Multi-source, heterogeneous traffic information from the new power system is preprocessed, converting this information into a unified, standardized format. This process includes detecting abnormal traffic, generating security alarm information containing key fields such as alarm ID, load, and timestamp, and constructing a globally unique event ID based on sensor identifiers.

[0026] The preprocessing of alarm information includes the following steps: Step 1.1, Anomaly detection of heterogeneous traffic: For the diverse communication methods of distributed resources such as photovoltaics, energy storage, and charging piles, existing detection mechanisms are used to perform deep packet inspection (DPI) on cross-domain heterogeneous traffic, identifying and extracting contextual features strongly related to power services; Step 1.2, Security alarm generation: Anomaly-prone heterogeneous traffic is converted into a standardized format containing power service semantics, generating security alarm information containing standard fields; To address the complex and variable characteristics of distributed resource services, heterogeneous traffic is converted into a unified standardized format containing power service semantics, generating security alarm information containing fields such as alarm ID, payload, payload type, timestamp, protocol information, device endpoint, and sensor ID; In the context data required for event association, security-related information mainly comes from distributed intrusion detection systems (IDS), which provide standardized alarm data and deviation information from normal system operation characteristics. Through analysis of traffic collected according to the detected domain, the standard fields for anomaly detection are shown in Table 3.

[0027] Table 3: Fields for Abnormal Traffic Detection

[0028] When the IDS detects abnormal traffic, it identifies it as a suspicious operation and generates security alert information as shown in Table 4.

[0029] Table 4: Security Alarm Information

[0030] Step 1.3, Event ID Construction and Route Classification: Construct globally unique event IDs using sensor identifiers and classify communication paths from the same source to form traversable information transmission routes.

[0031] Globally unique event IDs are constructed using sensor identifiers, and communication paths from the same source are classified to form a traversable information transmission route. Information preprocessing identifies event clusters sent within the same network domain by iterating through all data packet events, enabling the classification of individual alarms with the same information. The preprocessing component, based on unique sensor identifiers, constructs a unique event ID across the entire new power system network by combining the ID of each data packet with the ID of the corresponding sensor reporting the event. Therefore, step 1, by reconstructing the topology data packets and timestamps captured by monitoring into a topology-time and time mapping with the same structure, enables the classification of communication paths from the same source, ultimately forming a traversable information transmission route.

[0032] Step 1 of this invention utilizes existing detection mechanisms to identify abnormal traffic and extracts context fields related to the power system to describe suspicious operations; it generates structured alarms according to a fixed field system to ensure that data from all sources have a consistent format, facilitating matching in subsequent stages; it constructs globally unique event IDs based on sensor identifiers and classifies communication paths from the same source, thereby making events traceable in the power system topology.

[0033] Compared to conventional IDS that only output simple alarm records, this invention enforces a unified field specification during the preprocessing stage. This ensures that alarm information from different sensors and detection domains within the new power system has the same data structure, allowing it to be directly used by subsequent multi-source event correlators and attack sequence construction modules. Existing IDS generally use rules or features geared towards general IT networks and do not define this set of fields specifically for the periodic scheduling and topology of power systems. This invention establishes a clear interface between anomaly detection results and subsequent attack phase modeling through this set of fields. Conventional IDS alarms are often only valid within a single sensor or single session, lacking unified event identifiers across sensors and paths. This invention completes event ID construction and route classification during the preprocessing stage, enabling subsequent steps to perform behavior sequence, attack path, and kill chain phase matching based on "event clusters on the same path," rather than relying solely on isolated alarms.

[0034] Step 2, Multi-Source Event Correlation and Attack Behavior Sequence Construction: A multi-source event correlator is used to perform correlation analysis on the preprocessed alarm data to reveal potential attacker behavior patterns. This process aims to integrate scattered, isolated security events into a coherent sequence of attack behaviors and assign a trust index to each attack behavior.

[0035] The multi-source event correlation and attack behavior sequence construction process includes the following steps: Step 2.1, Behavior Analysis and Event Differentiation. Analyze the behaviors executed on the infected device, consider the communication behavior of the infected device, and integrate multi-source data such as system logs and real-time alerts to identify and differentiate recurring suspicious events. In the process of network intrusion into a new type of power system, attackers communicate with malicious components in the target system through a Command and Control (C2) server to achieve remote control of the infected device. To improve the accuracy of threat detection, the multi-source event correlator not only considers the communication behavior of the infected device but also integrates multi-source data such as system logs and real-time alerts.

[0036] Step 2.2, Kill Chain Mapping. Malicious behaviors are mapped to a network kill chain model to identify and define individual attacker operating patterns. Step 2.3, Indicator Dataset Collection. Access attempt indicator datasets are collected and systematically grouped according to time series and network host targets. The multi-source event correlator constructs a dataset containing access attempt indicators by collecting behavioral data such as network scanning, login attempts, and privilege escalation, and systematically groups them according to time series and network host targets. Step 2.4, Intrusion Indicator Key-Value Pair Generation. Based on detected intrusion indicators, a set of key-value pairs in the format {source address-target address} is formed to distinguish between local and remote access attempts. This dataset undergoes further filtering and processing, forming a set of key-value pairs in the format {source address-target address} based on detected Indicators of Compromise (IoC). The multi-source event correlator uses the length and endpoint attributes of these sets to distinguish between local and remote access attempts to identify the source and nature of the attack. Step 2.5, Assigning Behavioral Quality Functions to Potential Access Attempts.

[0037] Through behavioral quality function Each access attempt is assigned a trust index to quantify the likelihood that the alert corresponding to that access attempt represents a genuine attack. Let there be a set of hypotheses used to describe the state of access attempt behavior. . for A subset representing cases where the current access attempt constitutes a "real attack"; for A subset representing the "non-attack" scenario of the current access attempt; for A subset representing the current access attempt behavior in a "state of uncertainty," meaning it's impossible to distinguish between a "real attack" and a "non-attack." Behavioral quality function. The domain is set All subsets; of which This indicates the degree to which the current evidence supports the conclusion of a "real attack"; This indicates the degree to which the current evidence supports the judgment of "non-attack"; among which This indicates the degree of support for the uncertain state when the current evidence is insufficient to distinguish between a "real attack" and a "non-attack". All three values ​​are within the interval [0,1], and their sum is 1, used to satisfy the normalization constraints of the basic probability assignment. Function Based on the security alert attributes and their associated features extracted from the system logs and real-time alerts for this access attempt behavior, the associated features may include contextual information related to the behavior, alert severity, number of recurrences within a preset time window, degree of matching with intrusion indicators or attack technique templates, and degree of deviation of the communication pattern from the historical baseline.

[0038] It should be noted that this invention does not limit the mass function. The specific construction method and function form are as follows. Under the premise of satisfying the basic probability assignment constraints, different feature combinations, mapping functions, and parameter configurations can be selected according to different application scenarios. For example, linear mapping, piecewise functions, logarithmic functions, or other monotonic mapping methods can be used to combine the features used to describe access attempt behavior and obtain the corresponding quality function.

[0039] Based on the above analysis, the multi-source event correlator can further investigate whether infected devices exhibit lateral movement behavior. For example, it can monitor whether a device attempts to penetrate other systems to establish communication connections or to install malware. For devices that choose to remain hidden within the network to conduct reconnaissance or await instructions after installing malware, the multi-source event correlator will identify them by detecting whether the device establishes a connection with a server on a non-process network or whether a C2 server sends communication messages to it.

[0040] Step 3, Alert Trust Balancing and Uncertainty Handling: During the event correlation process, the behavioral quality functions obtained from multiple evidence sources in Step 2.5 undergo uncertainty handling and combination. Each evidence source provides numerical support for three categories of judgments: "real attack," "non-attack," and "uncertain state." Since different evidence sources may conflict or be inconsistent, these support values ​​need to be uniformly merged to form a comprehensive trust level for access attempt behavior.

[0041] For the behavioral quality functions of multiple evidence sources, they are combined recursively. Let the... The behavioral quality function given by each source of evidence is denoted as... It provides solutions for three scenarios: "real attack," "non-attack," and "uncertain state." , and In the combination up to the first When there is one source of evidence, the former The behavioral quality distribution obtained from the combination of evidence sources With the The behavioral quality distributions of the evidence sources are merged to obtain a new merged behavioral quality distribution. The merge operation calculates the support based on the matching relationship between the two sets of support and normalizes the conflict degree of contradictory parts so that the sum of the three supports obtained by merging is 1.

[0042] First, the behavioral quality function of the first source of evidence is denoted as the initial combination result: ;in, Values , or When the first When there is one source of evidence, the former The combined result of the evidence sources With the Behavioral quality function of each source of evidence Merge the results to obtain new combinations. For any state The comprehensive behavior quality distribution is obtained by merging the combined results of the first k evidence sources with the behavior quality function of the (k+1)th evidence source. satisfy:

[0043] in, This represents each possible behavioral state in the result obtained after merging the first k evidence sources; Let represent the various possible behavioral states provided by the (k+1)th source of evidence, and Values ​​and The values ​​are the same. (Set) This indicates that the state is considered after combination. All state pairs Its meaning can be specifically defined as: 1) When for hour, At least include ;2) When for hour, At least include ;3) The remaining combinations that do not have clear consistency between the two sources of evidence are considered by the system to be in an "uncertain state".

[0044] Conflict level Defined as:

[0045] in, This represents a set of contradictory pairs of states between two sources of evidence, including at least... and Through the analysis of The normalization process makes A new distribution of behavioral quality is obtained.

[0046] According to the above rules, from The process continues recursively until all evidence sources are merged, resulting in the comprehensive behavioral quality function. The comprehensive behavioral quality function is applicable to... , as well as The three states are given the fused support, which is used for subsequent threshold interpretation and attack phase judgment.

[0047] Step 4, Attack Scheme Association and Sequence Graph Generation: Based on the potential attack behavior sequence provided by the event correlator, an attack sequence graph is dynamically adjusted and generated. Specifically, based on the potential attack behaviors obtained in Step 2, a corresponding behavior node is generated for each attack behavior. Meta-information or semantic connections are added according to its attack scheme and kill chain stage, thus forming a graph structure containing overall scheme nodes, attack behavior nodes, and kill chain stage nodes. The nodes of this sequence graph contain semantic connections between the overall scheme, attack behavior, and kill chain stage. Its edges define the logical relationships between different attack behaviors, including common origin, common target, extension, etc., and an edge transition quality function value is set to reflect the confidence level of the transition.

[0048] The attack scheme association and sequence graph generation process includes the following steps: Step 4.1, attack sequence graph generation: Based on the potential attack behavior sequence generation in Step 2, an attack sequence graph based on real attack behaviors is generated. The nodes contain metadata, overall scheme, attack behaviors, and semantic connections of kill chain stages. Specifically, the potential attack behavior sequence arranged in chronological order is read, and a behavior node is generated for each attack behavior. The behavior node records at least the behavior identifier, timestamp, source host and target host, the corresponding {source address – target address} key-value pair, and the kill chain stage label determined in Step 2.2. For behavior sets that share the same attacker identifier, the same {source address – target address} key-value pair, or are in the same attack stage group, the corresponding overall scheme node is constructed, and semantic connections are established between the overall scheme node, behavior nodes, and kill chain stage nodes to obtain the initial attack sequence graph.

[0049] Step 4.2, Edge Connection Type Definition. Define the connection types of edges (directed edges) in the attack sequence graph, including same source, same target, extension, no overlap, and arbitrary overlap, to add semantic information; the possible connection types for directed edges from nodes of behavior A to nodes of behavior B are defined as shown in Table 6.

[0050] Table 6: Definition of Connection Types for Attack Behaviors

[0051] Based on the relationships between attack behaviors in the attack sequence graph, the connection types between two attack behaviors are obtained. These connection types add semantic information to the attack sequence graph, enabling it to describe attack behaviors and decision-making processes. Therefore, the edges in the attack sequence graph need to consider host connections and the behavioral transitions involved. By analyzing the paths and connection types in the attack sequence graph, the current stage of the kill chain in which the attacker's behavior occurs can be determined. Combined with event correlation alerts obtained in chronological order, the attack process can be reconstructed.

[0052] Step 4.3, Assigning Edge Transition Quality Functions. Set an edge transition quality function for each edge of the attack sequence graph. The semantics and probabilistic properties of the connections are defined. The quality distribution reflects the transition probability from one attack node to another and is a key component of the attack sequence graph structure, representing the likelihood of reconstructing the attack process. Its definition is shown in the formula:

[0053] in, This indicates that the attack step represented by this edge has "actually occurred". This indicates that the attack step represented by this edge has not occurred. This indicates an "uncertain" state where the attack steps represented by the edge cannot be distinguished as occurring or not occurring. To gauge the degree of direct support for "actual occurrence," To determine the direct support for "not having occurred", This indicates that the support level is left as uncertain. Parameter Used to characterize the initial confidence level of the attack transition represented by the edge that actually occurred.

[0054] Weighting factors The weight factor s can be determined based on features such as the connection type (same origin, same target, extension, etc.) of the edge, the quality function of the attacks at both ends, the edge's position in the attack sequence graph, and the historical alarm hit rate associated with that connection. It can also be configured according to different deployment environments. It should be noted that this invention does not limit the specific construction method and value rules of the weight factor s. Under the condition that… Under the premise of basic constraints and the ability to reflect the relative confidence of edges, different feature combinations, mapping functions, and parameter configurations can be selected according to different application scenarios. For example, linear mapping, piecewise functions, or other monotonic mapping methods can be used to combine the above features and obtain the corresponding... value.

[0055] Each edge of the attack sequence graph has an initial quality value. This value is correlated with the confidence level of the attack reconstruction path represented by the edge. A larger value indicates a higher confidence level. A higher value increases the system's initial trust in the corresponding transition, making paths involving that edge more likely to be identified as suspected attack paths, thus increasing the risk of false positives; a lower value... A lower initial trust value for the transition may lead to a lower overall confidence level for the actual attack path, thus increasing the risk of false negatives. Weighting factor Calibration can be performed during the trial operation phase according to the specific environment to adapt to different network and service scenarios.

[0056] The predefined attack sequence graph expands based on alerts by adding new edges that can explain undetected operations or detected anomalous behavior, dynamically interpreting the current attack state. The quality distribution depends on the shortest path length between the nodes in question and the quality of the edges along this shortest path, while assigning confidence to each path in the attack sequence graph.

[0057] Step 5, Attack Path Reconstruction and Reconstruction: Based on the attack sequence graph, all possible attack path schemes are traversed and reconstructed in chronological order. By adjusting the quality distribution of all behaviors in the path and evaluating the distribution of the path quality function, the potential attack path and attack sequence graph with the highest belief value exceeding a preset threshold are finally selected.

[0058] The attack path restoration and reconstruction process includes the following steps: Step 5.1, Path Traversal and Path Quality Function Modification: Traverse the nodes and their reachable edges in the attack sequence graph in chronological order to generate several candidate attack paths. Each candidate path consists of a set of consecutive directed edges. Composition. For each edge in the path. Its edge mass function is As defined in step 4.3, the support of this edge in the three states of "actually occurred", "did not occur", and "uncertain" is denoted as . , and To characterize the overall credibility of the entire path as a real attack path, a path quality function is constructed. Its definition and structure are consistent with the behavior quality function and edge quality function, still including three states: "actually occurred," "did not occur," and "uncertain." The path quality function is constructed by combining the edge quality functions of each edge on the path through a multiplicative chain. The specific definition is as follows.

[0059] The overall support for paths that are the actual attack paths is:

[0060] The overall support for paths that are not attack paths is

[0061] The remaining support levels that cannot be distinguished by the above two types of states are classified as uncertain states, defined as follows:

[0062] in This represents the number of edges in the path. This construction method allows the path confidence to reflect the coherence condition of all edges occurring consecutively on the path. If the "true occurrence" support of any edge in the path is low, then... This will decrease accordingly. Simultaneously, by modeling the cumulative support for the "not occurred" state... It can reflect the overall probability of path interruption. Under the above definition, , and The values ​​of all three are located in the interval [0,1], and the sum of the three is always 1.

[0063] Step 5.3, Optimal Path Selection: Among potential attack paths, the attack path with the highest confidence is selected by evaluating the path quality function distribution, and the corresponding attack sequence graph is constructed. Specifically, the confidence of the path is evaluated in the following ways: 1) Path Evaluation and Selection: Using the path quality function constructed in Step 5.1... Calculate the overall credibility of each potential attack path. For each path, a path quality function is used. , and These represent the support levels for the path as a real attack path, a non-attack path, and an uncertain state, respectively. The path with the highest trust level (i.e., ...) is selected. The highest path is considered the most likely attack path.

[0064] 2) Credibility Verification: In addition to evaluating the quality distribution of the paths, it is also necessary to ensure that the credibility of the selected path exceeds a preset threshold. This threshold can be dynamically adjusted based on factors such as the characteristics of the attack scenario, the length of the path, and the confidence level of the edges to ensure that the selected path has sufficient statistical credibility. By adjusting this threshold, the system can flexibly select the most probable attack path in different application scenarios.

[0065] 3) Path Selection and Update: During the attack path traversal, the credibility of each path is updated and a decision is made by detecting the quality function of each path. If the credibility of a path exceeds a preset credibility threshold, the path is considered a potential attack path and is retained; if the credibility of a path is below the threshold, the path is removed from the path list.

[0066] 4) Path Optimization and Output: Finally, the system will output the path with the highest confidence level, as the most likely attack path. This path will serve as a core part of the attack reconstruction model, helping to further reconstruct the attack behavior graph and providing a basis for subsequent attack behavior analysis.

[0067] Step 6, Attack Phase Matching and Result Output: Using the quality function value, the most credible attack sequence diagram and the most reasonable attack reconstruction path are found, thereby identifying the optimal result pair that best represents the attack activity. By taking the phase of the last attack node in the path as the current kill chain phase of the attacker, this method finally outputs the detected attack behavior, scheme, the phase the attacker is in, and a list of infected devices.

[0068] The attack phase matching and result output process includes the following steps: Step 6.1, Optimal Result Pair Identification. The system uses the quality function value to find the most credible attack sequence graph and the most reasonable attack reconstruction path, identifying the optimal result pair. Step 6.2, Minimum Confidence Level Control. The minimum confidence level of the result is controlled by comparing the quality function value with predefined thresholds and limits. Step 6.3, Kill Chain Phase Determination. In the selected result set, the system further checks whether the most credible path is included in the most credible graph to determine the final optimal solution. This optimal solution contains two key parts: the attack scheme, defined by the attack sequence graph in the optimal solution; and the attack phase matching, determined by the phase of the last attack node in the path, which represents the current kill chain phase of the attacker. If the system fails to find an optimal solution matching pair that meets the conditions, the association process will output two states: no attack or unknown attack. This judgment is based on whether the system detects an infected device or observes abnormal access behavior to other hosts within a specific time period. If no such abnormalities are detected, it is determined as no attack; otherwise, it is determined as an unknown attack.

[0069] Step 6.4, Output of Correlation Results. The output of the correlation process includes the detected attack behaviors and schemes, the last stage of the kill chain perceived by the attacker, and a list of detected infected devices. This information collectively helps us identify the time frame, development process, and schemes followed by the attack. By matching the detected attack schemes with the stages of the kill chain, accurate detection of the attack scenario is achieved.

[0070] As shown in Figure 2, this invention also proposes a threat detection system based on multi-stage attack process matching to achieve the aforementioned threat detection based on multi-stage attack process matching. This system includes: an alarm preprocessing module for formatting multi-source security alarm information and generating a standardized alarm containing alarm ID, timestamp, and sensor ID, while simultaneously constructing a globally unique event ID; an event association module, acting as a multi-source event correlator, for mapping alarm data to a network kill chain model, constructing an attack behavior sequence, and assigning a quality function to each potential attack behavior; and a trust balancing module, employing an evidence reasoning framework for handling incomplete knowledge, through robust combination... The system integrates dependency and conflict evidence with rule-based and log-robust combination rules, and introduces an environmental impact quality function to achieve a dynamic balance of alarm confidence. A scheme association module dynamically generates an attack sequence graph with semantically connected nodes and typified edges based on the attack behavior sequence, and assigns initial quality function values ​​to the edges. An attack path reconstruction module traverses the attack sequence graph in chronological order, reconstructs all possible attack scheme paths after adjusting the quality distribution, and selects the path with the highest belief value. An attack phase matching module identifies the optimal result pair formed by the most credible attack sequence graph and the most reasonable attack reconstruction path, thereby determining the current kill chain phase of the attacker.

[0071] The beneficial effects of this invention are that, compared with the prior art, this invention performs security analysis through multiple information sources, which can effectively integrate heterogeneous information and handle uncertainties, accurately identify and reconstruct attack schemes, and achieve complete attack situation awareness.

[0072] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0073] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0074] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0075] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A threat detection method based on multi-stage attack process matching, characterized in that, The process includes the following steps: Step 1, heterogeneous traffic detection is performed on the new power system, and the heterogeneous traffic is preprocessed to generate security alarm information, thus obtaining the information transmission route; Step 2, the preprocessed security alarm information is correlated using a multi-source event correlator to construct a sequence of potential attack behaviors, and a trust index is assigned to each attack behavior; Step 3, uncertainty processing and combination of the behavior quality functions from multiple evidence sources are performed to obtain a comprehensive behavior quality function; Step 4, based on the potential attack behavior sequence obtained in Step 2, corresponding behavior nodes are generated for each attack behavior, and metadata or semantic connections are added according to the attack scheme and kill chain stage to which the behavior node belongs. Based on the comprehensive behavior quality function, edge transition quality functions are assigned to each semantic connection to obtain an attack sequence graph; Step 5, based on the attack sequence graph and attack methods, potential attack paths are reconstructed through attack path restoration, and the quality of all behaviors in the path is assessed. The distribution is adjusted, and the distribution of the path quality function is evaluated. The potential attack path with the highest belief value and exceeding the preset threshold and the corresponding attack sequence graph are selected. Step 6: The potential attack path selected in step 5 and the corresponding attack sequence graph are structurally consistent to ensure that the path logic and the topological semantics of the graph are completely matched, thereby identifying the behavioral combination that best represents the attack activity and defining it as the optimal result pair. Then, the path quality function value in the optimal result pair is extracted and compared with the preset constraints to control the minimum confidence level of the result. The current state is determined as no attack or unknown attack based on the matching completeness of the optimal result pair. If the confidence level meets the standard and the match is successful, the kill chain stage is determined by mapping the behavioral attributes of the end node of the path in the optimal result pair. Based on the topological and temporal information covered by the optimal result pair, an association result report containing a list of infected devices, specific attack schemes and attack time ranges is output.

2. The threat detection method based on multi-stage attack process matching according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1, performing anomaly detection on heterogeneous traffic in the new power system, identifying and extracting contextual features strongly related to power business; Step 1.2, converting the heterogeneous traffic with anomalies into a standardized format containing power business semantics, generating security alarm information containing standard fields; Step 1.3, constructing globally unique event IDs through sensor identifiers, classifying same-source communication paths, and forming traversable information transmission routes.

3. The threat detection method based on multi-stage attack process matching according to claim 2, characterized in that, The security alarm information in the standard fields includes the following fields: alarm ID, load, load type, timestamp, protocol information, device endpoint, and sensor ID.

4. The threat detection method based on multi-stage attack process matching according to claim 2, characterized in that, Step 2 specifically includes: Step 2.1, Behavioral Analysis and Event Differentiation: Analyze the behavior executed on the device that issued the alarm signal, consider the communication behavior of the infected device, integrate system logs and security alarm information, and identify and extract recurring suspicious events; Step 2.2, Kill Chain Mapping: Obtain harmful behaviors based on security alarm information and suspicious events, map the harmful behaviors to the network kill chain model, and identify and define the single attacker's operation pattern; Step 2.3, Obtain an access attempt indicator dataset based on suspicious events and security alarm information, and systematically group it according to time series and network host targets to obtain a behavioral dataset arranged according to attack intent and time sequence; Step 2.4, Generate intrusion indicator key-value pairs based on the detected intrusion indicators. The format of the intrusion indicator key-value pairs is {source address-target address}. Local and remote access attempts are distinguished by the source address and target address; Step 2.5, Assign a behavioral quality function to the access attempt behavior based on the intrusion indicator key-value pairs, and then obtain the trust index corresponding to each access attempt behavior. The trust index serves as the probability that the alarm corresponding to each access attempt behavior represents a real attack behavior.

5. The threat detection method based on multi-stage attack process matching according to claim 4, characterized in that, Step 2.5 specifically includes: establishing a set of hypotheses for the access attempt behavior state. ,in, This indicates that the current access attempt is a genuine attack. This indicates that the current access attempt is a non-attack scenario; This indicates that the current access attempt is in a state of uncertainty, meaning it's impossible to distinguish between a real attack and a non-attack; Behavior Quality Function The domain is set All subsets of; among which, This indicates the degree to which the current evidence supports the conclusion that it was a real attack; This indicates the degree to which the current evidence supports the conclusion that there was no attack. This indicates the degree of support for uncertain states when the current access attempt is in a state of uncertainty. 、 and The values ​​of all values ​​belong to the interval [0,1] and satisfy: + + =1 Behavioral quality function Based on the security alert attributes and their associated features extracted from the system logs and real-time alerts for this access attempt behavior, the associated features include contextual information related to the behavior, alert severity, number of recurrences within a preset time window, degree of matching with intrusion indicators or attack technique templates, and degree of deviation of the communication pattern from the historical baseline.

6. The threat detection method based on multi-stage attack process matching according to claim 5, characterized in that, Step 3 specifically includes: Let the first The behavioral quality function corresponding to each source of evidence is: For any state Comprehensive behavioral quality distribution satisfy: in, This represents each possible behavioral state in the result obtained after merging the first k evidence sources; This represents the various possible behavioral states provided by the (k+1)th source of evidence introduced; Describing the degree of conflict, set This indicates that the state is considered after combination. All state pairs ;gather The meaning is defined as: when for hour, At least include ;when for hour, At least include The remaining combinations that do not have clear consistency between the two sources of evidence are considered by the system to be in an uncertain state.

7. The threat detection method based on multi-stage attack process matching according to claim 6, characterized in that, The degree of conflict The formula for calculation is: in, This represents a set of contradictory pairs of states between two sources of evidence, including at least... and 。 8. The threat detection method based on multi-stage attack process matching according to claim 1, characterized in that, Step 4 specifically includes: Step 4.1, generating an attack sequence graph, wherein the nodes of the attack sequence graph contain metadata, overall scheme, attack behavior, and semantic connections of the kill chain stages; Step 4.2, defining the edge connection types of the attack sequence graph, wherein the edge connection types include same source, same target, extended, non-overlapping, and arbitrary overlap; Step 4.3, setting an edge transition quality function for each edge of the attack sequence graph, wherein the edge transition quality function is used to define the semantic and probabilistic characteristics of the connection.

9. The threat detection method based on multi-stage attack process matching according to claim 8, characterized in that, Step 4.1 specifically includes: reading the sequence of potential attack behaviors arranged in chronological order, generating a behavior node for each attack behavior, wherein the behavior node records at least the behavior identifier, timestamp, source host and target host, the corresponding {source address – target address} key-value pair, and the kill chain stage label determined in step 2.2; for behavior sets that share the same attacker identifier, the same {source address – target address} key-value pair, or are in the same attack stage group, constructing the corresponding overall scheme node, and establishing semantic connections between the overall scheme node, behavior nodes, and kill chain stage nodes to obtain the initial attack sequence diagram.

10. The threat detection method based on multi-stage attack process matching according to claim 1, characterized in that, Step 5 specifically includes: Step 5.1, traversing nodes and their reachable edges in the attack sequence graph in chronological order to generate candidate attack paths; Step 5.2, using the edge quality function of each edge on the path as input, constructing the path quality function through multiplicative chain combination; Step 5.3, calculating the comprehensive credibility of each potential attack path based on the path quality function, and selecting the path with the highest comprehensive credibility. If the credibility of the selected path exceeds a preset threshold, then the path is considered a potential attack path, and the attack behavior graph is reconstructed using the potential attack paths.

11. The threat detection method based on multi-stage attack process matching according to claim 10, characterized in that, Step 6 specifically includes: Step 6.1, identifying the optimal result pair based on the potential attack paths and reconstructed attack behavior graph obtained in Step 5.3; Step 6.2, obtaining potential attack paths that meet the preset threshold limit by comparing the path quality function value corresponding to the potential attack path with a preset threshold; Step 6.3, taking the stage of the last attack node in the attack reconstruction path as the current kill chain stage of the attacker; Step 6.4, outputting the correlation process, including the detected attack behaviors and schemes, the kill chain stage last perceived by the attacker, and the list of detected infected devices.

12. A threat detection system based on multi-stage attack process matching, characterized in that, The method for implementing the threat detection method as described in any one of claims 1-11 includes: an alarm preprocessing module for formatting multi-source security alarm information and generating a normalized alarm containing an alarm ID, timestamp, and sensor ID, while constructing a globally unique event ID; an event association module, acting as a multi-source event correlator, for mapping alarm data to a network kill chain model, constructing an attack behavior sequence, and assigning a quality function to each potential attack behavior; and a trust balancing module that employs an evidence reasoning framework for handling incomplete knowledge, merging dependencies and conflicts through robust combination rules and logarithmic robust combination rules. The system employs several methods: 1) breaking through existing evidence and introducing an environmental impact quality function to achieve a dynamic balance of alarm confidence; 2) a scheme association module to dynamically generate an attack sequence graph with semantically connected nodes and typified edges based on the attack behavior sequence, and assign initial quality function values ​​to the edges; 3) an attack path reconstruction module to traverse the attack sequence graph in chronological order, reconstruct all possible attack scheme paths after adjusting the quality distribution, and select the path with the highest confidence value; and 4) an attack phase matching module to determine the current kill chain phase of the attacker by identifying the optimal result pair formed by the most credible attack sequence graph and the most reasonable attack reconstruction path.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the threat detection method according to any one of claims 1-11.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the threat detection method according to any one of claims 1-11.