A method for restoring network attack chains based on knowledge graphs

Through the cyberattack chain restoration method based on knowledge graph technology, using expert knowledge to enrich the semantics of the attack chain, the problems such as false alarms and dependency explosions in the recovery of the cyberattack chain in the existing technology are solved, and the traceability analysis results with higher accuracy and readability are achieved.

CN116112211BActive Publication Date: 2025-06-17ZHUHAI HENGQIN KUAJINGSHUO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211564513.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-07
Publication Date
2025-06-17
Estimated Expiration
2042-12-07

AI Technical Summary

Technical Problem

The existing technology cannot provide stable and effective detection capabilities, and false alarms, dependence explosions, disconnection and other problems are prone to occur in network attack chain restoration, and there is a lack of relevant technologies and methods for network attack chain restoration based on expert knowledge.

Method used

Using the network attack chain restoration method based on knowledge graph technology, through the steps of knowledge extraction, alarm recognition, security knowledge graph construction, local attack chain recognition, alarm false alarm elimination and local attack chain merging, expert knowledge is used to enrich the semantics of the attack chain to improve the restoration accuracy of the attack chain during the traceability process.

Benefits of technology

It achieves higher accuracy and readability traceability analysis results, which can effectively eliminate false positives, enhance the connectivity and interpretability of the attack chain, and provide more targeted protection measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116112211B_ABST
    Figure CN116112211B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for restoring a network attack chain based on a knowledge graph. The method includes the following steps: a knowledge extraction step of parsing system logs and intelligence texts and encapsulating them into JSON files in a unified format; an alarm recognition step of combining security devices such as EDR to identify alarm information in the logs; a security knowledge graph construction step of instantiating the security knowledge graph using log data combined with external intelligence information and annotating alarm information on the graph; a local attack chain recognition step of finding an attack entry point based on the alarm information and further identifying a local attack chain on the security knowledge graph; an alarm false alarm elimination step of scoring the alarm itself using external intelligence information, thereby scoring the local attack chain and filtering out false alarms according to an expert threshold; a local attack chain merging technique of merging existing attack chains according to overlapping entities of the local attack chains and simultaneously finding potential association information between different attack chains. Based on traditional attack traceability techniques, the present invention introduces knowledge graph technology to increase the semantic information of the interaction behavior between system entities in the underlying logs, realizing a more accurate and readable restoration of the attack chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security, and particularly to a method for restoring a network attack chain. Background Art

[0002] Network security, as the foundation for protecting the reliable and normal operation of computer and network systems, aims to prevent illegal intrusion and resource destruction by malicious groups such as hackers and is involved in important fields such as politics, economy, national defense, and military. With the continuous expansion of the scale of Internet users and the doubling of the number of networked terminal devices, network attack incidents from anonymous hackers or competitors around the world are becoming increasingly frequent, and network attack technologies are also constantly innovating.

[0003] APT attack (Advanced Persistent Threat) is a long-term and continuous network attack that uses advanced attack means to mine highly sensitive information in order to achieve attack purposes such as information stealing or resource control for specific targets. Different from traditional network attacks, APT has characteristics such as latency, long duration, and concealment, making it difficult to prevent and detect, which increases the difficulty of network security defense.

[0004] In recent years, relevant scholars have focused their defense efforts on attack tracing. Attack tracing analyzes and identifies attack behaviors by collecting information data such as traffic and logs, and can restore to a certain extent the attacker's attack path, technical strategy, and true intention, and discovers how the attacker completes the penetration of the system through what attack process.

[0005] During the process of attack tracing, one or more attack chains for describing the context information of the attacker's attack behavior will be generated. An attack chain usually appears as a path or subgraph in a graph structure. The nodes in the graph can represent various entities, such as system processes, files, ports, vulnerability intelligence, and alarm information, etc. The edges in the graph can represent the interaction relationships between entities, such as a process reading a file or communicating with a specific port, etc. When a security disaster occurs, the tracer will restore the attacker's attack link by mining the attack behavior, so as to detect and restore the process of APT attack penetrating the system, showing how the attacker gradually attacks the system and what kind of impact it has caused.

[0006] Currently, the more widely used attack chain restoration methods mainly use the traceability graph to identify the attack path. The traceability graph is a collection of system interaction events, which describes the data flow, control flow and timing relationship between system entities as a directed graph, effectively describes the underlying logical relationship between system entities, and retains the historical records of all system executions. Compared with traditional attack detection methods, the traceability graph preserves system behaviors irrelevant to attacks, has richer underlying information and execution history, makes attack detection more effective, solves the problem of narrow coverage of traditional methods, and is of great significance for the attack traceability of APT attacks.

[0007] The attack chain restoration technology is crucial for the entire attack traceability. By analyzing the attack chain, not only can the entire attack process of the attacker be obtained, but also more targeted protection or blocking measures can be formulated accordingly to achieve active defense. However, the current attack traceability methods cannot provide a stable and effective detection ability. There are problems such as false positives, dependency explosion, and disconnection in the restored attack chain, and a comprehensive attack chain restoration cannot be carried out. The main reasons include the following two points:

[0008] One is the non-comprehensiveness of collecting information such as system logs, incorrect data collection methods and graph construction strategies, or alarm forgetting caused by the overly long latency of APT attacks. Collecting system logs during long-term operation will bring expensive performance overhead and affect the normal operation of the original business system. Therefore, in order to meet the basic log volume requirements for traceability and evidence collection, the actual data probe collection scheme only collects some system logs. Due to the non-comprehensiveness of the logs, the necessary paths required for attack chain restoration may be missing in the traceability graph constructed by the traceability analysis method relying on the underlying logs, and the disconnected attack chains cannot be associated, and only multiple unconnected attack chains can be obtained.

[0009] The second is the lack of high-order semantics in the attack traceability process, which makes it impossible to depict high-order associations. High-order semantics includes some attack intelligence that describes attack intentions, and can specify the specific methods and goals to be achieved in this stage of the attack. However, the introduction of a knowledge base is lacking in the construction process of the traceability graph, and it only has the interaction information of the underlying system processes and files itself, so the identification of the attack path can only rely on the statistical information of nodes and edges, lacking the semantics required for analysis, and unable to depict the attacker's attack intention. Therefore, the interpretability of the traceability graph is weak, and the semantic guidance provided for security experts to conduct traceability reasoning is very limited.

[0010] Introducing the guidance of expert knowledge in the process of attack traceability is a technical method urgently needed in the industry. That is, there is currently a lack of technologies and methods related to the restoration of network attack chains based on expert knowledge. Summary of the Invention

[0011] To solve the problem of the current lack of relevant technologies and methods for restoring network attack chains based on expert knowledge, the present invention provides a method for restoring network attack chains based on knowledge graph technology, aiming to enrich the semantics of attack chains using expert knowledge, thereby improving the accuracy of restoring attack chains during the traceability process.

[0012] The present invention discloses a method for restoring network attack chains based on a knowledge graph, comprising the following steps:

[0013] Step 1, knowledge extraction, obtaining entities and relationships in log or intelligence texts, and using the ETW logs collected in the Windows system, the Auditd logs in the Linux system, and open-source intelligence and information;

[0014] Step 2, alert identification, using a suspicious information flow correlation identification method to identify alert events in the logs, through a detection rule library based on an expert model and an EDR tool;

[0015] Step 3, constructing a security knowledge graph, instantiating the graph in a NoSQL database, and using a Neo4j database;

[0016] Step 4, local attack chain identification, finding the earliest alert as the attack entry, and dividing a more accurate and smaller-scale attack path in the system entity network of the graph, thereby generating an attack subgraph;

[0017] Step 5, eliminating false alerts, using an abnormal path scoring method to score the identified local attack chains, distinguishing the danger level and suspicious level of different attack chains, and filtering out possible false alerts based on an expert experience threshold;

[0018] Step 6, merging local attack chains, merging attack chains with the same entity or implicit association relationships into one to obtain the restored attack chain;

[0019] In step 5, the score for a single alert is defined as:

[0020] ThreatScore(technique)=(a*SeverityScore)+(b*LikelihoodScore)

[0021] Where SeverityScore and LikelihoodScore are two risk assessment indicators included in the CAPEC (Common Attack Pattern Enumeration and Classification) intelligence, which are quantified to 1-5 according to the levels of Very Low, Low, Medium, High, and Very High;

[0022] In step 5, considering the partial order relationship between alarm events, the subsequence of the ADP (Alert Dependency Path) attack flow will be found according to the tactical map of UKC (Unified Kill Chain), and the one with the highest score in the longest subsequence that conforms to the UKC order will be selected;

[0023] The scoring of UKC is defined as:

[0024]

[0025] Among them, UKC i represents the i-th UKC sub-chain, represents the j-th alarm event on this sub-chain, represents the score of the j-th alarm event on this sub-chain, and len(UKC i ) represents the length of this sub-chain;

[0026] In step 5, a penalty factor is introduced to characterize the impact of potential false alarms on the ADP score, and the score is reduced by the number of repetitions of the UKC stage in the attack flow of ADP;

[0027] The definition of the penalty factor is:

[0028]

[0029] Among them, len(ADP) represents the length of ADP, and n i represents the number of alarms corresponding to the i-th UKC step in ADP;

[0030] In step 5, the score of ADP is:

[0031]

[0032] Among them, is the set of all longest subsequences obtained by dividing ADP according to the UKC model.

[0033] Furthermore, in step 2, the "Techniques" defined by ATT&CK will be mainly used to standardize the description of alarm events; if it is detected that there is an attempt in the log to destroy a large amount of data and files on a specific system or network, thereby interrupting the availability of system, service, and network resources, the corresponding relationship will be marked as the attack technique "Data Destruction" in ATT&CK in the security knowledge graph.

[0034] Further, in step 3, the ontology framework of the knowledge graph is mainly divided into two modules: one is the endogenous intelligence module, which will mainly read from the system logs. In this module, the processes, files, and communications, the three types of system entities in the log data, will be saved as nodes in the graph, and the system calls between entities will be summarized into five relationship types: CREATE, DELETE, EXECUTE, READ, and WRITE; the other is the external intelligence module, which will mainly read from the intelligence database and the knowledge database. In this module, the alarm nodes matched by combining the open-source knowledge base and the intelligence database will be stored, including the kill chain nodes designed by combining the 18 tactics defined by The Unified Kill Chain, and the threat intelligence nodes for scoring alarms designed by combining CAPEC; these nodes correspond to the alarm behaviors in the log, that is, a certain system call relationship in the endogenous intelligence module; the two modules are connected by the alarm information matched from the log, that is, the single-step attack behavior.

[0035] Further, in step 4, the earliest alarm event will be defined as the Attack Origin Point (AOP); the AOP corresponds to a system entity that satisfies the following two conditions: (1) the entity corresponds to a process that executes the alarm event; (2) there are no other alarm events in the reverse trace starting from this alarm event in the knowledge graph; the local attack chain ADP (Alert Dependency Path) represents the subgraph derived from the attack origin point AOP in the security graph, and each path on it has an alarm event, corresponding to a short-term and strongly coherent attack executed by the attacker; starting from the AOP as the entry point, perform a forward traversal in the graph, and add all the edges with alarm events found during the traversal process to the path; if there are no more alarm events after a certain edge, stop the traversal process at this edge; by continuously repeating the above process until all alarm events have been traversed at least once, all the ADPs in the graph can be identified.

[0036] The beneficial effects of the present invention are as follows: The present invention can solve the problem of the lack of related technologies and methods for restoring network attack chains based on expert knowledge. The method for restoring network attack chains based on the knowledge graph of the present invention can enrich the semantics of the attack chain by using expert knowledge, and achieve more accurate and more readable traceability analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the ontology framework of the security knowledge graph.

[0038] Figure 2 is the schematic flow diagram of the attack chain completion method based on the knowledge graph of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] The following will further elaborate on the technical solution of the invention in conjunction with Figure 2 to provide a more detailed description of the technical solution of the invention.

[0040] This figure provides a method for constructing a security knowledge graph. The ontology framework of the graph is as Figure 1 shown, and is used to implement the attack chain restoration task in multiple different scenarios, including the following steps:

[0041] Step 1, knowledge extraction, to obtain entities and relationships in log or intelligence texts, and it is possible to use ETW logs collected from Windows systems, Auditd logs from Linux systems, and open-source intelligence and information;

[0042] In this step, the Auditd log data collected from the attack and defense drill scenario is utilized, combined with relevant intelligence information on the CAPEC official website, and the 18 attack tactics of The Unified Kill Chain.

[0043] Step 2, alarm recognition, using the suspicious information flow association recognition method to identify alarm events in the logs, which can be achieved through a detection rule library based on an expert model and an EDR tool;

[0044] In this step, the Wazuh EDR tool is introduced to assist in identifying alarm behaviors existing in the Auditd logs.

[0045] Step 3, security knowledge graph construction, instantiating the graph in a NoSQL database, and it is possible to use the Neo4j database;

[0046] In this step, the Neo4j database is used to instantiate the knowledge graph.

[0047] Step 4, local attack chain identification, finding the earliest alarm as the attack entry, and dividing a more accurate and smaller-scale attack path in the system entity network of the graph to generate an attack subgraph;

[0048] In this step, the earliest occurring alarm event that has not been traversed yet is used as the AOP, and a forward DFS is issued to generate a subgraph rooted at this vertex.

[0049] Step 5, alarm false positive elimination, using the abnormal path scoring method to score the identified local attack chains, distinguish the danger level and suspicious level of different attack chains, and filter out possible false positives based on the expert experience threshold.

[0050] In this step, referring to the longest increasing subpath, the UKC subchain is found using a method combining dynamic programming and greed, and the abnormal path scoring method is used to score the identified local attack chains. The attack chain score is defined as:

[0051]

[0052] Among them, T is the set of all the longest subsequences divided by ADP according to the UKC model, and UKC i represents the i-th UKC sub-chain among them, represents the j-th alarm event on this sub-chain, and len(UKC i ) represents the length of this sub-chain.

[0053] Step 6, Merge local attack chains. Merge the attack chains with the same entity or implicit association relationship into one to obtain the restored attack chain.

[0054] In this step, the attack chains with the same entity or implicit association relationship are merged into one.

[0055] This embodiment also provides a network attack chain restoration system based on a knowledge graph, including: a knowledge extraction unit, an alarm recognition unit, a security knowledge graph construction unit, a local attack chain recognition unit, an alarm false alarm elimination unit, and a local attack chain merging unit

[0056] The knowledge extraction unit is used to parse system logs and intelligence texts and encapsulate them into a JSON file in a unified format;

[0057] The alarm recognition unit uses a suspicious information flow association recognition method to identify alarm events in the logs, and can use a detection rule library based on an expert model and an EDR tool;

[0058] The security knowledge graph construction unit instantiates the graph in a NoSQL database, and a Neo4j database can be used;

[0059] The local attack chain recognition unit finds the earliest alarm as the attack entry, divides a more accurate and smaller-scale attack path in the system entity network of the graph, and thus generates an attack sub-graph;

[0060] The alarm false alarm elimination unit uses an abnormal path scoring method to score the identified local attack chains, distinguish the danger level and suspicious level of different attack chains, and filter out possible false alarms based on the expert experience threshold.

[0061] The local attack chain merging unit merges the attack chains with the same entity or implicit association relationship into one to obtain the restored attack chain.

[0062] Although the present invention has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present invention. The protection scope of the present invention is defined by the appended claims and may include various variations, modifications, and equivalent solutions made to the invention without departing from the protection scope and spirit of the present invention.

Claims

1. A method for restoring a network attack chain based on a knowledge graph, characterized in that, It includes the following steps: Step 1, Knowledge extraction: Obtain entities and relationships in log or intelligence texts, using ETW logs in Windows systems, Auditd logs in Linux systems, and open-source intelligence and information for collection; Step 2, Alert identification: Use a suspicious information flow association and identification method to identify alert events in the logs, through a detection rule library based on an expert model and an EDR tool; Step 3, Security knowledge graph construction: Instantiate the graph in a NoSQL database, using the Neo4j database; Step 4, Local attack chain identification: Find the earliest alert as the attack entry point, and divide a more accurate and smaller-scale attack path in the system entity network of the graph to generate an attack subgraph; Step 5, False alert elimination: Use an abnormal path scoring method to score the identified local attack chains, distinguish the danger level and suspicious level of different attack chains, and filter out possible false alerts based on the expert experience threshold; Step 6, Local attack chain merging: Merge attack chains with the same entity or implicit association relationship into one to obtain the restored attack chain; In Step 5, the score for a single alert is defined as: ThreatScore(technique)=(a*SeverityScore)+(b*LikelihoodScore), where SeverityScore and LikelihoodScore are two risk assessment indicators included in CAPEC (Common Attack Pattern Enumeration and Classification) intelligence, and are quantified to 1-5 according to the levels of Very Low, Low, Medium, High, Very High; In Step 5, considering the partial order relationship between alert events, the tactical diagram based on UKC (Unified Kill Chain) will be used to find subsequences of the ADP (Alert Dependency Path) attack flow, and select the one with the highest score in the longest subsequence that conforms to the UKC order; The score definition of UKC is: Among them, UKC i represents the i-th UKC sub-chain, represents the j-th alarm event on this sub-chain, represents the score of the j-th alarm event on this sub-chain, and len(UKC i ) represents the length of this sub-chain; In Step 5, a penalty factor is introduced to describe the impact of potential false alerts on the ADP score, and the score is reduced by the number of repetitions of the UKC stage in the attack flow of ADP; The definition of the penalty factor is: where len(ADP) represents the length of ADP, and n i represents the number of alarms corresponding to the i-th UKC step in ADP; In Step 5, the score of ADP is: Among them, is the set of all the longest subsequences divided by ADP according to the UKC model.

2. The method for restoring a network attack chain based on a knowledge graph according to claim 1, characterized in that, In Step 2, the "techniques" defined by ATT&CK will be used to standardize the description of alert events; if it is detected that there is an attempt in the log to destroy a large amount of data and files on a specific system or network, thereby interrupting the availability of system, service, and network resources, the corresponding relationship will be marked as the attack technique "Data Destruction" in ATT&CK in the security knowledge graph.

3. The method for restoring a network attack chain based on a knowledge graph according to claim 1, characterized in that, In step 3, the ontology framework of the knowledge graph is divided into two modules: one is the endogenous intelligence module, which will be read from the system logs. In this module, the processes, files, and communications, which are three types of system entities in the log data, will be saved as nodes in the graph, and the system calls between entities will be classified into five relationship types: CREATE, DELETE, EXECUTE, READ, and WRITE. The other is the external intelligence module, which will be read from the intelligence database and the knowledge database. In this module, the alert nodes matched by combining the open-source knowledge base and the intelligence database will be stored, including the kill chain nodes designed by combining the 18 tactics defined by The Unified Kill Chain, and the threat intelligence nodes for scoring alerts designed by combining CAPEC. These nodes correspond to the alert behaviors in the log, that is, a certain system call relationship in the endogenous intelligence module. The two modules are connected by the alert information matched from the log, that is, the single-step attack behavior.

4. The method for restoring a network attack chain based on a knowledge graph according to claim 1, characterized in that, In step 4, the earliest alert event will be defined as the Attack Origin Point (AOP); the AOP corresponds to a system entity that meets the following two conditions: (1) the entity corresponds to a process that executes the alert event; (2) no other alert events are included in the reverse trace from this alert event in the knowledge graph; the local attack chain ADP (Alert Dependency Path) represents the subgraph derived from the attack origin point AOP in the security graph, where each path has an alert event, corresponding to a short-term and strongly coherent attack executed by the attacker. Starting from the AOP as the entry point, perform a forward traversal in the graph, and add all the edges with alert events found during the traversal to the path. If no alert event appears after a certain edge, stop the traversal at that edge. By continuously repeating the above process until all alert events have been traversed at least once, all the ADPs in the graph can be identified.

Citation Information

Patent Citations

  • Network attack path tracking method and device

    CN113783896A

  • Intelligent system for detecting multistage attacks

    US20200143052A1