Attack scene reconstruction method based on network threat clue causal mining

By constructing a causal map based on causal mining of network threat clues and integrating the marginal causal relationship of multi-source data, the problems of insufficient generalization capabilities and low efficiency of large-scale network data processing in complex attack scenarios are solved, and efficient and accurate attack scenario reconstruction and defense are achieved.

CN120263518APending Publication Date: 2025-07-04GUANGZHOU UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510545917.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-04

Smart Images

  • Figure CN120263518A_ABST
    Figure CN120263518A_ABST
Patent Text Reader

Abstract

The invention provides an attack scene reconstruction method based on network threat clue causal mining, and relates to the technical field of network security. The method comprises the following steps: performing feature extraction and representation on collected network data to obtain multi-source data; extracting open source data based on an open source threat intelligence platform to establish a marginal causal relationship of multi-source data; the method comprises the following steps: constructing a completely undirected graph based on multi-source data, initializing a condition set, testing whether adjacent node pairs are independent or not based on the condition set, removing edges without a direct causal relationship when the adjacent node pairs are independent, obtaining a skeleton graph, and expanding a marginal causal relationship based on the condition set at the moment; and performing edge orientation processing on the skeleton diagram based on the causal priori knowledge base to obtain a causal diagram, and obtaining a causal relationship between the network data based on the causal diagram and reconstructing an attack path. According to the method, a causal discovery process based on marginal causal priori knowledge extension is introduced, so that the generalization ability of the model is enhanced, and the adaptability to unknown attacks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to an attack scenario reconstruction method based on causal mining of network threat clues. Background Art

[0002] With the continuous improvement of informatization and digitalization levels, network security has long become an important issue for various industries and government agencies. The means of network attacks are becoming increasingly complex and diverse. In particular, advanced persistent threat (APT) attacks have become a huge challenge in modern network security protection. APT attacks are usually carried out by highly organized attackers, and the targets are often to obtain sensitive information, conduct espionage activities, or damage critical infrastructure, etc. These attacks often have the following characteristics: 1. Long-term and covert: Different from traditional hacker attacks, APT attackers often lurk in the target network for a long time and adopt covert means to avoid being detected. 2. High technical content and complexity: APT attacks often involve a variety of attack techniques and tools, including zero-day vulnerabilities, malware, social engineering attacks, etc. Attackers can flexibly respond to the monitoring and detection of defense systems.

[0003] In the field of network attack detection and defense, traditional intrusion detection / defense systems (IDS / IPS), security information and event management systems (SIEM), etc. have always been important components. These systems generate corresponding alarm messages for various suspicious behaviors. Naturally, alarm-based correlation analysis methods are used to analyze the attack path and reconstruct the attack scenario. Specifically, it mainly includes the following methods: 1. The alarm correlation method based on the probability statistical model uses probability theory and statistical analysis to identify the correlation relationship between alarm events. For example, the Bayesian network represents alarm events and their interdependent relationships by constructing a directed acyclic graph, uses prior probability and conditional probability to quantify the dependence relationship, and calculates the posterior probability of other related alarm events occurring under the condition that a certain alarm event occurs through Bayes' theorem, so as to determine the correlation degree between alarms, and combines alarm events with a relatively high correlation probability into possible attack links. The hidden Markov model (HMM) regards alarms as observed values of hidden states, assumes that the hidden state sequence controls the generation of alarms, analyzes the alarm sequence, and uses the transition and emission probability matrices to infer the most likely hidden state sequence, reveals the internal correlation of alarms, and completes the reconstruction and understanding of the attack link. 2. The alarm correlation method based on natural language processing (NLP) reconstructs the attack process through the semantic correlation of alarm information. First, the alarm information is preprocessed into a format suitable for the model, and then a deep learning model (such as LSTM, etc.) is used to learn the semantic features of the alarm information, and the semantic correlation degree between alarms is calculated accordingly. Then, relevant alarm events are connected in sequence to reconstruct the attack link, and finally, the attack scenario is reconstructed by combining information such as network topology, so as to deeply understand the semantic logic of alarm events and improve the accuracy and integrity of the attack link and scenario reconstruction.

[0004] Furthermore, with the rapid development of related technologies of the traceback graph, its application in the field of network security is becoming increasingly extensive and in-depth. As a directed relationship graph representing the interaction relationships between system objects (such as processes, files, network sockets), the traceback graph provides a new perspective and analysis means for the reconstruction of the attack path. Specifically, it mainly includes the following methods: 1. The method based on machine learning applies traditional machine learning techniques to the traceback graph. For example, after extracting features from elements such as nodes and edges, the community detection algorithm is used on the traceback graph. This method aims to identify malicious activities in the network by discovering communities formed by graph elements with closely related internal anomalies detected, and rely on the discovered communities to reconstruct the attack scenario, so as to achieve an in-depth investigation of the attack activities. 2. The method based on deep learning uses the powerful representation learning ability of neural networks to analyze the traceback graph. For example, through the graph neural network (GNN), the structural features of the traceback graph are learned to capture the complex patterns and relationships between graph elements. This method is beneficial to identifying intricate attack paths and can provide a finer-grained understanding of the attack link, thus improving the accuracy of the attack link reconstruction.

[0005] However, based on probability statistical models that rely on accurate prior probabilities, the changes in the network environment and the uncertainty of unknown attacks make it difficult to accurately obtain prior probabilities, resulting in the model being unable to effectively identify attack associations, leading to deviations in attack scenario reconstruction and insufficient generalization ability. Deep learning models based on natural language processing (such as LSTM for alarm correlation) and deep learning-based traceability graph methods (such as graph neural networks) have complex internal structures, and the decision-making process involves numerous parameters and complex calculations; this makes it difficult for security experts to understand how the model arrives at the results of attack link reconstruction and scenario reconstruction, that is, the interpretability of the model is poor. Moreover, with the expansion of the network scale and the increase in data generated by network activities, the computational complexity of traditional probability statistical models (such as Bayesian networks calculating posterior probabilities) and some deep learning models (such as traceability graph methods based on graph neural networks when storing and calculating graph structure features) grows exponentially, demanding huge computational resources. This results in these technologies often being less efficient when dealing with large-scale network security data.

[0006] Therefore, there is an urgent need to provide a solution to improve the above problems. Summary of the Invention

[0007] The purpose of the present invention is to provide an attack scenario reconstruction method based on causal mining of network threat clues, which improves the problems of insufficient generalization ability and low efficiency in processing large-scale network data in existing network security technologies when dealing with unknown attacks in complex attack scenarios.

[0008] An attack scenario reconstruction method based on causal mining of network threat clues provided by the present invention adopts the following technical solutions:

[0009] Collect and preprocess network traffic data, process behavior data, and system log data, respectively extract and represent features from the network traffic data, process behavior data, and system log data to obtain multi-source data;

[0010] Extract prior knowledge related to multi-source data features based on an open-source threat intelligence platform, continuously mine marginal causal prior knowledge based on the prior knowledge, and construct a set of marginal causal relationships between multi-source data features;

[0011] Take the multi-source data as variables and model them as multiple nodes. The multiple nodes are interconnected to form a complete undirected graph. Initialize the size of the conditional set, gradually expand the conditional set by adding 1 successively in each loop during the iterative process, test whether adjacent node pairs are independent based on the conditional set, remove the edges without direct causal relationships when the adjacent node pairs are independent to obtain a skeleton graph, and use the conditional set at this time as the minimum separation set to expand the marginal causal relationship set;

[0012] Orient the edges of adjacent node pairs in the skeleton graph based on the extended marginal causal relationship set to obtain a causal graph, obtain the causal relationship between network data based on the causal graph, and reconstruct the attack scenario based on the causal relationship.

[0013] Optionally, in the process of respectively extracting and representing features from network traffic data, process behavior data, and system log data, it includes:

[0014] Calculate the traffic size of network sessions, the duration of TCP connections and UDP flows, and the packet sending frequency in network traffic data, and obtain the TCP flag bit and UDP port distribution;

[0015] Statistically analyze the sensitive file operation features of processes in a time window. The sensitive file operation features include frequency, file type, and operation type encoding. Statistically analyze registry operations by location, type, and frequency, and statistically analyze mutex operations by acquisition and release frequency and holding duration, and represent them in a structured format;

[0016] Statistically analyze event frequencies, mine event combination patterns, and extract the process behavior characteristics of users.

[0017] Optionally, in the process of connecting multiple nodes to form a complete undirected graph, when multiple node datasets are {X1, X2, X3, …, X N}, there are N(N - 1) / 2 edges in the complete undirected graph connecting these nodes.

[0018] Optionally, in the process of testing whether adjacent node pairs are independent based on a conditional set, it includes:

[0019] Calculate the conditional mutual information of adjacent node pairs based on the conditional set, and calculate the independence probability value based on the conditional mutual information. When the independence probability value is greater than the pre-set variable significance level probability threshold, it is determined that the adjacent nodes are independent under the given conditions.

[0020] Optionally, the adjacent node pairs need to meet the condition: size(adjSet(X i , G)\{X j}) ≥ m, where adjSet(X i , G)\{X j} represents the set of nodes adjacent to node X i in the current graph G and excluding node X j , m represents the size of the conditional set, and size represents the number of nodes in the node set.

[0021] Optionally, in the process of using the current conditional set as the minimum separation set to expand the marginal causal relationship, it includes:

[0022] When for any two non - adjacent nodes X and Y in a complete undirected graph, if X is an ancestor of Y, that is, there exists a marginal causal relationship X→Y, then each variable in the minimum separating set S of X and Y is an ancestor of Y, and when and only when the number of nodes in the minimum separating set S is 1, this node is a descendant of X and an ancestor of Y;

[0023] When for any two non - adjacent nodes X and Y in a complete undirected graph, if X is not an ancestor of Y, that is, there is no marginal causal relationship then each variable in the minimum separating set S of X and Y is not a descendant of X.

[0024] Optionally, in the process of edge - orienting adjacent node pairs in the skeleton graph based on the causal prior knowledge base, it includes:

[0025] Edge - orient adjacent node pairs in the skeleton graph based on direct marginal causal relationships;

[0026] Edge - orient adjacent node pairs in the skeleton graph based on V - structures;

[0027] Finally, edge - orient the remaining unoriented adjacent node pairs in the skeleton graph based on Meek's rules.

[0028] Optionally, in the process of edge - orienting adjacent node pairs in the skeleton graph based on direct marginal causal relationships, it includes:

[0029] When there is an X i -X j in the skeleton graph, and there is a direct causal relationship X i →X j in the marginal causal relationship set, orient the edge in the skeleton graph as X i →X j , where the X i -X j represents that node X i is adjacent to node X j .

[0030] Optionally, in the process of edge - orienting adjacent node pairs in the skeleton graph based on V - structures, it includes:

[0031] Obtain an unshielded node triple X i -X j -X k in the skeleton graph. If X j is not in the separating set sep(X i and X k ), then for the triple X i ,X k ), orient the triple X i -X j -X kOriented as a V - collision structure: X i →X j ←X k wherein the X i -X j -X k is represented as node X j adjacent to node X i and node X k while node X i is not adjacent to node X k .

[0032] Optionally, during the process of finally orienting the remaining unoriented adjacent node pairs in the skeleton graph based on the Meek rule, it includes:

[0033] When there is a structure of X i →X j -X k in the skeleton graph, and node X i and node X k are not adjacent, then orient X j -X k as X j →X k ;

[0034] When there is a structure of X i →X j →X k in the skeleton graph, and node X i and node X k are adjacent, i.e., X i -X k , then orient X i -X k as X i →X k ;

[0035] When there is a structure of X i →X j ←X k in the skeleton graph, and node X i and node X k are adjacent, i.e., X i -X k , then orient X i -X k as X i →X k .

[0036] An attack scenario reconstruction method based on causal mining of network threat clues provided by the present invention has the beneficial effects that:

[0037] 1. The present invention integrates multi-source data such as network traffic, system logs, and process behaviors, constructs a causal graph by using conditional independence testing and marginal causal prior knowledge, accurately captures the causal relationships of attack behaviors, can accurately restore the complex network attack processes of multi-stage and multi-technology fusion, and greatly improves the accuracy of attack scenario reconstruction.

[0038] 2. The present invention constructs a marginal causal prior knowledge base, widely collects various attack correlation information, covering various attack scenarios, technical means, etc., lays a knowledge foundation for dealing with unknown attacks, and through the expansion operation of marginal causal prior knowledge in the causal graph construction process, identifies the potential causal logic in threat clues, and then effectively excavates potential attack paths or patterns, improves the generalization ability of dealing with unknown attacks, enhances the defense ability against unknown attacks, and makes up for the deficiencies of the prior art in the face of unknown attacks.

[0039] 3. The causal graph of the present invention can be visually displayed, and security experts can intuitively master the attack process and deeply analyze the attack intention. Based on the interpretable causal relationships, the security team can more reasonably evaluate risks, determine the priority of vulnerability repair, allocate defense resources, improve the efficiency of network security defense, and reduce attack losses.

[0040] 4. The present invention uses marginal causal prior knowledge and conditional independence testing to optimize calculations, reduces complexity, can efficiently process a large amount of data, and at the same time, can stably and efficiently operate in various large-scale network environments, accurately analyze attack behaviors, and strengthen the security protection ability of large-scale networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 FIG. is a schematic structural implementation diagram of an attack scenario reconstruction method based on causal mining of network threat clues provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein are intended to mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.

[0043] An embodiment of the present invention provides a method for reconstructing an attack scenario based on causal mining of network threat clues. Refer to Figure 1 , including:

[0044] S1. Collect and preprocess network traffic data, process behavior data, and system log data, respectively extract and represent features from the network traffic data, process behavior data, and system log data to obtain multi-source data;

[0045] S2. Extract prior knowledge related to the features of multi-source data based on an open-source threat intelligence platform, continuously mine marginal causal prior knowledge based on the prior knowledge, and construct a set of marginal causal relationships between the features of multi-source data;

[0046] S3. Take the multi-source data as variables and model them as multiple nodes. The multiple nodes are interconnected to form a complete undirected graph. Initialize the size of the conditional set, gradually expand the conditional set by adding 1 sequentially in each loop during the iterative process, test whether adjacent node pairs are independent based on the conditional set, remove the edges without direct causal relationships when the adjacent node pairs are independent to obtain a skeleton graph, and use the conditional set at this time as the minimum separation set to expand the marginal causal relationship set;

[0047] S4. Perform edge orientation processing on adjacent node pairs in the skeleton graph based on the expanded marginal causal relationship set to obtain a causal graph, obtain the causal relationship between network data based on the causal graph, and reconstruct the attack scenario based on the causal relationship.

[0048] In some embodiments, during the execution of step S1, it includes:

[0049] S1.1. Collect and preprocess network traffic data, process behavior data, and system log data;

[0050] S1.2. Respectively extract and represent features from the network traffic data, process behavior data, and system log data.

[0051] Specifically, during the execution of step S1.1, it includes:

[0052] S1.1.1. Collect and preprocess network traffic data;

[0053] S1.1.2. Collect and preprocess process behavior data;

[0054] S1.1.3. Collect and preprocess system log data.

[0055] Specifically, during the execution of step S1.1.1, it includes: deploying packet capture tools at key network nodes, capturing network data packets according to a certain time window and subnet rules, and recording information such as source IP, destination IP, port, protocol, length, timestamp, etc. The key network nodes include gateways and core switches.

[0056] Furthermore, parse the original data packets into a structured format, and filter out data packets with abnormal lengths, checksum errors, abnormal source or destination IPs, and duplicate data packets.

[0057] Specifically, during the execution of step S1.1.2, it includes: using a file monitoring tool to monitor the operations of processes on sensitive files, including read, write, modify, and delete operations. The sensitive files include system configuration files, user password files, and key business data files.

[0058] For example, monitor access to Windows registry files. For registry operations, through the system audit tool, monitor the modify, add, and delete operations of processes on registry items; for example, record the operations of key registry items. Abnormal operations on these registry items may affect the system startup items or the running of services, and may be behaviors of malware attempting to achieve self-start or persistence.

[0059] Furthermore, clean the process behavior data, remove the normal parts in the sensitive file operations, registry operations, and mutex operations data caused by normal system maintenance or normal user operations, and unify data from different sources and formats into a structured table form.

[0060] Specifically, during the execution of step S1.1.3, it includes: collecting operating system and application logs. In Windows, obtain system, application, and security logs through the Event Viewer, and in Linux, collect them using the syslog service, and record information such as login events, system service start and stop, and file access.

[0061] Furthermore, parse system logs and application logs in different formats into a unified structured format, extract key information, and remove records that do not conform to the format specifications, are incorrect, or are duplicates.

[0062] Specifically, during the execution of step S1.2, when performing feature extraction and representation on network traffic data, process behavior data, and system log data respectively, it includes:

[0063] Calculate the traffic size of network sessions, the durations of TCP connections and UDP flows, and the packet sending frequency in network traffic data, and obtain the TCP flag bit and UDP port distribution; count the sensitive file operation characteristics of processes in a time window, where the sensitive file operation characteristics include frequency, file type, and operation type encoding, count registry operations by location, type, and frequency, and count mutex operations by acquisition and release frequency and holding duration, and represent them in a structured format; count event frequencies, mine event combination patterns, and extract the process behavior characteristics of users.

[0064] In some embodiments, during the execution of step S2, it includes: S2.1. Extract prior knowledge related to multi-source data characteristics based on an open-source threat intelligence platform;

[0065] S2.2. Data cleaning and classification;

[0066] S2.3. Continuously mine marginal causal prior knowledge based on the prior knowledge, and construct a set of marginal causal relationships between multi-source data characteristics.

[0067] Specifically, during the execution of step S2.1, it includes: determining reliable open-source threat intelligence sources such as the MITRE ATT&CK knowledge base, security vendors, and professional threat intelligence platforms, as well as platforms such as Virustotal, and extracting open-source data related to multi-source data characteristics based on the platforms.

[0068] For example, obtain the association information between specific network port scanning behaviors and subsequent possible intrusion attempts from open-source intelligence, such as common port scans of ports 135 and 445, such as SMB protocol vulnerability attacks against Windows systems; for sensitive file operations, collect the access and tampering patterns of various malware families for specific types of sensitive files; for process behaviors, obtain the causal connection intelligence between the startup of abnormal processes and specific system failures or security events, such as system crashes and abnormal network connection interruptions.

[0069] Furthermore, execute step S2.2 to clean the collected data, removing duplicate, ambiguous, or obviously incorrect information. For example, if there are different and contradictory descriptions of the same attack behavior in different intelligence sources, it is necessary to further verify and remove the incorrect information.

[0070] Specifically, during the execution of step S2.3, it includes: establishing marginal causal relationships of multi-source data based on the open-source data. For example, for process behaviors, if it is observed that an unknown process frequently calls system-sensitive APIs and matches the behavior pattern of a certain type of malware in open-source intelligence, then establish a marginal causal relationship between the abnormal API calls of the unknown process and malware activities.

[0071] In some embodiments, during the execution of step S3, it includes:

[0072] S3.1. Construct a complete undirected graph;

[0073] S3.2. Construct a skeleton graph;

[0074] S3.3. Expand the marginal causal relationship set.

[0075] Specifically, during the execution of step S3.1 to construct an undirected graph, it includes: using the multi-source data as variables and modeling them as multiple nodes, and connecting these multiple nodes to form a complete undirected graph.

[0076] Furthermore, during the process of connecting these multiple nodes to form a complete undirected graph, when the multiple node datasets are {X1, X2, X3, …, X N}, there are N(N - 1) / 2 edges in the complete undirected graph connecting these nodes.

[0077] Specifically, during the execution of step S3.2 to construct a skeleton graph, it includes:

[0078] S3.2.1. Set the size of the conditioning set;

[0079] S3.2.2. Iteratively select adjacent node pairs to be tested;

[0080] S3.2.3. Remove edges based on conditional independence testing to obtain a skeleton graph.

[0081] Specifically, during the execution of step S3.2.1, it includes: initializing the size of the conditioning set, setting the initial value of the conditioning set to 0, and gradually expanding the conditioning set by adding 1 sequentially in each loop during the iterative process;

[0082] In fact, when m = 0, test whether nodes X and nodes X i are independent without introducing any mediating variables, i.e., the conditioning set; j When m = 1, test whether nodes X k and nodes X i are independent given the conditioning set C = {X j} when introducing a mediating variable node X k ; and so on. As m increases, the conditioning set gradually expands, thereby gradually considering the potential influence of more variables on the causal relationship between nodes.

[0083] Further, perform step S3.2.2. This process will select adjacent node pairs for subsequent conditional independence tests based on the size m of the conditional set and the marginal causal prior knowledge set PK. The adjacent node pairs need to meet the condition: size(adjSet(X i ,G)\{X j})≥m, where adjSet(X i ,G)\{X j} represents the set of nodes adjacent to node X i in the current graph G and excluding node X j . m represents the size of the conditional set, and size represents the number of nodes in the node set.

[0084] Actually, the setting of this condition is considered that only when the number of neighbor nodes is large enough, at least m, is it possible to select a conditional set of size m for conditional independence tests. And according to the marginal causal prior knowledge set PK, variable pairs with known causal information are skipped, that is, if X j →X i or then these variable pairs are skipped. Thus, through the constraint of prior knowledge, repeated calculations are avoided, and the construction efficiency of the causal graph is improved.

[0085] Specifically, in the process of performing step S3.2.3, it includes: testing whether adjacent node pairs are independent based on the conditional set, and removing edges without direct causal relationships when the adjacent node pairs are independent to obtain a skeleton graph.

[0086] Further, testing whether adjacent node pairs are independent based on the conditional set includes: calculating the conditional mutual information of adjacent node pairs based on the conditional set, and calculating the independence probability value based on the conditional mutual information. When the independence probability value is greater than the preset variable significance level probability threshold, it is determined that the adjacent nodes are independent under the given conditions.

[0087] Specifically, in the process of calculating the conditional mutual information of adjacent node pairs, the following formula is used:

[0088]

[0089] where I(X;Y∣Z) is the conditional mutual information of adjacent nodes X and Y under the given conditional set Z, p(x,y,z) is the joint probability density function of the variables, and p(x,y∣z), p(x∣z), and p(y∣z) are conditional probability density functions.

[0090] Further, the kernel density estimation formula of the joint probability distribution p(x,y,z) is as follows:

[0091]

[0092] Where n represents the number of samples for which the variables are sampled, h is the bandwidth parameter used to control the smoothness of the estimate, d is the total dimension of these variables, m is the number of variables included in the conditioning set Z, and K(·) is the kernel function. For example, the Gaussian kernel function is often used:

[0093] Furthermore, the conditional probability distributions p(x, y|z), p(x|z), and p(y|z) are obtained by calculating the quotient of the joint probability distribution and the marginal probability distribution. Here, p(x, y|z) is taken as an example: In the kernel density estimation framework, p(x, y, z) is calculated using the joint probability distribution formula given above, and the marginal probability distribution p(z) is obtained through a similar formula, that is:

[0094] In fact, based on the conditional mutual information calculated above, further comparison is needed to finally determine the conditional independence of two variables under a given conditioning set. The variable significance level is a pre-set probability threshold, which is usually taken as 0.05 in practice. The independence probability value can be calculated based on the conditional mutual information. For example, it can be calculated using the permutation test, and then the independence probability value is compared with the probability threshold to determine whether they are independent.

[0095] Furthermore, according to the above conditional independence test results, if X i and X j are independent under the given conditioning set C (i.e., X i ⊥X j |C), it indicates that there is no direct causal relationship between these two variables. Therefore, the edge directly connecting them is removed. And, the conditioning set C at this time is recorded as the separating set S = sep(X i , X j ). Since the size of the conditioning set increases iteratively from the initial value, the separating set S is the minimum separating set of the variables X i and X j . This set is used for the subsequent extension of causal relationships.

[0096] Specifically, in the process of executing step S3.3 to expand the marginal causal relationship set, it includes: when any two non-adjacent nodes X and Y in the complete undirected graph, if X is an ancestor of Y, that is, there is a marginal causal relationship X → Y, then each variable in the minimum separating set S of X and Y is an ancestor of Y, and when and only when the number of nodes in the minimum separating set S is 1, this node is a descendant of X and an ancestor of Y; when any two non-adjacent nodes X and Y in the complete undirected graph, if X is not an ancestor of Y, that is, there is no marginal causal relationship then each variable in the minimum separating set S of X and Y is not a descendant of X.

[0097] In some embodiments, during the execution of step S4, it includes:

[0098] S4.1. Orient the edges of adjacent node pairs in the skeleton graph based on the extended marginal causal relationship set to obtain a causal graph;

[0099] S4.2. Obtain the causal relationship between network data based on the causal graph, and reconstruct the attack scenario based on the causal relationship.

[0100] Specifically, during the execution of step S4.1, it includes:

[0101] S4.1.1. Orient the edges of adjacent node pairs in the skeleton graph based on the direct marginal causal relationship;

[0102] S4.1.2. Orient the edges of adjacent node pairs in the skeleton graph based on the V-structure;

[0103] S4.1.3. Finally, orient the edges of the remaining unoriented adjacent node pairs in the skeleton graph based on the Meek rule.

[0104] Specifically, during the execution of step S4.1.1, it includes: When there is X in the skeleton graph i -X j , and there is a direct causal relationship of X i →X j in the marginal causal relationship set, orient the edge in the skeleton graph as X i →X j , where the X i -X j represents that node X i is adjacent to node X j .

[0105] Furthermore, when executing step S4.1.2 and orienting the edges of adjacent node pairs in the skeleton graph based on the V-structure, it includes: Obtain an unshielded node triple X i -X j -X k in the skeleton graph. If X j is not in the separation set sep(X i , X k ) of X i and X k , orient the triple X i -X j -X k into a V-collider structure: X i →X j ←X k , where the X i -X j-X k Denoted as node X j With node X i , node X k Adjacent, node X i Not adjacent to node X k .

[0106] Furthermore, in the process of performing step S4.1.3 and finally orienting the edges of the remaining unoriented adjacent node pairs in the skeleton graph based on the Meek rule, it includes: when there is a structure of X i →X j -X k in the skeleton graph, and node X i and node X k are not adjacent, then orient X j -X k as X j →X k ; when there is a structure of X i →X j →X k in the skeleton graph, and node X i and node X k are adjacent, that is, X i -X k , then orient X i -X k as X i →X k ; when there is a structure of X i →X j ←X k in the skeleton graph, and node X i and node X k are adjacent, that is, X i -X k , then orient X i -X k as X i →X k .

[0107] Specifically, in the process of performing step S4.2, it includes: comprehensively and deeply analyzing the constructed causal graph to identify the key causal paths and core nodes therein. For example, identifying possible key hubs of attack propagation that have causal associations with many other nodes through analysis. In addition, through the visual display of the causal graph and combined with the knowledge of network security experts, further exploring potential causal relationships, attack paths, and patterns. For example, analyzing whether there are hidden causal chains, that is, seemingly unrelated events ultimately lead to serious security threats through indirect causal relationships.

[0108] Furthermore, data such as network traffic, system logs, and process behaviors are continuously collected according to a certain time window. With the evolution of the network environment, the causal graph is continuously updated in a timely manner to discover new causal relationships and security risks. For example, when a new abnormal traffic pattern is discovered, its possible causes and potential impacts are analyzed through the causal graph. In addition, by comparing the changes in the causal graph at different time periods, the evolution trend of the network security situation is explored, and potential security risks are predicted. This ensures that the protection always adapts to the actual security needs, thereby effectively defending against various security threats.

[0109] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.

Claims

1. An attack scenario reconstruction method based on causal mining of cyber threat clues, characterized in that, Including the following steps: Collect and preprocess network traffic data, process behavior data, and system log data, respectively extract and represent features from the network traffic data, process behavior data, and system log data to obtain multi-source data; Extract prior knowledge related to the features of the multi-source data based on an open-source threat intelligence platform, continuously mine marginal causal prior knowledge based on the prior knowledge, and construct a set of marginal causal relationships between the features of the multi-source data; Use the multi-source data as variables and model them as multiple nodes. The multiple nodes are interconnected to form a complete undirected graph. Initialize the size of the conditioning set, gradually expand the conditioning set by adding 1 successively in each loop during the iterative process. Based on the conditioning set, test whether adjacent node pairs are independent. When the adjacent node pairs are independent, remove the edges representing direct causal relationships between them to obtain a skeleton graph, and use the conditioning set at this time as the minimum separating set to expand the set of marginal causal relationships; Based on the expanded set of marginal causal relationships, perform edge orientation processing on adjacent node pairs in the skeleton graph to obtain a causal graph. Based on the causal graph, obtain the causal relationships between network data. Based on the causal relationships, mine key attack paths and core nodes, and reconstruct the attack scenario.

2. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 1, characterized in that During the process of respectively extracting and representing features from network traffic data, process behavior data, and system log data, it includes: Calculate the traffic size of network sessions in network traffic data, the duration of TCP connections and UDP flows, the packet sending frequency, and obtain the TCP flag bit and UDP port distribution; Statistically analyze the sensitive file operation characteristics of processes in a time window. The sensitive file operation characteristics include frequency, file type, and operation type encoding. Statistically analyze registry operations by location, type, and frequency, and statistically analyze mutex operations by acquisition and release frequency and holding duration, and represent them in a structured format; Statistically analyze event frequencies, mine event combination patterns, and extract the process behavior characteristics of users.

3. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 1, wherein In the process of connecting multiple nodes to form a complete undirected graph, when the multiple node datasets are {X1, X2, X3, …, X N}, there are N(N - 1) / 2 edges in the complete undirected graph connecting these nodes.

4. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 1, wherein During the process of testing whether adjacent node pairs are independent based on the conditioning set, it includes: Calculate the conditional mutual information of adjacent node pairs based on the conditioning set, and calculate the independence probability value based on the conditional mutual information. When the independence probability value is greater than the pre-set variable significance level probability threshold, it is determined that the adjacent nodes are independent under the given conditions.

5. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 4, characterized in that, The adjacent node pairs need to satisfy the condition: size(adjSet(X i ,G)\{X j})≥n, where adjSet(X i ,G)\{X j} represents the set of nodes adjacent to node X i in the current graph G and excluding node X j , n represents the size of the condition set, and size represents the number of nodes in the node set.

6. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 1, wherein, During the process of using the conditioning set at this time as the minimum separating set to expand the set of marginal causal relationships, it includes: For any two non-adjacent nodes X and Y in the complete undirected graph, if X is an ancestor of Y, that is, there is a marginal causal relationship X→Y, then each variable in the minimum separating set S of X and Y is an ancestor of Y, and when and only when the number of nodes in the minimum separating set S is 1, this node is a descendant of X and an ancestor of Y; When for any two non - adjacent nodes X and Y in a complete undirected graph, if X is not an ancestor of Y, that is, there is no marginal causal relationship then each variable in the minimum separating set S of X and Y is not a descendant of X.

7. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 1, wherein During the process of performing edge orientation processing on adjacent node pairs in the skeleton graph based on the expanded set of marginal causal relationships, it includes: Perform edge orientation on adjacent node pairs in the skeleton graph based on direct marginal causal relationships; Perform edge orientation on adjacent node pairs in the skeleton graph based on V-structures; Finally, perform edge orientation on the remaining unoriented adjacent node pairs in the skeleton graph based on the Meek rule.

8. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 7, characterized in that, During the process of performing edge orientation on adjacent node pairs in the skeleton graph based on direct marginal causal relationships, it includes: When X exists in the skeleton graph i -X j , and X exists in the marginal causal relationship set i →X j of the direct causal relationship, orient the edge in the skeleton graph as X i →X j , where the X i -X j represents the node X i adjacent to the node X j .

9. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 8, characterized in that, During the process of edge orientation for adjacent node pairs in the skeleton graph based on the V-structure, it includes: Obtain the unshielded node triple X in the skeleton graph i -X j -X k , if X j is not in the separation set sep(X i and X k ), the triple X i ,X k ) is oriented as a V - collider structure: X i -X j -X k is represented as node X i →X j ←X k , where the X i -X j -X k means that node X j is adjacent to node X i , and node X k is not adjacent to node X i and node X k .

10. The attack scenario reconstruction method based on causal mining of network threat clues according to claim 9, characterized in that, During the process of edge orientation for the remaining unoriented adjacent node pairs in the skeleton graph based on the Meek rule finally, it includes: When there is an X in the skeleton diagram i →X j -X k structure, and node X i and node X k are not adjacent, then X j -X k is oriented as X j →X k ; When there is an X in the skeleton diagram i →X j →X k structure, and node X i and node X k are adjacent, that is, X i -X k , then orient X i -X k as X i →X k ; When there is an X in the skeleton diagram i →X j ←X k structure, and node X i and node X k are adjacent, that is, X i -X k , then orient X i -X k as X i →X k .

Citation Information

Cited By

  • Response processing method and system for network security event

    CN120658497A