Heuristic APT attack tracing method based on flow detection and attack graph

Through a heuristic method based on traffic detection and attack graphs, anomaly detection and attack tracing are handled in stages, which solves the problem of high dependence on data integrity in existing technologies, realizes lightweight real-time APT attack tracing, and improves detection depth and adaptability.

CN120602118APending Publication Date: 2025-09-05UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510587288.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies rely heavily on the integrity of data and logs when tracing the source of APT attacks, resulting in high computational complexity and low processing efficiency. Effective tracing is particularly difficult in the absence of host logs.

Method used

Through heuristic methods based on traffic detection and attack graphs, anomaly detection and attack tracing are carried out in stages. Attack graphs are constructed using ATT&CK knowledge to identify nodes and security relationships in the APT attack process, reduce dependence on data integrity, and achieve lightweight real-time tracing.

Benefits of technology

It significantly improves the detection depth and tracing capabilities of APT attacks, reduces computational complexity, realizes lightweight real-time tracing, has strong adaptability, can identify abnormal behaviors and unknown attacks in encrypted traffic, and is suitable for a variety of network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602118A_ABST
    Figure CN120602118A_ABST
Patent Text Reader

Abstract

The invention discloses a heuristic APT attack tracing method based on flow detection and an attack graph, and relates to the field of network space security. The method comprises two stages: an exception detection stage and an attack tracing stage, in the exception detection stage, firstly, network flow data features easy to detect are collected from network flow to carry out exception detection on feature information, and a plurality of exception events are generated; in the attack tracing stage, according to ATTamp; and the CK knowledge is combined with an abnormal detection result, and an APT attack process is described through attack graph reasoning and heuristic traceability, so that nodes of all related parties in the APT attack process are identified, node roles are distinguished, a security relationship among the nodes is described, an APT attack chain is restored, and alarm information is generated. According to the method, the APT attack process can be traced without processing massive data packets and logs and a machine learning process, so that the calculation complexity is reduced, the lightweight real-time traceability is realized, the processing efficiency is improved, the calculation is efficient, and the data dependence is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cyberspace security, and in particular to a heuristic APT attack tracing method based on traffic detection and attack graphs. Background Art

[0002] APT attacks are characterized by their long duration, high concealment, multi-point coordination, and multi-hop nature, making them difficult to detect and trace. Traditional methods use host logs and network traffic as input, analyzing abnormal events and their correlations within these logs and traffic. Using methods such as cause-and-effect diagrams, they establish traceability models in the hope of tracing the complete APT attack process and attack chain. Discovering and tracing APTs is crucial for ensuring network security and detecting network failures.

[0003] There are various methods for tracing the source of an APT attack. The process generally involves a network early warning system detecting an attack and requesting tracking. The attack data stream is then tracked and located, and the network device or host that sent the attack data is analyzed to determine. After identifying the attacking host, the system analyzes its input and output information, including its system logs, to determine whether the device is controlled by a third party, leading to the generation of the attack data. Based on this information, the previous control node in the attack control chain is determined, and the process of tracing the source continues in this cycle. According to research, the following common methods are currently used for tracing the source.

[0004] (1) Data flow-based tracing methods, such as a network traffic-based APT attack detection and path reconstruction method disclosed in a Chinese patent document with publication number CN117914571A and publication date January 8, 2024. By collecting a large amount of network traffic data, reverse debugging along the attack data flow path to query its source, tracking data containing path information or collecting marked data, and using algorithms to analyze communication patterns to reconstruct the attack flow path to achieve APT attack tracing.

[0005] (2) Tracing methods based on system logs, such as the APT online detection method based on system logs and deep learning disclosed in the Chinese patent document with publication number CN116760604A and publication date September 15, 2023. Since the logs of system and network equipment provide a large amount of historical activity information, by summarizing various security events, unified analysis, correlating different events, and screening effective data, hidden security threats can be discovered and APT attack paths can be found.

[0006] (3) Tracing methods based on traceability graphs, such as the APT attack tracing method and system based on log graph representation method disclosed in the Chinese patent document with publication number CN119051923A and publication date August 9, 2024. By extracting abnormal events or security alarm events, the causal relationship is captured in the form of a traceability graph to generate a causal relationship graph, find the system's correlation and information flow, and finally restore the attack process.

[0007] The above methods each have their own advantages and can trace the source of APT attacks to a certain extent, but they all have the problems of high time and space complexity to varying degrees, as follows: (1) Data flow-based tracing methods require processing and storing a large number of network data packets, and then analyzing and tracking the sample data. This method has a long time span, a huge amount of data to be tracked, and poor detection of slow-execution attacks. In the face of encrypted communications, the increased difficulty of analysis and tracking also restricts the accuracy of tracing. (2) The tracing method based on system logs requires professional tools and skills to parse and correlate large amounts of log data. Purchasing, deploying, and maintaining a log security information system requires a large investment and constant adjustment of rules and policies to avoid false positives and missed positives. At the same time, it also faces the possibility that attackers may tamper with or delete logs. (3) The traceability method based on the traceability graph requires prior knowledge and a large number of high-quality multidimensional features to construct a training set and conduct offline machine learning. As the input feature dimension and noise data increase, the number of unknown entities increases, and the analysis and generation of causal relationships face problems such as poor accuracy and robustness, which to a certain extent reduces the performance of tracing and tracking.

[0008] The main drawback of these methods is their high reliance on data and log integrity, often requiring complete data chains or log information for tracing back the source. However, capturing all of this data and logs is often difficult, requiring the deployment of data and log collection modules on the host side, which consumes significant storage resources and bandwidth. Network traffic information is relatively inexpensive to obtain; simply applying a traffic collection strategy on the switch side yields a relatively complete picture. However, current research has rarely successfully traced APT attacks based solely on traffic information. This is because traffic information only provides a superficial representation of APT behavior and cannot capture user or host behavior at the same granularity as host logs. Therefore, tracing APT attacks solely through network traffic, in the absence of host logs, is a significant challenge. Summary of the Invention

[0009] In order to overcome the defects and shortcomings of the above-mentioned prior art, the present invention provides a heuristic APT attack tracing method based on traffic detection and attack graphs, which has low dependence on the integrity of data and logs. It can trace the APT attack process only through network traffic without processing massive data packets and logs, and without the need for machine learning. It reduces the computational complexity, realizes lightweight real-time tracing, improves processing efficiency, and is computationally efficient.

[0010] The present invention is achieved through the following technical solutions: A heuristic APT attack tracing method based on traffic detection and attack graphs, including the first and second stages: Phase 1: Anomaly detection. First, network flow feature information is collected from network traffic and anomaly detection is performed on the network flow feature information. Several anomaly events are generated and a list of anomaly events is obtained. The anomaly event includes the time when the anomaly event occurred, the IP address where the anomaly event occurred, the IP address of the other end of the anomaly event, and the type of anomaly event. Secondly, the network flow corresponding to the IP address where the anomaly event occurred is subjected to time sequence anomaly detection to obtain anomaly events in the host network activity space and anomaly events of newly added other end nodes on the host. The second stage: the attack tracing stage. Based on ATT&CK knowledge, combined with the abnormal events obtained in the first stage, abnormal events in the host network activity space, and abnormal events of newly added peer nodes in the host, the attack graph model is initialized and the attack graph is reasoned. Then, the attack graph is heuristically traced to characterize the APT attack process, thereby identifying the nodes in the APT attack process, distinguishing the node roles, and characterizing the security relationship between the nodes, restoring the APT attack chain and generating alarm information.

[0011] The specific steps of initializing the attack graph model and performing attack graph reasoning are as follows: Initialize the attack graph model. The attack graph contains four types of nodes: original attack nodes, auxiliary attack nodes, victim nodes, and zombie nodes, and four types of edges: attack edges, auxiliary attack edges, zombie attack edges, and remote control edges: Based on the abnormal events obtained in the first stage, the original attack node, victim node, zombie node, attack edge, and zombie attack edge are obtained, and duplicate edges are removed; Attack graph reasoning: Based on the host network activity space anomaly events and the host's newly added peer node anomaly events obtained in the first phase, the auxiliary attack nodes in the attack graph are completed. Based on the information of both parties in the host's newly added peer node anomaly events, the auxiliary attack edges and zombie attack edges are obtained.

[0012] The heuristic tracing of the attack graph specifically refers to: Traverse each edge in the current attack graph in turn, trace back the network flow feature information output in the first stage, find the active reverse connection initiated from the attacked to the attacker, and mark the edges from the victim node to other nodes as remote control edges.

[0013] The obtaining of the original attacking node, victim node, zombie node, attacking edge, and zombie attacking edge from the abnormal event obtained in the first stage specifically includes: Extract the attack source from the abnormal events output in the first stage, filter out the attack targets, and form the original attack node set; Extract attack targets from the abnormal events output in the first phase, filter out attack sources, and form a set of victim nodes; Extract dual-identity nodes that hit both the attack target and the attack source from the abnormal events output in the first stage to form a zombie node set; Relationships are extracted from the abnormal events output in the first stage. Each abnormal event corresponds to an edge from the source node to the target node. The edge from the attack node to the victim node or zombie node is marked as an attack edge, and the edge from the zombie node to the victim node or zombie node is marked as a zombie attack edge. Duplicate edges are deduplicated.

[0014] The method of completing the auxiliary attack nodes in the attack graph and obtaining the auxiliary attack edges and zombie attack edges based on the information of both parties in the abnormal event of the host adding a new peer node specifically includes: The time, source, and destination are extracted from each abnormal event output in the first stage. Then, fuzzy search is performed on the abnormal events in the host network activity space and the abnormal events of the host's newly added peer nodes based on the conditions of "source + time" and "destination + time", respectively. Auxiliary attack nodes that are synchronized with the attacker's behavior before and after the abnormal event are identified, and the auxiliary attack nodes are added to the attack graph. Based on the information of both parties in the abnormal events of the newly added peer nodes, the edges from the auxiliary attack nodes to the victim nodes or zombie nodes are marked as auxiliary attack edges, and the edges from the zombie nodes to the victim nodes are marked as zombie attack edges.

[0015] The network flow characteristic information includes at least: time, IP address, port, protocol, number of received messages, number of sent messages, duration and response time.

[0016] The warning information includes time information, attack source information, victim target information, attack chain information and supporting evidence.

[0017] The first stage includes the following steps: Step 11: Record the characteristic information of each network flow from the network traffic; Step 12: Behavior anomaly detection: For each host, perform anomaly detection on all network flow feature information within a specific learning cycle. Each feature in each network flow is divided into input features and output features according to the message inflow or outflow direction. The ratio of input features to output features is defined as the feature input-output ratio. The feature input-output ratio of each feature is obtained to form the feature input-output ratio set ri. The feature input-output ratio of each feature is multiplied to obtain the host behavior volume of the host within the specific learning cycle. Step 13: Set a threshold range for the host behavior volume. After the learning cycle ends, in subsequent detection cycles, the host behavior volume within the detection cycle is calculated in the same way. If it exceeds the threshold range, it is determined to be an abnormal event. Step 14: For the set of IP addresses where the abnormal event occurred, filter each network flow with it as the source address and segment it using a sliding time window of length t. Perform difference operations on the TCP session number series {Xt} and the data transmission volume series {Yt} to stationary series. Determine the optimal order of the ARIMA (p, d, q) model using the AIC or BIC criterion. After fitting the model, identify the set of mutation points {St} based on the 90% confidence interval. Extract the network flows containing the mutation points and mark the corresponding destination IP address set as the victim node. Based on this, generate the network activity space abnormal event. Step 15. For each node s in the set of peer IP addresses of the abnormal event, define the m time periods before the abnormality occurs as the learning period, and the target address set within the learning period as Ds. Within the time window t, count the number of TCP SYN packets sent by s to the new target address d that does not belong to the target address set Ds within the learning period, and calculate the dynamic threshold threshold based on the historical behavior of s within the learning period. If the current number of TCP SYN packets exceeds the threshold, it is determined to be a new abnormal event of the peer node.

[0018] The threshold range for setting the host behavior volume includes: Obtain the host behavior volume of the host in multiple learning cycles, take 3 times the maximum value as the upper bound threshold of the host behavior volume, and take 0.3 times the minimum value as the lower bound threshold of the host behavior volume.

[0019] The dynamic threshold value is obtained by the following formula: threshold=μ+3σ Where μ is the arithmetic mean of the number of TCP SYN packets, and σ is the standard deviation of the number of TCP SYN packets.

[0020] Compared with the prior art, the beneficial technical effects brought about by the present invention are as follows: 1. The present invention, through phased processing, based on ATT&CK knowledge, combines the abnormal events obtained in the first phase, the abnormal events of the host network activity space, and the abnormal events of the host's newly added peer nodes, and through multi-source data fusion, attack graph reasoning and heuristic tracing, significantly improves the detection depth and tracing capabilities of APT attacks. It monitors the traffic characteristics and behavioral characteristics of each host and traces the APT attack process only through network traffic. There is no need to process massive data packets and logs, and no machine learning process is required, which reduces computational complexity, realizes lightweight real-time tracing, improves processing efficiency, and is computationally efficient.

[0021] 2. The present invention mainly relies on network traffic information, which greatly reduces the requirements for data integrity. It can effectively trace the source even in the absence of host logs, reducing data dependence.

[0022] 3. The present invention can effectively identify abnormal behaviors in encrypted traffic through feature analysis and behavior detection, breaking through the technical bottleneck of encrypted data analysis, and can discover and identify unknown attack behaviors, thereby improving the detection rate of unknown anomalies and having a wide detection range.

[0023] 4. The present invention reduces the requirements for the environment and data. Since the protocols and roles of various hosts on the Internet are different, their behavior patterns have large individual differences. The present invention can accurately detect anomalies based on changes in the behavior patterns of individual hosts, with stronger adaptability, wider application scenarios, and strong compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of the technical process of the two stages in the method of the present invention; Figure 2 This is a partial schematic diagram of the multi-dimensional behavioral characteristics of the host in the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions of the present invention in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0026] ATT&CK Knowledge: The "Adversarial Tactics, Techniques, and Common Knowledge" framework. It is a public, community-driven knowledge base covering tactics and techniques used by attackers in cyberattacks.

[0027] Example 1 This embodiment discloses a heuristic APT attack tracing method based on traffic detection and attack graph, which includes the first and second stages: Phase 1: Anomaly detection phase. First, network flow feature information is collected from network traffic, including the following features within a fixed time window T: basic traffic features: including connection time, source IP, destination IP, source port, destination port, and protocol type; traffic statistical features: including packet length, response time, duration, average IP inbound and outbound traffic (bps), average TCP inbound and outbound traffic (bps), and average UDP inbound and outbound traffic (bps). Anomaly detection is then performed on the network flow feature information to generate several anomaly events and obtain an anomaly event list. The anomaly event includes the time when the anomaly event occurred, the IP address where the anomaly event occurred, the IP address of the other end of the anomaly event, and the type of the anomaly event. Secondly, the network flow corresponding to the IP address where the anomaly event occurred is subjected to time series anomaly detection to obtain anomaly events in the host network activity space and anomaly events in which the host adds a new opposite end node. The second stage: the attack tracing stage. Based on ATT&CK knowledge, combined with the abnormal events obtained in the first stage, abnormal events in the host network activity space, and abnormal events of newly added peer nodes in the host, the attack graph model is initialized and the attack graph is reasoned. Then, the attack graph is heuristically traced to characterize the APT attack process, thereby identifying the nodes in the APT attack process, distinguishing the node roles, and characterizing the security relationship between the nodes, restoring the APT attack chain and generating alarm information.

[0028] This embodiment monitors the traffic characteristics and behavioral characteristics of each host, does not require processing massive data packets and logs, does not require machine learning processes, reduces computational complexity, implements lightweight real-time tracing, improves processing efficiency, and is computationally efficient.

[0029] Example 2 This embodiment discloses a heuristic APT attack tracing method based on traffic detection and attack graph, which includes the first and second stages: Phase 1: Anomaly detection. First, network flow feature information is collected from network traffic and anomaly detection is performed on the network flow feature information. Several anomaly events are generated and a list of anomaly events is obtained. The anomaly event includes the time when the anomaly event occurred, the IP address where the anomaly event occurred, the IP address of the other end of the anomaly event, and the type of anomaly event. Secondly, the network flow corresponding to the IP address where the anomaly event occurred is subjected to time sequence anomaly detection to obtain anomaly events in the host network activity space and anomaly events of newly added other end nodes on the host. The second stage: the attack tracing stage. Based on ATT&CK knowledge, combined with the abnormal events obtained in the first stage, abnormal events in the host network activity space, and abnormal events of newly added peer nodes in the host, the attack graph model is initialized and the attack graph is reasoned. Then, the attack graph is heuristically traced to characterize the APT attack process, thereby identifying the nodes in the APT attack process, distinguishing the node roles, and characterizing the security relationship between the nodes, restoring the APT attack chain and generating alarm information.

[0030] The first phase specifically includes the following steps: Step 11: Record characteristic information of each network flow from the network traffic, including at least time, IP address, port, protocol, number of received messages, number of sent messages, duration, and response time; Step 12: Behavior anomaly detection. For each host, perform anomaly detection on all network flow feature information within a specific learning cycle T. Each feature in each network flow is divided into input features and output features according to the message inflow or outflow direction. The ratio of input features to output features is defined as the feature input-output ratio. The feature input-output ratio of each feature is obtained to form a feature input-output ratio set ri. The multiple feature input-output ratios of a host obtained within the learning cycle T are multiplied to obtain the host behavior volume V of the host within the learning cycle T. The calculation method is V = r1×r2×…×rn, where n is the number of features. Step 13: Obtain the host behavior volume for multiple learning cycles. Take 3 times the maximum value as the upper threshold for the host's behavior volume, and 0.3 times the minimum value as the lower threshold for the host's behavior volume. After the learning cycle, calculate the host behavior volume in the same way in subsequent detection cycles. If it exceeds the threshold, it is determined to be an abnormal event. Step 14: For the set of IP addresses where the abnormal event occurred, filter each network flow with it as the source address and segment it using a sliding time window of length T. Perform difference operations on the TCP session number series {Xt} and the data transmission volume series {Yt} to stationary series. Determine the optimal order of the ARIMA (p, d, q) model using the AIC or BIC criterion. After fitting the model, identify the set of mutation points {St} based on the 90% confidence interval. Extract the network flows containing the mutation points and mark the corresponding destination IP address set as the victim node. Based on this, generate the network activity space abnormal event. Step 15: For each node s in the set of peer IP addresses S of the abnormal event, define the m time periods before the abnormality occurs as the learning period. The set of target addresses within the learning period is Ds. Within the time window T, count the number of TCP SYN packets sent by s to the new target address d (d∉Ds). Based on the historical behavior of s within the learning period, calculate the dynamic threshold threshold = μ + 3σ (μ and σ are the arithmetic mean and standard deviation of the number of TCP SYN packets, respectively). If the current number of TCP SYN packets exceeds the threshold, it is determined to be a new abnormal event for the peer node.

[0031] In the second phase, the attack graph model is initialized and the attack graph reasoning is performed as follows: Initialize the attack graph model. The attack graph contains four types of nodes: original attack nodes, auxiliary attack nodes, victim nodes, and zombie nodes, and four types of edges: attack edges, auxiliary attack edges, zombie attack edges, and remote control edges: Extract attack sources from the abnormal events output in the first stage, filter out attack targets, and form a set of attack source nodes; Extract attack targets from the abnormal events output in the first phase, filter out attack sources, and form a set of victim nodes; Extract dual-identity nodes that hit both the attack target and the attack source from the abnormal events output in the first stage to form a zombie node set; Extract relationships from the abnormal events output in the first phase. Each abnormal event corresponds to an edge from the source node to the target node. The edge from the attack node to the victim node or zombie node is marked as an attack edge, and the edge from the zombie node to the victim node or zombie node is marked as a zombie attack edge. Duplicate edges are removed. Attack graph reasoning: Based on the host network activity space anomaly events and the host newly added peer node anomaly events obtained in the first phase, the auxiliary attack nodes in the attack graph are completed. Based on the information of both parties in the host newly added peer node anomaly events, auxiliary attack edges and zombie attack edges are obtained. Specifically, The time, source, and destination are extracted from each abnormal event output in the first phase. Fuzzy searches are then performed on the host network activity space abnormal events and the host's newly added peer node abnormal events, using "source + time" and "destination + time" as conditions, respectively. Auxiliary attack nodes that are synchronized with the attacker's behavior before and after the abnormal event are identified and added to the attack graph. Based on the information of both parties in the newly added peer node abnormal event, the edges from the auxiliary attack node to the victim node or zombie node are marked as auxiliary attack edges, and the edges from the zombie node to the victim node are marked as zombie attack edges. In the second phase, heuristic tracing of the attack graph specifically involves: Traverse each edge in the current attack graph in turn, trace back the network flow feature information output in the first phase, find the active reverse connection initiated from the attacked node to the attacker, and mark the edges from the victim node to other nodes as remote control edges; Restoring the APT attack chain and generating alarm information specifically includes: restoring the APT attack chain based on the results of the above steps and generating alarm information, wherein the alarm information includes time information, attack source information, victim target information, attack chain information and supporting evidence.

[0032] This embodiment monitors the traffic characteristics and behavioral characteristics of each host. There is no need to process massive data packets and logs, and no machine learning process is required. This reduces computational complexity, implements lightweight real-time tracing, improves processing efficiency, and achieves high computational efficiency. It also greatly reduces the requirements for data integrity and allows for effective tracing even in the absence of host logs, reducing data dependency.

[0033] Example 3 This embodiment discloses a heuristic APT attack source tracing method based on traffic detection and attack graphs. First, a traffic collection device is deployed. Using traffic collection technologies such as Sniffer, SNMP, NetFlow, and SFlow, traffic is collected from various network devices, including switches, routers, and host ports. The data is recorded by time, protocol type, TCP / UDP, port number, application layer protocol, and data volume. The collected traffic data is then analyzed to obtain multi-dimensional network flow characteristics and host behavior features.

[0034] Phase 1: By deploying collection devices at key network nodes, network traffic characteristics of network hosts are collected in m consecutive periods T, and the following characteristics are recorded: SRC_IP, DST_IP, SRC_PORT, DST_PORT, PROTOCOL, RESPONSE_TIME, IP_INBPS, IP_OUTBPS, TCP_INBPS, TCP_OUTBPS, TCP_SYN, etc. Taking the host behavior feature "number of IP packets" as an example, first capture the host's traffic data within the time period T1 = 1 minute, analyze it to obtain the "IP packet output" IP_OUTBPS and the "IP packet input" IP_INBPS, then the feature input-output ratio r1 = (IP_OUTBPS+1) / (IP_INBPS+1). Similarly, obtain the feature input-output ratios r1, r2, ..., rn of other behavior features. Using the feature input-output ratio set as a parameter, according to the host's behavior volume calculation formula: V = r1 × r2 × ... × rn, multiply the multiple feature input-output ratios obtained within the time period T1 = 1 minute to obtain the host's behavior volume V1 within the T1 period. By collecting traffic from network hosts for m consecutive periods T, the same operation is performed on the traffic data of each period T, and finally multiple behavior volumes V = {V1, V2, ..., Vm} are calculated. The upper limit threshold of the host's behavior volume is Vmax = 3×max(V1, V2, ..., Vm), and the lower limit threshold of the host's behavior volume is Vmin = 0.3×min(V1, V2, ..., Vm). This method is used to obtain the behavior volume threshold of each host. Within the same time period T = 1min, the behavior volume Vt of each host in each period is calculated, and Vt is compared with Vmax and Vmin. If Vt is greater than the upper limit threshold Vmax or less than the lower limit threshold Vmin, a host behavior abnormality warning is output, and a list of abnormal events is output; Extract all attack sources from abnormal events to obtain the attack source node set Original_Attacker. For each network flow with a source IP address as an element of the set, count the number of TCP sessions for m consecutive periods T, using a time period of 1 minute as a slice, to obtain the sequence {Xt}, and the statistical transmission volume to obtain the sequence {Yt}. Perform a stationarity test on each sequence {Xt} and sequence {Yt}, and calculate the optimal autoregressive term, moving average term, and difference number combination (Xp, Xq, Xd) and (Yp, Yq, Yd) corresponding to the stationary sequence. Then, fit the two sequences using an ARIMA model, and identify mutation points based on a 90% confidence interval. If a mutation point exists in the network flow, mark the destination IP address of the network flow as the victim node, and add the time, source IP address, and destination IP address to the list of abnormal events in the network activity space. All attack targets are extracted from the anomaly event to obtain the victim node set Victim_Host. For each node s in the victim node set Victim_Host, its network flow is sliced ​​and analyzed based on a 1-minute time period T. The m periods before the anomaly occur as a learning period, and the historical target address set Victim_Host_Ds of s is obtained. Within a 3-minute detection window, if the number of TCP SYN packets dt sent by s to a new target d (d∉Victim_Host_Ds) exceeds the dynamic threshold threshold = μ + 3σ (μ and σ are the mean and standard deviation of the number of SYN packets during the learning period), this behavior is marked as an abnormal new peer node event, and the time, source IP address, and destination IP address are added to the abnormal new peer node event list.

[0035] Phase 2: The attack graph model is constructed, and the node types include original attack nodes, auxiliary attack nodes, victim nodes, and zombie nodes. First, the set of all attack source nodes (Original_Attacker) and the set of all victim nodes (Victim_Host) are obtained from the abnormal event list. Nodes with dual identities that hit both the attack target and the attack source are extracted to form the zombie node set (Zombie_Machine). Relationships are extracted from abnormal events. Each abnormal event corresponds to an edge from the source node to the target node. The edge from the attack node to the victim node or zombie node is marked as an attack edge, and the edge from the zombie node to the victim node or zombie node is marked as a zombie attack edge. Duplicate edges are deduplicated. To complete the auxiliary attack nodes, we perform fuzzy searches based on "source + time" and "destination + time" in spatial anomaly events and host-added peer node anomaly events, respectively. We identify auxiliary attack nodes that are synchronized with the attacker's behavior before and after the anomaly event and add them to the graph. Based on the information of both parties in the host-added peer node anomaly event, we mark the edges from the auxiliary attack node to the victim node or zombie node as auxiliary attack edges, and the edges from the zombie node to the victim node as zombie attack edges. Traverse each edge in the current attack graph in turn, trace back the network flow feature information of the abnormal event, find the active reverse connection initiated from the attacked to the attacker, identify persistence and data theft behaviors, and mark the edges from the victim node to other nodes as remote control edges; The generated alarm information contains the following elements: time information, attack source information, victim target information, attack chain information, and supporting evidence.

Claims

1. A heuristic APT attack tracing method based on traffic detection and attack graph, characterized in that: Including the first and second phases: Phase 1: Anomaly detection. First, network flow feature information is collected from network traffic and anomaly detection is performed on the network flow feature information. Several anomaly events are generated and a list of anomaly events is obtained. The anomaly event includes the time when the anomaly event occurred, the IP address where the anomaly event occurred, the IP address of the other end of the anomaly event, and the type of anomaly event. Secondly, the network flow corresponding to the IP address where the anomaly event occurred is subjected to time sequence anomaly detection to obtain anomaly events in the host network activity space and anomaly events of newly added other end nodes on the host. The second stage: the attack tracing stage. Based on ATT&CK knowledge, combined with the abnormal events obtained in the first stage, abnormal events in the host network activity space, and abnormal events of newly added peer nodes in the host, the attack graph model is initialized and the attack graph is reasoned. Then, the attack graph is heuristically traced and the APT attack process is characterized. In this way, the nodes of the APT attack process are identified, the node roles are distinguished, and the security relationship between the nodes is characterized. The APT attack chain is restored and alarm information is generated.

2. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 1 is characterized by: The specific steps of initializing the attack graph model and performing attack graph reasoning are as follows: Initialize the attack graph model. The attack graph contains four types of nodes: original attack nodes, auxiliary attack nodes, victim nodes, and zombie nodes, and four types of edges: attack edges, auxiliary attack edges, zombie attack edges, and remote control edges: Based on the abnormal events obtained in the first stage, the original attack node, victim node, zombie node, attack edge, and zombie attack edge are obtained, and duplicate edges are removed; Attack graph reasoning: Based on the host network activity space anomaly events and the host's newly added peer node anomaly events obtained in the first phase, the auxiliary attack nodes in the attack graph are completed. Based on the information of both parties in the host's newly added peer node anomaly events, the auxiliary attack edges and zombie attack edges are obtained.

3. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 2 is characterized by: The heuristic tracing of the attack graph specifically refers to: Traverse each edge in the current attack graph in turn, trace back the network flow feature information output in the first stage, find the active reverse connection initiated from the attacked to the attacker, and mark the edges from the victim node to other nodes as remote control edges.

4. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 3 is characterized by: The obtaining of the original attacking node, victim node, zombie node, attacking edge, and zombie attacking edge from the abnormal event obtained in the first stage specifically includes: Extract the attack source from the abnormal events output in the first stage, filter out the attack targets, and form the original attack node set; Extract attack targets from the abnormal events output in the first phase, filter out attack sources, and form a set of victim nodes; Extract dual-identity nodes that hit both the attack target and the attack source from the abnormal events output in the first stage to form a zombie node set; Relationships are extracted from the abnormal events output in the first stage. Each abnormal event corresponds to an edge from the source node to the target node. The edge from the attack node to the victim node or zombie node is marked as an attack edge, and the edge from the zombie node to the victim node or zombie node is marked as a zombie attack edge. Duplicate edges are deduplicated.

5. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 4 is characterized in that: The method of completing the auxiliary attack nodes in the attack graph and obtaining the auxiliary attack edges and zombie attack edges based on the information of both parties in the abnormal event of the host adding a new peer node specifically includes: The time, source, and destination are extracted from each abnormal event output in the first phase. Fuzzy searches are then performed on the host network activity space abnormal events and the host's newly added peer node abnormal events using "source + time" and "destination + time" as conditions, respectively. Auxiliary attack nodes that are synchronized with the attacker's behavior before and after the abnormal event are identified and added to the attack graph. Based on the information of both parties in the newly added peer node abnormal event, the edges from the auxiliary attack node to the victim node or zombie node are marked as auxiliary attack edges, and the edges from the zombie node to the victim node are marked as zombie attack edges.

6. A heuristic APT attack tracing method based on traffic detection and attack graph according to any one of claims 1 to 5, characterized in that: The network flow characteristic information includes at least: time, IP address, port, protocol, number of received messages, number of sent messages, duration and response time.

7. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 6 is characterized in that: The warning information includes time information, attack source information, victim target information, attack chain information and supporting evidence.

8. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 7 is characterized in that: The first stage includes the following steps: Step 11: Record the characteristic information of each network flow from the network traffic; Step 12: Behavior anomaly detection: For each host, perform anomaly detection on all network flow feature information within a specific learning cycle. Each feature in each network flow is divided into input features and output features according to the message inflow or outflow direction. The ratio of input features to output features is defined as the feature input-output ratio. The feature input-output ratio of each feature is obtained to form the feature input-output ratio set ri. The feature input-output ratio of each feature is multiplied to obtain the host behavior volume of the host within the specific learning cycle. Step 13: Set a threshold range for the host behavior volume. After the learning cycle ends, in subsequent detection cycles, the host behavior volume within the detection cycle is calculated in the same way. If it exceeds the threshold range, it is determined to be an abnormal event. Step 14: For the set of IP addresses where the abnormal event occurred, filter each network flow with it as the source address and segment it using a sliding time window of length t. Perform difference operations on the TCP session number series {Xt} and the data transmission volume series {Yt} to stationary series. Determine the optimal order of the ARIMA (p, d, q) model using the AIC or BIC criterion. After fitting the model, identify the set of mutation points {St} based on the 90% confidence interval. Extract the network flows containing the mutation points and mark the corresponding destination IP address set as the victim node. Based on this, generate the network activity space abnormal event. Step 15. For each node s in the set of peer IP addresses of the abnormal event, define the m time periods before the abnormality occurs as the learning period, and the target address set within the learning period as Ds. Within the time window t, count the number of TCP SYN packets sent by s to the new target address d that does not belong to the target address set Ds within the learning period, and calculate the dynamic threshold threshold based on the historical behavior of s within the learning period. If the current number of TCP SYN packets exceeds the threshold, it is determined to be a new abnormal event of the peer node.

9. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 8 is characterized by: The threshold range for setting the host behavior volume includes: Obtain the host behavior volume of the host in multiple learning cycles, take 3 times the maximum value as the upper bound threshold of the host behavior volume, and take 0.3 times the minimum value as the lower bound threshold of the host behavior volume.

10. The heuristic APT attack tracing method based on traffic detection and attack graph according to claim 9, characterized in that: The dynamic threshold value is obtained by the following formula: threshold=μ+3σ Where μ is the arithmetic mean of the number of TCP SYN packets, and σ is the standard deviation of the number of TCP SYN packets.

Citation Information

Patent Citations

  • APT online detection method based on system log and deep learning

    CN116760604A

  • APT attack detection and path reconstruction method based on network traffic

    CN117914571A

  • APT attack tracing method and system based on log graph representation method

    CN119051923A