Automatic fine-grained log labeling method

By collecting and analyzing a variety of log information in the attack scenario, building a traceability map and identifying anchor points to form an attack sub-map, the problems of rough granularity and high dependence of the existing log annotation methods are solved, and fine-grained log annotation with high accuracy are achieved.

CN119995938AActive Publication Date: 2025-05-13TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202411995678.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing automatic logging annotation methods have problems with rough granularity, insufficient coverage of different log types, and excessive dependence on manual work and domain expertise.

Method used

By running preset attack scenarios, we collect log information including unmarked logs, alignment information logs, separator logs, attack logos and attack logs, build an initial traceability map, and use this information to identify and connect anchor points to form a refined attack sub-map to achieve labeling.

Benefits of technology

It realizes fine-grained annotation of logs, reduces manual workload, reduces dependence on domain expertise, ensures the accuracy and completeness of the annotation, and can label traffic, audit and application logs at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995938A_ABST
    Figure CN119995938A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic fine-grained log labeling method which comprises the following steps: running a preset attack scene, and collecting log information including unlabeled logs, alignment information logs, separator logs, attack marks and attack logs based on a running result; constructing an initial traceability graph according to an audit log in the alignment information log, associating the application log and the traffic log with the audit log by using the alignment information log, and identifying an anchor point in the initial traceability graph by using an attack mark and an attack log based on an association result; and dividing nodes meeting a preset segmentation condition in the initial traceability graph into a plurality of execution units by using a separator log to obtain a refined traceability graph, connecting all anchor points based on the refined traceability graph to form an attack sub-graph, and obtaining a labeled log based on the attack sub-graph. Therefore, manual annotation workload is reduced, traffic, audit and application logs can be annotated at the same time, accurate annotation of individual log entries is achieved, and annotation accuracy and completeness are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to an automated fine-grained log annotation method. Background Art

[0002] In recent years, Advanced Persistent Threat (APT) has become a major concern in the field of network security. These persistent, targeted attacks are usually planned and carried out by well-funded and well-organized groups, and are known for their long-term and covert characteristics. Attackers usually lurk in the victim's network for a long time, using intermediate hosts to expand the scope of the attack or steal sensitive data.

[0003] Log analysis techniques are critical in detecting, blocking, and tracking APT attacks. For example, network-based intrusion detection systems (NIDS) detect anomalies at the network layer by analyzing network traffic logs. System-level audit logs are used to create traceability graphs for post-attack anomaly detection and forensic analysis. Application logs, such as web server logs, provide high-level insights into attack methods. These techniques all rely on large log datasets that are critical for training models, evaluating performance, and benchmarking. The effectiveness of these processes is highly dependent on the quality of log annotations.

[0004] Existing automatic log annotation methods usually fail to meet these requirements:

[0005] Time window-based method: This method uniformly marks all log entries within a specific time window as attack behaviors. Since APT attacks are often disguised as normal activities, the annotation granularity of this method is too coarse, resulting in insufficient accuracy of the results.

[0006] The annotation method based on the behavioral model BP (Behavioral Profiles): automatically generates normal traffic and attack traffic by simulating network behavior characteristics, so annotation can be performed directly when the traffic is generated. However, this method cannot annotate real traffic and cannot be applied to audit logs and application logs.

[0007] Detection-based methods: Use network security tools, such as sniffers and honeypot systems, to identify and annotate attack behaviors. However, the accuracy and completeness of annotations in this method are highly dependent on the performance of the detection tools, and it is impossible to ensure that all attack-related logs are accurately identified and recorded.

[0008] Rule-matching-based method: Traditional methods annotate logs by matching rules such as IP addresses line by line. The Kyoushi method optimizes this by combining the context and sequence of events to improve the matching accuracy. However, the Kyoushi method still requires a lot of manual input: for each trace left by the APT attack in each data source, a matching rule needs to be written separately, which makes the number of rules very large and highly dependent on domain knowledge; at the same time, it is difficult to prove the accuracy and completeness of the annotation rules. Summary of the invention

[0009] The present application provides an automated fine-grained log annotation method to address the problems of current automatic log annotation methods, such as coarse granularity, insufficient coverage of different log types, and excessive reliance on manual work and domain expertise.

[0010] The first aspect of the present application provides an automated fine-grained log annotation method, comprising the following steps: running a preset attack scenario, and collecting log information including unlabeled logs, alignment information logs, separator logs, attack flags, and attack logs based on the running results; constructing an initial traceability graph based on the audit log in the alignment information log, associating the application log and the traffic log with the audit log using the alignment information log, and identifying anchor points in the initial traceability graph based on the association results using the attack flag and the attack log; using the separator log to divide the nodes in the initial traceability graph that meet the preset segmentation conditions into multiple execution units to obtain a refined traceability graph, connecting all anchor points based on the refined traceability graph to form an attack subgraph, so as to obtain annotated logs based on the attack subgraph.

[0011] Optionally, the use of the alignment information log to associate the application log and the traffic log with the audit log includes: disabling the application's buffer, injecting the timestamp and thread ID in the alignment information log into the application log to obtain a modified application and a modified application log; executing the modified application to regenerate the alignment information log, and using the timestamp and thread ID of the regenerated alignment information log to associate the modified application log with the audit log.

[0012] Optionally, the use of the alignment information log to associate the application log and the traffic log with the audit log also includes: obtaining each TCP connection, UDP session timestamp and four-tuple information in the traffic log, and obtaining the four-tuple information in the audit log; based on the UDP session timestamp and four-tuple information in the traffic log, searching the audit log for the network system call that is closest to the UDP session timestamp and the four-tuple information in the traffic log; within the life cycle of the traffic file descriptor corresponding to the closest network system call, batch matching all system calls with the same four-tuple to establish an association between the traffic log and the audit log.

[0013] Optionally, the use of the attack flag and the attack log to identify the anchor point in the initial traceability graph includes: using a preset attack flag scanner to detect all IP addresses in the attack scenario and addresses corresponding to the scenario reversal; identifying connections in which the reserved bit in the IP address header of the traffic log is set to 1, and marking the connections in which the reserved bit in the IP address header is set to 1 as anchor points in the traffic log.

[0014] Optionally, the use of the attack flag and the attack log to identify the anchor point in the initial traceability graph also includes: integrating a preset kernel module to hook into the system call and injecting the attack flag into the attack payload; using a preset attack flag scanner to intercept the system call, restore the attack payload to a state before the attack, and execute the restored attack payload; identifying the system call parameters with the attack flag in the audit log, and marking the corresponding log entries in the audit log as anchor points based on the system call parameters with the attack flag.

[0015] Optionally, when an anchor point in the initial tracing graph is not identified, the method includes: recording a behavior log of the attacker, wherein the behavior log includes a timestamp, attack parameters and preset matching rules; using the timestamp to refine the scope of the preset matching rules, and using the attack parameters and the preset matching rules, taking into account event combinations, contexts and attack parameters, to detect attack log signatures; and identifying the anchor point in the initial tracing graph based on the attack log signatures.

[0016] Optionally, when connecting all anchor points based on the refined traceability graph to form an attack subgraph, it includes: constructing a handler derivation relationship tree; determining handlers with shared derivation relationships through the handler derivation relationship tree, and classifying system calls in the handlers with shared derivation relationships as belonging to the same execution unit.

[0017] Optionally, when connecting all anchor points based on the refined traceability graph to form an attack subgraph, it includes: constructing an initial handler tree; evaluating the number of accept calls in each node subtree of the initial handler tree, and deleting nodes whose accept call times are greater than or equal to a preset number, to obtain a target handler tree, so as to define different execution units according to the target handler tree.

[0018] Optionally, the method of connecting all anchor points based on the refined provenance graph to form an attack subgraph to obtain labeled logs based on the attack subgraph includes: based on the refined provenance graph, processing in reverse topological order from nodes without outgoing edges; maintaining a set for each node and updating the set for each new node, obtaining all marked path edges by taking the set union of all direct successor nodes and subtracting all intersections, wherein, if the node is a starting point, all paths starting from the starting point are marked, and the set of each node is adjusted to obtain all marked path edges; connecting all marked path edges to form the attack subgraph.

[0019] Optionally, after forming the attack subgraph, the method includes: checking whether all anchor points in the refined traceability graph have been connected to form an attack path; if there are unconnected anchor points, determining the positions of the unconnected anchor points, and prompting the user to add anchor points.

[0020] In the above implementation, a preset attack scenario is run, and log information including unlabeled logs, alignment information logs, separator logs, attack flags and attack logs is collected based on the running results. An initial traceability graph is constructed based on the audit log in the alignment information log, and the application log and traffic log are associated with the audit log using the alignment information log. Based on the association results, the attack flags and attack logs are used to identify anchor points in the initial traceability graph, and the separator log is used to divide the nodes in the initial traceability graph that meet the preset segmentation conditions into multiple execution units to obtain a refined traceability graph. All anchor points are connected based on the refined traceability graph to form an attack subgraph, so as to obtain labeled logs based on the attack subgraph. As a result, the problems of current automatic log annotation methods, such as coarse granularity, insufficient coverage of different log types, and excessive reliance on manual work and domain expertise, are solved. Manual workload is reduced, human involvement is minimized, and reliance on domain expertise is reduced. Multiple data sources are covered, and traffic, audit, and application logs can be annotated simultaneously. Fine-grained labels are provided: precise annotation of individual log entries is achieved, specific connections, audit events, and application log entries are identified, accuracy and completeness of annotations are ensured, entries not related to attacks are avoided, and missing attack logs are minimized.

[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0023] Figure 1 A flowchart of an automated fine-grained log annotation method provided according to an embodiment of the present application;

[0024] Figure 2 This is a flowchart of an automated fine-grained log annotation method according to an embodiment of the present application;

[0025] Figure 3 A code schematic diagram of log matching according to Algorithm 1 of one embodiment of the present application;

[0026] Figure 4 A schematic diagram of matching traffic logs and audit logs according to an embodiment of the present application;

[0027] Figure 5 A schematic diagram of a system call included in a matching traffic log according to an embodiment of the present application;

[0028] Figure 6 A code schematic diagram of attack flag scanning according to Algorithm 2 of one embodiment of the present application;

[0029] Figure 7 A schematic diagram of an example of an execution model according to an embodiment of the present application;

[0030] Figure 8 A schematic diagram of a service application used in an experiment according to an embodiment of the present application;

[0031] Fig. 9 A schematic diagram of APP-audit log matching accuracy according to an embodiment of the present application;

[0032] Fig.10 A schematic diagram of APP audit log matching accuracy according to an embodiment of the present application;

[0033] Fig.11 A schematic diagram of a traffic-audit log matching success rate according to an embodiment of the present application;

[0034] Fig.12 A schematic diagram of traffic log anchor point verification according to an embodiment of the present application;

[0035] Fig.13A schematic diagram of audit log anchor point identification according to an embodiment of the present application;

[0036] Fig.14 A schematic diagram of a server application audit log unit partitioning experiment according to an embodiment of the present application;

[0037] Fig.15 A schematic diagram of static analysis time according to an embodiment of the present application;

[0038] Fig.16 is a schematic diagram of runtime load and occupied load according to an embodiment of the present application;

[0039] Fig.17 A schematic diagram of audit log space overhead according to an embodiment of the present application;

[0040] Fig.18 A schematic diagram of an attack tracing diagram according to an embodiment of the present application;

[0041] Fig.19 A schematic diagram of an insertion point in Python according to an embodiment of the present application;

[0042] Fig. 20 A schematic diagram of an insertion point in NODE.JS according to an embodiment of the present application;

[0043] Fig.21 A schematic diagram of an instrumentation point in NGINX according to an embodiment of the present application;

[0044] Fig. 22 A schematic diagram of an insertion point in REDIS according to an embodiment of the present application;

[0045] Fig.23 A schematic diagram of attack marker injection in a traffic log according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0047] The following describes the automated fine-grained log annotation method of the embodiment of the present application with reference to the accompanying drawings. In view of the problems that the current automated log annotation method mentioned in the above background technology has coarse granularity, insufficient coverage of different log types, and excessive reliance on manual work and domain expertise, the present application provides an automated fine-grained log annotation method, in which a preset attack scenario is run, and log information including unlabeled logs, alignment information logs, separator logs, attack flags, and attack logs is collected based on the running results, an initial traceability graph is constructed based on the audit log in the alignment information log, the application log and the traffic log are associated with the audit log using the alignment information log, and the anchor points in the initial traceability graph are identified based on the association results using the attack flag and the attack log, the nodes in the initial traceability graph that meet the preset segmentation conditions are divided into multiple execution units using the separator log to obtain a refined traceability graph, all anchor points are connected based on the refined traceability graph to form an attack subgraph, and the annotated log is obtained based on the attack subgraph. As a result, the problems of current automatic log annotation methods, such as coarse granularity, insufficient coverage of different log types, and excessive reliance on manual work and domain expertise, are solved. Manual workload is reduced, human involvement is minimized, and reliance on domain expertise is reduced. Multiple data sources are covered, and traffic, audit, and application logs can be annotated simultaneously. Fine-grained labels are provided: precise annotation of individual log entries is achieved, specific connections, audit events, and application log entries are identified, accuracy and completeness of annotations are ensured, entries not related to attacks are avoided, and missing attack logs are minimized.

[0048] Given that APT attacks can last for months and generate a large amount of log data from multiple sources, manual log annotation becomes very time-consuming and labor-intensive. Therefore, there is an urgent need for an automatic log annotation method to generate a benchmark dataset that meets the following requirements:

[0049] (1) Reduce manual workload: Minimize manual involvement and reduce reliance on domain expertise;

[0050] (2) Covering multiple data sources: Ability to annotate traffic, audit, and application logs simultaneously;

[0051] (3) Provide fine-grained labeling: Enable precise labeling of individual log entries, identifying specific connections, audit events, and application log entries;

[0052] (4) Ensure the accuracy and completeness of annotations: avoid including entries that are not related to the attack and minimize missing attack logs.

[0053] Current status of traffic log annotation: Multiple researchers reviewed 28 traffic datasets using various annotation methods: time window (TW), behavior profile (BP), and detection tool (DT). Some datasets use a combination of these techniques. The TW method is used in 12 datasets, BP is used in 12 datasets, and DT is used in 19 datasets, which is the most common method.

[0054] Current status of audit log annotation: Currently, the public audit log benchmark datasets used for training and performance evaluation include the DARPATC series, CERT series, LANL series, StreamSpot series, Unicorn series, Alchemist series, and Atlas series.

[0055] The DARPA TC series dataset only provides coarse-grained attack scenarios and does not provide precise log annotations at the audit log entry level.

[0056] CERT and LANL datasets define malicious activities based on user identity, which is a limitation in APT scenarios because it is usually impossible to distinguish between legitimate and malicious users under APT attacks.

[0057] The StreamSpot and Unicorn datasets have log annotations at the scene level, but do not provide precise log annotations at the audit log entry level.

[0058] Datasets like Atlas and Alchemist provide manual log annotations by authors based on domain knowledge, but this manual approach requires a lot of manpower and lacks scalability.

[0059] Many studies use private audit log datasets, which are not publicly available and only provide narrative attack descriptions. These internal datasets rely on manual annotation by experts, which is expensive and difficult to replicate.

[0060] The current status of multi-source log annotation: parsing logs into semi-structured formats and matching them using rule-based methods to annotate audit, traffic, and application logs simultaneously. However, it requires many domain-specific rules to cover all attack-related logs, making it time-consuming and difficult to ensure the accuracy and completeness of the rules.

[0061] The idea of ​​the technical solution of the embodiment of the present application is:

[0062] A provenance graph is a directed graph used to represent the relationship between topics (such as processes, threads, etc.) and objects (such as files, network sockets, etc.) in a system. The direction of the edge represents the flow of data. The provenance graph is mainly constructed from audit logs, where each edge in the graph may correspond to multiple audit log entries. At present, attack investigation methods based on provenance graphs have been widely proposed. These methods generate provenance graphs based on audit logs, and then start from symptom events (i.e., attack events) to analyze and extract subgraphs related to the attack in the provenance graph.

[0063] Inspired by this, a new method is proposed to solve the problem of automatic log annotation, which simplifies it into the problem of automatically obtaining an accurate and complete attack subgraph in the traceability graph. Once the attack subgraph is identified, the relevant logs can be annotated according to the correspondence between the edges and log entries in the traceability graph, but this method faces the following challenges:

[0064] Challenge 1: In addition to audit logs, application logs and traffic logs need to be annotated. Therefore, it is a challenge to establish the accurate relationship between these logs and build a unified traceability graph to integrate multi-source logs.

[0065] Challenge 2: A prerequisite for extracting attack subgraphs is to have an accurate set of symptom events. A key challenge is how to automatically locate the edges (called anchor points) corresponding to these symptom events in the unified provenance graph.

[0066] Challenge 3: In the process of building a connected attack subgraph starting from the anchor point, we will inevitably encounter the classic dependency explosion problem, which will cause edges irrelevant to the attack to be included in the attack subgraph, thereby reducing the accuracy of the attack subgraph. It is a challenge to automatically, systematically, and universally solve the dependency explosion problem of log annotation tasks.

[0067] The dependency explosion problem is a classic problem in attack investigation in traceability graphs. A long-running process usually performs many input and output operations, resulting in a large number of inbound and outbound edges connected to the corresponding nodes, greatly complicating the attack investigation process.

[0068] BEEP first proposed dividing long-running applications into execution units to obtain more precise attack attribution. MPI, Omegalog, and TeSec further advanced this process, dealing with asynchronous models such as multi-process, multi-threading, thread pools, and task queues.

[0069] MPI can track thread pools by annotating data structures to track dependencies. However, its implementation complexity and the deep code knowledge required may limit its practical application, especially in applications using modern asynchronous models.

[0070] Omegalog can automatically handle multi-process and multi-thread scenarios, but it has difficulty dealing with asynchronous task queues.

[0071] TeSec proposes a solution that can accommodate various execution models, including asynchronous task queues and thread pools. Although it provides improvements, TeSec does not fully address the complexity introduced when these models are combined, such as using both asynchronous task queues and thread pools in applications such as Node.js.

[0072] In automatic annotation, the accuracy of unit partitioning is crucial as it determines whether the log is relevant to an attack. Currently, there is no practical general technique to automatically partition units across various modern asynchronous execution models and their combinations.

[0073] This application focuses on automatic log annotation in the context of advanced persistent threats (APTs), where attackers typically launch attacks through applications on one server and then expand their activities to multiple servers. We assume that the target is a Linux server with a built-in audit logging mechanism (such as sysdig) that is able to capture the start and end time of syscalls, as well as the caller PID / TID information. The system also allows the insertion of custom kernel modules to intercept and modify network and syscalls, without involving a client application with a graphical interface.

[0074] Furthermore, it is assumed that all applications involved in the scenarios are open source server applications, which allows source code analysis and instrument modifications and ensures the reproducibility of the dataset.

[0075] On top of this, we assume that AUTOLABEL users are able to collect network traffic logs and application logs of relevant applications. All audit logs, application logs, network traffic logs, and logs generated by AUTOLABEL must be protected from tampering by attackers. This assumption is widely required by log analysis methods. In addition, the integrity of the control flow must be maintained to ensure the validity of the audit logs.

[0076] In this article, the present application embodiment proposes AUTOLABEL, a system designed to automatically perform fine-grained log annotation. The workflow of AUTOLABEL can be divided into three stages:

[0077] 1) Preparation phase: AUTOLABEL modified the experimental environment and applications related to the attack scenario to generate necessary auxiliary information during the attack execution phase. In addition, AUTOLABEL adjusted the attack plan and tools used by the attacker to inject attack signs to identify the attack.

[0078] 2) Attack execution phase: When the attacker executes the attack, AUTOLABEL collects various logs, including auxiliary information. These logs include unlabeled logs, alignment information logs (used to establish the relationship between different logs), separator logs (used for unit partitioning), and attack logs and attack flags (used to identify anchor points in the traceability graph).

[0079] 3) Annotation phase: This is the core phase of AUTOLABEL and includes five steps:

[0080] (1) Constructing a traceability graph: AUTOLABEL uses audit logs to construct a traceability graph.

[0081] (2) Log correlation: AUTOLABEL uses alignment information logs to correlate application logs and traffic logs with audit logs to solve Challenge 1.

[0082] (3) Locating anchor points: AUTOLABEL uses attack flags and attack logs to identify anchor points in the traceability graph and solves Challenge 2.

[0083] (4) Unit segmentation: Using the separator log, AUTOLABEL splits the nodes in the traceability graph that may cause the dependency explosion problem into multiple units, thus obtaining a refined traceability graph. This partially solves Challenge 3.

[0084] (5) Extracting attack subgraph: AUTOLABEL extracts a complete and precise attack subgraph starting from the anchor point, ensuring that all attack-related logs (i.e., traffic, audit, and application logs) are accurately and comprehensively labeled, thus solving Challenge 3.

[0085] The embodiment of the present application is also evaluated through experiments to evaluate the performance of AUTOLABEL in various attack scenarios, covering different types of server applications, attack tools and vulnerabilities. The experimental results show that for 6 server applications, as long as the audit log is complete, the accuracy of associating application / traffic logs to audit logs is always 100%. And 8 common attack tools and 4 typical vulnerabilities were used for testing. In all cases, AUTOLABEL successfully identified anchors in audit and traffic logs. For the dependency explosion problem, AUTOLABEL has been verified in 7 server applications. For a complete attack story involving multi-source logs, AUTOLABEL has achieved accurate and comprehensive log tagging. In terms of performance overhead, the additional load introduced by AUTOLABEL is acceptable, with a maximum overhead of 8.5% (runtime overhead) and 53.7% (space overhead).

[0086] As a scenario example of an embodiment of the present application:

[0087] Alice is a security expert whose goal is to create a log dataset containing APT attacks to train her security model based on log analysis. She built a network with multiple servers and applications, made the entry point public, and invited security expert Bob to conduct a seven-day penetration test. At the same time, she organized multiple users to use the application normally.

[0088] During the attack, Bob recorded his attack methods. Alice collected traffic logs between hosts, system audit logs for each host, and application logs. The following is an example of an attack step recorded by Bob: 1) A malicious request is sent. 2) The Nginx proxy processes and forwards it to the backend server. 3) The Flask application on the backend processes and forks a child process to execute the malicious command.

[0089] Labeling using existing methods. After the attack, Alice collected millions of traffic logs, billions of bytes of audit logs, and billions of bytes of application logs. In order to train and evaluate security models, Alice needs to accurately label these logs at the log entry level. Given the huge number of logs, manual labeling is obviously impractical.

[0090] Alice attempts to use a rule-matching approach for log annotation, writing rules across different data sources for each attack step. Relevant logs include network logs, Nginx logs, application logs, and audit logs of Flask and fork processes. However, she faces two difficulties: (1) The large amount of log traces makes it difficult to ensure the integrity of log annotation. For a simple attack step, Alice needs to write matching rules for five different types of logs. Even after writing these rules, it is difficult to determine whether all relevant logs are covered. (2) It is difficult to write robust matching rules. API logs only record the requested URL without parameters, providing limited information; audit logs lack advanced readable information about network connections; at the same time, advanced attack techniques make it particularly complex to write rules for traffic logs.

[0091] To use AutoLabel to annotate a dataset, the following preparations are required:

[0092] For Alice: First, she uses the AUTOLABEL tool to automatically detect Nginx and Flask applications. Then, she installs the AUTOLABEL kernel module in the experimental environment. Finally, she modifies the startup commands of Nginx and Flask applications to load the dynamic libraries provided by AUTOLABEL.

[0093] For Bob: He slightly adjusted his attack strategy based on the tools and guidance provided by AUTOLABEL, ensuring that the malicious request and forked child process were injected with the attack flag.

[0094] After the attack, Alice does not need to write any rules. AUTOLABEL will automatically create a traceability graph from the collected logs. By finding the injected attack flags in the traceability graph to locate the anchor point, AUTOLABEL outputs an attack subgraph and automatically annotates all related log entries. This process not only simplifies the annotation work, but also ensures the integrity and accuracy of the annotated logs.

[0095] Specifically, Figure 1 A flowchart of an automated fine-grained log annotation method provided in an embodiment of the present application.

[0096] like Figure 1 As shown, the automatic fine-grained log annotation method includes the following steps:

[0097] In step S101, a preset attack scenario is run, and log information including unlabeled logs, alignment information logs, separator logs, attack flags, and attack logs is collected based on the running results.

[0098] Specifically, during the preparation phase, AUTOLABEL performs two types of modifications:

[0099] (1) AUTOLABEL modifies the environment and applications in the attack scenario. These modifications enable the environment to generate auxiliary information required for automatic annotation during the attack execution phase.

[0100] (2) AUTOLABEL makes slight modifications to the attacker’s attack plan and the attack tools used, allowing attack flags to be injected into the attack process and leaving attack traces in the log.

[0101] The auxiliary information generated from these two types of modifications will be used in different steps of the annotation phase. Specifically, the modifications include the following components:

[0102] Modify the application in the attack scenario:

[0103] a) Alignment Information Logger: Dynamic link hooks and kernel module registration are used to weave the functionality of the alignment information logger into the target applications, which enables them to generate alignment information logs to correlate application logs and audit logs while recording application logs. This information will be used in the log correlation step of the annotation phase.

[0104] b) Delimiter logger: Automatic analysis of the asynchronous execution model of the target application. Using instrumentation technology, AUTOLABEL weaves the functionality of the delimiter logger into the target applications, enabling them to automatically generate delimiter logs, including information about unit partitions, unit derivation relationships, and context switches. This information will be used in the unit partitioning step of the annotation phase.

[0105] Modify the attack plan and attack tools: Slightly modify the attack process to ensure that the traffic packets and system calls generated during the attack contain attack signatures that can be identified by the attack signature scanner. These adjustments are lightweight and provide methods to modify the attack tools for common attack patterns. These methods are transparent to the attacker.

[0106] Modify the attacked environment: Due to the adjustment of the attack plan and tools, the attack flags are injected into the traffic and system calls, resulting in the failure of the attack. Therefore, it is necessary to restore the injected traffic and system calls. To this end, the embodiment of the present application implements an attack flag scanner as a kernel module, whose main goals are: 1) During the attack execution phase, scan and extract attack flags from traffic and audit logs, providing information that can accurately locate the annotation phase of key events; 2) In order to prevent the attack from failing or leaving irreversible traces in the log, the attack flags are also deleted from the attack payload, restoring the payload to the state before modification.

[0107] In step S102, an initial traceability graph is constructed based on the audit log in the alignment information log, the application log and the traffic log are associated with the audit log using the alignment information log, and the anchor points in the initial traceability graph are identified using attack flags and attack logs based on the association results.

[0108] Optionally, in some embodiments, the application log and the traffic log are associated with the audit log using the alignment information log, including: disabling the application's buffer, injecting the timestamp and thread ID in the alignment information log into the application log to obtain a modified application and a modified application log; executing the modified application to regenerate the alignment information log, and associating the modified application log with the audit log using the timestamp and thread ID of the regenerated alignment information log.

[0109] Optionally, in some embodiments, the application log and the traffic log are associated with the audit log using the alignment information log, and also include: obtaining each TCP connection, UDP session timestamp and four-tuple information in the traffic log, and obtaining the four-tuple information in the audit log; based on the UDP session timestamp and four-tuple information in the traffic log, searching the audit log for the network system call that is closest to the UDP session timestamp and four-tuple information in the traffic log; within the life cycle of the traffic file descriptor corresponding to the closest network system call, batch matching all system calls with the same four-tuple to establish an association between the traffic log and the audit log.

[0110] Optionally, in some embodiments, attack flags and attack logs are used to identify anchor points in the initial traceability graph, including: using a preset attack flag scanner to detect all IP addresses in the attack scenario and addresses corresponding to the scenario reversal; identifying connections in which the reserved bit in the IP address header of the traffic log is set to 1, and marking the connections in which the reserved bit in the IP address header is set to 1 as anchor points in the traffic log.

[0111] Optionally, in some embodiments, using attack flags and attack logs to identify anchor points in the initial traceability graph also includes: integrating a preset kernel module to hook into the system call and injecting the attack flag into the attack payload; using a preset attack flag scanner to intercept the system call, restore the attack payload to a state before the attack, and execute the restored attack payload; identifying system call parameters with attack flags in the audit log, and marking corresponding log entries in the audit log as anchor points based on the system call parameters with the attack flags.

[0112] Optionally, in some embodiments, when an anchor point in the initial traceability graph is not identified, the method includes: recording a behavior log of the attacker, wherein the behavior log includes a timestamp, attack parameters, and preset matching rules; using the timestamp to refine the scope of the preset matching rules, and using the attack parameters and the preset matching rules, taking into account event combinations, contexts, and attack parameters to detect attack log signatures; and identifying the anchor point in the initial traceability graph based on the attack log signature.

[0113] During the attack execution phase, the attack scenario begins to run, the attacker launches the attack, and AUTOLABEL collects various logs. The generated information includes:

[0114] Unlabeled logs are logs that need to be labeled;

[0115] The alignment information log is used to correlate application logs and audit logs, and is used in the log correlation step of the annotation phase;

[0116] The separator log is used for unit partitioning and is used in the unit partitioning step of the annotation phase;

[0117] The attack flag is a key attack indicator and is used in the anchor point location step of the annotation phase;

[0118] Attack logs are records of the attack process and are used in the anchor point location step of the annotation phase. These logs include the timestamp range of each attack step and some log records that can be used for rule matching.

[0119] In the annotation phase, there are five steps:

[0120] (1) Constructing a traceability graph: AUTOLABEL uses the collected audit logs to construct a traceability graph, where each edge corresponds to multiple audit log entries.

[0121] (2) Log association (addressing challenge 1): AUTOLABEL maps traffic logs and application logs to audit logs using aligned information logs. Therefore, the edges in the provenance graph are attached not only to audit logs but also to traffic logs and application logs.

[0122] (3) Anchor point location (solving challenge 2): By analyzing attack flags and optional annotation rules in attack logs, AUTOLABEL identifies attack-related edges in the provenance graph, called anchor points.

[0123] (4) Unit partitioning (solving challenge 3): AUTOLABEL uses delimiter logs to perform unit partitioning on applications that may cause dependency explosion problems, splitting nodes related to dependency explosion into multiple unit nodes.

[0124] (5) Attack subgraph extraction (solving challenge 3). Finally, AUTOLABEL connects the anchor points to form a complete attack subgraph.

[0125] To achieve annotation across three different log sources, we need to associate application logs and traffic logs with audit logs. The following describes how to associate application logs with traffic logs.

[0126] The specific method of associating application logs with audit logs is:

[0127] The main difficulty in matching application logs with audit logs is the lack of matching information. In addition, it is impossible to ensure a one-to-one correspondence between application log entries and audit log entries. To solve this problem, AUTOLABEL performs the following three phases of tasks:

[0128] 1) Preparation: (1) To ensure that each API call for writing to the application log (generating a log entry) corresponds to a file write system call, the application's buffer needs to be disabled. (2) To accurately match each application log entry with the underlying file write system call, we need to enhance the application log by injecting the system call's timestamp and thread ID (i.e., alignment information log) into the application log.

[0129] 2) Attack execution phase: During the attack, the logs generated by the application include alignment information logs.

[0130] 3) Log correlation step in the annotation phase: Use the alignment information log to match the application log with the corresponding audit log.

[0131] Specifically, disable buffering mechanisms: The following is a review of the server program logging application data. The program uses specific logging functions to write logs, such as printf in C, logging.info in Python, or logger.info in Java's Log4j, which usually use buffering mechanisms, accumulating data until a threshold is reached or a newline character is encountered, and then flushing to the output device. Finally, the write system call is triggered, writing the data to the file, creating a new application log entry and a corresponding audit log entry.

[0132] Buffering mechanisms vary for different programming languages ​​and frameworks. For basic languages, such as C / C++, functions such as printf follow the implementation of the standard library. High-level languages, such as Python and Java, have internal buffering mechanisms in their logging libraries.

[0133] If buffering is enabled, it is theoretically impossible to directly match log entries to system calls. We take the following steps in preparation to disable buffering:

[0134] For basic languages, such as C / C++, buffering can be disabled by hooking the dynamic link library, for example, using setbuf in the glibc library to disable buffering after the file handle acquisition function, such as fdopen.

[0135] For high-level languages, global buffering settings can be easily adjusted. For example, in Python logging, buffering can be set to 0 by monkey-patching, and in Java's Log4j, buffering can be disabled by configuring it in the log4j.properties file.

[0136] Further insert alignment information log:

[0137] Precise matching between application and audit logs is achieved by using accurate timestamps and thread IDs. Although long-running applications may use a multi-process or multi-threaded model, system calls from a single thread are sequential. By inserting a kernel module to hook the write system call and configuring it to identify log-related file writes, each log entry can be tagged with a thread ID and a precise timestamp (e.g., using ktime_get_real_ts64). This tagging is applied to each write operation during the attack execution phase.

[0138] Further matching logs in the annotation phase:

[0139] In the log-related steps of the annotation phase, Algorithm 1 is as follows Figure 3 As shown, the audit logs are sorted by thread ID and the application logs are matched with the most recent system call timestamp, thus ensuring that the matching process is efficient with a time complexity of O(n).

[0140] The specific steps to associate traffic logs with audit logs are:

[0141] The goal is to match each TCP connection and UDP session in the traffic log with the corresponding network-related system call in the audit log.

[0142] The audit log already contains detailed information for easy matching: the four-tuple of IP address and port number at both ends of the network traffic can be mapped one-to-one with a TCP connection or UDP session.

[0143] The matching process is described below and is Figure 4 (The system calls involved in this process are described in Figure 5 Listed in ):

[0144] (1) Match the closest system call: Identify the network system call whose timestamp and quadruple are closest to each entry in the traffic log.

[0145] (2) Locating the life cycle of traffic file descriptors: Track the creation and destruction of file descriptors used for system calls in the process audit log to establish matching boundaries.

[0146] (3) Batch matching within an interval: All system calls sharing the same quadruple are considered as matches to the traffic log.

[0147] In order to associate traffic logs and audit logs within a time complexity of O(n), an embodiment of the present application develops a two-step scanning algorithm that can be used in the log association step of the annotation phase. First, we scan the audit log to track the life cycle of each network file descriptor (fd) associated with the network system call. Subsequently, we scan the traffic log to determine the audit log entry e with the closest timestamp to the current quadruple match. For this log entry e, we locate all related audit log entries that share the same quadruple during the fd life cycle, so that it can be integrated with the traffic log.

[0148] Furthermore, anchor points are identified in traffic logs and audit logs:

[0149] Anchor points are log entries that represent attack activities, corresponding to the edges in the provenance graph. These points can be classified according to their type: traffic logs and audit logs.

[0150] There are some difficulties when trying to write rules to find anchors:

[0151] For traffic logs: attackers may use tools such as Burp Suite to intercept and modify browser requests to make them look like regular user activity, with only small but critical API changes. In addition, attackers can bypass rule-based intrusion detection systems (IDS), which require deep knowledge of the domain to accurately match these complex logs.

[0152] For audit logs: attackers may manipulate normal business processes, such as changing configuration files, to perform malicious actions, which often mimic legitimate activities, produce identical system calls and log entries, and complicate rule-based log matching.

[0153] To address these issues, this application modifies attack tools and strategies to inject attack flags into the payload, which can be directly associated with entries in traffic logs and audit logs to help identify anchor points. However, this raises two issues: 1) How to ensure the effectiveness of the attack: the attack payload, such as parameters in the traffic, is usually finely tuned. The payload needs to be modified without jeopardizing the success of the attack. 2) How to ensure the integrity of the log: irreversible changes cannot be introduced into the log.

[0154] To address these issues, during the preparation phase, an attack sign scanner was designed as a kernel module in the target environment. The scanner intercepts all traffic and system calls, detects attack signs, and restores them to their original state, which ensures the success of the attack without interruption and without leaving traces in the logs.

[0155] Further discussion of methods for identifying anchor points in traffic logs and audit logs.

[0156] The specific method of marking anchor points in traffic logs is as follows: In order to mark anchor points in traffic logs, the destination address of the traffic is used to inject attack markers. The destination address is usually under the control of the attacker and can be manipulated in scenarios such as external attacks or lateral movement within the network. Specifically, AUTOLABEL reverses the destination IP address (an IP address is a 32-bit number reversed through a logical NOT operation to create a new IP address. The reversed IP address avoids conflicts with network destinations and ensures uniqueness and concealment. Subnet consistency remains unchanged, which helps tools like nmap to scan and attack settings more easily).

[0157] In the preparation phase, the scanner records all IP addresses in the scenario and their reversed counterparts. In the attack execution phase, once the scanner detects a reversed IP address, it leaves an attack flag and then sets the reserved bit in the IP header to 1 (the unused reserved bit in the IP packet header is set to 0, carrying non-destructive information. During transmission, this bit remains unchanged, retaining the tracking function. Once the anchor point is found, the bit is reset to 0 and the packet is restored). In the marking phase, during the identification of the anchor point, the connection with the reserved bit in the IP header set to 1 is marked as an anchor point.

[0158] In Description A, the embodiment of the present application details how AUTOLABEL ensures that IP address modifications remain transparent to attackers.

[0159] The approach to designing an attack sign scanner is as follows: To ensure that the attack can proceed smoothly, even after the destination address is modified, we designed and implemented a kernel module, an attack sign scanner based on the Netfilter framework to intercept and modify the traffic. This interception and modification needs to be applied to both inbound and outbound traffic, because the upper-layer application that generates the forged traffic is unaware that the underlying destination address has changed. If the destination address of the response traffic is incorrect, it may cause the application to behave abnormally.

[0160] Algorithm 2 is as follows Figure 6 As shown, outbound traffic and inbound traffic are handled separately:

[0161] For outbound traffic:

[0162] 1. Traffic is intercepted at the NF_INET_LOCAL_OUT hook. If the destination is not an inverted IP address, it is forwarded directly. Otherwise, proceed to the next step.

[0163] 2. The reserved bit is set to 1, and the destination address is restored to its original value. The new source address is calculated based on the routing table to ensure correct routing, and this address is set as the new source.

[0164] 3. The actual destination address and source port number are recorded in hack_connection so that they can be adjusted in the response traffic.

[0165] For inbound traffic:

[0166] 1. Traffic is intercepted at the NF_INET_LOCAL_IN hook. If the source address and destination port do not match any entry in hack_connection, it is forwarded directly, otherwise, proceed to the next step.

[0167] 2. The reserved bit is set to 1, and the source address is changed to prevent application-level anomalies. The new destination address is calculated based on the routing table to ensure correct delivery.

[0168] The specific method of marking the anchor points in the audit log is as follows: To mark the anchor points in the audit log, pay attention to the main injection points of file names and file contents in file read / write operations, and command parameters in command execution. These elements are usually under the control of the attacker, such as writing malicious trojans, accessing sensitive files, or executing harmful commands. Include these attack signs in the attack payload and show their effects through disk operations and system calls related to command execution.

[0169] In the preparation phase, an embodiment of the present application integrates a kernel module to hook into the system call and inject an attack flag into the attack payload. For example, when writing a malicious Trojan, a flag is added to the file content. In the attack execution phase, the attack flag scanner intercepts system calls, restores them to their pre-attack state, and then executes the payload. For example, in the example of a malicious Trojan, during the writing of a system call, the file content is restored to its original form. In the anchor positioning step of the annotation phase, logs related to the attack are easily identified because the audit log captures the system call parameters with the attack flag. After identification, the attack flag is removed and the log is restored.

[0170] The embodiments of the present application apply specific methods to insert and remove attack flags in system calls, related file operations and command execution:

[0171] File name in file operations: For the openat system call, a fixed-length identifier is prepended to the file name. The pointer (rdi) is then adjusted forward by that length to access the original file name.

[0172] File content is written in the write operation: In the write system call, a fixed-length identifier is added before the content. The content pointer (rsi) is moved forward by that length to restore the original content.

[0173] Command parameters in execution: For execve, a special identifier is inserted as a new parameter. Detection of this identifier prompts the subsequent parameters to move forward one position to restore the original command state. After execution, the parameters are moved backward to maintain the consistency of the user space program.

[0174] A similar approach applies to variants of these system calls.

[0175] In Note B, it is detailed how AUTOLABEL can modify existing attack tools to subtly change command execution, file names, and file contents without alerting the attacker.

[0176] If the key anchor points cannot be identified using conventional methods, AUTOLABEL will adopt the rule matching strategy proposed by Kyoushi.

[0177] During the attack execution phase, attackers may choose to record their actions in the attack log, including:

[0178] Timestamps mark the start and end of an attack: Timestamps help refine the scope of rule matches, improving accuracy and minimizing false positives.

[0179] Attack parameters: Record specific details such as IP addresses or traffic parameters for accurate matching.

[0180] Specific matching rules: Leverage Kyoushi’s nested rule approach to detect log signatures by considering event combinations, contexts, and attack parameters.

[0181] During the anchor localization step of the annotation phase, AUTOLABEL integrates attack logs to accurately locate anchors, thus addressing any deficiencies left by earlier methods.

[0182] In step S103, the nodes in the initial traceability graph that meet the preset segmentation conditions are divided into multiple execution units using the delimiter log to obtain a refined traceability graph, and all anchor points are connected based on the refined traceability graph to form an attack subgraph, so as to obtain a labeled log based on the attack subgraph.

[0183] Optionally, in some embodiments, when connecting all anchor points based on the refined traceability graph to form an attack subgraph, it includes: constructing a handler derivation relationship tree; determining handlers that share a derivation relationship through the handler derivation relationship tree, and classifying system calls in the handlers that share a derivation relationship as belonging to the same execution unit.

[0184] Optionally, in some embodiments, when connecting all anchor points based on the refined traceability graph to form an attack subgraph, it includes: building an initial handler tree; evaluating the number of accept calls in each node subtree of the initial handler tree, and deleting nodes whose accept call times are greater than or equal to a preset number, to obtain a target handler tree, so as to define different execution units according to the target handler tree.

[0185] Optionally, in some embodiments, all anchor points are connected based on a refined provenance graph to form an attack subgraph, so as to obtain labeled logs based on the attack subgraph, including: based on the refined provenance graph, processing in reverse topological order from nodes with no outgoing edges; maintaining a set for each node, and updating the set for each new node, by taking the set union of all direct successor nodes and subtracting all intersections to obtain all marked path edges, wherein, if the node is a starting point, all paths starting from the starting point are marked, and the set of each node is adjusted to obtain all marked path edges; connecting all marked path edges to form an attack subgraph.

[0186] Optionally, in some embodiments, after forming the attack subgraph, it includes: checking whether all anchor points in the refined traceability graph have been connected into an attack path; if there are unconnected anchor points, determining the positions of the unconnected anchor points, and prompting the user to add anchor points.

[0187] The specific method to solve the dependency explosion problem of attack subgraph extraction is as follows:

[0188] 1) Divide the execution units and refine the traceability graph;

[0189] 2) Expand from the anchor points in the refined traceability graph and construct an attack subgraph.

[0190] Unit division principle: Before dividing the execution unit, you need to define what the execution unit is. In AUTOLABEL, the execution unit refers to all operations performed by the application when processing a single request. Because the application may use an asynchronous execution model, these operations may occur across different threads and times, but their common goal is to complete the same task.

[0191] In the asynchronous execution model, the completion of a task is split into multiple handlers, each of which handles a part of the task. These handlers have a derivative relationship, which means that the derived handlers are still processing the same task.

[0192] like Figure 7 As shown, here are some specific examples of execution units for four common execution models when a request arrives:

[0193] 1) Sequential processing tasks: Each handler directly corresponds to a request, thus forming an execution unit.

[0194] 2) Asynchronous task queue: One request is processed by multiple handlers. For example, request 1 is initially processed by handler 1, which then spawns handlers 3 and 6. Handlers 1, 3, and 6 together form an execution unit.

[0195] 3) Create additional threads: When a request is accepted by a handler on the main thread, it creates a new thread to service the request. The handler on the main thread and the handler on the thread it created together form an execution unit.

[0196] 4) Thread pool: When a request is accepted by a handler on the main thread, an idle worker thread in the thread pool is scheduled to service the request. The worker threads process different tasks in sequence, and each segment corresponds to a handler. Therefore, the combination of the handler on the main thread and the handler corresponding to the derived task forms an execution unit.

[0197] This illustration only shows operations using a single execution model. However, modern applications often use a combination of multiple execution models. For example, a common pattern is for the main thread to use an asynchronous event queue, while long-running tasks are submitted to a background thread pool for maintenance.

[0198] Based on the above discussion, the embodiment of the present application takes the following two steps to partition the execution unit:

[0199] 1) Partition all system calls in a thread into different handlers. Record the handler ID at the start of each handler so that you know when to switch and which handler to start running. For example, in an asynchronous task queue, recording occurs where each event callback is executed; in a thread pool, recording occurs where each task is executed.

[0200] 2) Construct a handler derivation tree (called handler tree). The derivation is recorded where the new handler is created, recording the new handler's ID. The current handler can be connected to the newly derived handler. For example, in an asynchronous task queue, the recording is done at the event registration point; in a thread pool, it happens where the task is submitted.

[0201] By following these steps, all system calls of an application can be partitioned into different handlers and a tree can be built to capture their derivation relationships. Therefore, system calls in handlers that share a derivation relationship belong to the same execution unit.

[0202] The embodiment of the present application has manually injected the official asynchronous libraries of Python, Node.js and Java to enable execution partitioning, which is a one-time task. The details are described in Note D.

[0203] Unit Partitioning: Automated Instrumentation. Although only one manual instrumentation is required to meet the unit partitioning requirements of server applications using popular asynchronous libraries for high-level languages ​​and frameworks, for server applications developed in C language, instrumentation points still need to be manually identified, which requires a lot of domain knowledge and effort. To solve this problem, the present application embodiment proposes an automated method to analyze and instrument C language projects to facilitate unit partitioning.

[0204] The main difficulties currently faced include:

[0205] 1) Automatically locate processor execution and derivative points.

[0206] 2) Identify the top-level processor that accepts the request and aggregate its system calls and its descendants into a single execution unit.

[0207] For the first difficulty, the embodiment of the present application takes the following steps:

[0208] 1) Static analysis: We analyze C code features and use static matching abstract syntax trees (ASTs) to preliminarily identify and instrument code executed and registered by processors. Specifically, in C, processors are usually implemented as function pointers in a class structure, similar to event → process (event). Using tools such as clangd, generate and traverse ASTs to identify all function pointer calls and assignments, which correspond to processor executions and derivatives. The embodiment of the present application uses a macro wrapper to record the event addresses of these points.

[0209] 2) Dynamic analysis: Accurately identify processor activity based on the dynamic behavior of the code. Specifically, after instrumentation, recompile and run the software to collect logs, including event addresses. The present application embodiment applies three rules to further improve:

[0210] Mutable event addresses: Event addresses should show mutability, indicating dynamic assignment rather than a static function pointer.

[0211] Repeated handler calls: Evaluate the execution interval of each function pointer, record their first and last execution, and then merge any overlapping intervals. The longest intervals identify function pointers active in the event loop, while shorter or non-overlapping intervals indicate function pointers used before the event loop started or after it ended.

[0212] Processor is at the top of the call stack: The processor should mainly be at the top of the call stack during execution.

[0213] This precise set of function pointers marks where the processor executes and spawns, and this approach has been successfully implemented in projects such as Nginx, Haproxy, Redis, and Libuv.

[0214] For the second difficulty, this embodiment constructs a handler tree and evaluates the number of accept calls in each node subtree. Nodes with more than two accept calls may not correspond to a single request. By removing these nodes, we refine the handler tree into multiple independent trees, each of which is rooted at the handler directly linked to the TCP connection, thereby defining different execution units.

[0215] This automated process simplifies code injection and unit partitioning for long-running C-based server applications.

[0216] The steps of attack subgraph extraction are:

[0217] After identifying anchor points in the traceability graph, the goal of the embodiment of the present application is to connect them into an attack subgraph. The embodiment of the present application only marks the path edge between two anchor points if there is a unique path between them. In order to complete this process efficiently, an O(kn) algorithm based on topological sorting is designed, where k is the number of anchor points. The steps are as follows:

[0218] 1) Define the starting nodes of the edges corresponding to all anchor points as “starting points”.

[0219] 2) Using the structure of the traceability graph, i.e., a directed acyclic graph, we start from the nodes without outgoing edges and process them in reverse topological order.

[0220] 3) For each node i, maintain a set Si containing all the starting points that can be reached from i through a unique path.

[0221] 4) Update Si for each new node i by taking the set union of all direct successor nodes and then subtracting any intersections to ensure that the path from i remains unique.

[0222] This step requires careful handling of various complex scenarios, such as distinguishing between pointer and variable usage, and embedded function pointer calls.

[0223] 5) If i is a starting point, mark all paths starting from i and adjust Si to include only i, because the paths to other points in Si must pass through i.

[0224] If there are still unconnected anchors after this process, it means that the current anchors are not enough to cover all attack paths. In this case, AUTOLABEL will point out the locations of these unconnected anchors and prompt the user to add more anchors or write rules to enhance the attack subgraph construction.

[0225] After solving the three challenges, the embodiment of the present application obtains an attack subgraph containing attack-related nodes and edges. In view of the association between application, network and audit logs, the multi-source logs related to each edge connected to the attack-related nodes in the attack subgraph are further labeled, and this process produces a fully labeled log dataset.

[0226] The Linux server used for the experiment has an Intel 13th generation Core i9-13900H CPU 2.60GHz to 5.40GHz and 32GBRAM, running Ubuntu 22.04.3LTS.

[0227] Environment setup: Docker was used to create a network topology and containers that mimic real-world hosts.

[0228] Advantages of using containers include:

[0229] 1) Easily manage starting and stopping through automation.

[0230] 2) Native support for setting the LD_PRELOAD environment variable, which is essential for AUTOLABEL's dynamic linking hooks.

[0231] 3) Convenient kernel module insertion on all hosts. It only needs to be implemented once on the master server to affect all containers. The kernel module distinguishes containers by identifying the namespace of the current process.

[0232] Log collection: Use sysdig to collect audit logs, use tcpdump to collect traffic logs by monitoring Docker's virtual network devices, and direct application logs to a specified directory.

[0233] Attack launch: The attack was launched from a host running Ubuntu Desktop with custom attack tools and a configured terminal that automatically sets the necessary environment variables. We also registered the required kernel modules on this machine.

[0234] Log matching accuracy assessment: The log matching process that combines application logs with audit logs and traffic logs with audit logs was evaluated. Figure 8 As shown, the experiments of the embodiments of the present application involve six server applications, representing common categories, generating a large amount of traffic, application and audit logs.

[0235] Application log to audit log: Tested whether AutoLabel can accurately map application logs to corresponding audit logs.

[0236] One-to-one correspondence evaluation: The alignment information log of each application log line was checked to ensure that different entries are not mixed in the same buffer.

[0237] Match Accuracy Assessment: Using Sysdig, we confirmed that the prefix of the content written by the write system call matched the prefix in the application log.

[0238] The buffering mechanism varies between applications:

[0239] Nginx and Python Logging use custom log buffering, which is configurable and disabled by default.

[0240] Apache, Morgan (Node.js), Redis, and MySQL Server (excluding MySQL binlog) rely on the C standard library's buffering.

[0241] Results (such as Fig. 9 and Fig.10 The following table shows successful one-to-one log mapping for all applications, except for Apache where audit log entries are lost at higher concurrency levels (over 50). This issue can be addressed by using a more reliable audit logging framework.

[0242] Traffic log to audit log: records both traffic log and audit log while accessing the above applications. Fig.11 As shown, all traffic log entries have found corresponding quads in the audit log, which confirms the successful matching of all entries.

[0243] Evaluate the usefulness of anchor positioning:

[0244] AUTOLABEL uses attack signatures to identify attack points, ensuring accurate anchor point detection. The feasibility of this method in traffic and audit logs was tested by modifying the attack tools and plans.

[0245] 1) Locating anchor points in traffic logs: To demonstrate the feasibility of this method, Fig.12 As shown, 8 experiments were conducted, using common attack tools to inject attack flags, as described in Note E. The results show that all attack flags were injected and all anchors were successfully identified.

[0246] 2) Locating anchor points in the audit log: Four scenarios involving arbitrary command execution and file operations were reproduced, and attack flags were injected at key points. All anchor points were successfully identified by AUTOLABEL, such as Fig.13 Record.

[0247] Effectiveness in solving the dependency explosion problem

[0248] exist Fig.14 In , cell partitioning is implemented in the audit log for various applications, and the consistency and correctness of log entries within a cell are evaluated:

[0249] For Nginx and Haproxy, multiple reverse proxy rules are set up. The assessment checks whether the writing of audit logs and network system calls are matched within the unit.

[0250] For Redis, the standard is the consistency of the four-tuple of network system calls within each cell.

[0251] For Express.js and Fast API applications, we evaluated whether system calls to write audit logs for API requests and file writes were consistent across cells.

[0252] For Apache using the Prefork MPM, unit partitioning is based on audit log sequences with the accept system call as a delimiter, using a similar evaluation method as Nginx and Haproxy.

[0253] For MySQL Server, in a configuration where each connection is managed by a unique thread, the evaluation verifies that all SQL statements in the log within the cell are identical.

[0254] The experimental results show that all evaluation criteria are met and the execution units are correctly partitioned (e.g. Fig.14 ).

[0255] Evaluate performance on the entire dataset:

[0256] We evaluate the labeling effectiveness of AUTOLABEL in a real network environment involving a 10-step attack between 4 hosts with normal user activities as background. Manual verification confirms the accuracy and completeness of the labeling logs.

[0257] To demonstrate the capabilities of AUTOLABEL, selected attack steps on a host are described in detail in Note C, showing how AUTOLABEL extracts the attack graph and labels the logs. In this detailed description, we outline the manual work required for AUTOLABEL for this complete example and explain how each log is labeled.

[0258] Evaluate performance overhead:

[0259] In the preparation phase, automated annotation analysis for C language projects mainly involves static analysis, which is time-consuming because it is necessary to traverse the project's abstract syntax tree (AST) and identify annotation points. The time for this analysis depends on the complexity of the project. Fig.15 The static analysis time of HAProxy is the longest, reaching 90 seconds, while Redis, Nginx and libuv are relatively short.

[0260] During the attack execution phase: various applications were tested for concurrent access with a concurrency of 255 to evaluate the runtime overhead and space overhead.

[0261] During the attack execution phase, the runtime overhead includes two aspects:

[0262] Application Execution: For log correlation, kernel modules are inserted to print alignment information logs and hooks are applied to dynamic link libraries to disable buffering; to identify attack points, kernel modules intercept system calls; for execution unit segmentation, delimited logs are printed at specific locations.

[0263] Traffic Transmission: To identify attack points, kernel modules are inserted to intercept traffic.

[0264] Fig.16 The blue part in the figure records the change in request access time of each application before and after the attack. It can be seen that Redis has the highest runtime overhead, reaching 8.5%.

[0265] During the attack execution phase, the space overhead mainly includes:

[0266] The attack flag is injected into the sysdig log. Since the system calls involved in APT attacks only account for a small part of the audit log, this load can be ignored.

[0267] The alignment information log is injected into the application log. Since each line of the application log corresponds to an alignment information log entry, this part will generate a significant space load. Fig.16 The red part has calculated the changes in application logs, and the maximum increase in application log space is 53.7%.

[0268] Delimited logs: Since the goal of delimited logs is to segment the audit log, this increase is considered "audit space overhead". Fig.17 As shown, the largest overhead in the audit log occurs in the Haproxy application, which does not exceed 8%.

[0269] During the annotation phase:

[0270] The time complexity of each step in the annotation phase is linear with the number of log entries, which is equivalent to scanning the log multiple times, so the time consumption is acceptable. In addition, since the annotation phase is a post-analysis step, there is enough time. In our hardware environment, the speed of traversing the three types of logs using Python is about 105 entries per second, which is considered acceptable.

[0271] Limitation and Discussion

[0272] 1) AUTOLABEL support for GUI applications. AUTOLABEL annotation has difficulties on GUI applications because complex UI events such as malicious clicks or keystrokes are not fully captured in system logs. The multifaceted nature of GUI applications makes it complicated to identify logs related to malicious activities. Improving automatic annotation of GUI applications is a key future challenge.

[0273] 2) Injecting attack markers into audit logs requires domain knowledge. AUTOLABEL relies on the attacker injecting attack markers to identify anchor points in the audit log, which requires a deep understanding of the attack. Although modifying the attack tool is discussed in Note B, it is not always feasible, especially with custom PoC scripts. Although modifying the attack payload may seem simple to a knowledgeable attacker, more research is needed to find more subtle ways to inject attack markers. Automatically identifying the insertion points in the attack tool code can simplify and improve the annotation process.

[0274] 3) Static analysis in unit partitioning needs to consider special cases. Our static analysis of C projects automatically identifies and macro-encapsulates function pointer calls and assignments. However, function pointer calls hidden in macros or complex expressions require advanced handling. We have added extensive logic to handle these issues, and future work aims to refine this approach and cover more special cases and advanced C features.

[0275] Related Work

[0276] Automatically annotating multi-source logs is relatively less explored. Kyoushi is notable in that it parses multi-source logs into a semi-structured format and applies rule templates such as query and sequence rules. However, it has difficulties in writing large numbers of rules, integrating between different log types, and identifying complex attacks.

[0277] In terms of traffic log annotation, time window-based methods utilize specific time frames to distinguish normal traffic from malicious traffic, but encounter difficulties when granularity is required in complex attacks. Behavioral model-based methods generate and annotate network traffic through summaries of normal and malicious behaviors. Although GANs and tools like the CREME toolbox simulate real traffic to enhance datasets, they are mainly applicable to synthetic traffic and cannot be directly applied to real-world traffic. Detection tool-based methods use NIDS and statistical tools to improve annotation. However, their effectiveness depends largely on the detection tool used.

[0278] Approaches to address the dependency explosion problem. To address the dependency explosion problem, BEEP proposes to partition long-running applications into several execution units to improve attack traceability. However, BEEP's direct approach is to train and analyze loop and memory dependencies, which may lead to incorrect partitioning. MPI improves partitioning by using data structure annotations and tracking indicator variables, although this approach is complex and difficult to apply in complex asynchronous codes. Omegalog and Alchemist use application logs to partition audit logs, but they rely on specific log content and format, which limits their applicability. TeSec, which focuses on web applications, addresses different asynchronous models but fails to consider their combined use in frameworks like Express.js, resulting in inaccurate unit partitioning.

[0279] Unit partitioning methods based on taint analysis have been proposed to achieve more fine-grained attack traceability by merging low-level data flow information beyond system calls. However, these methods cannot ensure accurate taint propagation. Log annotation tasks require deterministic algorithms and 100% accuracy, so they are not suitable for log annotation.

[0280] Conclusion

[0281] The present application embodiment introduces AUTOLABEL, an automatic log annotation system. AUTOLABEL merges audit, traffic, and application logs to create a traceability graph, identifying and annotating APT attack-related entries. By subtly adjusting the attack tool, AUTOLABEL injects identifiers during the attack to achieve automatic annotation, eliminating manual intervention. It reduces the need for manual annotation and improves accuracy and completeness. In testing, AUTOLABEL achieved 100% annotation accuracy, and the running time and space overhead were within 8.5% and 53.7% respectively, demonstrating its practicality.

[0282] Description A

[0283] Methods to covertly modify the destination address of traffic

[0284] The destination address of network traffic is modified covertly by dynamically hooking into the link layer function, thereby injecting attack markers into the traffic log without the attacker's knowledge.

[0285] This application involves intercepting the connect and sendto functions in the glibc library, which are crucial for network operations in Linux:

[0286] connect function: starts a connection request with the server.

[0287] sendto function: Sends a data packet to a specific address in a connectionless network.

[0288] A custom library is loaded using the LD PRELOAD trick to modify these functions. When the application uses these functions, our library first verifies if the destination IP is part of the target cluster. If yes, it modifies the address and forwards it to the original library function.

[0289] The modification is transparent and automatic, perfectly integrated into the attacker's toolkit:

[0290] Setting the LD PRELOAD variable in the attacker's endpoint ensures that all network applications inherit this modification without additional setup.

[0291] Setting up LD PRELOAD in a compromised internal network ensures that all network tools covertly change their communication addresses.

[0292] An attacker who is not familiar with the environment configuration can use pre-configured tools or endpoints where the necessary settings have already been applied.

[0293] Description B

[0294] A method to covertly modify attack payloads in attack tools

[0295] Various attack tools have been improved to secretly inject attack flags into payloads, affecting file and command related system calls without the attacker's knowledge.

[0296] Many advanced attack tools provide backdoor management capabilities after establishing a foothold:

[0297] Tools like Antsword, Godzilla, and Behinder feature graphical user interfaces for file and command management, including drag-and-drop and file renaming capabilities. The Metasploit Framework provides a Meterpreter session for similar purposes.

[0298] Examining the source code of these tools revealed common file management and command execution functions, such as the file manager and terminal modules in Antsword, and meterpreter / extensions in Metasploit Framework.

[0299] By modifying these unified functions to automatically insert attack flags into parameter inputs, we can change the payload without the attacker knowing.

[0300] An attacker using an integrated backdoor management tool can modify similar functions by himself, or use the pre-modified versions of Antsword and Metasploit Framework provided in the embodiments of the present application.

[0301] Description C

[0302] Complete flow example of AUTOLABEL

[0303] Fig.18 Depicted is a provenance graph containing red nodes and edges marking anchors, blue represents audit logs, and yellow edges represent connections during attack subgraph extraction.

[0304] Attack scenario description: In order to penetrate the host, the attacker previously implanted frpc in the intranet and established a socks5 tunnel. Through this tunnel, the attacker can directly attack the target host. At the same time, the target host has a web application developed using the FastAPI framework.

[0305] The attack involved:

[0306] Attack step 1: Exploiting SQL injection to retrieve administrator credentials.

[0307] Second attack step: Use these credentials to upload a malicious backup file through a file upload vulnerability and set it to execute as a cron job.

[0308] The third attack step: execute commands through the cron job to detect the host.

[0309] The fourth attack step: Use nc to establish a reverse shell and gain control of the host.

[0310] Marking method:

[0311] 1) Identify anchor points in traffic logs: First, the address that frpc accesses frps from is reversed, and the target address that the attacker accesses is reversed. This allows the attack flag to be injected directly into the proxy's traffic. Additionally, when nc executes a reverse shell, the target address filled in is reversed, allowing the attack flag to be injected into the reverse shell traffic. This step marks all red nodes and edges.

[0312] 2) Identify anchor points in the audit log: The command that the attacker fills in the backup file includes the attack flag in the parameter. This step marks all blue nodes and edges.

[0313] 3) Divide the units in the FastAPI application and connect the paths: After identifying the anchors, the logs that are still not considered include:

[0314] Other information within the FastAPI unit, such as malicious SQL traffic sent and API records in application logs.

[0315] The process forking and file read / write relationship between the backup file writing and the malicious command execution.

[0316] By partitioning the cells and connecting the anchors, we obtain a connected attack subgraph which marks all yellow nodes and edges.

[0317] Through these three steps, we achieved accurate and complete labeling of attack-related logs in this attack scenario.

[0318] Description D

[0319] Example of environmental unit division

[0320] Example 1: Python language. Python language includes the official asynchronous library Asyncio and the official thread pool ThreadPoolExecutor. Fig.19 The instrumentation points are shown in . For FastAPI applications developed in Python, the subtree formed by the handlers of the \_accept_connection2 function can be treated as a unit because this function has a one-to-one correspondence with the accept system call.

[0321] Example 2: Node.js language. Node.js includes the libuv library, which provides an event loop. For file reading and writing operations, Node.js uses the thread pool provided by libuv. There are seven types of handlers in the libuv event loop; it also implements the uv__work class to encapsulate the handlers of the thread pool. Fig. 20The instrumentation points are indicated in . For Express.js applications developed using Node.js, each TCPWRAP handler directly corresponds to the accept system call, and its subtree can be considered as a unit.

[0322] Example 3: Java. Java provides the official thread pool ThreadPoolExecutor.

[0323] The execution of the handler is implemented in the beforeExecute method.

[0324] · The derivation of the handler is implemented in the execute method.

[0325] Example 4: Nginx. Nginx implements an event loop. There are three types of handlers: IO type (usually implemented by epoll on Linux); Posted type; and Timer type. Fig.21 The insertion point is indicated in .

[0326] Example 5: Redis. Redis implements an event loop. There are two types of handlers: file event handlers and time event handlers. Fig. 22 The implementation point is indicated in .

[0327] Example 6: Haproxy. Haproxy implements a task-based event loop, represented by struct task.

[0328] The execution of the handler is implemented in run_tasks_from_lists.

[0329] The derivation of the handler is implemented in the task_queue function (task queue).

[0330] Description E

[0331] Description of the attack signature injection method in attack tools and traffic logs

[0332] In order to verify the practicality of locating anchor points in traffic logs, e.g. Fig.23 As shown in the , experiments were conducted using common attack tools. The following are their descriptions and the methods used to inject attack flags:

[0333] Nmap is a network scanning tool used to discover devices and services on the network. By directly modifying the scanned target address, you can inject attack flags.

[0334] Burpsuite is an integrated platform for testing the security of web applications. Changing the IP address that Burpsuite accesses can inject attack flags.

[0335] Antsword is a web shell management tool. When connecting to the backdoor, specify the backdoor's modified IP address.

[0336] Bash Reverse Shell allows remote command execution and changes the IP address when the IP address of the jump target is specified.

[0337] DNS Shell uses the DNS protocol for covert command and control communications. When the DNS server address controlled by the attacker is specified, the IP address of the DNS server is changed to the modified IP address.

[0338] Frpc&Frps is a network penetration tool that modifies the IP address when Frpc specifies the Frps IP address.

[0339] SSH Tunnel & Proxychains4 can establish a communication tunnel using a socks5 proxy. The settings include the SSH target address, the Proxychains4 proxy address, and the internal network address accessed through the socks5 proxy. Modifying all of these addresses can inject attack flags into all traffic.

[0340] Metasploit Framework is a versatile penetration testing framework that is suitable for all stages of penetration. Combined with modifying the target and rebound address, it can inject attack flags.

[0341] According to the automated fine-grained log annotation method proposed in the embodiment of the present application, a preset attack scenario is run, and log information including unlabeled logs, alignment information logs, separator logs, attack flags and attack logs is collected based on the running results, an initial traceability graph is constructed based on the audit logs in the alignment information logs, the application logs and traffic logs are associated with the audit logs using the alignment information logs, and the anchor points in the initial traceability graph are identified using the attack flags and attack logs based on the association results, the nodes in the initial traceability graph that meet the preset segmentation conditions are divided into multiple execution units using the separator logs to obtain a refined traceability graph, all anchor points are connected based on the refined traceability graph to form an attack subgraph, so as to obtain annotated logs based on the attack subgraph. As a result, the problems of current automatic log annotation methods, such as coarse granularity, insufficient coverage of different log types, and excessive reliance on manual work and domain expertise, are solved. Manual workload is reduced, human involvement is minimized, and reliance on domain expertise is reduced. Multiple data sources are covered, and traffic, audit, and application logs can be annotated simultaneously. Fine-grained labels are provided: precise annotation of individual log entries is achieved, specific connections, audit events, and application log entries are identified, accuracy and completeness of annotations are ensured, entries not related to attacks are avoided, and missing attack logs are minimized.

[0342] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0343] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0344] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

Claims

1. An automated fine-grained log annotation method, characterized in that: The following steps are involved: Run the preset attack scenario and collect log information including unlabeled logs, alignment information logs, delimiter logs, attack flags, and attack logs based on the running results; Constructing an initial traceability graph according to the audit log in the alignment information log, associating the application log and the traffic log with the audit log using the alignment information log, and identifying anchor points in the initial traceability graph using the attack flag and the attack log based on the association result; The delimiter log is used to divide the nodes in the initial traceability graph that meet the preset segmentation conditions into multiple execution units to obtain a refined traceability graph, and all anchor points are connected based on the refined traceability graph to form an attack subgraph, so as to obtain a labeled log based on the attack subgraph.

2. The method according to claim 1, characterized in that The step of associating the application log and the traffic log with the audit log using the alignment information log includes: disabling a buffer of the application, injecting the timestamp and thread ID in the alignment information log into the application log, and obtaining a modified application and a modified application log; The modified application is executed to regenerate an alignment information log, and the modified application log is associated with the audit log using a timestamp and a thread ID of the regenerated alignment information log.

3. The method according to claim 1, characterized in that The utilizing the alignment information log to associate the application log and the traffic log with the audit log further includes: Obtaining each TCP connection, UDP session timestamp and four-tuple information in the traffic log, and obtaining the four-tuple information in the audit log; According to the UDP session timestamp and the four-tuple information in the traffic log, searching in the audit log for a network system call that is closest to the UDP session timestamp and the four-tuple information in the traffic log; Within the life cycle of the traffic file descriptor corresponding to the closest network system call, all system calls with the same four-tuple are matched in batches to establish an association between the traffic log and the audit log.

4. The method according to claim 1, characterized in that: The using the attack flag and the attack log to identify the anchor point in the initial tracing graph includes: Use the preset attack sign scanner to detect all IP addresses in the attack scenario and the addresses corresponding to the scenario reversal; A connection whose reserved bit in the IP address header of the traffic log is set to 1 is identified, and the connection whose reserved bit in the IP address header is set to 1 is marked as an anchor point in the traffic log.

5. The method according to claim 1, characterized in that The identifying the anchor point in the initial tracing graph by using the attack mark and the attack log also includes: Integrate a preset kernel module to hook into the system call and inject the attack flag into the attack payload; intercepting the system call using a preset attack signature scanner, restoring the attack payload to a state before the attack, and executing the restored attack payload; A system call parameter with an attack flag in the audit log is identified, and a corresponding log entry in the audit log is marked as an anchor point based on the system call parameter with the attack flag.

6. The method according to claim 1, characterized in that When the anchor point in the initial traceability graph is not identified, the following steps are included: Recording the attacker's behavior log, wherein the behavior log includes a timestamp, attack parameters, and preset matching rules; Refining the scope of the preset matching rule by using the timestamp, and detecting the attack log signature by using the attack parameter and the preset matching rule, taking into account the event combination, context and attack parameter; Anchor points in the initial traceability graph are identified according to the attack log signature.

7. The method according to claim 1, characterized in that When all anchor points are connected based on the refined traceability graph to form an attack subgraph, it includes: Construct a handler derivation tree; The processing programs that share a derivation relationship are determined through the processing program derivation relationship tree, and the system calls in the processing programs that share a derivation relationship are classified as belonging to the same execution unit.

8. The method according to claim 1, characterized in that When all anchor points are connected based on the refined traceability graph to form an attack subgraph, it includes: Build the initial handler tree; The number of accept calls in each node subtree of the initial processing program tree is evaluated, and nodes whose number of accept calls is greater than or equal to a preset number are deleted to obtain a target processing program tree, so as to define different execution units according to the target processing program tree.

9. The method according to claim 1, characterized in that: The step of connecting all anchor points based on the refined traceability graph to form an attack subgraph, and obtaining annotated logs based on the attack subgraph, includes: Based on the refined traceability graph, the nodes without outgoing edges are processed in reverse topological order; Maintain a set for each node and update the set for each new node by taking the set union of all direct successor nodes and subtracting all intersections to obtain all marked path edges. If the node is a starting point, mark all paths starting from the starting point and adjust the set of each node to obtain all marked path edges. Connecting all the marked path edges forms the attack subgraph.

10. The method according to claim 9, characterized in that After forming the attack subgraph, it includes: Check whether all anchor points in the refined traceability graph have been connected into an attack path; If there are any unconnected anchor points, the positions of the unconnected anchor points are determined, and the user is prompted to add additional anchor points.

Citation Information

Patent Citations

  • Big data analysis method applied to information security field

    CN109660526A

  • Attack tracing method and device based on log association analysis

    CN114615063A

  • Attack tracing method and device based on log association analysis

    CN115396137A

  • Method and device for adjusting NGINX connection number, electronic equipment and storage medium

    CN115883580A

  • Log analysis and APT attack tracing method based on unsupervised learning

    CN118611962A

Cited By

  • Heterogeneous event association representation model construction method, system and equipment based on attack stage semantic alignment

    CN120811708A