Agile robust APT detection method based on traceability graph multi-view comparative learning
By employing a multi-view comparative learning method based on a continuous-time dynamic graph on the source graph, combined with η-BFS and -DFS sampling strategies, the adaptability and agility issues of existing APT detection methods in complex scenarios are solved, achieving efficient and robust APT attack detection, suitable for cybersecurity protection in key areas such as government, military, and finance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-28
AI Technical Summary
Existing APT detection methods based on source maps suffer from poor adaptability to fixed encoding methods, weak robustness, and insufficient detection agility due to a single sampling perspective when facing unknown, covert, and complex attacks. They are difficult to adapt to complex and ever-changing APT attack scenarios and cannot accurately locate key attack time points and core affected areas.
We adopt a multi-view contrastive learning method based on continuous-time dynamic graphs. We use a hybrid subgraph sampling strategy of η-BFS and -DFS to sample from both temporal and structural perspectives. By using the multi-view contrastive learning framework, we can capture important changes in the source graph and improve detection robustness and agility.
It achieves high-precision detection of complex and ever-changing APT attacks, improves the model's generalization ability and robustness against unknown attacks, significantly shortens attack response time, and reduces deployment and maintenance costs. It is suitable for cybersecurity protection in key areas such as government, military, and finance.
Smart Images

Figure CN121940177A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to an agile and robust APT detection method based on multi-view comparative learning of source graphs. Background Technology
[0002] Advanced Persistent Threats (APTs) have become a prevalent multi-dimensional adversarial mode in cyberspace. [1-2] APTs are characterized by high complexity and multiple execution stages, making them the preferred method for elite attackers to target high-value assets such as government agencies, military facilities, and financial institutions. Therefore, developing timely and accurate APT detection technologies has become a key task in contemporary network defense research.
[0003] Traditional intrusion detection systems primarily focus on native system log analysis or rely on malware signatures. [3-4] However, APT attackers typically employ stealthy and persistent attack patterns: they establish a foothold in target systems by exploiting zero-day vulnerabilities and evade detection for extended periods. This makes traditional detection methods either unable to identify new vulnerabilities or limited to log-related event analysis, failing to effectively counter APT threats.
[0004] To address these challenges, an increasing number of researchers are dedicated to transforming system logs into source graphs that can characterize the flow of event information. [5] A source graph is a collection of system entities and system events, typically presented as a heterogeneous graph: system entities (such as files, processes, and sockets) are represented as nodes, and system events (such as read and write operations) are represented as edges. This heterogeneous source graph can uniformly model various entities and their complex relationships involved in normal and attack events, providing rich information support for attack tracing and behavior analysis.
[0005] APT detection methods based on source graphs can be divided into three categories: rule-based methods, statistical methods, and deep learning-based methods. Rule-based methods... [6-7] Heuristic rules are built based on known attacks (such as the MITRE ATT&CK framework) to identify threats. These methods have low false positive rates, but suffer from challenges in rule design, reliance on specialized knowledge, and difficulty in detecting unknown or potential attacks due to limitations imposed by specific conditions. Statistical methods... [8-9] Anomaly scores are determined by recording the frequency of historical events and basing them on the rarity of edges; some methods are also based on statistics.
[10] Histograms are transformed into fixed-dimensional vectors. However, fixed vector dimensions lead to significant information loss, resulting in both false positives and false negatives. Deep learning-based methods address this issue. [11-14]Extracting node features from a source graph using graph neural networks can be further divided into supervised learning and unsupervised learning. Among these, supervised methods...
[11] It can detect graph-level or node-level anomalies, but requires a large amount of labeled attack data (which is usually scarce), and lacks robustness to zero-day vulnerabilities; unsupervised methods [12-13] Anomaly detection is achieved by modeling normal behavior, but most methods rely on static source graphs, making it difficult to capture subtle dynamic changes. To address this, researchers have proposed a novel method based on time-based source graph modeling. [14-16] This type of method can capture continuous and subtle temporal changes in the source map more comprehensively and accurately, thereby improving the APT detection effect.
[0006] Although existing APT detection methods based on source graphs have made some progress, they still have limitations. First, while some studies model source graphs as continuous-time dynamic graphs, they rely solely on graph autoencoders to achieve spatiotemporal information fusion. Because graph autoencoders have a fixed encoding method, they lack flexibility and struggle to adapt to complex and ever-changing APT attack scenarios, leading to decreased generalization ability and reduced robustness under different attack conditions. Second, existing methods fail to adequately sample and deeply analyze source graphs from both temporal and structural perspectives. APT attacks are often accompanied by abnormal fluctuations and structural abrupt changes in the time series of source graphs. However, limited by sampling strategies, existing methods may miss key time points or structural details, failing to accurately locate critical time points and core impact areas of the attack, resulting in a lack of agility in quickly responding to attack behaviors.
[0007] To address this, this invention proposes a multi-view contrastive APT attack detection method based on continuous-time dynamic graphs, used to identify unknown, covert, and complex fine-grained attack behaviors. This model employs a self-supervised contrastive learning framework to improve detection robustness and performs thorough sampling from both temporal and structural perspectives to capture as many important changes as possible in the source graph, thereby enhancing detection agility. By performing contrastive learning on the dual-view subgraphs, this invention helps learn more generalized representations, further enhancing detection robustness.
[0008] References: [1]M. Zipperle, F. Gottwalt, E. Chang, and T. Dillon, “Provenance-based intrusion detection systems: A survey,” ACM Computing Surveys, vol. 55, no. 7, pp. 1–36, 2022. [2]M. Zhong, M. Lin, C. Zhang, and Z. Xu, “A survey on graph neuralnetworks for intrusion detection systems: Methods, trends and challenges,”Computers & Security, p. 103821, 2024. [3] H. Dornhackl, K. Kadletz, R. Luh, and P. Tavolato, “Maliciousbehavior patterns,” in 2014 IEEE 8th international symposium on serviceoriented system engineering. IEEE, 2014, pp. 384–389. [4]M. Wagner, F. Fischer, R. Luh, A. Haberson, A. Rind, D. A. Keim,and W. Aigner, “A survey of visualization systems for malware analysis,”2015. [5]J. Ren and R. Geng, “Provenance-based apt campaigns detection viamasked graph representation learning,” Computers & Security, vol. 148, p.104159, 2025. [6]M. N. Hossain, S. M. Milajerdi, J. Wang, B. Eshete, R. Gjomemo, R.Sekar, S. Stoller, and V. Venkatakrishnan, “Sleuth: Real-time attack scenarioreconstruction from cots audit data,” in 26th USENIX Security Symposium(USENIX Security 17), 2017, pp. 487–504. [7]S. M. Milajerdi, R. Gjomemo, B. Eshete, R. Sekar, and V.Venkatakrishnan, “Holmes: real-time apt detection through correlation ofsuspicious information flows,” in 2019 IEEE Symposium on Security and Privacy(SP). IEEE, 2019, pp. 1137–1152. [8]W. U. Hassan, S. Guo, D. Li, Z. Chen, K. Jee, Z. Li, and A. Bates,“Nodoze: Combatting threat alert fatigue with automated provenance triage,”in network and distributed systems security symposium, 2019. [9]Y. Liu, M. Zhang, D. Li, K. Jee, Z. Li, Z. Wu, J. Rhee, and P.Mittal, “Towards a timely causality analysis for enterprise security.” inNDSS, vol. 24, 2018, p. 141.
[10] X. Han, T. Pasquier, A. Bates, J. Mickens, and M. Seltzer,“Unicorn: Runtime provenance-based detector for advanced persistent threats,”arXiv preprint arXiv:2001.01525, 2020.
[11] T. Chen, C. Dong, M. Lv, Q. Song, H. Liu, T. Zhu, K. Xu, L. Chen,S. Ji, and Y. Fan, “Apt-kgl: An intelligent apt detection system based onthreat knowledge and heterogeneous provenance graph learning,” IEEETransactions on Dependable and Secure Computing, 2022.
[12] S. Wang, Z. Wang, T. Zhou, H. Sun, X. Yin, D. Han, H. Zhang, X.Shi, and J. Yang, “Threatrace: Detecting and tracing host-based threats innode level through provenance graph learning,” IEEE Transactions onInformation Forensics and Security, vol. 17, pp. 3972–3987, 2022
[13] Z. Jia, Y. Xiong, Y. Nan, Y. Zhang, J. Zhao, and M. Wen, “MAGIC:Detecting advanced persistent threats via masked graph representationlearning,” in 33rd USENIX Security Symposium (USENIX Security 24), 2024, pp.5197–5214.
[14] MU Rehman, H. Ahmadi, and WU Hassan, “Flash: Acomprehensive approach to intrusion detection via provenance graphrepresentation learning,” in 2024 IEEE Symposium on Security and Privacy(SP). IEEE Computer Society, 2024, pp. 139–139.
[15] Z. Cheng, Q. Lv, J. Liang, Y. Wang, D. Sun, T. Pasquier, and X.Han, “Kairos: Practical intrusion detection and investigation using whole-system provenance,” in 2024 IEEE Symposium on Security and Privacy (SP), 2024, pp. 3533–3551.
[16] L. Wang, L. Fang, and Y. Hu, “A dynamic provenance graph-based detector for advanced persistent threats,” Expert Systems with Applications, vol. 265, p. 125877, 2025. Summary of the Invention
[0009] In view of the above-mentioned deficiencies of the prior art, the present invention at least solves the following technical problems: 1. Poor adaptability of fixed encoding methods and weak robustness to unknown APT attacks: Existing APT detection methods based on continuous-time dynamic source graphs rely solely on fixed graph autoencoders to achieve spatiotemporal information fusion. This lack of flexibility makes them ill-suited for complex and ever-changing APT attack scenarios, resulting in poor model generalization ability and weak robustness to unknown attacks. This is a core and critical problem in the field of APT detection that urgently needs to be solved, directly impacting the reliability of detection technology in real-world attack scenarios. 2. Blind spots exist in the single sampling perspective, resulting in insufficient APT detection agility: Existing APT detection methods do not fully sample the source map from both temporal and structural perspectives. Exploring only from a single perspective can easily miss key time nodes or structural details, making it impossible to accurately locate key attack time points and core affected areas. This leads to insufficient detection agility and can easily expand the scope of attack impact due to response delays.
[0010] This invention discloses an agile and robust APT detection method based on multi-view contrastive learning of source graphs, the method comprising the following steps: S1: Obtain the kernel audit logs of the target system, filter core fields, and filter invalid logs; S2: Construct a continuous-time dynamic traceability graph based on the core fields, and divide the time window into lengths. A continuous interval of minutes; S3: Using an η-BFS sampler and - The DFS sampler performs hybrid subgraph sampling on the continuous-time dynamic source graph to obtain a time subgraph and a structure subgraph, both of which contain positive sample subgraphs and negative sample subgraphs; S4: Perform time comparison training on the time subgraph and structural comparison training on the structure subgraph to obtain a generalized source graph representation through multi-view comparison learning; S5: Based on the generalized source graph representation, calculate the edge reconstruction loss and node deviation score, construct a queue of suspicious nodes and calculate the queue anomaly score. When the queue anomaly score exceeds a preset threshold, it is determined that there is an APT attack. Furthermore, the construction of the continuous-time dynamic source graph in step S2 specifically includes: Define the node type as process, file, or socket, and the edge type as process-file read / write event or process-socket connection event; Hierarchical feature hashing is used to encode node attributes. File nodes and socket byte points are split into substrings respectively. The substrings of node attributes are mapped to a low-dimensional feature space through a hash function. The node attribute vector is obtained by summing the feature vectors of the substrings. Furthermore, the sampling of the η-BFS sampler in step S3 specifically includes: Using the core node in the continuous-time dynamic source graph as the root node, extract the first-order neighbor set of the root node. With event time set ; The time-normalized value is calculated using the formula:
[0011] in, For neighboring nodes The time of the event For the current time, For time window Inner root node The minimum timestamp of the associated event; The sampling probability of neighboring nodes is calculated based on the chronological probability function, expressed by the formula:
[0012] in, Given the temperature parameter, random sampling is used to generate time-positive sample subplots based on this probability; The inverse-time normalized value is calculated using the following formula: ; The sampling probability is calculated based on the inverse chronological probability function, expressed by the formula:
[0013] in, Given the temperature parameter, random sampling is performed according to this probability to generate time-negative sample subplots; Furthermore, the steps described in step S3 - The sampling of the DFS sampler specifically includes: Sort the first-order neighbors of the root node in ascending order of timestamps and filter out a preset number of recently interacted nodes; Explore the node interaction path in depth, and prioritize the generation of positive sample subgraphs with more recent timestamps; Randomly select a non-root node as the new root node and perform the above operation to generate a structural negative sample subgraph. Furthermore, the multi-view comparison learning described in step S4 specifically includes: The Readout mean pooling function is used to aggregate the node states of the temporal subgraph and the structural subgraph into a subgraph embedding; Temporal comparison training uses the InfoNCE loss function to focus on short-term time dependence and fluctuations, making the similarity between recent positive sample subgraph embeddings and root node embeddings higher than that between outdated negative sample subgraph embeddings; structural comparison training uses the InfoNCE loss function to monitor long-term topological deviations, making the similarity between target node subgraph embeddings and root node embeddings higher than that between random node subgraph embeddings. Furthermore, the time-contrast training employs the InfoNCE loss function, expressed by the formula:
[0014] in, For time window The set of nodes within, root node In the time window Embedded, For positive sample subplots within the time window Embedded For time negative sample subplots in the time window Embedding; The structural contrastive training uses the InfoNCE loss function, expressed by the formula:
[0015] in, For time window The set of nodes within, root node In the time window Embedded, For the structure of the positive sample subgraph in the time window Embedded, For the negative sample subplot of the structure in the time window Embedding; Specifically, this invention addresses the suddenness and stealth of APT attacks by establishing a collaborative logic in the technical chain design of "source graph modeling - subgraph sampling - comparative training": First, the source graph is modeled as a continuous-time dynamic graph, and node attributes are comprehensively encoded and contextual information is preserved through multi-dimensional feature extraction methods such as hierarchical feature hashing, laying the foundation for subsequent accurate sampling and feature learning; Second, a breadth-first search (BFS) sampler (i.e., η-BFS sampler) and a depth-first search (DFS) sampler (i.e., ... -DFS Sampler): BFS focuses on the direct neighbors of nodes, capturing close causal relationships and current key factors in a short period of time through time-aware probability priority, avoiding the omission of short-term attack traces; DFS explores deeply along the node interaction path, filtering recent nodes in chronological order, revealing long-term structural changes and long-distance path evolution, and fully presenting the life cycle of the attack event. The two work together to cover the time anomaly fluctuations and structural mutations in the source graph as much as possible, significantly improving detection agility; finally, based on the time subgraph and structure subgraph obtained by the above sampling, time evolution patterns and structural patterns are captured through time comparison training and structure comparison training, respectively, so that the model can not only identify attack signs in short-term time fluctuations, but also monitor anomalies in long-term topological deviations, further strengthening the ability to identify APT attacks, providing a more generalizable graph representation foundation for subsequent pre-training and anomaly detection, and ensuring that it can still maintain excellent performance in complex adversarial scenarios.
[0016] Furthermore, it also includes a pre-training step for edge type prediction, specifically including: Embed the two-end nodes of the edges in the continuous-time dynamic tracing graph, and output the edge type prediction probability through MLP and sigmoid function; The cross-entropy loss for edge type prediction is expressed by the formula:
[0017] in, ,and Indicates time Time node and nodes Does a certain type of edge exist between them? For time window The set of edges inside; The total pre-training loss is expressed by the formula:
[0018] in, To compare training loss over time, For structural contrast training loss, For cross-entropy loss, The weighted coefficients are used to determine the optimal values through grid search. Furthermore, the determination of the existence of an APT attack in step S5 specifically includes: The edge reconstruction loss is the squared L2 norm of the original feature vector and the reconstructed feature vector of the edge, and its decision threshold is set as the mean of all edge reconstruction losses within the corresponding time window. Node deviation score according to formula Calculation, where This represents the total number of nodes within the time window. For nodes The number of associated nodes; Queue anomalies are categorized as the product of the mean edge reconstruction losses for each time window within the suspicious node queue, expressed by the formula:
[0019] in, For window The mean of the edge reconstruction loss, For queue The number of time windows included, when If the threshold for anomalies is exceeded, an APT attack is detected. Furthermore, the core fields mentioned in step S1 include entity ID, event type, timestamp, and attribute information; Furthermore, in the continuous-time dynamic tracing graph described in step S2, nodes represent system entities, including process nodes, file nodes, and socket byte nodes; edges represent interaction events between system entities, including process-file read edges, process-file write edges, and process-socket interaction edges.
[0020] This invention achieves at least the following technical effects: 1. By combining a multi-view contrastive learning framework with edge type prediction pre-training, this approach overcomes the rigid limitations of traditional fixed graph autoencoders and effectively adapts to complex and ever-changing APT attack scenarios. On the StreamSpot dataset, it achieves a detection accuracy and F1 score of 1.0; on the DARPA E3-CADETS dataset, it achieves an accuracy of 1.0, precision of 0.9997, and recall of 0.9983; and on the DARPA E3-THEIA dataset, it achieves an accuracy of 1.0 and an F1 score of 0.9956. These performances significantly outperform HOLMES (CADETS dataset F1 score 0.47), Unicorn (THEIA dataset accuracy 0.83), and FLASH (accuracy drops by 30%+ in adversarial scenarios). Furthermore, in adversarial scenarios with 1-16 benign events injected, the anomaly score remains stable, demonstrating outstanding robustness against unknown APT attacks. 2. Through η-BFS and - The DFS hybrid subgraph sampling strategy collaboratively covers "short-term to long-term" and "local to global" source graph information, while η-BFS quickly captures short-term causal relationships. -DFS reveals long-term structural changes and can accurately locate key attack time points and core impact areas, avoiding information omissions caused by a single sampling perspective. Ablation experiments verify that removing this sampling strategy will lead to a significant decrease in attack tracking accuracy, effectively shortening the APT attack response time. 3. Employing a self-supervised mode, pre-training can be completed solely based on benign system logs, eliminating the need for scarce labeled attack samples; it supports dual-mode deployment in both cloud and local environments, can interface with SIEM systems or integrate into existing IDS / IPS systems without requiring reconstruction of the original security architecture, and the deployment cycle is controlled within 1-2 weeks; it can run stably on CentOS 7.9 system (Intel Xeon E5-2630 v4 CPU, 12 cores 2.2GHz), without requiring a dedicated GPU, significantly reducing enterprise deployment and maintenance costs; 4. This invention eliminates the need for customized coding methods for different attack scenarios, reducing model iteration and long-term maintenance costs; it also eliminates the need to deploy multiple sampling tools, reducing hardware and software resource consumption and simplifying the detection system architecture. 5. This invention enhances the ability to identify covert and unknown APT attacks, effectively protecting the core systems and sensitive data security of key sectors such as government, military, and finance, and reducing the risk of major security incidents; it shortens the response time to APT attacks, reducing economic losses and social impact caused by attacks; at the same time, it can assist in locating key attack nodes and paths, providing support for post-attack tracing and defense strategy optimization, and making up for the shortcomings of traditional detection methods that only issue alarms and are difficult to trace the source.
[0021] This invention addresses the core need for detecting Advanced Persistent Threats (APTs) in cyberspace, focusing on revealing covert, long-term, and complex APT attack behaviors through continuous-time dynamic attribution graph modeling and multi-view comparative learning. The solution combines publicly available authoritative datasets (StreamSpot, DARPA E3-CADETS, DARPA E3-THEIA) with dynamic graph learning technology to construct an APT detection mechanism centered on "continuous-time dynamic attribution graph construction, hybrid subgraph sampling, and multi-view comparative learning," accurately identifying APT attacks with covert and multi-stage characteristics. This solution boasts advantages such as strong detection robustness, high agility, and wide scenario adaptability, providing crucial support for cybersecurity protection in key sectors such as government, military, and finance. It possesses significant practical value and promising prospects for widespread adoption, specifically in the following aspects: 1. Technological advantages 1.1 Multi-view Comparative Learning Based on Continuous-Time Dynamic Source Graph This scheme proposes a multi-view comparative learning method based on continuous-time dynamic source graph, overcoming the rigid limitations of traditional APT detection methods that rely on fixed graph autoencoders. First, the system log is modeled as a continuous-time dynamic source graph (nodes represent entities such as files, processes, and sockets, while edges represent events such as read / write and interaction). Hierarchical feature hashing technology is used to encode node attributes (such as file paths and IP addresses) in multiple dimensions, comprehensively preserving entity semantic information. Based on this, temporal comparison training and structural comparison training are conducted for sampled subgraphs—temporal comparison focuses on short-term temporal dependencies and fluctuations (comparing recent positive sample subgraphs with outdated negative sample subgraphs), while structural comparison monitors long-term topological deviations (comparing DFS subgraphs of target nodes and random nodes), learning a highly generalizable graph representation in a self-supervised mode. This approach effectively avoids the problem of fixed encoding methods being difficult to adapt to complex APT scenarios, while also addressing the weakness of traditional methods in robustness to unknown attacks, providing a more reliable feature foundation for APT detection.
[0022] 1.2η-BFS and To improve the agility of APT detection, this scheme designs η-BFS (time-aware) and hybrid subgraph sampling in collaboration with DFS. - A hybrid subgraph sampling strategy that combines structure-aware (DFS) and η-BFS sampling. The η-BFS sampler uses a time-aware probability function to focus on the direct neighbors and nearest neighbor regions of nodes, quickly capturing causal relationships and key attack factors within a short period of time, avoiding the omission of short-term attack traces; The DFS sampler filters recent interaction nodes chronologically, explores along the path in depth, reveals long-term structural changes and the evolution of long-distance attack paths, and presents a complete lifecycle of attack events. Together, they cover "short-term to long-term" and "local to global" source map information, effectively avoiding the omission of key time nodes or structural details caused by a single sampling perspective, improving attack localization efficiency, and providing support for rapid response to APT attacks.
[0023] 2. Performance indicators
[0024] This approach was evaluated on publicly available authoritative datasets, demonstrating significantly superior detection accuracy and stability compared to existing mainstream methods: On the StreamSpot dataset, it achieves an accuracy of 1.0 and an F1 score of 1.0; on the DARPAE3-CADETS dataset, it achieves an accuracy of 1.0, a precision of 0.9997, and a recall of 0.9983; and on the DARPAE3-THEIA dataset, it achieves an accuracy of 1.0 and an F1 score of 0.9956. Comparative experiments show that its performance far surpasses HOLMES (CADETS dataset F1 score 0.47), Unicorn (THEIA dataset accuracy 0.83), and FLASH (accuracy drops by over 30% in adversarial scenarios). Furthermore, ablation experiments verify that removing the contrastive learning module reduces the accuracy on the CATETS dataset from 1.0 to 0.4545, further demonstrating the effectiveness of the core mechanism.
[0025] 3. Production Implementation Aspects
[0026] 3.1 System Architecture This solution adopts a highly modular development model to ensure high system availability and scalability. Core modules include: ① Source graph construction module (log parsing, node attribute encoding, dynamic graph generation); ② Hybrid subgraph sampling module (η-BFS sampling, ... - DFS sampling and subgraph selection); ③ Multi-view comparison learning module (time comparison training, structure comparison training, edge type prediction pre-training); ④ Anomaly detection module (edge reconstruction loss calculation, node deviation scoring, time window-level anomaly judgment). The modules are loosely coupled to facilitate subsequent function iteration and customization.
[0027] 3.2 Model training was performed using publicly available authoritative datasets (StreamSpot, DARPA E3-CADETS, DARPA E3-THEIA): a self-supervised mode was adopted, relying solely on benign system logs to complete pre-training (without the need for scarce labeled attack samples); hyperparameters were determined to be optimally configured through grid search, improving the model's adaptability to different APT scenarios.
[0028] 3.3 The system supports both cloud and local deployment modes to meet the needs of different scenarios: cloud deployment can connect to enterprise cloud security platforms to achieve large-scale cross-domain APT detection; local deployment is suitable for scenarios such as internal core servers and classified networks, ensuring data privacy and security. Furthermore, the system provides standard API interfaces, allowing direct integration into existing network defense systems (such as IDS / IPS systems) without reconstructing the original security architecture, and the deployment cycle can be controlled within 1-2 weeks; hardware requirements are low, requiring only CentOS 7.9 system (Intel Xeon E5-2630 v4 CPU, 12 cores 2.2GHz) for stable operation, eliminating the need for a dedicated GPU and reducing enterprise deployment costs. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating the agile and robust APT detection method based on multi-view comparative learning of source graphs according to the present invention. Figure 2 This is a schematic diagram showing the verification (ablation experiment) results of the core innovation of the agile and robust APT detection method based on source graph multi-view comparative learning in this invention; Figure 3 This is a schematic diagram showing the robustness verification results (comparison of abnormal scores in a benign event injection scenario) of the agile and robust APT detection method based on source graph multi-view comparative learning of the present invention. Detailed Implementation
[0030] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0031] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components is appropriately exaggerated in the drawings.
[0032] This invention proposes an agile and robust APT detection method based on multi-view contrastive learning of source graphs, such as... Figure 1As shown, the method includes the following steps: S1: Obtain the kernel audit logs of the target system, filter core fields, and filter invalid logs; S2: Construct a continuous-time dynamic traceability graph based on the core fields, and divide the time window into continuous intervals of length tw minutes; S3: Using an η-BFS sampler and - The DFS sampler performs hybrid subgraph sampling on the continuous-time dynamic source graph to obtain a time subgraph and a structure subgraph, both of which contain positive sample subgraphs and negative sample subgraphs; S4: Perform time comparison training on the time subgraph and structural comparison training on the structure subgraph to obtain a generalized source graph representation through multi-view comparison learning; S5: Based on the generalized source graph representation, calculate the edge reconstruction loss and node deviation score, construct a queue of suspicious nodes and calculate the queue anomaly score. When the queue anomaly score exceeds a preset threshold, it is determined that there is an APT attack.
[0033] In specific implementations, the study targets two scenarios: controlled laboratory simulations and real enterprise networks. It focuses on detecting stealthy download attacks and multi-vulnerability APT attacks, corresponding to the core scenarios of the StreamSpot and DARPAE3 datasets. In the laboratory environment, the attacks cover five benign behaviors: viewing Gmail, browsing CNN.com, downloading files, watching YouTube videos, and playing video games, as well as malicious URL attacks exploiting Flash vulnerabilities to gain root privileges. In the enterprise network environment, the attacks involve red teams exploiting vulnerabilities in web servers, mail servers, and SSH servers to steal sensitive information, while simultaneously incorporating benign behaviors such as web browsing, email viewing, and SSH logins to conceal their activities. Blue teams, on the other hand, use source capture systems such as CADETS and THEIA to record all system activity. The study requires accurate differentiation between malicious and normal activities in different scenarios to verify the detection capabilities of the method against single attack patterns and complex, covert APT attacks. The specific detection methods include: 1. Complete steps and functions
[0034] 1.1 System Log Preprocessing and Construction of Continuous-Time Dynamic Source Graph
[0035] 1.1.1 Extract kernel audit logs from the dataset and filter core fields: entity ID (process PID, file path, socket IP:port), event type (read / write / connect / fork), timestamp (accurate to second), attribute information (process command line, file path, IP address), and filter out invalid logs such as duplicates and empty attributes; 1.1.2 Define node and edge types: Nodes are divided into subject (processes, such as nginx, bash), fileObject (files, such as / etc / passwd, / var / log / nginx.log), and netflowObject (sockets, such as 81.49.200.166:44623); edges are represented by triples (source node ID, target node ID, timestamp), corresponding to interactions such as process-file (read / write) and process-socket (connect); 1.1.3 Hierarchical Feature Hash Encoding of Node Attributes (Innovation Supporting Steps): Substrings are split for file nodes and socket byte nodes; characters are mapped to a low-dimensional feature space using a hash function h(s_j). Mapped to {±1}, calculate the substring feature vector. The final node attribute vector ; 1.1.4 Dividing the time window: Set the window length to tw minutes (experimental optimization parameter) and divide the data into multiple continuous windows.
[0036] The purpose of this step is to transform unstructured logs into structured, time-series dynamic source graphs. Hierarchical feature hashing enables low-dimensional lossless encoding of high-dimensional attributes, providing a high-quality data foundation for subsequent sampling and feature learning, and solving the problem that traditional static graphs cannot capture temporal dynamics.
[0037] 1.2 η-BFS and -DFS hybrid subgraph sampling
[0038] 1.2.1 η-BFS Sampling (Time-Aware): Select the core node of the window (such as the Nginx process, sensitive file access node) as the root node, and extract the first-order neighbor set of the root node. With event time set ; Calculate the sampling probability using the chronological probability function:
[0039] The above formula This is the time-aware probability function of the η-BFS sampler. By using this function to prioritize the sampling weights of recently interacting nodes, the causal relationships within a short period of time can be accurately captured. in, For neighboring nodes The time of the event For the current time, For time window Inner root node Minimum timestamp of related events Given the temperature parameter, a time-positive sample subplot is generated by randomly sampling according to this probability.
[0040] Randomly sample multiple first-order neighbors; recursively sample to generate a second-order positive sample subgraph. ; Negative sample subgraphs are generated based on inverse chronological probability sampling. The calculation method is as follows:
[0041] in, Given the temperature parameter, a time-negative sample subplot is generated by randomly sampling according to this probability.
[0042] 1.2.2 -DFS sampling (structure-aware): Sort the first-order neighbors of the root node in ascending order of timestamps and filter the nodes with the most recent interactions; explore the depth along the node path, prioritizing edges with more recent timestamps to generate a second-order positive sample subgraph. Randomly select non-root nodes implement -DFS sampling generates negative sample subgraphs Each root node generates a corresponding... , One set of each.
[0043] η-BFS focuses on the nearest neighbor region of a node and captures short-term causal relationships by sampling recently interacting nodes with time-aware probability priority. - DFS explores along the path in depth, combined with time sorting to filter key nodes and reveal long-term structural changes; the two work together to cover "short-term-long-term" and "local-global" information; it solves the problem of missing key time nodes or structural details in existing single sampling, and provides comprehensive subgraph samples for subsequent comparative learning.
[0044] 1.3 Multi-view comparison learning
[0045] 1.3.1 Time-based comparison training: Using the Readout function (mean pooling) to... Node status Aggregation into subgraph embeddings ; Then, calculate the InfoNCE loss:
[0046] in, For time window The set of nodes within, root node In the time window Embedded, For positive sample subplots within the time window Embedded For time negative sample subplots in the time window Embedding; 1.3.2 Structural Comparison Training: Similarity Convergence get ; Then, calculate the InfoNCE loss:
[0047] in, For time window The set of nodes within, For root node i in the time window Embedded, For the structure of the positive sample subgraph in the time window Embedded, For the negative sample subplot of the structure in the time window Embedded.
[0048] Temporal comparison enhances the capture of short-term temporal fluctuations by learning the difference between positive and negative subgraph embeddings; structural comparison focuses on topological deviations, learning long-term structural features through the difference between the subgraph embeddings of target nodes and random nodes; both are optimized based on InfoNCE loss, making positive sample embeddings more similar and negative sample embeddings more distant. This approach overcomes the rigid limitations of traditional fixed graph autoencoders, learns graph representations with strong generalization capabilities, and improves robustness against unknown attacks.
[0049] 1.4 Self-supervised pre-training and anomaly detection
[0050] 1.4.1 Edge type prediction pre-training: 1) Opposite side splicing node embedding Probability prediction using MLP and sigmoid output edge type ; 2) Calculate the cross-entropy loss:
[0051] in, ,and Indicates time Time node and nodes Does a certain type of edge exist between them? For time window The set of edges inside.
[0052] 3) Calculate the total pre-training loss:
[0053] in, To compare training loss over time, For structural contrast training loss, For cross-entropy loss, The weighting coefficients are used to determine the optimal values through grid search.
[0054] 1.4.2 Anomaly Detection: 1) Calculate the edge reconstruction loss The threshold is set within the window. Mean; 2) Calculate the node deviation score ; 3) Construct a queue of suspicious nodes Q, and calculate the queue anomaly score:
[0055] in, For window The mean of R(e), For queue If the score of the included time window exceeds the threshold, it is considered an APT attack.
[0056] Edge type prediction pre-training improves the quality of graph representation, and anomaly detection combines edge reconstruction loss and node deviation score to realize a complete detection process of "fine-grained recognition-queue aggregation judgment", which solves the problem of high false alarm rate of traditional unsupervised methods.
[0057] This invention uses the StreamSpot and DARPAE3 datasets to validate the method: StreamSpot is a simulation dataset publicly released by the research team of the same name, collected in a controlled laboratory environment using the SystemTap auditing system. The dataset contains 600 source images from six activity scenarios. Five of these are benign scenarios, covering activities such as checking Gmail emails, browsing CNN.com, downloading files, watching YouTube videos, and playing video games; the remaining one is an attack scenario simulating a "drive-by-download attack"—in which a malicious URL exploits a Flash vulnerability, allowing the attacker to gain root access to the host. The DARPA E3 dataset is a core component of the DARPA Transparent Computing program, collected from adversarial interaction scenarios within enterprise network environments. These scenarios primarily involve two teams: the Red Team and the Blue Team. The Red Team's mission is to launch Advanced Persistent Threat (APT) attacks: they exploit various vulnerabilities in the network to steal sensitive information from within the enterprise, targeting security-critical services such as web servers, mail servers, and SSH servers. Simultaneously, to conceal their activities and reduce the probability of detection, the Red Team also performs seemingly benign activities such as browsing the web, checking emails, and SSH logins. The Blue Team is responsible for recording all host activity across the entire system to identify such APT attacks. They deploy multiple source capture systems (including CADETS and THEIA) on different platforms to comprehensively record all activities occurring within the network.
[0058] This invention compares with six advanced source-information-based APT detection baseline methods, covering both statistical and deep learning approaches. Among them, HOLMES defines rule-matching detection based on an APT lifecycle model; Unicorn utilizes historical information and summarization techniques to model normal behavior; Threatrace uses GraphSAGE to learn node representations and identify biases; FLASH improves detection performance by integrating multiple techniques; KAIROS employs a novel graph neural network architecture to address specific challenges; and STGAN combines spatiotemporal graphs and attention mechanisms to cope with complex attacks.
[0059] This invention, referencing existing research evaluation criteria, first clarifies the basic definitions of graph classification results: successfully identified attack graphs are defined as true positives (TP), and undetected attack graphs are defined as false negatives (FN); correctly identified non-attack graphs are defined as true negatives (TN), and misidentified benign graphs are defined as false positives (FP). Based on this, the study calculates five evaluation metrics: accuracy, precision, recall, F1 score, and AUC. This approach maintains consistency with existing mature methods and provides a clear analytical framework for evaluating model performance.
[0060] This invention is based on CentOS 7.9 system, Intel Xeon E5-2630 v4 CPU (12 cores 2.2GHz), PyTorch 2.0 framework, and has no GPU dependency. To comprehensively evaluate the technical advancement and practical value of this invention, the following verification will be conducted from three dimensions: overall detection capability, core innovative value, and anti-interference capability. The positioning and objectives of each verification module are as follows: 1. Validation Experiment: Verifying the overall detection capability of the ARAD method The effectiveness experiments are fundamental verifications to evaluate whether the ARAD method can solve the core problem of APT detection. They aim to verify ARAD's comprehensive detection performance in real-world attack scenarios by comparing it with various baseline methods across multiple scenario datasets. The experiments select three authoritative public datasets: StreamSpot, DARPA E3-CADETS, and DARPA E3-THEIA, covering typical APT scenarios such as browser plugin attacks and server vulnerability attacks, as well as the two mainstream operating system environments: FreeBSD and Ubuntu. Simultaneously, six mainstream baseline methods—rule-based, statistical, and deep learning-based—are introduced to comprehensively measure ARAD's performance in the detection phase, focusing on core metrics such as accuracy, precision, recall, F1 score, and AUC. This provides an overall performance benchmark for subsequent innovation point verification and robustness verification.
[0061] Table 1: Comparison of Experimental Performance of This Method with Existing Baseline Methods
[0062] The experimental results are shown in Table 1 above. ARAD performs exceptionally well across all evaluation metrics on all datasets of StreamSpot, CADETS, and THEIA, consistently achieving superior scores. This fully demonstrates the effectiveness and superiority of our approach of treating the source graph as a continuous-time dynamic graph. ARAD, leveraging a multi-view contrastive learning paradigm, generates subgraphs through BFS and DFS sampling, capturing long- and short-term causal dependencies on the source graph. Furthermore, by utilizing structure and temporal contrastive learning, it successfully overcomes the shortcomings of existing methods in terms of insufficient robustness of graph structure when detecting unknown APT attacks in real-world environments based on full log files. Simultaneously, by employing temporal and structure samplers, it accurately characterizes the evolution of events or states on the source graph, improving detection agility.
[0063] 2. Innovation Point Validation: Verify the necessity and effectiveness of the core innovation points.
[0064] Innovation validation is crucial for revealing the source of ARAD's technological advantages. This validation focuses on ARAD's two core innovations—η-BFS and... - DFS hybrid subgraph sampling and temporal-structural multi-view comparative learning, through ablation experiments (removing innovative points one by one and comparing performance changes), verify the contribution of each innovative point to the detection effect. The experiments use the DARPA E3-CADETS dataset as a benchmark to quantitatively analyze the necessity of a single innovative point and the synergistic effect of multiple innovative points, clarifying ARAD's technical breakthrough in solving the sampling blind spots of existing methods and overcoming the limitations of fixed encoding generalization, supporting the core argument for the method's creativity.
[0065] The experimental results are as follows Figure 2 As shown, (a) and (b) correspond to the precision and recall of the CADETS-E3 dataset, respectively, while (c) and (d) correspond to the precision and F1 score of the THEIA-E3 dataset, respectively. To further explore the impact of each component on the overall performance, this invention conducts ablation studies on the DARPA CADETS-E3 and THEIA-E3 datasets, removing the contrastive learning module, temporal contrastive analysis, and structural contrastive analysis to analyze their roles in APT detection. Contrastive learning extracts short- and long-term information from the source graph through temporal and structural samplers, generating positive and negative subgraphs, and then making the positive subgraph representation more similar to the original representation and the negative subgraph representation more distant. After removing the contrastive learning module, the model continuously declined on all evaluation metrics of both datasets. Taking the CADETS-E3 dataset as an example, the precision dropped from 1 to 0.4545, and the recall dropped from 1 to 0.8333. The metrics of the THEIA-E3 dataset also declined significantly, demonstrating that contrastive learning is crucial for enhancing the robustness and generalization of the representation. Temporal contrast is used to discover short-term causal dependencies in the temporal dimension of continuous-time dynamic graphs. Removing temporal contrast significantly degrades model performance across all datasets and metrics. Precision on the CADETS-E3 dataset drops from 1 to 0.9736, and the F1 score drops from 1 to 0.5882. A similar decline is observed on the THEIA-E3 dataset, indicating that short-term causal dependencies are indispensable for extracting local semantic information from the source graph and revealing attack temporal evolution patterns. Structural contrast focuses on long-term graph features, monitoring subtle changes in the topology of the APT source graph in real time. Removing structural contrast significantly degrades model performance on both datasets. Precision on the THEIA-E3 dataset drops from 1 to 0.7, and recall drops from 1 to 0.875, demonstrating the crucial importance of long-term features in identifying long-term temporal relationships and potential patterns, ensuring robustness and agility in APT attack detection.
[0066] 3. Robustness Verification: Verify ARAD's ability to resist interference scenarios.
[0067] Robustness verification is a crucial step in evaluating the reliability of ARAD in real-world complex environments. It aims to simulate mimicry attack scenarios commonly used by APT attackers, where benign events are injected to mask malicious behavior, thus verifying ARAD's ability to resist interference and maintain stable detection. Experiments are conducted using attack windows on the DARPA E3-CADETS dataset. Scenarios with varying interference intensities are constructed by injecting 1-16 benign events. Anomaly score stability is used as the core metric to compare the performance of ARAD with mainstream dynamic graph detection methods. This verifies ARAD's detection stability when malicious behavior is masked by benign behavior, ensuring the method can effectively identify threats even under the covert characteristics of real-world attacks, and providing experimental support for reliable industrial deployment.
[0068] Experimental results are as follows Figure 3As shown, to evaluate the robustness of ARAD, this invention conducted an adversarial attack experiment by adding benign features to the source map (simulating the coexistence of normal activities and potential threats in a real-world scenario) and compared the anomaly scores generated by FLASH and ARAD. The horizontal axis represents the number of added benign events, and the vertical axis represents the anomaly score. Blue dots represent FLASH, red dots represent ARAD, and a threshold line is also provided for reference. When the number of added benign events increased from 1 to 16, the anomaly score of FLASH fluctuated significantly: the score was high when a few benign events were added, but dropped sharply as the number increased, even falling below the threshold when a moderate number were added; while the anomaly score of ARAD remained stable, always well above the threshold regardless of the number of benign events added. It is evident that regardless of the number of simulated normal behaviors introduced, ARAD maintains a stable and high level of anomaly detection capability, demonstrating outstanding robustness; conversely, the anomaly score of FLASH decreases significantly with the increase of benign events, indicating a weakening detection performance. ARAD's superior robustness stems primarily from contrastive learning, which enables the model to better distinguish between real malicious patterns and simulated benign features, resist interference from added benign events, and maintain stable and reliable detection performance, significantly outperforming FLASH.
[0069] Experimental results on three different datasets demonstrate that this method can effectively identify APT attacks by accurately capturing the continuous temporal dynamics and complex structural details in the source graph, and outperforms existing state-of-the-art methods in terms of detection accuracy, agility, and robustness.
[0070] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. An agile and robust APT detection method based on multi-view comparative learning of source graphs, characterized in that, The method includes the following steps: S1: Obtain the kernel audit logs of the target system, filter core fields, and filter invalid logs; S2: Construct a continuous-time dynamic traceability graph based on the core fields, and divide the time window into lengths. A continuous interval of minutes; S3: Using an η-BFS sampler and - The DFS sampler performs hybrid subgraph sampling on the continuous-time dynamic source graph to obtain a time subgraph and a structure subgraph, both of which contain positive sample subgraphs and negative sample subgraphs; S4: Perform time comparison training on the time subgraph and structural comparison training on the structure subgraph to obtain a generalized source graph representation through multi-view comparison learning; S5: Based on the generalized source graph representation, calculate the edge reconstruction loss and node deviation score, construct a queue of suspicious nodes and calculate the queue anomaly score. When the queue anomaly score exceeds a preset threshold, it is determined that there is an APT attack.
2. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, The construction of the continuous-time dynamic source graph in step S2 specifically includes: Define the node type as process, file, or socket, and the edge type as process-file read / write event or process-socket connection event; Hierarchical feature hashing is used to encode node attributes. File nodes and socket byte points are split into substrings respectively. The substrings of node attributes are mapped to a low-dimensional feature space through a hash function. The node attribute vector is obtained by summing the feature vectors of the substrings.
3. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, The sampling of the η-BFS sampler in step S3 specifically includes: Using the core node in the continuous-time dynamic source graph as the root node, extract the first-order neighbor set of the root node. With event time set ; The time-normalized value is calculated using the formula: in, For neighboring nodes The time of the event For the current time, For time window Inner root node The minimum timestamp of the associated event; The sampling probability of neighboring nodes is calculated based on the chronological probability function, expressed by the formula: in, Given the temperature parameter, random sampling is used to generate time-positive sample subplots based on this probability; The inverse-time normalized value is calculated using the following formula: ; The sampling probability is calculated based on the inverse chronological probability function, expressed by the formula: in, Given the temperature parameter, a time-negative sample subplot is generated by randomly sampling according to this probability.
4. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, The step S3 described above - The sampling of the DFS sampler specifically includes: Sort the first-order neighbors of the root node in ascending order of timestamps and filter out a preset number of recently interacted nodes; Explore the node interaction path in depth, and prioritize the generation of positive sample subgraphs with more recent timestamps; Randomly select a non-root node as the new root node and perform the above operation to generate a structural negative sample subgraph.
5. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, The multi-view comparison learning described in step S4 specifically includes: The Readout mean pooling function is used to aggregate the node states of the temporal subgraph and the structural subgraph into a subgraph embedding; Temporal comparison training uses the InfoNCE loss function, focusing on short-term time dependence and fluctuations, making the similarity between recent positive sample subgraph embeddings and root node embeddings higher than that between outdated negative sample subgraph embeddings; structural comparison training uses the InfoNCE loss function, monitoring long-term topological deviations, making the similarity between target node subgraph embeddings and root node embeddings higher than that between random node subgraph embeddings.
6. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 5, characterized in that, The time-comparison training uses the InfoNCE loss function, expressed by the formula: in, For time window The set of nodes within, root node In the time window Embedded, For positive sample subplots within the time window Embedded For time negative sample subplots in the time window Embedding; The structural contrastive training uses the InfoNCE loss function, expressed by the formula: in, For time window The set of nodes within, root node In the time window Embedded, For the structure of the positive sample subgraph in the time window Embedded, The embedding of the negative sample subgraph of the structure within the time window t.
7. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, It also includes a pre-training step for edge type prediction, specifically including: Embed the two-end nodes of the edges in the continuous-time dynamic tracing graph, and output the edge type prediction probability through MLP and sigmoid function; The cross-entropy loss for edge type prediction is expressed by the formula: in, ,and Indicates time Time node and nodes Does a certain type of edge exist between them? The set of edges within the time window t; The total pre-training loss is expressed by the formula: in, To compare training loss over time, For structural contrast training loss, For cross-entropy loss, The weighting coefficients are used to determine the optimal values through grid search.
8. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, Step S5, which determines the existence of an APT attack, specifically includes: The edge reconstruction loss is the squared L2 norm of the original feature vector and the reconstructed feature vector of the edge, and its decision threshold is set as the mean of all edge reconstruction losses within the corresponding time window. Node deviation score according to formula Calculation, where This represents the total number of nodes within the time window. For nodes The number of associated nodes; Queue anomalies are categorized as the product of the mean edge reconstruction losses for each time window within the suspicious node queue, expressed by the formula: in, For window The mean of the edge reconstruction loss, For queue The number of time windows included, when If the threshold for anomalies is exceeded, an APT attack is detected.
9. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, The core fields mentioned in step S1 include entity ID, event type, timestamp, and attribute information.
10. The agile and robust APT detection method based on multi-view contrastive learning of source graphs as described in claim 1, characterized in that, In the continuous-time dynamic tracing graph described in step S2, nodes represent system entities, including process nodes, file nodes, and socket byte nodes; edges represent interaction events between system entities, including process-file read edges, process-file write edges, and process-socket interaction edges.