APT attack detection method fusing comparative learning and cross-domain recommendation
By integrating the methods of comparative learning and cross-domain recommendation, the APT attack detection model is constructed and optimized, and the problems of high false alarm rate and sparse data of APT attack detection in the existing technology are solved, and higher detection accuracy and model adaptability are achieved.
Patent Information
- Application Number
- CN202510118165.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has problems in APT attack detection with high false positive rate, time-consuming rule formulation and error prone, and insufficient training samples and diversity, resulting in poor detection performance.
Using the method of fusion contrast learning and cross-domain recommendation, a traceability map and two-part graph of the source domain and the target domain are constructed, and the r-ego network is generated, and the graph encoder is pre-trained using a self-supervised learning scheme, and information is transferred from the data-rich source domain to the target domain through cross-domain recommendation, and the matrix decomposition model is fine-tuned to identify potential attack threats.
It improves the accuracy and generalization ability of APT attack detection, reduces the false positive rate, enhances the adaptability of the model and the ability to perform against unseen system behaviors, and alleviates the problem of data sparseness.
Smart Images

Figure CN120185849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to APT attack detection and machine learning technology, and in particular to an APT attack detection method integrating contrastive learning and cross-domain recommendation. Background Art
[0002] The development of Internet information technology has also spawned many internal and external security risks and threats. In recent years, the threat of Advanced Persistent Threat (APT) attacks has been gradually escalating. With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, the situation of network attack and defense has become increasingly complex in the environment of human-machine-object integration, and the network attack surface has also expanded infinitely. APT attacks pose a huge threat to information security, mainly manifested in the theft of important data and the destruction of system integrity. Compared with traditional attack modes, APT attacks have the characteristics of long duration, long attack chain, high concealment, diverse means, and strong harmfulness, so they are more difficult to be detected by security methods.
[0003] In order to facilitate the investigation of APT attacks in large-scale host records, researchers use data tracing technology to guide APT attack detection by describing the tracing graph of system execution history. The detection based on this is mainly divided into three categories: statistical-based detection, rule-based detection, and learning-based detection. Although existing solutions have shown good detection performance, they still have some inherent shortcomings. Statistical-based detection is prone to produce a large number of false positives for rare but normal system activities; although rule-based detection is very effective in the face of known attacks, the formulation of such heuristic rules is often time-consuming and error-prone; although learning-based detection can more accurately identify hidden and fine-grained APT attack patterns, due to the low probability, difficulty in reproduction, and difficulty in labeling of complex network attacks, the number and diversity of relevant training samples are extremely low, which cannot support the training of attack detection models with high generalization capabilities, and most learning methods only generate detection signals at a coarse-grained level. Summary of the invention
[0004] The purpose of the present invention is to solve the problems existing in the prior art and propose an APT attack detection method that integrates contrastive learning and cross-domain recommendation, maps the network security concept of system entity context interaction to the recommendation concept of user-item interaction, and uses the side information of system entities to form high-order connectivity to predict the possibility of interaction between entities. At the same time, a cross-domain recommendation method is used to transfer information from the data-rich source domain to the target domain, improving the recommendation performance on the target domain while alleviating the problem of sparse APT attack data.
[0005] In order to achieve the above object, the technical solution provided by the present invention is:
[0006] An APT attack detection method that combines contrastive learning and cross-domain recommendation, including:
[0007] Construct a source domain traceability graph based on the source domain dataset, construct a source domain bipartite graph according to the source domain traceability graph, construct a target domain traceability graph based on the target domain dataset, and construct a target domain bipartite graph according to the target domain traceability graph;
[0008] Generate r-ego networks for all nodes in the source domain bipartite graph;
[0009] Perform two random walks on the r-ego network of a node in the source domain bipartite graph to generate the first subgraph positive pair g of the current node q And the second subgraph positive pair g k As positive samples, perform random walks on the r-ego networks of each other node except the current node to generate subgraphs of each other node as negative samples, and generate positive and negative samples corresponding to each node in the source domain bipartite graph;
[0010] Use the positive and negative samples corresponding to each node in the source domain bipartite graph to pre-train the first graph encoder and the second graph encoder, obtain the pre-trained first graph encoder, and transfer the pre-trained first graph encoder to the target domain;
[0011] Generate r-ego networks for all nodes in the target domain bipartite graph, and perform random walks on the r-ego networks of the nodes in each target domain bipartite graph to generate target domain node subgraphs of the nodes in each target domain bipartite graph;
[0012] Input the target domain node subgraph into the pre-trained first graph encoder to obtain the initial embedding of the nodes in each target domain bipartite graph;
[0013] Use the initial embedding to fine-tune the matrix factorization model;
[0014] Input the initial embedding of the nodes in each target domain bipartite graph into the fine-tuned matrix factorization model to obtain the final embedding of the nodes in each target domain bipartite graph;
[0015] Calculate the inner product according to the final embeddings of every two nodes, calculate the probability of interaction between two nodes according to the inner product. If the probability is greater than the preset threshold, it is considered that the two nodes interact and are marked as potential attack threats. If the probability is less than or equal to the preset threshold, it is considered that the two nodes do not interact.
[0016] Furthermore, the source domain dataset is the LANL dataset and the target domain dataset is the DAPRA dataset.
[0017] Furthermore, the generating r-ego networks for all nodes in the source domain bipartite graph includes:
[0018] Taking a node in the source domain bipartite graph as the central node, a set S of r-hop neighbor nodes of the central node is obtained v ={v': d(v, v') ≤ r}, where v represents the central node, v' represents the r-hop neighbor node of the central node, d(v, v') represents the shortest path distance between the central node v and the r-hop neighbor node v', and r ∈ [1, 2];
[0019] The subgraph generated according to the central node and the r-hop neighbor nodes is the r-ego network of the current node;
[0020] Taking each node in the source domain bipartite graph as the central node to generate an r-ego network respectively.
[0021] Furthermore, the pre-training of the first graph encoder and the second graph encoder by using the positive samples and negative samples corresponding to each node in the source domain bipartite graph includes:
[0022] Feeding the first subgraph positive pair g q into the first graph encoder f q for encoding to generate a first vector, feeding the second subgraph positive pair g k into the second graph encoder f k for encoding to generate a second vector, and feeding each negative sample into the second graph encoder f k for encoding to generate a third vector for each negative sample;
[0023] Using the contrastive loss InfoNCE to perform self-supervised optimization on the first graph encoder and the second graph encoder, which is expressed by the formula as follows:
[0024]
[0025] where L InFoNCE represents the contrastive loss InfoNCE, e q represents the first vector, represents the transpose of the first vector, e k represents the second vector, e i represents the third vector generated by the i-th negative sample, n represents the number of negative samples, and τ is the temperature hyperparameter.
[0026] Furthermore, the fine-tuning of the matrix factorization model by using the initialization embedding includes:
[0027] Feeding the initial embedding into the matrix factorization model;
[0028] Using the Bayesian personalized ranking loss function to optimize the matrix factorization model;
[0029] Obtaining the fine-tuned matrix factorization model.
[0030] Compared with the prior art, the remarkable advantages of the present invention are as follows: 1. The similarity between network threat detection and task recommendation is discovered, the network security concept of system entity context interaction is mapped to the recommendation concept of user-item interaction, and the side information of system entities is used to form high-order connectivity to predict their interaction. 2. A cross-domain recommendation method is used to transfer information from a data-rich source domain to a target domain, improving the recommendation performance on the target domain while alleviating the data sparsity problem of APT attack data. 3. A self-supervised learning scheme is adopted to train the graph encoder, reducing the prediction bias from the source domain and enhancing the generalization of the model. 4. A design result feedback mechanism is designed to improve the model adaptability and the performance ability for unseen system behaviors. Brief Description of the Drawings
[0031] Figure 1 It is a flowchart of an APT attack detection method integrating contrastive learning and cross-domain recommendation according to the present invention;
[0032] Figure 2 It is a schematic diagram of the recommendation model according to the present invention. Detailed Embodiments
[0033] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0034] The present invention reveals fine-grained network threats by determining the possibility of interaction between one entity and another entity in the system, and a similar problem has been explored in the field of recommendation, whose main goal is to predict the possibility of users consuming goods. The cross-domain recommendation method can transfer information from a data-rich network traffic dataset and a multi-source network security event dataset to the APT attack field with sparse target domain data, which can effectively make up for the problem of data sparsity and expand the recommendation scope. As Figure 1 shown, an APT attack detection method integrating contrastive learning and cross-domain recommendation includes the following steps:
[0035] (1) Graph construction based on log data: Graph construction is respectively performed on the datasets of the source domain and the target domain. First, the log data records are converted into a traceability graph, and a bipartite graph is extracted based on system entity interaction.
[0036] (1-1) Traceability graph construction: In the traceability graph, nodes represent system entities involved in log data, and edges represent information flows between system entities. Based on node types and edge types, a traceability graph G=(V, R) is constructed, where V is the set of nodes and R is the set of edges. The node types include processes, files, users, hosts, and sockets. Relationship types are defined according to the interaction behaviors between nodes, including the derivation relationship between processes, the login relationship between users and hosts, the creation relationship between processes and files, the communication relationship between hosts, the connection relationship between processes and sockets, etc.
[0037] (1-2) Bipartite graph construction: The interactions between system entities in the traceability graph reflect causal relationships. In the recommendation scenario, user-item interactions are usually presented in the form of a bipartite graph to maintain collaborative filtering signals. Therefore, next, a bipartite graph G B ={(v, yvv', v')|v, v'∈V} is defined according to system entity interactions, where G B represents the bipartite graph, and yvv' represents the link between node v and node v'. The link yvv' = 1 indicates that there is an interaction between node v and node v', while the link yvv' = 0 indicates no interaction. Here, the source domain dataset is set to the multi-source network security event log dataset LANL dataset with rich data volume, and the target domain dataset is set to the APT attack-related dataset DARPA dataset. That is, a source domain traceability graph is constructed based on the LANL dataset, and a source domain bipartite graph is constructed according to the source domain traceability graph A target domain traceability graph is constructed based on the DAPRA dataset, and a target domain bipartite graph is constructed according to the target domain traceability graph
[0038] (1-3) Defining r-ego network to extract context relationships: The r-ego network is a local network composed of a central node and its neighbor nodes. The r-hop neighbors of node v are defined as S v ={v': d(v, v')≤r}, where v represents the central node, v' represents the r-hop neighbor node of the central node, d(v, v') represents the shortest path distance between the central node v and the r-hop neighbor node v', and r∈[1, 2]. The r-ego network of node v is denoted as G s , which is a subgraph composed of the central node, the r-hop neighbor nodes of the central node, and the edges between these nodes.
[0039] (2) Generating a cross-domain recommendation model: The source domain dataset is enhanced, pre-trained using a self-supervised scheme on the source domain, and then transferred to the target domain for adjustment to generate a cross-domain recommendation model.
[0040] (2-1) Data augmentation: Since the number of malicious samples in the source domain dataset is much smaller than that of benign data, subgraphs of a node are constructed to perform data augmentation on positive and negative samples. First, two random walks are performed on the r-ego network G of node v to generate two subgraph positive pairs of a node (the first subgraph positive pair g s and the second subgraph positive pair g q ) as positive samples. The subgraphs generated by the r-ego networks of other nodes except node v are used as negative samples (g1 in k is the first negative sample) for processing. The subgraph positive pair is a subgraph sampled from the random walk path. Figure 2
[0041] (2-2) Pre-training on the source domain: After generating positive and negative samples, they are fed into two graph encoders (the first graph encoder f q and the second graph encoder f k ). The first graph encoder f q is used to encode the first subgraph positive pair g q to generate the first vector e q , while the second graph encoder f k is used to encode other subgraphs to generate the second vector e k for the second subgraph positive pair g k , and the third vector e i is generated for each negative sample. The first vector e q , the second vector e k , and the third vector e i are low-dimensional representative vectors. In this embodiment, the graph attention network (GAT) is selected as the graph encoder. The contrastive loss InfoNCE is used to perform self-supervised optimization on the graph encoder to maximize the consistency between the two subgraph positive pairs while maintaining the difference between the positive and negative sample pairs. The self-supervised optimization refers to the reference “A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” CoRR, vol. abs / 1807.03748, 2018”. The formula for the contrastive loss InfoNCE is as follows:
[0042]
[0043] where L InFoNCE represents the contrastive loss InfoNCE, e q represents the first vector, represents the transpose of the first vector, e k represents the second vector, and e iIt represents the third vector generated by the i-th negative sample, n represents the number of negative samples, and τ is the temperature hyperparameter.
[0044] (2-3) Model adjustment on the target domain: After obtaining the pre-trained first graph encoder from the source domain, transfer the pre-trained first graph encoder to the target domain. Use the pre-trained first graph encoder to initialize the node embeddings in the target domain, and obtain the initialized embeddings of each node in the target domain bipartite graph. Specifically, generate r-ego networks for all nodes in the target domain bipartite graph, perform random walks on the r-ego networks of each node in the target domain bipartite graph, and generate the target domain node subgraphs of each node in the target domain bipartite graph. Input the target domain node subgraphs into the pre-trained first graph encoder to obtain the initialized embeddings of each node in the target domain bipartite graph.
[0045] Fine-tune a matrix factorization (MF) model, optimize it using the Bayesian personalized ranking (BPR) loss function, and at the same time use the labeled malicious samples in the target domain dataset as supervision signals to further fine-tune the matrix factorization model so that it can identify abnormal behaviors. The process of fine-tuning the matrix factorization model in this embodiment is the process described in the literature "S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme, "BPR: bayesian personalized ranking from implicit feedback," in UAI 2009, Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, June 18-21, 20". Finally, input the initialized embeddings of each node in the target domain bipartite graph into the fine-tuned matrix factorization model to obtain the final embeddings of each node in the target domain bipartite graph as the central nodes.
[0046] As Figure 2 shown, the pre-trained first graph encoder and the fine-tuned matrix factorization model together constitute a recommendation model. In the figure, circles represent processes and squares represent files, and different colors represent different processes and files.
[0047] (3) Threat detection: Obtain each node embedding through learning by the recommendation model, apply the inner product to the representations of two system entities to predict the possibility of interaction between the two system entities, thereby finding potential threats, and design a feedback mechanism to improve the model adaptability.
[0048] (3-1) Threat detection: After optimization in the fine-tuning stage, the final embedding of each entity node is obtained. For the interaction between every two entities, the inner product is calculated using the final embedding, and the softmax function is used to convert the inner product value to [0, 1], representing the probability that an entity node v does not interact with another entity node v'. If the predicted probability is greater than the preset threshold, the interaction is marked as a potential attack threat. If the probability is less than or equal to the preset threshold, it is considered that the two nodes do not interact.
[0049] (3-2) Model adaptability: When faced with benign but previously unseen system behaviors, false positives may occur. In this case, if these new false positive or false negative behaviors are screened and reviewed manually in the real scenario, the new result feedback is used as an additional label to retrain and correct the recommendation model to improve the performance of the recommendation model for unseen system behaviors.
[0050] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. An APT attack detection method integrating contrastive learning and cross-domain recommendation, characterized in that: The APT attack detection method integrating contrastive learning and cross-domain recommendation includes: Based on the source domain data set, a source domain traceability graph is constructed, and a source domain bipartite graph is constructed according to the source domain traceability graph. Based on the target domain data set, a target domain traceability graph is constructed, and a target domain bipartite graph is constructed according to the target domain traceability graph. Generate r-ego network for all nodes in the source domain bipartite graph; Perform two random walks on the r-ego network of a node in the source domain bipartite graph to generate the first subgraph of the current node g q Opposite to the second subgraph g k As positive samples, perform random walks on the r-ego network of each other node except the current node, generate subgraphs of each other node as negative samples, and generate positive and negative samples corresponding to each node in the source domain bipartite graph; Pre-training the first graph encoder and the second graph encoder using positive samples and negative samples corresponding to each node in the bipartite graph of the source domain to obtain a pre-trained first graph encoder, and transferring the pre-trained first graph encoder to the target domain; Generate an r-ego network for all nodes in the target domain bipartite graph, perform random walks on the r-ego network of each node in the target domain bipartite graph, and generate a target domain node subgraph for each node in the target domain bipartite graph; Input the target domain node subgraph into the pre-trained first graph encoder to obtain the initialization embedding of each node in the target domain bipartite graph; Fine-tune the matrix factorization model using the initialized embeddings; Input the initial embedding of the nodes in each target domain bipartite graph into the fine-tuned matrix factorization model to obtain the final embedding of the nodes in each target domain bipartite graph; The inner product is calculated based on the final embedding of every two nodes, and the probability of the two nodes interacting is calculated based on the inner product. If the probability is greater than the preset threshold, the two nodes are considered to interact and are marked as potential attack threats. If the probability is less than or equal to the preset threshold, the two nodes are considered not to interact.
2. The APT attack detection method integrating contrastive learning and cross-domain recommendation according to claim 1 is characterized in that: The source domain dataset is the LANL dataset and the target domain dataset is the DAPRA dataset.
3. The APT attack detection method integrating contrastive learning and cross-domain recommendation according to claim 1 is characterized in that: The generating of the r-ego network for all nodes in the source domain bipartite graph includes: Take a node in the source domain bipartite graph as the central node and obtain the set S of r-hop neighbor nodes of the central node v ={v':d(v,v')≤r}, where v represents the central node, v' represents the r-hop neighbor node of the central node, d(v,v') represents the shortest path distance between the central node v and the r-hop neighbor node v', r∈[1,2]; The subgraph generated based on the central node and r-hop neighbor nodes is the r-ego network of the current node; Each node in the source domain bipartite graph is used as the central node to generate the r-ego network.
4. The APT attack detection method integrating contrastive learning and cross-domain recommendation according to claim 1 is characterized in that: The method of pre-training the first graph encoder and the second graph encoder using positive samples and negative samples corresponding to each node in the source domain bipartite graph includes: Align the first subgraph to g q Feed to the first image encoder f q Encode and generate the first vector, and put the second subgraph directly on g k Feed to the second image encoder f k Encode to generate a second vector, feed each negative sample to the second graph encoder f k Encode and generate a third vector for each negative sample; The contrast loss InfoNCE is used to perform self-supervisory optimization on the first image encoder and the second image encoder, which can be expressed as follows: Among them, L InFoNCE Denotes the contrast loss InfoNCE, e q represents the first vector, represents the transpose of the first vector, e k represents the second vector, e i represents the third vector generated by the i-th negative sample, n represents the number of negative samples, and τ is the temperature hyperparameter.
5. The APT attack detection method integrating contrastive learning and cross-domain recommendation according to claim 1 is characterized in that: The method of fine-tuning the matrix decomposition model by using initialization embedding includes: Feed the initial embedding into the matrix factorization model; The Bayesian personalized ranking loss function is used to optimize the matrix factorization model; Get a fine-tuned matrix factorization model.
Citation Information
Cited By
Practical cross-system shilling attack method with limited data access
CN115859286A
Practical cross-system attack method with restricted data access
CN115859286B