APT detection method and system based on multi-branch graph neural network
Through multi-branch graph neural network (MBGNN) combined with GAT, GIN and GraphSAGE, the detection accuracy and efficiency of existing APT detection methods in large-scale and complex network environments are solved, and efficient and accurate APT detection is achieved.
Patent Information
- Application Number
- CN202510859530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing APT detection method based on traceability maps has low detection accuracy and low computing efficiency in large-scale complex network environments, making it difficult to comprehensively model the multi-dimensional structural and semantic relationships in the graph, and lacks the coordinated utilization of neighbor information, resulting in missed detection or misjudgment.
Multi-branch graph neural network (MBGNN) is used to combine Graph Attention Network (GAT), Graph Isomorphism Network (GIN) and Graph Sample and Aggregation (GraphSAGE), node features are extracted through attention mechanism, structure-sensitive aggregation and neighbor sampling, and embedding reuse optimization strategy and neighbor perception confidence fusion mechanism are introduced to optimize computing efficiency and detection accuracy.
The accuracy and recall rate of APT detection are improved, the calculation overhead is reduced, and the adaptability and detection efficiency are improved in complex scenarios. The experimental results show that the accuracy is increased by 2%-15%, the recall rate is increased by 1%-8%, the F1 value is increased by 2%-8%, and the calculation time is reduced by 20%-25%.
Smart Images

Figure CN120378226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of advanced persistent threat (APT) detection, and particularly to an APT detection system and method based on a multi-branch graph neural network. Background Art
[0002] With the continuous evolution of cyber attack technologies, advanced persistent threat (APT) has become an important challenge in the field of network security due to its characteristics such as strong concealment, long attack cycle, and strong targetability. APT attacks usually achieve the penetration and persistent control of the target system through multi-stage and complex behavior chains (such as vulnerability exploitation, privilege escalation, data theft, etc.). Traditional intrusion detection systems (IDSes) based on signature matching or rules often face problems such as low detection accuracy, high false alarm rate, and inability to effectively identify unknown attack patterns when dealing with such attacks. In recent years, the detection method based on the traceability graph has gradually become an important technical path for APT detection because it can comprehensively capture the context dependencies between system behaviors by parsing system audit logs and constructing event causal relationship graphs. The traceability graph models system behaviors through nodes (representing entities such as processes, files, network flows, etc.) and edges (representing the causal relationships between events), providing rich structured information for anomaly detection and showing significant advantages in improving detection accuracy and interpretability.
[0003] However, the existing APT detection methods based on the traceability graph still face multiple challenges in practical applications. First of all, the traceability graph usually has the characteristics of large scale and complex node relationships. Existing methods mostly use a single graph neural network (such as GraphSAGE or GAT) for node feature learning, making it difficult to comprehensively model the multi-dimensional structural, semantic, and adjacency relationships in the graph, resulting in insufficient key information mining and limited detection accuracy. Secondly, in confidence calculation, existing methods usually only rely on the own features of the target node and ignore the influence of neighbor nodes. However, the APT attack path often spreads in a chain, and the abnormal state of a node is closely related to its neighbors. The lack of collaborative utilization of neighbor information is likely to lead to missed detections or misjudgments. In addition, there are a large number of nodes with similar semantics and structures in the large-scale traceability graph. Existing methods calculate the embedding representation for each node one by one, causing redundant calculations in the training and inference stages, significantly increasing the computational overhead, and limiting the application efficiency of the method in large-scale scenarios.
[0004] To address the above problems, some studies have attempted to improve detection performance by optimizing the structure of graph neural networks or introducing embedding storage mechanisms. For example, some methods enhance the ability to capture attack behaviors through subgraph matching or path analysis. However, these methods often rely on predefined attack patterns and are difficult to adapt to unknown or mutated attacks. Other methods store node embeddings in a database to reduce computational overhead but do not fully consider the dynamic reuse of node similarity and neighbor relationships, resulting in limited optimization effects. Therefore, there is an urgent need for an APT detection method that can comprehensively utilize multi-dimensional graph features, fuse neighbor information, and optimize computational efficiency to improve detection accuracy and recall rate and meet the requirements of efficient APT detection in complex network environments. Summary of the Invention
[0005] Aiming at the deficiencies of existing APT detection methods in insufficient mining of traceability graph features, low detection efficiency, and low accuracy, the present invention discloses an APT detection system and method based on a multi-branch graph neural network.
[0006] The detection method first constructs a traceability graph by parsing system audit logs, modeling the system execution process as a structured graph containing process nodes, object nodes, and event edges. The edges between nodes carry timestamps and causal relationship labels to comprehensively capture the context dependencies between events. Subsequently, the Word2Vec technique is used to semantically encode the nodes to generate preliminary node feature embeddings as the input for subsequent processing. At the same time, a multi-branch graph neural network model (Multi-Branch Network Architecture - MBGNN) is designed, aggregating three graph neural networks, namely Graph Attention Network (GAT), Graph Isomorphism Network (GIN), and Graph Sample and Aggregation (GraphSAGE), into a unified model to extract node features from three dimensions: neighbor importance, structural distinctiveness, and local feature aggregation. Among them, GAT dynamically models the importance differences between nodes and neighbors through an attention mechanism to strengthen the contributions of key neighbors; GIN enhances the discriminative ability of local topological structures through a structure-sensitive aggregation function to capture similar attack patterns; GraphSAGE efficiently extracts global and local features through neighbor sampling and feature aggregation to reduce computational overhead. The outputs of the three branches are fused through weight parameters to generate a comprehensive node embedding vector. During the training process, a multi-class cross-entropy loss function with class weights is used to alleviate the class imbalance problem and improve the ability to identify malicious nodes.
[0007] To optimize the computational efficiency, this method proposes an embedding reuse optimization strategy. The system uses the SQLite database to store the embedding vectors and neighbor relationships of nodes, and measures the similarity of neighbor sets between nodes through the Overlap coefficient. When the similarity exceeds the set threshold, the existing embeddings in the database are preferentially reused, and new embeddings are only generated for unmatched nodes. At the same time, strategies such as neighbor set deduplication, bidirectional deduplication, and consistency checking are adopted to optimize the storage efficiency, significantly reducing redundant calculations in the training and inference phases.
[0008] In addition, this method introduces a neighbor-aware confidence fusion mechanism, which uses the GAT model to weighted aggregate the confidence of the target node and its neighbors, enhancing the context awareness ability of complex attack chains. The confidence fusion combines the classification confidence of the node itself (based on the difference between the maximum and the second-largest category scores) and the weighted result of neighbor confidence, and balances the contributions of the two through adjustable parameters to improve the detection sensitivity of rare attack signals.
[0009] The technical solution of the present invention is as follows: An APT detection method based on a multi-branch graph neural network, including: First, construct a traceability graph by parsing system logs; Then, use Word2Vec to perform semantic embedding on the nodes in the traceability graph to capture the context semantics of the nodes; After the embedding generation stage, the semantic information of the nodes is input into the trained multi-branch network model MBGNN to extract node representations and perform APT detection from three perspectives: neighbor importance modeling, structural discriminability, and local structure aggregation; The multi-branch network model includes GAT, GraphSAGE, and GIN; The GAT branch uses the attention mechanism to model the importance difference between nodes and neighbors, strengthening key neighbor features; The GIN branch improves the discriminability of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes feature aggregation through neighbor sampling; In the detection process, a neighbor-aware confidence fusion mechanism is introduced to fuse the confidence of the node itself with the confidence of its neighbor nodes through GAT; and an embedding reuse optimization strategy is introduced to calculate the node similarity and reuse the embedding representations of similar nodes.
[0010] According to the preference of the present invention, constructing a traceability graph by parsing system logs includes: First, extract event logs from the log data of multiple hosts, and screen and filter the event log entries to extract valid data, including source IP address, target IP address, and event timestamp; Subsequently, the filtered log entries are converted into JSON format; Finally, the structured log data, i.e., the log entries converted into JSON format, are processed in a batch manner to generate a traceability graph; the traceability graph includes process nodes and object nodes, where the object nodes represent entities in the system, and the process nodes represent the occurrence of events and their related operations; the nodes are connected by edges, and the edges carry event type labels and timestamps, which are used to represent the causal relationships between the nodes; each node contains rich context information.
[0011] According to the preferred embodiment of the present invention, the GraphSAGE branch updates the representation of the target node by sampling neighbor nodes and aggregating their features; The GraphSAGE branch includes two layers of convolutional operations; for the target node vi, its initial feature representation is , and the neighbor node set is ; First, the first layer of convolutional operation sage1 aggregates the neighbor features and updates the node representation: ; Among them, is the mean aggregation function, defined as: ; is the weight matrix of the first layer of convolutional operation, and the RELU activation function introduces nonlinearity, and then Dropout is applied to prevent overfitting; is the representation of the first layer of convolutional operation of node i; Then, the second layer of convolutional operation sage2 further updates the node representation: ; Among them, is the weight matrix of the second layer of convolutional operation, and finally the node representation of the GraphSAGE branch is obtained .
[0012] According to the preferred embodiment of the present invention, GIN calculates the node representation by weighted summing of neighbor node features and performing a nonlinear transformation using a multi-layer perceptron; The GIN branch includes two layers of convolutional operations; for the target node vi, the first layer of convolutional operation gin1 first performs a weighted sum of neighbor features and transforms through an MLP, as follows: ; Among them, ; ; , is the weight matrix of the MLP. After ReLU activation, Dropout is applied. MLP represents a multi-layer perceptron, including two layers of linear transformation, with ReLU activation function in between. Then, the second convolutional operation gin2 continues to update the node representation: ; Among them, is a two-layer fully connected network: ; , is the weight matrix, and finally the node representation of the GIN branch is obtained .
[0013] According to the preference of the present invention, GAT assigns different weights to each neighbor node through the self-attention mechanism to capture the relative importance of neighbor nodes to the target node. The GAT branch includes two layers of convolutional operations, and each layer uses the multi-head attention mechanism. For the target node vi, the first convolutional operation gat1 uses 3 attention heads: ; Among them, is the weight matrix of the k-th attention head, is the attention weight calculated by the k-th attention head, which is dynamically generated through the self-attention mechanism: ; Among them, is the attention parameter, and || represents vector concatenation. After concatenating the outputs of the three attention heads, ReLU activation is used, and then Dropout is applied; Then, the second convolutional operation gat2 also uses 3 attention heads, but takes the average of the multi-head outputs: ; Among them, is the weight matrix, and finally the node representation of the GAT branch is obtained .
[0014] According to the preference of the present invention, the multi-branch network model aggregates the outputs of the GAT branch, GIN branch, and GraphSAGE branch by introducing weights , and ; The final node embedding vector is expressed as: ; The aggregated node representation Subsequently, classification is performed through a linear layer and a softmax function.
[0015] Preferably according to the present invention, node V i cross-entropy loss is defined as: ; wherein: ; N is the total number of all nodes, C is the total number of categories, is the number of samples of category c, represents the predicted probability that sample i belongs to category c by the multi-branch network model, represents the true label, which is 1 if sample i belongs to category c, otherwise 0, represents the weight of category c; taking the average of the losses of all samples, the total loss of the multi-branch network model is expressed as: .
[0016] Preferably according to the present invention, during the training and inference process of the multi-branch network model MBGNN, the system generates the embedding vector of each node and stores it in the database; including: 1) Check whether there is a table for storing embedding vectors; if not, create a table structure including the following fields: node label: used to uniquely identify the node; embedding vector: representing the feature embedding of the node; neighbor set: storing the neighbor node information of the node in JSON format, representing the relationship between nodes; 2) Perform deduplication through the following deduplication strategies: Neighbor set deduplication: Deduplicate the neighbor set through set operations to ensure the uniqueness of each neighbor set; Bidirectional deduplication: If the neighbor set of node A1 already includes node B1, then node B1 is not stored repeatedly when storing node A1; Neighbor set consistency check: Before inserting node information, query whether there is a record with the same neighbor set in the database. If so, reuse the existing embedding.
[0017] Preferably according to the present invention, an embedding matching mechanism based on the Overlap coefficient is introduced to measure the similarity between nodes; the calculation formula of the Overlap coefficient is as follows: ; wherein, A and B respectively represent the neighbor sets of the current node and the node stored in the database; The specific process of embedding matching is as follows: The system extracts the embedding vectors of all stored nodes and their neighbor sets from the database; For each input node, calculate the Overlap coefficient between the set of its neighbors and the set of neighbors of the stored nodes in the database to evaluate node similarity; If the Overlap coefficient is higher than the preset threshold, the system directly reuses the corresponding embedding vector in the database as the feature representation of the current node; otherwise, a new embedding vector is generated for the node.
[0018] According to the preference of the present invention, a neighbor-aware confidence fusion mechanism is introduced during the detection process, and the confidence of the node itself and the confidence of its neighbor nodes are fused through GAT; including: Calculate the local confidence of the target node, which is: based on the relative difference between the maximum and the second-largest category scores in the classification result; Aggregate the confidence information of neighbor nodes through the GAT branch to generate neighbor-weighted confidence; Use adjustable parameters to balance the contributions of the local confidence of the target node and the neighbor-weighted confidence to generate the final fused confidence; The relevant formulas are as follows: ; Among them, ; ; Among them, is the confidence calculated based on the own information of node i, is the confidence calculated based on the own information of neighbor j of node i, is the result after weighted aggregation of the confidence of all neighbors through the GAT branch, α is a parameter that controls the fusion ratio of the own information and the neighbor information, usually represents node v i For node v j the attention weight of, used to measure v j to v i importance; Finally, based on the fused confidence judge whether the node is abnormal, and the node with a value lower than the threshold is marked as abnormal.
[0019] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned APT detection method based on a multi-branch graph neural network are implemented.
[0020] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned APT detection method based on a multi-branch graph neural network are implemented.
[0021] An APT detection system based on a multi-branch graph neural network includes: The traceability graph construction module is configured to: construct a traceability graph by parsing system logs; The semantic embedding module is configured to: perform semantic embedding on the nodes in the traceability graph using Word2Vec to capture the context semantics of the nodes, and then input them into the trained multi-branch network model MBGNN to extract node representations from three perspectives of neighbor importance modeling, structural distinctiveness, and local structure aggregation respectively, and perform classification through a linear layer and a softmax function; The multi-branch network model includes GAT, GraphSAGE, and GIN; The GAT branch uses the attention mechanism to model the importance difference between nodes and their neighbors, and strengthens the key neighbor features; The GIN branch improves the discriminability of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes feature aggregation through neighbor sampling; During the detection process, a neighbor-aware confidence fusion mechanism is introduced to fuse the confidence of the node itself with the confidence of its neighbor nodes through GAT; and an embedding reuse optimization strategy is introduced to calculate the node similarity and reuse the embedding representations of similar nodes.
[0022] The beneficial effects of the present invention are as follows: 1. The present invention proposes an efficient APT detection method based on multi-branch graph neural network aggregation: The present invention proposes an APT detection framework based on multi-branch graph neural network aggregation, MBGNN. Aiming at the problem of insufficient utilization of traceability graph feature information in existing APT detection methods, MBGNN dynamically aggregates three graph neural networks, namely Graph Attention Network (GAT), Graph Isomorphism Network (GIN), and Graph Sample and Aggregation (GraphSAGE). Among them, GAT identifies key neighbor nodes through the attention mechanism, GIN captures the node behaviors of similar attack patterns by enhancing graph isomorphism and local structure learning, and GraphSAGE aggregates the features of node neighbors through random sampling, which is suitable for capturing the global topological information and the dependencies between nodes in a large-scale graph. In the aggregation step, weight parameters are set to fuse the outputs of the above three graph neural networks, realizing the efficient learning of the global information of the large-scale traceability graph, and improving the comprehensive recognition ability of the model for APT behaviors and the adaptability in complex scenarios.
[0023] 2. The present invention proposes an embedding reuse optimization strategy: The present invention introduces an embedding reuse optimization strategy. By storing the embedding representations and neighbor relationships of nodes in a database, existing embeddings in the library are directly reused through similarity matching, and new embeddings are only generated for unmatched nodes, thus significantly improving the computational efficiency.
[0024] 3. The present invention proposes a neighbor-aware confidence fusion mechanism: The present invention proposes a neighbor-aware confidence fusion mechanism. In the APT detection stage, the confidence information of neighbor nodes is first introduced during the inference process and fused with the confidence of the target node. Different from existing methods that only rely on the confidence of the target node itself for anomaly inference, this method enhances the ability to identify complex attack behaviors by considering the influence of neighbor nodes and significantly improves the detection efficiency.
[0025] 4. During the detection process, classification is performed based on the node embedding vectors generated by the multi-branch network model (MBGNN) to predict the category and confidence of each node. Combining the neighbor-aware confidence fusion mechanism, it is determined whether a node has abnormal behavior. Experimental results show that in six scenarios of three public datasets (StreamSpot, DARPA TC #3, DARPA OpTC), the precision is improved by 2% - 15%, the recall rate is improved by 1% - 8%, and the F1 value is improved by 2% - 8%. The F1 value in five scenarios exceeds 95%, and the recall rate in four scenarios exceeds 99%. In terms of time efficiency, the execution time in the node embedding generation and inference stages is reduced by approximately 20% and 25% respectively, which is better than existing advanced methods. Brief Description of the Drawings
[0026] Figure 1 It is the architecture diagram of an APT detection method based on a multi-branch graph neural network of the present invention; Figure 2 It is the architecture diagram of the multi-branch network model. Detailed Embodiments
[0027] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.
[0028] Embodiment 1 Term Explanation: 1. Advanced Persistent Threat (APT): APT is a form of network attack with strong concealment, a long attack cycle, and a clear target. It is usually carried out by organized attackers to penetrate and persistently control the target system through multi-stage and complex behavior chains (such as vulnerability exploitation, privilege escalation, data theft, etc.).
[0029] 2. Traceability Graph: A traceability graph is a structured graphical representation constructed by parsing system audit logs, where nodes represent system entities (such as processes, files, network flows, etc.), edges represent the causal relationships between events, and nodes and edges contain rich context information (such as timestamps, event types).
[0030] 3. Graph Neural Network: A graph neural network is a class of deep learning models based on graph structures. By aggregating the features of nodes and their neighbors, it learns the representations of nodes or subgraphs in the graph and is widely applied to tasks such as graph data classification, prediction, and anomaly detection.
[0031] 4. Attention Mechanism: The attention mechanism is a neural network technique that highlights important information by assigning dynamic weights to different inputs. In graph neural networks (such as GAT), it is used to measure the relative importance of a node and its neighbors.
[0032] 5. Node Embedding: Node embedding is a representation form that maps the features of nodes in a graph (including semantic information, structural information, etc.) to a low-dimensional vector space. It is usually generated by graph neural networks or word embedding techniques (such as Word2Vec) and is used for subsequent classification or detection tasks.
[0033] An APT detection method based on a multi-branch graph neural network, as Figure 1 shown, includes: First, construct a traceability graph by parsing system logs; Then, use Word2Vec to perform semantic embedding on the nodes in the traceability graph to capture the context semantics of the nodes; After the embedding generation stage, the semantic information of the nodes is input into the trained multi-branch network model MBGNN to extract node representations from three perspectives: neighbor importance modeling, structural discriminability, and local structure aggregation, and classification is performed through a linear layer and a softmax function; The multi-branch network model includes GAT, GraphSAGE, and GIN; The GAT branch uses the attention mechanism to model the importance difference between nodes and their neighbors and strengthen the key neighbor features; The GIN branch improves the discriminability of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes efficient feature aggregation through neighbor sampling; while reducing resource consumption, it enhances the global feature mining ability of the model; To further improve the judgment ability of the multi-branch network model (MBGNN) in the abnormal node recognition stage, a neighbor-aware confidence fusion mechanism is introduced during the detection process. The confidence of the node itself is fused with the confidence of its neighbor nodes through GAT, thereby effectively improving the accuracy of anomaly detection. This method makes full use of the confidence information of neighbor nodes and enhances the response ability to complex attack scenarios. At the same time, to improve the computational efficiency of the multi-branch network model (MBGNN) on large-scale graphs, an embedding reuse optimization strategy is introduced to solve the redundant calculation problem of node embeddings. This mechanism calculates node similarity and reuses the embedding representations of similar nodes, significantly improving the computational efficiency.
[0034] In the present invention, a method based on the multi-branch graph neural network (MBGNN) is adopted to detect APT. During detection, first, the system extracts relevant data from system logs. Subsequently, the node feature representation is calculated so that the node representation can fully integrate local features and graph structure characteristics, thereby better supporting anomaly detection. Then, based on the embedding representation of the node, it is input into the trained multi-branch network model MBGNN to extract node representations from three perspectives: neighbor importance modeling, structural discriminability, and local structure aggregation, and classification is performed through a linear layer and a softmax function to predict the category of each node, and whether the node has abnormal behavior is judged according to the classification result and confidence. For example, when the confidence of the node classification prediction is low, the node may be judged as abnormal.
[0035] Embodiment 2 A method for detecting APT based on the multi-branch graph neural network according to Embodiment 1, wherein: A traceability graph is constructed by parsing system logs, including: First, event logs are extracted from the log data of multiple hosts, and the event log entries are screened and filtered to extract valid data that meets the analysis requirements, including source IP address, target IP address, and event timestamp; Subsequently, the screened log entries are converted into JSON format; the JSON format is a structured data format suitable for subsequent processing, and these logs record detailed event information of the system, including but not limited to process execution, file operations, and network traffic, etc.
[0036] Finally, the structured log data, which has been converted into JSON - formatted log entries, is processed in a batch manner to generate a traceability graph. The traceability graph includes process nodes and object nodes. Among them, the object nodes represent entities in the system, such as files, processes, or network flows; the process nodes represent the occurrence of events and their related operations; the nodes are connected by edges, and the edges are labeled with event types (such as system calls, data flow directions) and timestamps, which are used to represent the causal relationships between nodes; each node contains rich context information, such as process names, file paths, IP addresses, and port numbers, etc., to enhance the analysis ability of the relationships between events and potential attack patterns. Through the above steps, the traceability graph constructed by the present invention can comprehensively capture the causal dependencies and context information during the system execution process, providing a reliable data basis for subsequent APT detection.
[0037] To achieve accurate detection of APT, the present invention uses Word2Vec to perform semantic encoding on the nodes in the traceability graph as the input of the graph neural network model.
[0038] On this basis, the method designs a multi - branch network model (Multi - Branch Network Architecture - MBGNN), aiming to improve the node feature mining ability through multi - branch design. As Figure 2 shown, the multi - branch network model (MBGNN) aggregates three classic graph neural networks: GraphSAGE, GIN, and GAT, and extracts node representations from three perspectives: local structure aggregation, structural discriminability, and neighbor importance modeling.
[0039] The GraphSAGE branch updates the representation of the target node by sampling neighbor nodes and aggregating their features; In the multi - branch network model (MBGNN), the GraphSAGE branch includes two - layer convolutional operations; for the target node vi, its initial feature representation is , and the set of neighbor nodes is ; First, the first - layer convolutional operation sage1 aggregates neighbor features and updates the node representation: ; Among them, is the mean aggregation function, defined as: ; is the weight matrix of the first - layer convolutional operation, and the RELU activation function introduces non - linearity. Then, Dropout is applied to prevent overfitting; is the representation of the first - layer convolutional operation of node i; Then, the second - layer convolutional operation sage2 further updates the node representation: ; Among them, is the weight matrix of the second-layer convolution operation, and finally the node representation of the GraphSAGE branch is obtained .
[0040] GIN calculates the node representation by weighted summation of neighbor node features and non-linear transformation using a multi-layer perceptron (MLP), which is particularly suitable for capturing the structural distinctiveness of the graph.
[0041] In the multi-branch network model (MBGNN), the GIN branch includes two layers of convolution operations; for the target node vi, the first layer of convolution operation gin1 first performs weighted summation on neighbor features and transforms through MLP as follows: ; Among them, ; ; , is the weight matrix of the MLP. After ReLU activation, Dropout is applied; MLP represents a multi-layer perceptron (MLP), including two layers of linear transformation, and the two layers of linear transformation are represented by two weight matrices and with ReLU activation function in between; Then, the second layer of convolution operation gin2 continues to update the node representation: ; Among them, is also a two-layer fully connected network: ; , is the weight matrix, and finally the node representation of the GIN branch is obtained .
[0042] GAT assigns different weights to each neighbor node through the self-attention mechanism to capture the relative importance of neighbor nodes to the target node; in the multi-branch network model (MBGNN), the GAT branch includes two layers of convolution operations, and each layer uses the multi-head attention mechanism; for the target node vi, the first layer of convolution operation gat1 uses 3 attention heads (heads = 3): ; Among them, is the weight matrix of the k-th attention head, is the attention weight calculated by the k-th attention head, dynamically generated through the self-attention mechanism: ; Among them, are attention parameters, || represents vector concatenation; the outputs of the three attention heads are concatenated and then activated using RELU, and then Dropout is applied; Then, the second-layer convolutional operation gat2 also uses 3 attention heads, but takes the average of the multi-head outputs: ; Among them, is the weight matrix, and finally the node representation of the GAT branch is obtained .
[0043] The multi-branch network model aggregates the outputs of the GAT branch, GIN branch, and GraphSAGE branch by introducing weights , and ; the weights , and are set manually, and the sum of the three weights is 1; the final node embedding vector is expressed as: ; The aggregated node representation is then classified through a linear layer and a softmax function.
[0044] The method of the present invention uses a multi-class cross-entropy loss function with class weights as the training objective to improve the recognition ability of minority classes and further improve the detection ability of the model in complex APT scenarios. Specifically, the cross-entropy loss of node Vi is defined as:
[0045] ; Among them: ; N is the total number of all nodes, C is the total number of classes, is the number of samples of class c, represents the predicted probability that the multi-branch network model assigns sample i to class c, represents the true label, which is 1 if sample i belongs to class c and 0 otherwise, represents the weight of class c; in order to obtain stable gradient updates in the batch, the losses of all samples are averaged, and the total loss of the multi-branch network model is expressed as: .
[0046] By introducing class weights , this loss function can effectively alleviate the class imbalance problem, enabling the multi-branch network model (MBGNN) to pay more attention to the correct classification of minority classes (such as malicious nodes) during the training process, thereby improving the detection performance for complex APTs.
[0047] Aiming at the problems of waste of computing resources and low inference efficiency caused by repeated calculation of node embeddings when existing APT detection methods handle large-scale traceability graphs, this method proposes an embedding reuse optimization strategy. Through two key steps of database storage and embedding matching, this method realizes efficient node embedding management and fast generation, significantly improving the computational efficiency of the training and inference processes.
[0048] To efficiently manage node embeddings, this method adopts a storage scheme based on SQLite. During the training and inference processes of the multi-branch network model MBGNN, the system generates the embedding vectors of each node and stores them in the database, including:
[0049] 1) Check whether there is a table for storing embedding vectors; if not, create a table structure including the following fields: Node label: used to uniquely identify the node; Embedding vector: representing the feature embedding of the node; Neighbor set: storing the neighbor node information of the node in JSON format, representing the relationship between nodes; 2) To optimize storage efficiency and reduce redundancy, the following deduplication strategies are used for deduplication: Neighbor set deduplication: Before inserting the node embedding and its neighbor information, perform deduplication on the neighbor set through set operations to ensure the uniqueness of each neighbor set; Bidirectional deduplication: Avoid bidirectional redundant storage. If the neighbor set of node A1 already includes node B1, then node B1 is not stored repeatedly when storing node A1; vice versa.
[0050] Neighbor set consistency check: Before inserting node information, query whether there is a record with the same neighbor set in the database. If so, reuse the existing embedding to avoid repeated calculation.
[0051] Through the above deduplication strategies, unnecessary repeated calculations are avoided, and the training and inference speeds are improved.
[0052] To further optimize feature reuse, the present invention introduces an embedding matching mechanism based on the Overlap coefficient to measure the similarity between nodes; compared with other similarity metrics used in traditional methods, the Overlap coefficient pays more attention to the number of shared neighbors and uses the minimum degree in the node pair as the normalization factor to reduce the bias caused by node degree differences, and is applicable to traceability graphs with uneven node degree distributions. The calculation formula of the Overlap coefficient is as follows:
[0053] ; Wherein, A and B respectively represent the neighbor sets of the current node and the stored nodes in the database; The specific process of embedding matching is as follows: The system extracts the embedding vectors of all stored nodes and their neighbor sets from the database; For each input node, calculate the Overlap coefficient between its neighbor set and the neighbor sets of the stored nodes in the database to evaluate the node similarity; If the Overlap coefficient is higher than the preset threshold (manually set by oneself, and the threshold I set is 1), the system directly reuses the corresponding embedding vector in the database as the feature representation of the current node; otherwise, generate a new embedding vector for the node. Through the above mechanism, this method can efficiently reuse existing embeddings among similar nodes, significantly reduce the amount of repeated calculations, and thus accelerate the training and inference processes of the model.
[0054] This detection method includes the following steps when performing detection: First, the system extracts relevant data from the system log to construct a traceability graph structure containing nodes and edges, where nodes represent processes or entities, and edges represent the causal relationships between events. Subsequently, in the embedding matching stage, if the embedding matching is successful, the node embedding is directly obtained; if it is unsuccessful, it is calculated using a multi-branch network architecture. After obtaining the node embedding, the node is classified, and the classification confidence is calculated. If the confidence of the node classification is lower than the preset threshold, the node is determined to be a potential abnormal node.
[0055] Aiming at the defect that the existing APT detection methods only rely on the information of the target node itself and ignore the influence of neighbor nodes when calculating the confidence, this method innovatively proposes a neighbor-aware confidence fusion mechanism. This mechanism realizes the weighted fusion of the confidence information of the target node and its neighbor nodes through the Graph Attention Network (GAT). GAT dynamically calculates the importance of each neighbor node to the target node through the attention mechanism, and automatically assigns higher weights to key neighbors, thereby enhancing the model's ability to capture the relationships between nodes. This mechanism is particularly suitable for detecting rare or hidden attack events and can significantly improve the sensitivity to weak attack signals.
[0056] In the detection process, a neighbor-aware confidence fusion mechanism is introduced to fuse the confidence of the node itself and the confidence of its neighbor nodes through GAT; including: Calculate the local confidence of the target node, which is: based on the relative difference between the maximum and the second-largest category scores in the classification result; reflecting the classification certainty of the node itself.
[0057] Aggregate the confidence information of neighbor nodes through the GAT branch to generate neighbor weighted confidence; Balance the contributions of the local confidence of the target node and the weighted confidence of neighbors using adjustable parameters to generate the final fused confidence; The relevant formula is as follows: ; where, ; ; conf local Measures the relative difference between the scores of the predicted maximum class and the second - largest class, and is used to evaluate the classification confidence degree of the model at the current node. is the confidence calculated based on the self - information of node i (i.e., the features of the node itself), is the confidence calculated based on the self - information of neighbor j of node i (i.e., the features of the neighbor node), is the result after weighted aggregation of the confidences of all neighbors through the GAT branch. α is a parameter that controls the fusion ratio of self - information and neighbor information, usually represents the node v i to node v j 's attention weight, which is used to measure the importance of v j to v i ;
[0058] Finally, according to the fused confidence judge whether the node is abnormal. Nodes below the threshold (manually set) are marked as abnormal. Through the above - mentioned method, this method not only fully utilizes the feature information of the target node, but also enhances the context - awareness ability through the confidence fusion of neighbor nodes, significantly improving the detection accuracy in complex attack scenarios.
[0059] This embodiment uses the Streamspot dataset and the Engagement 3 and OpTC datasets in the DARPA TC project. The Streamspot dataset contains 6 groups, with 100 information - flow graphs in each group, covering five normal scenarios and one attack scenario, and each scenario runs 100 times. The normal scenarios include viewing Gmail, browsing CNN, downloading files, watching YouTube, and playing games. The attack scenario simulates a drive - by download attack where the victim host accesses a malicious URL and exploits a Flash vulnerability to obtain root privileges.
[0060] The DARPA TC #3 dataset was released by DARPA in 2019 to support the transparent computing project. This dataset is widely used in APT detection research to verify the effectiveness of various detection methods. The dataset contains node and edge information from multiple execution teams (such as Theia, Trace, etc.). To match real - world APT activities, the attack team developed various tools to simulate different stages of APT while generating normal data.
[0061] The DARPA OpTC dataset is the latest achievement released by DARPA in recent years, with a total of more than 17 billion system event records. This dataset collected system audit logs from approximately 1,000 hosts running the Windows operating system over seven consecutive days. During the last three days of data collection, the attack team simulated three APT scenarios on multiple hosts. In this example experiment, the data of host number 501 in the second attack scenario (data exfiltration) and the data of host number 051 in the third attack scenario (malware upgrade) were selected as malicious samples. In the experimental analysis, to ensure consistency of expression, the attack data are all identified by the scenario name.
[0062] To evaluate the performance of the proposed detection method, a systematic comparison was made with the representative detection methods THREATRACE and FLASH in the current field (the method of the present invention is abbreviated as MBGNN), as shown in Table 1 and Table 2.
[0063] Table 1 Experimental results one on the StreamSpot, DARPA TC #3, and DARPA OpTC datasets;
[0064] Table 2 Experimental results two on the StreamSpot, DARPA TC #3, and DARPA OpTC datasets;
[0065] On the StreamSpot dataset, MBGNN achieved 1.00, 0.95, and 0.97 in terms of precision, recall, and F1-score metrics, respectively, being on par with or slightly better than FLASH and THREATRACE, which demonstrates excellent detection accuracy and recall rate. On the DARPA TCE3 dataset, MBGNN performed outstandingly in the Cadets and Trace scenarios, with F1-scores both reaching 0.99, far higher than FLASH and THREATRACE, which reflects the adaptability of the proposed method in different scenarios. In the Theia scenario, MBGNN also achieved an F1-score of 0.95, significantly better than FLASH (0.77) and comparable to THREATRACE (0.93). In the two attack scenarios (Custom Powershell and Malicious Upgrade) of the OpTC dataset, MBGNN outperformed the comparison methods in most metrics. Especially in the "Custom Powershell" scenario, the F1-score reached 0.97, significantly higher than FLASH (0.92) and THREATRACE (0.86). In the more challenging "Malicious Upgrade" scenario, although the overall metrics decreased, MBGNN still achieved an F1-score of 0.89, maintaining a significant performance advantage. Overall, MBGNN achieved high detection performance in multiple complex datasets and different attack types, verifying its effectiveness and strong competitiveness in the APT detection task.
[0066] Example 3 A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the APT detection method based on a multi-branch graph neural network described in Example 1 or 2 are implemented.
[0067] Example 4 A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the APT detection method based on a multi-branch graph neural network described in Example 1 or 2 are implemented.
[0068] Example 5 An APT detection system based on a multi-branch graph neural network, comprising: A traceability graph construction module, configured to: construct a traceability graph by parsing system logs; A semantic embedding module, configured to: perform semantic embedding on the nodes in the traceability graph using Word2Vec to capture the context semantics of the nodes, and then input them into the trained multi-branch network model MBGNN to extract node representations from three perspectives: neighbor importance modeling, structural distinctiveness, and local structure aggregation, and perform classification through a linear layer and a softmax function; The multi-branch network model includes GAT, GraphSAGE, and GIN; The GAT branch uses the attention mechanism to model the importance difference between nodes and their neighbors, strengthening the features of key neighbors; The GIN branch improves the discriminative power of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes feature aggregation through neighbor sampling; During the detection process, a neighbor-aware confidence fusion mechanism is introduced to fuse the confidence of a node itself with the confidence of its neighbor nodes through GAT; and an embedding reuse optimization strategy is introduced to calculate node similarity and reuse the embedding representations of similar nodes.
Claims
1. An APT detection method based on a multi-branch graph neural network, characterized in that Including: First, construct a traceability graph by parsing system logs; Then, use Word2Vec to perform semantic embedding on the nodes in the traceability graph to capture the context semantics of the nodes; After the embedding generation stage, the semantic information of the nodes is input into the trained multi-branch network model MBGNN to extract node representations and perform APT detection from three perspectives: neighbor importance modeling, structural discriminability, and local structure aggregation; The multi-branch network model includes GAT, GraphSAGE, and GIN; The GAT branch uses the attention mechanism to model the importance difference between nodes and their neighbors, strengthening key neighbor features; The GIN branch improves the discriminability of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes feature aggregation through neighbor sampling; During the detection process, introduce a neighbor-aware confidence fusion mechanism, and fuse the confidence of the node itself with the confidence of its neighbor nodes through GAT; And introduce an embedding reuse optimization strategy to calculate node similarity and reuse the embedding representations of similar nodes.
2. The APT detection method based on a multi-branch graph neural network according to claim 1, wherein Construct a traceability graph by parsing system logs; including: First, extract event logs from the log data of multiple hosts, and screen and filter the event log entries to extract valid data, including source IP address, target IP address, and event timestamp; Subsequently, convert the filtered log entries into JSON format; Finally, process the structured log data, that is, the log entries converted into JSON format, in a batch manner to generate a traceability graph; the traceability graph includes process nodes and object nodes. Among them, the object nodes represent entities in the system, and the process nodes represent the occurrence of events and their related operations; the nodes are connected by edges, and the edges carry event type labels and timestamps, which are used to represent the causal relationship between nodes; each node contains rich context information.
3. The APT detection method based on a multi-branch graph neural network according to claim 1, wherein The GraphSAGE branch updates the representation of the target node by sampling neighbor nodes and aggregating their features; The GraphSAGE branch includes two layers of convolutional operations; for the target node vi, its initial feature representation is , and the set of neighbor nodes is ; First, the first-layer convolutional operation sage1 aggregates neighbor features and updates the node representation: ; Among them, is the mean aggregation function, defined as: ; is the weight matrix of the first-layer convolution operation. The RELU activation function introduces non-linearity, and then Dropout is applied to prevent overfitting; is the representation of the first-layer convolution operation of node i; Then, the second-layer convolutional operation sage2 further updates the node representation: ; Among them, is the weight matrix of the second-layer convolution operation, and finally the node representation of the GraphSAGE branch is obtained .
4. The APT detection method based on a multi-branch graph neural network according to claim 1, wherein GIN calculates the node representation by weighted summing the neighbor node features and performing a non-linear transformation using a multi-layer perceptron; The GIN branch includes two layers of convolutional operations; for the target node vi, the first-layer convolutional operation gin1 first performs a weighted sum of the neighbor features and transforms through the MLP, as follows: ; where, ; ; , is the weight matrix of the MLP. After ReLU activation, Dropout is applied. MLP represents a multi-layer perceptron, including two layers of linear transformation, with a ReLU activation function in the middle. Then, the second-layer convolutional operation gin2 continues to update the node representation: ; Among them, is a two-layer fully connected network: ; , is the weight matrix, and finally the node representation of the GIN branch is obtained .
5. The APT detection method based on a multi-branch graph neural network according to claim 1, characterized in that, GAT assigns different weights to each neighbor node through the self-attention mechanism to capture the relative importance of the neighbor nodes to the target node; the GAT branch includes two layers of convolutional operations, and each layer uses the multi-head attention mechanism; for the target node vi, the first-layer convolutional operation gat1 uses 3 attention heads: ; Among them, is the weight matrix of the k-th attention head, is the attention weight calculated by the k-th attention head, which is dynamically generated through the self-attention mechanism: ; Among them, is the attention parameter, and || represents vector concatenation; the outputs of the three attention heads are concatenated and then activated using RELU, followed by applying Dropout; Then, the second-layer convolutional operation gat2 also uses 3 attention heads, but takes the average of the multi-head outputs: ; Among them, is the weight matrix, and finally the node representation of the GAT branch is obtained .
6. The APT detection method based on a multi-branch graph neural network according to claim 1, characterized in that, The multi-branch network model aggregates the outputs of the GAT branch, GIN branch, and GraphSAGE branch by introducing weights , and ; the final node embedding vector is expressed as: ; Aggregated node representation Subsequently, classification is performed through a linear layer and a softmax function; Node V i 's cross-entropy loss is defined as: ; where: ; N is the total number of all nodes, and C is the total number of categories. is the number of samples in category c. represents the predicted probability that sample i belongs to category c by the multi-branch network model. represents the true label, which is 1 if sample i belongs to category c, otherwise 0. represents the weight of category c. Taking the average of the losses for all samples, the total loss of the multi-branch network model is expressed as: 。 7. A method for APT detection based on a multi-branch graph neural network according to claim 1, characterized in that During the training and inference process of the multi-branch network model MBGNN, the system generates an embedding vector for each node and stores it in the database; including: 1) Check whether there is a table for storing embedding vectors; if not, create a table structure including the following fields: Node label: used to uniquely identify a node; Embedding vector: representing the feature embedding of the node; Neighbor set: storing the neighbor node information of the node in JSON format, representing the relationship between nodes; 2) Perform deduplication through the following deduplication strategies: Neighbor set deduplication: Deduplicate the neighbor set through set operations to ensure the uniqueness of each neighbor set; Bidirectional deduplication: If the neighbor set of node A1 already includes node B1, then node B1 will not be stored repeatedly when storing node A1; Neighbor set consistency check: Before inserting node information, query whether there is a record with the same neighbor set in the database. If so, reuse the existing embedding.
8. The APT detection method based on a multi-branch graph neural network according to claim 1, wherein Introduce an embedding matching mechanism based on the Overlap coefficient to measure the similarity between nodes; the calculation formula of the Overlap coefficient is as follows: ; where A and B respectively represent the neighbor sets of the current node and the nodes stored in the database; The specific process of embedding matching is as follows: The system extracts the embedding vectors and their neighbor sets of all stored nodes from the database; For each input node, calculate the Overlap coefficient between its neighbor set and the neighbor sets of the nodes stored in the database to evaluate the node similarity; If the Overlap coefficient is higher than the preset threshold, the system directly reuses the corresponding embedding vector in the database as the feature representation of the current node; otherwise, generate a new embedding vector for the node.
9. A method for APT detection based on a multi-branch graph neural network according to any one of claims 1-8, characterized in that, A neighbor-aware confidence fusion mechanism is introduced during the detection process, and the confidence of the node itself is fused with the confidence of its neighbor nodes through GAT; including: Calculate the local confidence of the target node, which is: based on the relative difference between the maximum and the second-largest category scores in the classification result; Aggregate the confidence information of neighbor nodes through the GAT branch to generate neighbor-weighted confidence; Use adjustable parameters to balance the contributions of the local confidence of the target node and the neighbor-weighted confidence to generate the final fused confidence; The relevant formulas are as follows: ; Among them, ; ; Among them, is the confidence calculated based on the self - information of node i, is the confidence calculated based on the self - information of neighbor j of node i, is the result after weighted aggregation of the confidences of all neighbors through the GAT branch. α is a parameter that controls the fusion ratio of self - information and neighbor information, usually represents the attention weight of node v i to node v j for measuring the importance of v j to v i ; Finally, based on the fusion confidence determine whether the node is abnormal, and the nodes below the threshold are marked as abnormal.
10. An APT detection system based on a multi-branch graph neural network, characterized in that, Including: A traceability graph construction module, configured to: construct a traceability graph by parsing system logs; A semantic embedding module, configured to: perform semantic embedding on the nodes in the traceability graph using Word2Vec to capture the context semantics of the nodes, and then input them into the trained multi-branch network model MBGNN to extract node representations from three perspectives of neighbor importance modeling, structural discriminability, and local structure aggregation respectively, and perform classification through a linear layer and a softmax function; The multi-branch network model includes GAT, GraphSAGE, GIN; The GAT branch uses the attention mechanism to model the importance differences between nodes and their neighbors, and strengthens the key neighbor features; The GIN branch improves the discriminability of local connection patterns through a structure-sensitive aggregation function; The GraphSAGE branch realizes feature aggregation through neighbor sampling; A neighbor-aware confidence fusion mechanism is introduced during the detection process, and the confidence of the node itself is fused with the confidence of its neighbor nodes through GAT; And introduce an embedding reuse optimization strategy to calculate node similarity and reuse the embedding representations of similar nodes.
Citation Information
Patent Citations
Intrusion detection method based on gating time convolutional network and graph
CN117579324A
APT attack detection method based on multi-dimensional edge optimization traceability graph
CN118827222A
Metadata-based abnormal data intelligent monitoring method and system
CN120123960A
Detection of adverserial attacks on graphs and graph subsets
US20210034737A1
Access point neighbor discovery for radio resource management
US20240365214A1
Cited By
Complex network anomaly detection method and system based on dynamic graph neural network
CN121644163A