Malicious attack behavior detection method based on Transform and GNN
By combining Transformer and GNN networks, the local and global dependencies of the graph are extracted and captured, and the problems of handling large-scale dynamic graphs, capturing long-range dependencies and lacking global perspectives in APT attack detection in the prior art are solved, and efficient and accurate APT attack detection is achieved.
Patent Information
- Application Number
- CN202510225577.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing technology is difficult to deal with large-scale dynamic graphs, unable to effectively capture long-range dependencies and lack of global perspectives, resulting in inefficiency and insufficient accuracy in APT attack detection.
The malicious attack behavior detection method based on Transformer and GNN is adopted to extract the local dependencies of the graph through the GNN network, combine the Transformer encoder to capture the global dependencies of the graph, and use the classification layer to determine the node type to achieve the generation of detection results.
This method can efficiently process large-scale dynamic graphs, capture long-range dependencies and provide a global perspective, significantly improving the accuracy and real-time nature of APT attack detection, and reducing the risk of false positives and underreports.
Smart Images

Figure CN119996022A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a malicious attack behavior detection method based on Transformer and GNN. Background Art
[0002] APT (Advanced Persistent Threat) attacks are complex and persistent network attacks, usually launched by organized groups, with the goal of invading network systems for a long time without being discovered. Currently, APT attack detection is performed by using graph neural networks and Transformer-based graph models, but the following problems still exist in this process:
[0003] Difficulty in processing large-scale dynamic graphs: Traditional GNN and GAT models usually assume that graphs are static and of moderate size. In practical applications, the networks involved in APT attacks are dynamically changing, and the scale of the graphs may be very large. Existing graph models are inefficient in processing dynamic large-scale graphs, making it difficult to achieve real-time detection.
[0004] Inability to effectively capture long-range dependencies: APT attacks usually involve multiple steps and long periods of latency, and there may be complex global dependencies between nodes. Traditional graph neural network models mainly focus on node information within a local neighborhood, and it is difficult to effectively capture the dependencies between distant nodes in the graph, which may miss key attack behaviors.
[0005] Lack of global perspective: APT attacks may be carried out simultaneously in multiple parts of the network. Existing graph neural network models mainly rely on local features when aggregating node information, lacking a global perspective of the entire graph. This may lead to insufficient overall understanding of attack behaviors and difficulty in accurately detecting attack activities distributed in different areas.
[0006] Therefore, there is an urgent need for an APT attack detection method that can handle large-scale dynamic graphs, capture long-range dependencies, and provide a global perspective. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a malicious attack behavior detection method based on Transformer and GNN.
[0008] In order to achieve the above object, the present invention adopts the following technical scheme.
[0009] In a first aspect, the present invention provides a malicious attack behavior detection method based on Transformer and GNN, comprising:
[0010] Obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph;
[0011] Performing word embedding generation processing on the tracing graph to obtain word embedding;
[0012] Inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result;
[0013] Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
[0014] Furthermore, the malicious attack behavior detection model includes a GNN network, a Transformer encoder, and a classification layer;
[0015] The step of inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result includes:
[0016] Using a GNN network to capture local dependencies of the graph from the traceability graph and the word embedding;
[0017] Inputting the local dependencies of the graph, the traceability graph, and the word embedding into a Transformer encoder to capture the global dependencies of the graph;
[0018] The local dependency relationship of the graph and the global dependency relationship of the graph are input into the classification layer to obtain the classification score of each node in the graph, and the node type is determined according to the classification score. The detection result includes the node type of each node.
[0019] Furthermore, the method of using the GNN network to capture the local dependency relationship of the graph from the traceability graph and the word embedding includes:
[0020] Based on the multi-head attention mechanism, the local dependency of the graph is captured from the traceability graph and the word embedding using the GNN network:
[0021]
[0022] In the formula, h i and h j are the feature vectors of node i and neighbor node j, respectively, extracted through the GNN network, W is the learnable weight matrix applied to each feature vector, a is the attention vector that determines the importance of each neighbor, α ij is the attention coefficient between node i and neighbor node j, || represents the connection operation, and k represents each neighbor node of node i, that is, is the set of neighbor nodes of node i, LeakyReLU is a nonlinear activation function, h′ i is the eigenvector h i The corresponding local dependencies, m∈M, M is the total number of attention heads, σ(·) is the nonlinear activation function;
[0023] The inputting the local dependency of the graph, the traceability graph, and the word embedding into the Transformer encoder to capture the global dependency of the graph includes:
[0024]
[0025] Among them, h″ i is the global dependency, W v is a learnable weight matrix for value projection, β ij is the attention weight, N is the total number of neighbor nodes in the neighbor node set, e ij is the attention score between node i and node j, W q and W k are the learnable projection matrices for query and key vectors, respectively, and d k is the dimension of the key vector.
[0026] Furthermore, constructing a traceability graph based on the system log data to obtain a traceability graph includes:
[0027] Nodes are created at least based on execution commands and paths in the system log data, and edges between nodes are determined at least based on data flows and network connections in the system log data, thereby obtaining a traceability graph including nodes and edges.
[0028] Furthermore, the performing word embedding generation processing on the source graph to obtain word embedding includes:
[0029] The Word2Vec model is used to perform word embedding generation processing on the traceability graph to capture the semantics and contextual information of the node interactions in the traceability graph to obtain word embedding.
[0030] Further, after determining the node type according to the classification score, the method further includes:
[0031] The Get_Adjacent function is used to analyze the current node according to its neighbor relationship, and the detection result is optimized according to the node analysis results of all nodes.
[0032] Furthermore, the malicious attack behavior detection model is based on the training log data and the corresponding label information, and is trained using the FocalLoss function, where the FocalLoss function is:
[0033]
[0034] Where α is a scaling factor used to balance the classes, and p i is the predicted probability of the true class label and γ is a focusing parameter.
[0035] In a second aspect, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the program.
[0036] In a third aspect, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.
[0037] In a fourth aspect, the present invention further provides a computer program product, comprising a computer program, wherein the computer program implements the above method when executed by a processor.
[0038] Beneficial effects of the present invention: The malicious attack behavior detection method based on Transformer and GNN provided by the present invention uses GNN to extract local features in the graph structure, such as neighbor information and local topological structure of nodes, and uses Transformer to capture global dependencies in the graph. This method can effectively fuse local and global feature information to ensure that key graph features are not missed when detecting APT attacks. Through this combination, the present invention can not only accurately identify local attack behaviors, but also discover complex attack patterns distributed throughout the network from a global graph perspective. In addition, the malicious attack behavior detection model proposed by the present invention can learn and model complex global dependencies, thereby more accurately detecting APT attacks. By using the Word2Vec semantic encoder to capture basic semantic attributes (such as process names and file paths), and combining the timing information of the graph structure, the model can identify subtle signs of attackers lurking and moving in the network. This significantly improves the accuracy of APT attack detection and reduces the risk of false positives and false negatives. Finally, the APT attack detection method based on graph neural network and Transformer of the present invention effectively improves the accuracy and real-time performance of APT attack detection by combining local and global graph features, enhances the adaptability and scalability of the system, and has significant practical value.
[0039] Additional aspects and advantages of the present invention will be given in part in the following description, which will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0041] Figure 1 One of the flow charts of a malicious attack behavior detection method based on Transformer and GNN provided in an embodiment of the present invention;
[0042] Figure 2 The second flowchart of the malicious attack behavior detection method based on Transformer and GNN provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0044] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "an" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0045] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.
[0046] Provenance-based intrusion detection devices (IDSes) mainly work by learning benign behavior models from provenance graphs. Anomalies are detected whenever there are deviations from these established models. For example, ProGrapher uses Graph2Vec and TextRCNN embeddings to identify anomalies at the graph level. Similarly, StreamSpot extracts graph features to build benign models and uses clustering techniques to identify anomalous graphs. On the other hand, Unicorn uses a graph similarity matching method to locate anomalous graphs. It simplifies provenance graph features into histograms and forms fixed-size graph sketches, which are then clustered to detect anomalies. ProvDetector adopts a path-based malware detection strategy. It can identify uncommon paths in the graph, uses path embedding generation, and adopts a local outlier factor method for anomaly detection. ThreaTrace uses GNN to perform node-level anomaly detection, learns the structural information of nodes in the benign dataset, and identifies anomalies based on deviations from the learned behavior. Finally, ShadeWatcher uses GNN to describe the preferences of system entities for interacting entities, prompting adversarial interactions through edge-level anomaly detection. Despite these advances, current attribution-based IDS still have some limitations that affect their practical application in the real world.
[0047] Attribution-based intrusion detection devices (IDSes) have gained popularity due to their potential in detecting sophisticated advanced persistent threat (APT) attacks. These IDSes employ attribution graphs created from system logs to identify potentially malicious activities. Despite their potential, they face challenges in accuracy, practicality, and scalability, especially when dealing with large attribution graphs.
[0048] In view of the problems existing in the APT attack detection method in the prior art, the present invention proposes a TransGNN model by combining the local feature extraction capability of GNN and the global dependency modeling capability of Transformer. The model has the following advantages:
[0049] Processing large-scale dynamic graphs: Through improved data loading and processing mechanisms, such as using neighbor sampling technology (NeighborLoader), TransGNN can efficiently process large-scale dynamic graphs and achieve real-time APT attack detection.
[0050] Capturing long-range dependencies: By leveraging the Transformer’s self-attention mechanism, TransGNN can capture global dependencies between nodes globally, thereby more accurately identifying complex APT attack patterns.
[0051] Providing a global perspective: The TransGNN model combines local feature extraction and global relationship modeling capabilities, enabling it to obtain a more comprehensive perspective in complex networks, thereby improving the accuracy and reliability of APT attack detection.
[0052] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0053] Example 1
[0054] See also Figure 1 and Figure 2 , a malicious attack behavior detection method based on Transformer and GNN, comprising the following steps:
[0055] S101, obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph.
[0056] In this step, nodes are created based at least on the execution commands and paths in the system log data, and edges between nodes are determined based at least on the data flows and network connections in the system log data, thereby obtaining a traceability graph including nodes and edges.
[0057] This step converts the original system log data into node features, edges, and corresponding labels to obtain a traceability graph.
[0058] S102, performing word embedding generation processing on the source graph to obtain word embedding.
[0059] Specifically, the Word2Vec model is used to perform word embedding generation processing on the traceability graph to capture the semantics and contextual information of the node interactions in the traceability graph to obtain word embedding.
[0060] S103, inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result; wherein the malicious attack behavior detection model is a model based on Transformer and GNN.
[0061] Specifically, the malicious attack behavior detection model includes a GNN network, a Transformer encoder and a classification layer. Further, S103 includes the following sub-steps:
[0062] Using a GNN network to capture local dependencies of the graph from the traceability graph and the word embedding;
[0063] Inputting the local dependencies of the graph, the traceability graph, and the word embedding into a Transformer encoder to capture the global dependencies of the graph;
[0064] The local dependency of the graph and the global dependency of the graph are input into the classification layer to obtain the classification score of each node in the graph, and the node type is determined according to the classification score. The detection result includes the node type of each node. The node type is divided into normal and abnormal. When the misclassified node is identified, the prediction result is compared with the true label, and the status flag of each node is updated.
[0065] More specifically, the GNN network consists of multiple graph attention layers (GATConv), each of which receives input node features and edge indices to generate new node representations. The GNN network contains a module with an adjustable number of layers (4 layers by default), each of which is implemented by GATConv to generate new node embeddings. Through the multi-head attention mechanism, these layers can effectively capture local and global information in the graph structure.
[0066] The GNN network also includes activation layers and dropout layers. Specifically, after each graph convolution layer, the ELU activation function is used and dropout is applied to prevent overfitting.
[0067] After calculation through the GNN network, the local dependencies of the generated graph, the traceability graph and the word embedding input are passed to the Transformer encoder. The main purpose of the Transformer encoder is to further learn the sequential dependencies between node embeddings, especially when processing large-scale graph data, Transformer can capture long-range dependencies.
[0068] The Transformer encoder consists of several layers of EncoderLayer, each with a multi-head attention mechanism and a feedforward neural network. These layers can capture complex patterns of node features.
[0069] In some embodiments of the present invention, the method of capturing the local dependency of the graph from the source graph and the word embedding using a GNN network includes:
[0070] Based on the multi-head attention mechanism, the local dependency of the graph is captured from the traceability graph and the word embedding using the GNN network:
[0071]
[0072] In the formula, h i and h j are the feature vectors of node i and neighbor node j, respectively, extracted through the GNN network, W is the learnable weight matrix applied to each feature vector, a is the attention vector that determines the importance of each neighbor, α ijis the attention coefficient between node i and neighbor node j, || represents the connection operation, and k represents each neighbor node of node i, that is, is the set of neighbor nodes of node i, LeakyReLU is a nonlinear activation function, h′ i is the eigenvector h i The corresponding local dependencies, m∈M, M is the total number of attention heads, and σ(·) is the nonlinear activation function.
[0073] The inputting the local dependency of the graph, the traceability graph, and the word embedding into the Transformer encoder to capture the global dependency of the graph includes:
[0074] h″ i =∑j=1 N β ij (W v h′ j )
[0075]
[0076] Among them, h″ i is the global dependency, W v is a learnable weight matrix for value projection, β ij is the attention weight, N is the total number of neighbor nodes in the neighbor node set, e ij is the attention score between node i and node j, W q and W k are the learnable projection matrices for query and key vectors, respectively, and d k is the dimension of the key vector, used for scaling.
[0077] Specifically, this application captures local interactions and long-term dependencies between nodes by combining graph neural networks (GNNs) with Transformer layers to identify advanced persistent threats (APTs). The method provided by this application overcomes the limitations of traditional GNNs, which are usually effective in local message passing but have difficulty capturing global dependencies in large, interconnected graphs. The following details each stage of the representation learning process:
[0078] 1. Locally Dependent Graph Neural Network (GNN) Layer
[0079] The first stage of graph representation learning is to apply graph neural network (GNN) layers, specifically graph attention network (GAT), to model local dependencies within the neighborhood of each node. In the GAT layer, a node aggregates information from its neighboring nodes and weights them through an attention mechanism to determine the importance of each neighboring node.
[0080] For a feature vector h i The node i and its neighbor node set The attention coefficient α between node i and neighbor node j ij The calculation is as follows:
[0081]
[0082] Among them, h i and h j are the feature vectors of node i and neighbor node j, respectively, extracted through the GNN network, W is the learnable weight matrix applied to each feature vector, a is the attention vector that determines the importance of each neighbor, α ij is the attention coefficient between node i and neighbor node j, || represents the connection operation, and k represents each neighbor node of node i, that is, The denominator of the above formula is The sum of all neighbor nodes k in α is calculated. ij Normalize so that the attention coefficients of all neighbors form a probability distribution, is the set of neighbor nodes of node i, and LeakyReLU is a nonlinear activation function.
[0083] Attention coefficient α ij reflects the relevance of node j to node i. The final representation of node i is calculated as the weighted sum of neighbor features:
[0084]
[0085] Among them, σ is a nonlinear activation function such as ReLU, which allows each node to integrate information from its neighbors in a weighted manner according to its relevance, thereby facilitating the task of identifying suspicious activities. Through this localized attention mechanism, GNNFormer is able to focus on important interactions within the neighborhood of each node, capturing local patterns that may indicate attacks or abnormal behavior.
[0086] 2. Multi-head attention mechanism to enhance expressiveness
[0087] In order to enhance the expressiveness of the GAT layer, a multi-head attention mechanism is adopted. In this setting, multiple attention heads work in parallel, and each head learns different aspects of local interactions. For each attention head m, the coefficient and node feature h′ i Computed independently, the outputs of all heads are then concatenated:
[0088]
[0089] Where M is the number of attention heads and || represents the connection operation. Multi-head attention enhances the model's ability to capture complex dependencies by allowing different parts of the network to focus on different aspects of the local graph structure. The aggregation of multiple heads ensures that GNN can capture diverse signals, making it more robust to changes in node behavior and interactions.
[0090] 3. Globally Dependent Transformer Layer
[0091] Although GNNs excel in capturing local dependencies, they are less effective in modeling long-range dependencies across the entire graph. To this end, we introduce a Transformer layer after the GNN layer to capture the global context. The self-attention mechanism in the Transformer enables each node to directly pay attention to all other nodes regardless of their distance in the graph, which is critical for understanding distributed and multi-stage attack patterns.
[0092] The Transformer layer calculates the attention score e between node i and node j using the following formula ij :
[0093]
[0094] Among them, W q and W k is the learnable projection matrix for query and key vectors, d k is the dimension of the key vector, used for scaling.
[0095] Use the softmax function to score e ij Normalize to generate attention weights β ij :
[0096]
[0097] The output of each node i is the weighted sum of all node features:
[0098]
[0099] Among them, W v is a learnable weight matrix for value projection. This operation enables GNNFormer to integrate information from nearby and distant nodes, providing a global view of the graph.
[0100] 4. Combination of local and global representations
[0101] By combining GNN layers with Transformer layers, GNNFormer is able to gain both local and global perspectives from the graph. The GNN layer provides a localized representation that highlights the immediate neighborhood and captures structural details. The Transformer layer captures long-range dependencies and context, which is critical for understanding multi-stage attacks across parts of the graph. This hybrid approach enables GNNFormer to effectively model both fine-grained and high-level dependencies, making it particularly suitable for detecting complex and subtle APT patterns.
[0102] 5. Hierarchical Attention Mechanism
[0103] To further refine the representation, GNNFormer uses a hierarchical attention mechanism to prioritize key nodes and edges. At the node level, the attention score highlights highly relevant nodes that may be associated with attacks, such as nodes associated with suspicious processes or abnormal network activities. At the graph level, the attention mechanism focuses on specific subgraphs that exhibit abnormal patterns, emphasizing interactions that indicate coordinated malicious behavior. This hierarchical structure enables GNNFormer to filter out irrelevant information and focus computational resources on high-risk areas in the graph.
[0104] 6. Final Representation and Classification
[0105] After processing through multiple GNN and Transformer layers, each node is represented by a feature vector that encodes local and global dependencies. These final node representations are used for classification, and a scoring function is used to assign a probability to each node or subgraph, indicating the likelihood of malicious activity. This final classification step leverages the context-aware representations learned by the model to accurately distinguish between benign and malicious behavior.
[0106] In some embodiments of the present invention, the attack detection phase in GNNFormer focuses on node-level classification within the provenance graph. This phase evaluates each node individually to determine its likelihood of being involved in malicious activity. This detection phase combines confidence-based screening and anomaly scoring to enhance the model's focus on high-risk nodes.
[0107] 1. Node-level classification
[0108] Each node embedding h obtained through the GNN and Transformer layers i , classified as benign or malicious based on the softmax probability score. This is done by feeding the embedding into a fully connected layer with a softmax activation:
[0109] P(malicious|h i )=σ(W 分 *h i +b)
[0110] Among them, W 分 is the weight matrix of the classification layer, b is the bias term, σ is the softmax function, and the output probability is between 0 and 1. P(malicious|h i ) is the probability that node i is classified as malicious.
[0111] Nodes exceeding a defined threshold are marked as potentially malicious. Threshold-based classification helps filter out benign nodes and focus the analysis on nodes with a higher probability of malicious behavior.
[0112] 2. Use FocalLoss to deal with category imbalance
[0113] To manage the inherent class imbalance between benign and malicious nodes, Focal Loss is used. This loss function mitigates the impact of a large number of benign nodes and enables the model to focus on a small number of malicious classes by increasing the weight of difficult-to-classify samples. FocalLoss is calculated as follows:
[0114]
[0115] Where α is a scaling factor used to balance the classes, and p i is the predicted probability of the true category label, and γ is the focusing parameter, which is used to reduce the weight of easy-to-classify samples and highlight the misclassified nodes.
[0116] Using Focal Loss helps GNNFormer to prioritize detecting malicious nodes, thereby improving accuracy and reducing false negatives in APT detection.
[0117] 3. Confidence-based screening
[0118] Each node is assigned a confidence score to assess the certainty of its classification. This confidence score is calculated by comparing the softmax scores of the top two classes:
[0119]
[0120] Among them, P top1 and P top2 are the highest and second highest probabilities. This normalized confidence ranges between 0 and 1 and reflects how certain the model is about its prediction.
[0121] Nodes with confidence scores above a certain threshold are retained as possible candidates for malicious activity, which enables the model to filter out low-confidence predictions, thereby reducing the likelihood of false positives.
[0122] 4. Anomaly scoring based on node centrality
[0123] In addition to classification confidence, GNNFormer assigns an anomaly score to each node based on the classification probability and the importance of the node in the graph. This score reflects not only the possibility of malicious behavior, but also the role of the node in the traceability graph, which is crucial for identifying key or central nodes with higher risks.
[0124] The anomaly score S of node i i The calculation is as follows:
[0125] Si=P(malicious|h i )×NodeImportance(i)
[0126] Among them, P(malicious|h i ) is the probability that node i is classified as malicious, and NodeImportance(i) is a measure of the centrality of node i in the graph (e.g. PageRank or eigenvector centrality).
[0127] Nodes with high anomaly scores are prioritized for further inspection, as these scores indicate high classification confidence and significant influence in the graph structure.
[0128] 5. Iterative Misclassification Reduction
[0129] To further improve classification accuracy, the implementation uses an iterative optimization process where nodes that were misclassified in the previous cycle are masked out and re-evaluated in subsequent iterations. This optimization helps improve the overall classification accuracy by focusing on the initially misclassified nodes.
[0130] In some embodiments of the present invention, after determining the node type according to the classification score, the method further includes:
[0131] The Get_Adjacent function is used to analyze the current node according to its neighbor relationship, and the detection result is optimized according to the node analysis results of all nodes.
[0132] In addition, the malicious attack behavior detection model is based on training log data and corresponding label information, and is trained using the FocalLoss function. The FocalLoss function is particularly suitable for dealing with class imbalance problems, especially for anomaly detection tasks where there are relatively few positive samples.
[0133] During the training process, in each iteration, NeighborLoader is used to batch the training traceability graph and sample subgraphs. The malicious attack behavior detection model is also trained through the back-propagation algorithm, and the gradient is clipped to stabilize the training process. After each epoch, the misclassified nodes are tracked and processed. When the trained malicious attack behavior detection model is obtained, its performance indicators such as precision, recall rate, F1 value, false alarm rate and true positive rate are calculated. These indicators provide a comprehensive evaluation of the model effect. Auxiliary functions are used to analyze the misclassification situation, and the true positive and false positive sets are updated through the information of two-hop neighbors to improve the prediction results.
[0134] In addition, after obtaining the system log data to be detected, the system log data will be preprocessed, and a traceability graph will be constructed based on the preprocessed system log data to obtain a traceability graph.
[0135] Specifically, the system log data to be detected is generally stored in JSON format, and the UUID and node type are extracted by using regular expressions. These UUIDs represent nodes in the graph, and edges are created based on the interactions between these nodes. After reading the system log data to be detected in JSON format, irrelevant data is filtered out and the UUID is mapped to the corresponding node type. The edge represents the interaction between nodes, such as data flow or communication event. Finally, the preprocessed system log data is obtained.
[0136] In other embodiments of the present invention, the malicious attack behavior detection model is a model based on LSTM and GNN. The long short-term memory network LSTM performs well in processing time series data and can effectively capture the temporal dependencies in the data. In APT attack detection, the LSTM network can be used to process the time series features of graph nodes or edges. Combining LSTM with GNN can enhance the ability to capture time dependencies and better deal with long-range dependency problems in APT attacks.
[0137] The malicious attack behavior detection method based on Transformer and GNN provided in an embodiment of the present invention uses GNN to extract local features in the graph structure, such as neighbor information and local topological structure of nodes, and uses Transformer to capture global dependencies in the graph. This method can effectively fuse local and global feature information to ensure that key graph features are not missed when detecting APT attacks. Through this combination, the present invention can not only accurately identify local attack behaviors, but also discover complex attack patterns distributed throughout the network from a global graph perspective. In addition, the malicious attack behavior detection model proposed in the present invention can learn and model complex global dependencies, thereby more accurately detecting APT attacks. By using the Word2Vec semantic encoder to capture basic semantic attributes (such as process names and file paths), and combining the timing information of the graph structure, the model can identify subtle signs of attackers lurking and moving in the network. This significantly improves the accuracy of APT attack detection and reduces the risk of false positives and false negatives. Finally, the APT attack detection method based on graph neural network and Transformer of the present invention effectively improves the accuracy and real-time performance of APT attack detection by combining local and global graph features, enhances the adaptability and scalability of the system, and has significant practical value.
[0138] Example 2
[0139] On the basis of Example 1, this Example 2 provides a malicious attack behavior detection device based on Transformer and GNN, which corresponds to the malicious attack behavior detection based on Transformer and GNN, and specifically includes:
[0140] A data acquisition module is used to acquire system log data to be detected, and to construct a traceability graph based on the system log data to obtain a traceability graph;
[0141] A word embedding processing module, used for performing word embedding generation processing on the source tracing graph to obtain word embedding;
[0142] A detection module, used for inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result;
[0143] Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
[0144] For details, please refer to the description of the malicious attack behavior detection method based on Transformer and GNN, which will not be repeated here.
[0145] Example 3
[0146] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a malicious attack behavior detection method based on Transformer and GNN, the method including the following process steps:
[0147] Obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph;
[0148] Performing word embedding generation processing on the tracing graph to obtain word embedding;
[0149] Inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result;
[0150] Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
[0151] Example 4
[0152] Embodiment 4 of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, a malicious attack behavior detection method based on Transformer and GNN is implemented. The method includes the following process steps:
[0153] Obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph;
[0154] Performing word embedding generation processing on the tracing graph to obtain word embedding;
[0155] Inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result;
[0156] Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
[0157] Example 5
[0158] Embodiment 5 of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, a malicious attack behavior detection method based on Transformer and GNN is implemented. The method includes the following process steps:
[0159] Obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph;
[0160] Performing word embedding generation processing on the tracing graph to obtain word embedding;
[0161] Inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result;
[0162] Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
[0163] In summary, the Transformer- and GNN-based malicious attack behavior detection methods, devices, electronic devices, media, and products provided in the embodiments of the present invention can be applied to products and projects in multiple network security fields, and can efficiently detect and defend against complex network attacks, thereby improving the security of network systems. In addition, it is suitable for network traffic monitoring and attack detection within an enterprise, helping enterprises identify and defend against complex APT attacks, and protecting the sensitive data and assets of the enterprise. In addition, it can also continuously monitor network traffic and host behavior, identify potential APT attacks through real-time analysis and graph neural network models, and automatically generate security alerts to help security teams respond and handle threats in a timely manner.
[0164] The malicious attack behavior detection model of the present invention can efficiently process large-scale network graph data, realize real-time detection and response to APT attacks, and ensure the security of the system. Through the combination of graph neural network and Transformer, it is possible to capture long-range dependencies and complex relationships in the network, and effectively respond to distributed, multi-stage APT attacks. In addition, it can also adapt to network environments of different scales and complexities, and has good scalability. Through the multi-task learning method, the detection performance of different types of attacks can be optimized at the same time, and the robustness of detection can be enhanced.
[0165] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0166] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the method or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The method and system embodiments described above are merely schematic, in which the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0167] The above are only preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A malicious attack behavior detection method based on Transformer and GNN, characterized in that: include: Obtaining system log data to be detected, and constructing a traceability graph based on the system log data to obtain a traceability graph; Performing word embedding generation processing on the tracing graph to obtain word embedding; Inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result; Among them, the malicious attack behavior detection model is a model based on Transformer and GNN.
2. The method according to claim 1, characterized in that The malicious attack behavior detection model includes a GNN network, a Transformer encoder and a classification layer; The step of inputting the traceability graph and the word embedding into a pre-trained malicious attack behavior detection model to obtain a detection result includes: Using a GNN network to capture local dependencies of the graph from the traceability graph and the word embedding; Inputting the local dependencies of the graph, the traceability graph, and the word embedding into a Transformer encoder to capture the global dependencies of the graph; The local dependency relationship of the graph and the global dependency relationship of the graph are input into the classification layer to obtain the classification score of each node in the graph, and the node type is determined according to the classification score. The detection result includes the node type of each node.
3. The method according to claim 1, characterized in that The method of using a GNN network to capture local dependencies of the graph from the source graph and the word embedding includes: Based on the multi-head attention mechanism, the local dependency of the graph is captured from the traceability graph and the word embedding using the GNN network: In the formula, h i and h j are the feature vectors of node i and neighbor node j, respectively, extracted through the GNN network, W is the learnable weight matrix applied to each feature vector, a is the attention vector that determines the importance of each neighbor, α ij is the attention coefficient between node i and neighbor node j, || represents the connection operation, and k represents each neighbor node of node i, that is, is the set of neighbor nodes of node i, LeakyReLU is a nonlinear activation function, h′ i is the eigenvector h i The corresponding local dependencies, m∈M, M is the total number of attention heads, σ(·) is the nonlinear activation function; The inputting the local dependency of the graph, the traceability graph, and the word embedding into the Transformer encoder to capture the global dependency of the graph includes: Among them, h" i " is a global dependency, W v is a learnable weight matrix for value projection, β ij is the attention weight, N is the total number of neighbor nodes in the neighbor node set, e ij is the attention score between node i and node j, W q and W k are the learnable projection matrices for query and key vectors, respectively, and d k is the dimension of the key vector.
4. The method according to claim 1, characterized in that: The constructing a traceability graph based on the system log data to obtain a traceability graph includes: Nodes are created at least based on execution commands and paths in the system log data, and edges between nodes are determined at least based on data flows and network connections in the system log data, thereby obtaining a traceability graph including nodes and edges.
5. The method according to claim 1, characterized in that The performing word embedding generation processing on the source tracing graph to obtain word embedding includes: The Word2Vec model is used to perform word embedding generation processing on the traceability graph to capture the semantics and contextual information of the node interactions in the traceability graph to obtain word embedding.
6. The method according to claim 2, characterized in that After determining the node type according to the classification score, the method further includes: The Get_Adjacent function is used to analyze the current node according to its neighbor relationship, and the detection result is optimized according to the node analysis results of all nodes.
7. The method according to claim 1, characterized in that The malicious attack behavior detection model is based on training log data and corresponding label information, and is trained using the FocalLoss function.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium, characterized in that: The device stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Advanced continuous threat detection method based on traceability graph and heterogeneous graph neural network
CN114900364A
APT attack traceability analysis method based on bidirectional long and short time memory network
CN115567306A
APT attack detection method based on traceability graph and self-attention mechanism
CN116192421A
Attack detection method and device, electronic equipment and storage medium
CN117375998A
Network threat deduction system and method based on Transform and graph attention network model
CN119449452A
Cited By
Malicious deletion traceability method of distributed file system based on block chain
CN120386768A
APT attack path traceability method and device based on traceability graph
CN121309047A