APT attack behavior pattern recognition method based on MPRSAGE model
Through the multi-channel design MPRSAGE model, combined with causal weight and attention mechanism, the problem of underutilization of causal relationships in the existing technology is solved, and high-precision identification and defense of APT attack behavior is achieved.
Patent Information
- Application Number
- CN202510660816.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
The existing APT attack prediction method based on graph neural networks fails to fully consider the causal relationship in attack behavior, and there are problems of information loss and model performance limitation in the comprehensive processing of structural diagrams and feature diagrams.
The MPRSAGE model with multi-channel design is adopted, and the GraphSAGE layer of causal weights, attention mechanism and residual connection blocks are combined with position coding to enhance the correlation between the structural diagram and the feature diagram to achieve comprehensive modeling of multi-dimensional features.
It significantly improves the recognition accuracy of APT attack behavior patterns and the stability of the model, breaks through the limitations of a single graph structure or node characteristics, and provides a more reliable defense method.
Smart Images

Figure CN120561685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graph neural network model application technology in network security technology, and more particularly to a method for recognizing APT attack behavior patterns based on the MPRSAGE (a multi-channel causal GraphSAGE with Position Encoding and Residual Connections in Attention, MPRSAGE) model. Background Art
[0002] As cyberattacks become increasingly complex, APT attacks, due to their persistence, stealth, and diversity, have become a major challenge in network security. Traditional attack prediction methods rely on rules or shallow feature analysis, making it difficult to effectively identify the complex behavioral patterns of APT attacks. In recent years, graph neural network technology has been gradually applied to network threat detection due to its ability to effectively process structured data. However, most existing APT attack prediction methods based on graph neural networks fail to fully consider the causal relationships in attack behavior patterns and lack the comprehensive processing of structural and feature graphs.
[0003] To this end, this paper proposes a method for identifying APT attack behavior patterns based on the MPRSAGE model. This method addresses the shortcomings of existing technologies by designing multi-channel processing structure graphs and feature graphs within a GraphSAGE model that incorporates causal weights and incorporating an improved attention mechanism, thereby improving the effectiveness of APT attack behavior pattern recognition. Summary of the Invention
[0004] In order to solve the problems that traditional APT attack prediction methods have poor recognition effect when dealing with complex and hidden attack behavior patterns, the existing graph neural network model fails to fully explore the causal relationship in attack behavior, and there is information loss and limited model performance when processing the fusion of structure graphs and feature graphs, the present invention provides an APT attack behavior pattern recognition method based on the MPRSAGE model.
[0005] A method for identifying APT attack behavior patterns based on the MPRSAGE model is implemented by the following steps:
[0006] Step 1: Obtain the feature matrix X and structure graph A of APT attack data, and construct the feature graph F;
[0007] Step 2: Construct an MPRSAGE model, which consists of three layers of GraphSAGE layers based on causal weights, a positional encoding and residual connection block based on the attention mechanism, and an MLP classifier;
[0008] Step 3: Initialize the MPRSAGE model and set the number of epochs; divide the APT attack dataset into a training set and a test set; and use the training set to train the MPRSAGE model.
[0009] Step 4: Determine whether the current epoch is less than or equal to the total number of times. If so, go to step 5; otherwise, go to step 9.
[0010] Step 5: Input the feature matrix X and the structure graph A of step 1 into the first layer of GraphSAGE based on causal weights to learn the embedded information Z T ;
[0011] Input the structure graph A and feature graph F described in step 1 into the second layer of GraphSAGE based on causal weights to learn the embedded feature Z CT and embedding feature Z CF , and embed the feature Z CT and embedding feature Z CF Average to obtain common information Z C ;
[0012] Input the feature matrix X and feature graph F described in step 1 into the third layer of GraphSAGE based on causal weights to learn the embedded information Z F ;
[0013] Step 6: embed the information Z obtained in step 5 T , common information Z C and embedded information Z F Input the position encoding based on the attention mechanism and the residual connection block to perform position addition, residual connection and weighted summation calculation to obtain output features;
[0014] Step 7: Input the output features into the MLP classifier to perform classification processing, obtain node probabilities, and complete the training of the MPRSAGE model;
[0015] Step 8: Increase the number of loops epoch by 1 and return to step 4.
[0016] Step 9: Use the test set to test the trained MPRSAGE model to achieve pattern recognition of APT attack data.
[0017] Beneficial effects of the present invention:
[0018] The identification method described in this invention addresses many challenges faced by existing methods in APT classification and prediction. These include a lack of comprehensive extraction and effective modeling of multidimensional features within APT attack sequences; or existing methods often rely on single graph structures or node features for training, failing to fully integrate and integrate information from structural and feature graphs. By proposing the MPRSAGE model, this invention provides an innovative solution to key issues in identifying APT attack behavior patterns.
[0019] The MPRSAGE model described in the present invention adopts a multi-channel design. The layer of each channel is a GraphSAGE layer designed based on causal weights. It performs feature sampling and aggregation representation on the structural graph and feature graph of the APT attack respectively, and effectively establishes the correlation and similarity between the structural graph and the feature graph through a common channel, breaking through the limitations of traditional methods that rely only on a single graph structure or node features. At the same time, the introduction of the attention mechanism of position encoding and residual connection further enhances the model's ability to capture potential correlations between APT attack elements, and significantly improves the stability of model training and the gradient transfer effect. This design idea realizes the comprehensive modeling of multi-dimensional features in APT attack behavior, effectively improves the accuracy of classification and prediction, and provides a new technical path for APT attack defense.
[0020] In summary, the detection method described in the present invention is reliable and provides a theoretical basis and broad application prospects for the identification and defense of APT attack behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of an APT attack behavior pattern recognition method based on the MPRSAGE model described in the present invention;
[0022] Figure 2 Flowchart of the GraphSAGE layer based on causal weights in the present invention;
[0023] Figure 3 This is a framework diagram of an APT attack behavior pattern recognition method based on the MPRSAGE model described in the present invention;
[0024] Figure 4 This is a comparison diagram of the MPRSAGE model described in the present invention and other existing models. DETAILED DESCRIPTION
[0025] Specific implementation method 1. Combination Figures 1 to 4 This embodiment describes an APT attack behavior pattern recognition method based on the MPRSAGE model. The method is implemented by the following steps:
[0026] Step 1: Obtain APT attack data and calculate the feature matrix X and structure diagram A; construct a feature diagram F based on the feature matrix X; divide the node CVE in the feature matrix X into the first 80% as a training set and the last 20% as a test set;
[0027] In this embodiment, each row in the feature matrix X represents a piece of attack data, which is a combination of vulnerabilities (CVEs), software weaknesses (CWEs), attack patterns (CAPECs), threat indicators and advanced threats (IOCs), and APTs. Each row of CWEs, CAPECs, and IOCs represents a node CVE, i.e., a feature vector for a vulnerability (CVE), with APT being the node CVE category.
[0028] The relationships between node CVEs form structure graph A. Structure graph A is a graph consisting of CVEs and CVEs, representing the relationships between CVEs. For example, in structure graph A, if nodes CVE-2017-11882 and CVE-2017-12824 are found to belong to the same APT-C-01 category during crawling, then an edge exists between nodes CVE-2017-11882 and CVE-2017-12824, indicating that node CVE-2017-12824 is a neighbor of node CVE-2017-11882.
[0029] In this embodiment, according to the feature matrix X, the Jaccard similarity method is used to calculate the top two node CVEs with the highest similarity scores for each node CVE in the feature matrix X. For example, the similarity score between the node CVE-2021-27065 and the node CVE-2021-40444 is 0.75, the similarity score between CVE-2021-27065 and CVE-2018-20250 is 0.33, and the similarity score between CVE-2021-27065 and CVE-2 The similarity score of 018-14847 is 0.16. Similarly, the similarity scores of CVE-2021-27065 and all nodes are calculated. The similarity scores of CVE-2021-27065 with CVE-2021-40444 and CVE-2018-20250 are the top two, indicating that CVE-2021-27065 is related to CVE-2021-40444, and CVE-2021-27065 is related to CVE-2018-20250. Calculate the similarity of each node CVE in the feature matrix X to form the feature graph F;
[0030] In this embodiment, the APT attack data records the behavioral characteristics, attack methods, affected systems, etc. To obtain data, common attack knowledge such as software vulnerabilities, weaknesses, attack patterns, APT organizations, and IOC information is crawled from major open source network security knowledge bases using crawler technology.
[0031] Step 2: Construct an MPRSAGE model; the MPRSAGE model consists of three layers of GraphSAGE layers based on causal weights, position encoding and residual connection blocks based on the attention mechanism, and an MLP classifier;
[0032] Initialize the MPRSAGE model, set the number of epochs of the MPRSAGE model, and train the MPRSAGE model. The overall task of the MPRSAGE model is to identify the APT category to which the CVE of the node in the test set belongs.
[0033] Step 3: Determine whether the current epoch is less than or equal to the total number of times. If so, go to step 4; otherwise, go to step 8.
[0034] Step 4: Input the feature matrix X and the structure graph A into the first layer of GraphSAGE layer (GraphSAGE1) based on causal weights to learn the embedded feature Z T ;
[0035] Input the structure graph A and feature graph F into the second layer of GraphSAGE layer (GraphSAGE2) based on causal weights to learn the feature Z CT and feature Z CF , the feature Z CF and feature Z CF Average to get the common embedding feature Z C ;
[0036] Input the feature matrix X and feature graph F into the third layer of GraphSAGE layer (GraphSAGE3) based on causal weights, and learn the embedded feature Z F ;
[0037] Step 5: Get the feature Z T 、Z C and Z F Input into the position encoding and residual connection block based on the attention mechanism to obtain the final output value inal_output; the specific process is:
[0038] Step 51: Add positional information positional_encoding to the input features. The calculation method is:
[0039] z′i =X new(i) +positional_encoding
[0040] Where, X new(i) is the feature of the i-th attack data without adding position encoding, z' i Add position encoding features to the i-th attack data. Figure 3 The attention mechanism section lists several CVEs. For example, the position encoding of CVE-1 is Positional Encoding 1, so the node CVE-1 adds position information Positional Encoding 1.
[0041] Step 52: Calculate the attention weight; obtain the attention weight w and the weight w of the i-th attack data through two-layer linear transformation and activation function. i The calculation method is:
[0042] w i =Linear2(tanh(Linear1(z′ i )))
[0043] Where Linear1 is the first-layer linear transformation, tanh is the activation function, and Linear2 is the second-layer linear transformation, and we get w i ;
[0044] Step 53: Use the softmax function to normalize the weights and calculate the weighted sum of each attack data. Finally, use the residual connection to add the input and output of the self-attention mechanism to obtain the final output value final_output. The calculation method is:
[0045]
[0046] Where, is the weighted sum of the attack data through the softmax function, E represents the exponential function, ∑ i X new(i) It is a simple sum of the original features of the attack data;
[0047] Step 6: Input the final output value (final_output) from Step 5 into the MLP classifier. The fully connected layer and the LogSoftmax function output the logarithmic probability for classification. For example, if the node CVE-2014-6352 is connected to CWE-94, CAPEC-77, and IOC-1, this attack data will be classified as APT-C-09 using the MPRSAGE model.
[0048] Step 7: Increase the number of loops epoch by 1 and return to step 3.
[0049] Step 8: Use the test set to test the trained MPRSAGE model to achieve pattern recognition of APT attack data.
[0050] like Figure 2 As shown, in this embodiment, the specific implementation steps of one of the three GraphSAGE layers based on causal weights described in step 2 (the implementation process of each GraphSAGE layer based on causal weights is the same) are:
[0051] Step 21: Use mean aggregation to calculate the initial aggregation matrix of neighbor features using the feature matrix X and the structure diagram A. The specific calculation method is:
[0052]
[0053] Based on the nonlinear transformation of neighbor features and the feature mean, the causal weight c_w of each node is calculated. The specific calculation method is:
[0054]
[0055] Where mean() is the average function, S is the Sigmoid function, dim is the feature dimension, which is the number of columns in the feature matrix X, and ∈ is a small constant to prevent the denominator from being zero;
[0056] Step 22: Determine whether each node CVE in the feature matrix X has a causal weight. If not, use the node CVE's own features and neighbor features to initialize the aggregation matrix. Splicing to get splicing features Execute step 23; otherwise, aggregate the initial mean matrix of neighbor features Perform causal weight calculation to obtain a new weighted matrix Then it is spliced with the node CVE's own features to obtain the spliced features Go to step 23;
[0057] In this embodiment, the weighting matrix The calculation method is:
[0058]
[0059] Step 23: Generate a new feature representation of the node through linear transformation and then nonlinear activation. After the causal weight judgment in step 2, the new feature representation matrix of the node CVE may have two kinds: or The calculation method is:
[0060]
[0061] In the formula, Leaky Relu() is used as the activation function, W is the weight matrix of the activation function, and b is the bias term.
[0062] The final learned embedding feature obtained in the above step 4 is the new feature representation matrix of each node CVE, which maps each node to a vector space.
[0063] Specific implementation method 2: Figure 4 This embodiment is described as a verification example of the APT attack behavior pattern recognition method based on the MPRSAGE model described in the first embodiment;
[0064] To verify the accuracy of the experiment, we compared the MPRSAGE model described in this paper with five other models suitable for APT attack data: GraphSAGE, UGCN, GAN_LSTM, SR2APT, and BiADG. This experiment provides a comprehensive evaluation of the model's overall performance. Table 1 shows the experimental results for different models.
[0065] Table 1
[0066] Models Precision Recall F1Score Acc MPRSAGE 0.987 0.9765 0.9609 0.986 GraphSAGE 0.9725 0.9438 0.9523 0.9738 UGCN 0.9404 0.9633 0.9493 0.9633 GAN-LSTM 0.9375 0.9161 0.9482 0.9633 SR2APT 0.9616 0.9231 0.9464 0.9695 BiADG 0.9503 0.9345 0.952 0.9521
[0067] As can be seen from Table 1, the MPRSAGE model described in this embodiment is compared with the graph neural network-based models GraphSAGE and UGCN. In Table 1, the four indicators of the MPRSAGE model all perform best, indicating that it has strong comprehensive performance in classification and recognition tasks. The recall rate (Recall) is 0.0327 higher than that of the GraphSAGE model, but the precision (Precision) and accuracy (Acc) are excellent, which is suitable for processing general graph structure data. UGCN is slightly inferior in precision (Precision) and F1 score (F1Score), which are 0.047 and 0.011 lower than the MPRSAGE model respectively, and may be suitable for tasks where recall is more important. The GAN_LSTM model, a method based on traditional machine learning, is slightly lower in precision (Precision) and recall rate (Recall), but performs relatively well in accuracy (Acc), which may be affected by the integration of generative adversarial networks and LSTM.
[0068] like Figure 4As shown in the figure, in this implementation, the MPRSAGE model is compared with the SR2APT and BiADG models. SR2APT and BiADG have relatively balanced overall performance, with similar accuracy (Acc) and F1 score (F1Score). SR2APT performs slightly better, especially in precision and recall, demonstrating its advantages in specific scenarios. Overall, MPRSAGE is the most comprehensive method.
[0069] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0070] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. The APT attack behavior pattern recognition method based on the MPRSAGE model is characterized by: The method is implemented by the following steps: Step 1: Obtain the feature matrix X and structure graph A of APT attack data, and construct the feature graph F; Step 2: Construct an MPRSAGE model, which consists of three layers of GraphSAGE layers based on causal weights, a positional encoding and residual connection block based on the attention mechanism, and an MLP classifier; Step 3: Initialize the MPRSAGE model and set the number of epochs; divide the APT attack dataset into a training set and a test set; and use the training set to train the MPRSAGE model. Step 4: Determine whether the current epoch is less than or equal to the total number of times. If so, proceed to step 5. Otherwise, go to step nine; Step 5: Input the feature matrix X and the structure graph A of step 1 into the first layer of GraphSAGE based on causal weights to learn the embedded information Z T ; Input the structure graph A and feature graph F described in step 1 into the second layer of GraphSAGE based on causal weights to learn the embedded feature Z CT and embedding feature Z CF , and embed the feature Z CT and embedding feature Z CF Average to obtain common information Z C ; Input the feature matrix X and feature graph F described in step 1 into the third layer of GraphSAGE based on causal weights to learn the embedded information Z F ; Step 6: embed the information Z obtained in step 5 T , common information Z C and embedded information Z F Input the position encoding based on the attention mechanism and the residual connection block to perform position addition, residual connection and weighted summation calculation to obtain output features; Step 7: Input the output features into the MLP classifier to perform classification processing, obtain node probabilities, and complete the training of the MPRSAGE model; Step 8: Increase the number of loops epoch by 1 and return to step 4. Step 9: Use the test set to test the trained MPRSAGE model to achieve pattern recognition of APT attack data.
2. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 1 is characterized in that: In step 2, the specific implementation process of the GraphSAGE layer based on causal weights is as follows: Step 2.1: Use mean aggregation to calculate the initial aggregation results of the adjacent features of the feature matrix X and the structure diagram A of the APT attack data; and calculate the causal weight of each node based on the nonlinear transformation of the adjacent features and the feature mean; Step 22: Determine whether each node has causal weight; If not, execute steps 2 and 3; if yes, perform causal weight adjustment on the initial mean aggregation result of the neighbor feature to obtain the weighted neighbor feature, and execute steps 2 and 3; Steps 2 and 3: Use the own features and neighbor features to splice, and generate updated node features through linear transformation and then nonlinear activation.
3. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 2 is characterized in that: In step 21, the initial aggregation result of the neighbor features The specific calculation formula is: Based on the nonlinear transformation of neighbor features and the feature mean, the causal weight c_w of each node is calculated using the formula: Where mean() is the average function, S is the Sigmoid function, dim is the feature dimension, which is the number of columns of the feature matrix X, and ∈ is a small constant to prevent the denominator from being zero.
4. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 2, characterized in that: In step 22, the calculation formula of the weighted neighbor feature is:
5. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 2 is characterized in that: In steps 2 and 3, the calculation formula for the updated node features is: Where Leaky Relu() is the activation function, W is the weight matrix of the activation function, b is the bias term, and X new It is the updated feature of the attack data obtained after linear transformation and activation function.
6. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 1 is characterized in that: In step 6, the specific implementation process of the position encoding and residual connection block based on the attention mechanism is as follows: Step 6.1: Add positional information positional_encoding to the input features. Step 6.2: Calculate the attention weight. Obtain the attention weight through two layers of linear transformation and activation function in the model. Step 6.3: Use the softmax function to normalize the attention weights and calculate the weighted sum of each attack data; finally, use the residual connection to add the input and output of the position encoding based on the attention mechanism and the residual connection block as the final output feature.
7. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 6 is characterized in that: The attention weight calculation method in step 62 is: w i =Linear2(tanh(Linear1(z′ i ))) Where Linear1 is the first-layer linear transformation, tanh is the activation function, Linear2 is the second linear transformation, and w i is the weight of the i-th attack data.
8. The APT attack behavior pattern recognition method based on the MPRSAGE model according to claim 6 is characterized in that: In step 63, the calculation formula for the final output feature is: Where, is the weighted sum of attack data through the softmax function, E is the exponential function, ∑ i z i is the simple sum of the original features of the attack data, w i is the weight of the i-th attack data, w j is the weight of the j-th attack data.