TTPs intelligent learning analysis method for hidden features and weak correlation features
Through random segmentation and hierarchical clustering of log files, combined with TTPs knowledge graph, the problem of identifying hidden features and weakly related features is solved, and the rapid and accurate detection and traceability of APT attacks are achieved.
Patent Information
- Application Number
- CN202510380810.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively identify and detect TTPs with hidden features and weakly related features, resulting in insufficient detection capabilities for advanced attacks, especially in complex attack scenarios, which is difficult to achieve accurate abnormal detection and traceability.
By extracting important features of the log file, random segmentation and recursive segmentation are performed to generate data points, and the vector representation embedded in hierarchical clustering and attack graphs are used to match them with the TTPs knowledge graph to identify exception points and attack chains.
It realizes rapid identification of hidden features and weakly related features, reduces the cost of data processing time, can discover potential correlation behaviors in the attack chain, and accurately identify APT attacks.
Smart Images

Figure CN120263455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to an intelligent learning and analysis method for TTPs for hidden features and weakly related features. Background Art
[0002] The main drawback of log analysis is that it often requires processing a vast amount of data, which poses a huge pressure on computing resources and storage space. With the increase in network scale and system complexity, the amount of log data has grown rapidly, making the analysis process burdensome and time-consuming. Especially when the log quality is uneven or there is noisy data, misjudgments are more likely to occur. In addition, log analysis usually relies on manually configured rules and templates, and its ability to detect unknown threats is relatively weak. This defect stems from the dependence of log analysis on historical data and established patterns, making it difficult to cope with emerging attack patterns.
[0003] The main drawback of TTPs recognition is its high dependence on threat intelligence. However, solely relying on threat intelligence is difficult to effectively identify those TTPs with hidden features or weakly related features. Many advanced attacks will hide key behavioral features through means such as obfuscation and delay, or adopt more indirect attack paths, making their features less obvious and difficult to fully match known patterns. In addition, attackers often use scattered and low-correlation operations to avoid detection, resulting in the difficulty of traditional threat-intelligence-based TTPs recognition in discovering these complex attack techniques, limiting its effectiveness in dealing with highly concealed and complex attacks.
[0004] Existing methods usually have difficulty in fully utilizing network structure information and behavior sequence information in anomaly detection and TTPs recognition, resulting in a lack of effective learning of system structure features and operation sequence features. In addition, these methods often ignore the causal relationship between abnormal behaviors and are difficult to achieve comprehensive detection and accurate traceability of APT attack chains under weakly related features. On the other hand, when dealing with TTPs deliberately hidden or dispersed by advanced attackers, the recognition accuracy of these methods for hidden features is significantly reduced, affecting the comprehensive detection of complex attack techniques and thus increasing the recognition difficulty.
[0005] Therefore, there is an urgent need to provide a solution to improve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent learning and analysis method for TTPs for hidden features and weakly related features, to improve the problem of insufficient accuracy in the anomaly detection and TTPs recognition of existing methods for hidden features and weakly related features.
[0007] An intelligent learning and analysis method for TTPs for hidden features and weakly related features provided by the present invention includes the following steps:
[0008] Extract and preprocess the log files in the network. The log files consist of process operations, File files, and sockets. Among them, the important features of the File files include file path, operating user, time feature, file type, and operation type;
[0009] Randomly select one feature from the important features as the feature to be cut, and randomly select a split point on the feature to be cut for recursive splitting to generate multiple data points;
[0010] Obtain the expected path length based on the data points, calculate the anomaly score of the data points based on the average path length and the expected path length, and determine whether it is an anomaly point based on the anomaly score and the first threshold;
[0011] Perform hierarchical clustering on the anomaly points to obtain multiple anomaly point clusters, generate an attack graph based on the clusters, randomly sample the neighbor nodes of the target nodes in the attack graph, and integrate the features of the neighbor nodes based on the mean aggregation function to obtain the vector representation of the attack graph embedding;
[0012] Pre-train the TTPs knowledge graph to generate the vector representation of the TTPs embedding, obtain the similarity score based on the cosine similarity, and compare the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold to obtain the comparison result.
[0013] Optionally, in the process of randomly selecting a split point on the feature to be cut for recursive splitting to generate multiple data points, it includes:
[0014] Randomly select a split point on the feature to be cut, divide the data into two subsets, select the repeated features of each subset and perform splitting based on recursion until the stop condition is met to generate multiple data points; the stop condition includes that the number of data points is 1 or the preset maximum tree depth is reached.
[0015] Optionally, in the process of obtaining the expected path length based on the data points, the following formula is used for calculation:
[0016]
[0017] Among them, c(n) is the expected path length, n represents the number of dataset samples, H(i) is the i-th harmonic number, and the mathematical expression is:
[0018] H(i) = ln(i) + γ;
[0019] Among them, γ is the Euler constant.
[0020] Optionally, in the process of obtaining the anomaly score of the data points based on the average path length and the expected path length, the following formula is used for calculation:
[0021]
[0022] Among them, s(x,n) represents the anomaly score of each data point, c(n) represents the expected path length, and h(x) represents the average path length.
[0023] Optionally, in the process of determining whether it is an anomaly point based on the anomaly score and the first threshold, it includes: when the anomaly score is greater than the first threshold, it is determined as an anomaly point, and the first threshold is set to 0.7 or 0.8.
[0024] Optionally, in the process of hierarchically clustering the anomaly points to obtain multiple anomaly point clusters, it includes:
[0025] Calculate the distance between anomaly points based on the distance metric method, regard each anomaly point as a separate cluster, and use the bottom-up aggregation method of hierarchical clustering to recursively merge the clusters with the closest distance until the set level is reached to obtain multiple anomaly point clusters.
[0026] Optionally, in the process of generating an attack graph based on the clusters, it includes:
[0027] Take each anomaly point in the cluster as a node of the attack graph, construct connection edges based on the chronological order, user association or file dependency between the nodes, and represent the directionality of the edges based on the process of the attack chain;
[0028] Assign the mean attribute of the clustering feature to each node, set weights for each edge based on the operation frequency, association strength or time interval to reflect the association strength between anomaly points, identify the critical path in the attack chain based on the association strength, and generate one or more attack paths.
[0029] Optionally, the process of randomly sampling the neighbor nodes of the target node in the attack graph includes:
[0030] Obtain the set of neighbor nodes of the target node, randomly set the sampling quantity of the neighbor nodes. When the sampling quantity is greater than or equal to the number of nodes in the set, select the entire set of neighbor nodes. Otherwise, randomly sample within the set of neighbor nodes.
[0031] Optionally, the process of integrating the features of the neighbor nodes based on the mean aggregation function includes:
[0032] Obtain the aggregated neighborhood feature vector based on the feature vector of the target node and the sampled set of neighbor nodes. Among them, the mathematical expression for calculating the neighborhood feature vector is:
[0033]
[0034] Among them, hN′(v) is the neighborhood feature vector, h u is the feature vector of node u, N′(v) is the set of neighbor nodes, and ∣N′(v)∣ is the number of sampled neighbor nodes; h u is the feature vector of neighbor node u;
[0035] Combines the self - features of the target node v with the neighborhood features to generate an updated feature vector;
[0036]
[0037] Among them, is the updated feature vector, σ is the activation function, W is the trainable weight matrix, CONCAT represents the vector concatenation operation, is the node feature vector of the current layer.
[0038] Optionally, in the process of comparing the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold, it includes:
[0039] If the similarity score exceeds the second threshold, it is determined that the current attack graph conforms to the TTPs pattern, marked as a potential APT attack behavior, and the name and description of the matching TTPs pattern are output, where the second threshold is set between 0.7 and 0.9.
[0040] A TTPs intelligent learning and analysis method for hidden features and weakly - related features provided by the present invention has the beneficial effects that:
[0041] 1. By randomly splitting features, the present invention can quickly identify outliers with sparse distribution in high - dimensional data, directly output high - risk potential anomalies, greatly reducing the time cost of data processing, avoiding the inefficient mode of analyzing each log one by one, and making the anomaly detection link faster and more accurate;
[0042] 2. By hierarchical clustering to reveal the weakly - related features in the outliers and identify the potential associated behaviors in the attack chain, the system can discover the hidden behavior associations in APT attacks, thus constructing a more complete attack chain;
[0043] 3. By matching the generated embedding representation with the TTPs knowledge graph, the present invention realizes the accurate identification of APT hidden features. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the flowchart of the TTPs intelligent learning and analysis method provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention pertains. The words such as "including" used herein mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects.
[0046] The embodiments of the present invention provide a TTPs intelligent learning and analysis method for hidden features and weakly correlated features. Refer to Figure 1 , including:
[0047] S1. Extract and preprocess the log files in the network. The log files are composed of process operations, File files, and sockets. Among them, the important features of the File files include file path, operating user, time feature, file type, and operation type;
[0048] S2. Randomly select one feature from the important features as the feature to be cut, and randomly select a split point on the feature to be cut for recursive splitting to generate multiple data points;
[0049] S3. Obtain the expected path length based on the data points, calculate the anomaly score of the data points based on the average path length and the expected path length, and determine whether it is an anomaly point based on the anomaly score and the first threshold;
[0050] S4. Perform hierarchical clustering on the anomaly points to obtain multiple anomaly point clusters. Generate an attack graph based on the clusters, randomly sample the neighbor nodes of the target nodes in the attack graph, and integrate the features of the neighbor nodes based on the mean aggregation function to obtain the vector representation of the attack graph embedding;
[0051] S5. Pre-train the TTPs knowledge graph to generate the vector representation of the TTPs embedding. Obtain the similarity score based on the cosine similarity, and compare the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold to obtain the comparison result.
[0052] In some embodiments, during the execution of step S1, it includes:
[0053] S1.1. Extract the log files;
[0054] S1.2. Preprocess the log files.
[0055] Specifically, in the process of executing step S1.1 to extract the log file, it includes: the system extracts a large amount of raw data from the network log, and these data can be divided into three types: Process process operations, File files, and Socket sockets.
[0056] Furthermore, in step S1.2, feature extraction, missing value filling, and normalization processing are performed on the extracted files. The important features extracted include file path, operating user, time feature, file type, and operation type.
[0057] In some embodiments, in the process of executing step S2, it includes:
[0058] S2.1. Randomly select one feature from the important features as the feature to be cut;
[0059] S2.2. Randomly select a split point on the feature to be cut and perform recursive splitting to generate multiple data points.
[0060] Specifically, in the process of executing step S2.1, an improved cutting method is obtained based on the traditional isolation forest algorithm, including: evaluating the importance of each feature and assigning priorities to each cut accordingly; for example, when analyzing the log, the file path, operating user, time feature, file type, and operation type have different degrees of contribution to anomaly detection. By calculating the information gain or similarity measure of these features, the most discriminative features can be more accurately identified for splitting.
[0061] Furthermore, comprehensive analysis is combined with multiple features; for example, when judging whether an operation is abnormal, multiple dimensions of data such as the identity of the operating user, the time of the operation, and the file type involved are considered simultaneously, which can effectively capture abnormal behaviors under complex patterns, rather than relying solely on a single feature.
[0062] Furthermore, a self-learning mechanism of the model is designed so that the model can automatically adjust the feature selection strategy according to newly emerging log data. Through the accumulation of data, the system will better understand the interaction between different features and their impact on anomaly detection, thereby continuously optimizing the feature cutting process.
[0063] Furthermore, various parameters of the isolation forest are customized and configured for specific application scenarios, such as the number of trees, sample size, etc., to achieve the best performance. This flexibility can ensure the efficient operation of the algorithm in different environments.
[0064] Specifically, during the execution of step S2.2, it includes: randomly selecting a splitting point on the feature to be cut, dividing the data into two subsets, selecting the repeated features of each subset and performing splitting based on recursion until the stopping condition is met, generating multiple data points; the stopping condition includes that the number of data points is 1 or the preset maximum tree depth is reached.
[0065] In fact, to improve the splitting efficiency and accuracy, before each splitting, the selection order of features will be determined according to the importance of features and information gain to ensure that the most effective features participate in the splitting first; during the splitting process, the feature data in multiple dimensions will be comprehensively considered to comprehensively capture potential abnormal patterns, ensuring that the splitting decision is more accurate and can effectively identify complex abnormal behaviors.
[0066] In some embodiments, during the execution of step S3, it includes:
[0067] S3.1. Obtain the expected path length based on the data points;
[0068] S3.2. Calculate the anomaly score of the data points based on the average path length and the expected path length;
[0069] S3.3. Determine whether it is an anomaly point based on the anomaly score and the first threshold.
[0070] Specifically, during the execution of step S3.1, the following formula is used for calculation:
[0071]
[0072] Among them, c(n) is the expected path length, n represents the number of dataset samples, H(i) is the i-th harmonic number, and the mathematical expression is:
[0073] H(i) = ln(i) + γ;
[0074] Among them, γ is the Euler constant.
[0075] Furthermore, when executing step S3.2 and obtaining the anomaly score of the data points based on the average path length and the expected path length, the following formula is used for calculation:
[0076]
[0077] Among them, s(x,n) represents the anomaly score of each data point, c(n) represents the expected path length, and h(x) represents the average path length.
[0078] Further, in the process of performing step S3.3 to determine whether it is an abnormal point based on the abnormal score and the first threshold, it includes: when the abnormal score is greater than the first threshold, it is determined as an abnormal point, and the first threshold is set to 0.7 or 0.8.
[0079] Actually, when the abnormal score is close to 1, it means that the data point is more likely to be an abnormal point, while when the score is close to 0, it means that the data point is more likely to be a normal point; that is, when the abnormal score s(x)≈1, the data point x is very likely to be an abnormal point; and when the abnormal score s(x)≈0, it means that x belongs to a normal point. Usually, a first threshold is set. When the abnormal score exceeds the first threshold, the data point is determined as an abnormal point. The recommended value of the first threshold is 0.7 or 0.8, which can be further adjusted according to the application scenario.
[0080] In some embodiments, in the process of performing step S4, it includes:
[0081] S4.1. Perform hierarchical clustering on the abnormal points to obtain multiple abnormal point clustering clusters;
[0082] S4.2. Generate an attack graph based on the clustering clusters;
[0083] S4.3. Randomly sample the neighbor nodes of the target node in the attack graph;
[0084] S4.4. Integrate the features of the neighbor nodes based on the mean aggregation function to obtain the vector representation of the attack graph embedding.
[0085] Specifically, in the process of performing step S4.1 to perform hierarchical clustering on the abnormal points to obtain multiple abnormal point clustering clusters, it includes: calculating the distance between abnormal points based on the distance metric method, regarding each abnormal point as a separate cluster, and using the bottom-up aggregation method of hierarchical clustering to recursively merge the clusters with the closest distance until the set level is reached to obtain multiple abnormal point clustering clusters.
[0086] Actually, in order to more accurately reflect the similarity between abnormal points and enhance the clustering effect, in the process of calculating the distance between abnormal points based on the distance metric method, the distance metric method of the present invention uses the Euclidean distance to calculate the similarity of abnormal points. Let the feature vectors of two abnormal points be x=(x1,x1,…,x n ) and y=(y1,y1,…,y n ), where n represents the number of feature dimensions.
[0087] Further, aggregate the abnormal point features from five feature dimensions, that is, based on the five dimensions of file path, operating user, time feature, file type, and operation type. Therefore, the number of feature dimensions n is selected as 5, and the distance metric formula is:
[0088]
[0089] Among them, d(x, y) is the distance between the abnormal point x and the abnormal point y, S is the covariance matrix of the feature vectors, S -1 is the inverse matrix of the covariance matrix, and (x - y) is the difference between the two feature vectors.
[0090] Specifically, in the process of performing step S4.2 to generate an attack graph based on the clustering clusters, it includes: taking each abnormal point in the clustering cluster as a node of the attack graph, constructing connection edges based on the chronological order, user association or file dependency between the nodes, and representing the directionality of the edges based on the process of the attack chain; assigning the mean attribute of the clustering features to each node, and setting weights for each edge based on the operation frequency, association strength or time interval to reflect the association strength between the abnormal points, identifying the critical path in the attack chain based on the association strength, and generating one or more attack paths.
[0091] Further, perform step S4.3 to randomly sample the neighbor nodes of the target node in the attack graph, including: obtaining the set of neighbor nodes of the target node, randomly setting the sampling quantity of the neighbor nodes, and when the sampling quantity is greater than or equal to the number of nodes in the set, selecting the entire set of neighbor nodes, otherwise randomly sampling within the set of neighbor nodes.
[0092] Specifically, in the process of performing step S4.4 to integrate the features of the neighbor nodes based on the mean aggregation function, it includes:
[0093] Obtaining the aggregated neighborhood feature vector based on the feature vector of the target node and the sampled set of neighbor nodes, where the mathematical expression for calculating the neighborhood feature vector is:
[0094]
[0095] where h N′(v) is the neighborhood feature vector, h u is the feature vector of node u, N′(v) is the set of neighbor nodes, ∣N′(v)∣ is the number of sampled neighbor nodes; h u is the feature vector of the neighbor node u;
[0096] Combining the self - feature of the target node v with the neighborhood feature to generate an updated feature vector;
[0097]
[0098] where, is the updated feature vector, σ is the activation function, W is the trainable weight matrix, CONCAT represents the vector concatenation operation, is the node feature vector of the current layer.
[0099] In some embodiments, during the execution of step S5, it includes:
[0100] S5.1. Pre-train the TTPs knowledge graph to generate a vector representation of the TTPs embedding;
[0101] S5.2. Obtain the similarity score based on the cosine similarity;
[0102] S5.3. Compare the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold to obtain a comparison result.
[0103] Specifically, during the execution of step S5.1, it includes: Using algorithms such as GraphSAGE to pre-train the TTPs graph, generating the embedding representation of each node and the overall TTPs pattern embedding, and constructing a comparable TTPs pattern library.
[0104] Furthermore, when executing step S5.2, the following formula is used for calculation:
[0105] S(G1,G2) = α·S struct (G1,G2)+(1 - α)S feat (G1,G2);
[0106] Wherein, S(G1,G2) is the similarity between the attack graph embedding and the TTPs pattern embedding, S struct (G1,G2) is the similarity metric based on the graph structure information, and S feat (G1,G2) is the similarity metric based on the node and edge features, and α is a weight parameter.
[0107] In fact, the present invention not only considers the graph structure information but also combines the features of nodes and edges, thereby providing a more comprehensive and accurate graph similarity metric. Through this comprehensive similarity calculation method that combines graph structure information and node and edge features, the present invention can more comprehensively reflect the similarity between different TTPs patterns, thereby significantly improving the accuracy and practicality of the graph similarity metric.
[0108] Furthermore, during the process of executing step S5.3 to compare the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold, it includes: If the similarity score exceeds the second threshold, it is determined that the current attack graph conforms to the TTPs pattern, marked as a potential APT attack behavior, and the name and description of the matching TTPs pattern are output, wherein the second threshold is set between 0.7 and 0.9.
[0109] Although the embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A TTPs intelligent learning and analysis method for hidden features and weakly correlated features, characterized in that Including the following steps: Extract and preprocess the log files in the network. The log files are composed of process operations, File files, and sockets. Among them, the important features of the File files include file path, operating user, time feature, file type, and operation type; Randomly select one feature from the important features as the feature to be cut, and randomly select a split point on the feature to be cut for recursive splitting to generate multiple data points; Obtain the expected path length based on the data points, calculate the anomaly score of the data points based on the average path length and the expected path length, and determine whether it is an anomaly point based on the anomaly score and the first threshold; Perform hierarchical clustering on the anomaly points to obtain multiple anomaly point clusters. Generate an attack graph based on the clusters. Randomly sample the neighbor nodes of the target nodes in the attack graph, and integrate the features of the neighbor nodes based on the mean aggregation function to obtain the vector representation of the attack graph embedding; Pre-train the TTPs knowledge graph to generate the vector representation of the TTPs embedding. Obtain the similarity score based on the cosine similarity. Compare the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold to obtain the comparison result.
2. The TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, wherein During the process of randomly selecting a split point on the feature to be cut for recursive splitting to generate multiple data points, it includes: Randomly select a split point on the feature to be cut, divide the data into two subsets, select the repeated features of each subset and perform splitting based on recursion until the stopping condition is met to generate multiple data points; The stopping condition includes that the number of data points is 1 or the preset maximum tree depth is reached.
3. A TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, characterized in that During the process of obtaining the expected path length based on the data points, the following formula is used for calculation: where c(n) is the expected path length, n represents the number of dataset samples, H(o) is the o-th harmonic number, and the mathematical expression is: H(i) = ln(i) + γ; where γ is the Euler constant.
4. The TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 3, wherein During the process of obtaining the anomaly score of the data points based on the average path length and the expected path length, the following formula is used for calculation: where s(x,n) represents the anomaly score of each data point, c(n) represents the expected path length, and h(x) represents the average path length.
5. A TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, characterized in that, During the process of determining whether it is an anomaly point based on the anomaly score and the first threshold, it includes: when the anomaly score is greater than the first threshold, it is determined as an anomaly point, and the first threshold is set to 0.7 or 0.
8.
6. The TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, characterized in that During the process of performing hierarchical clustering on the anomaly points to obtain multiple anomaly point clusters, it includes: Calculate the distance between the anomaly points based on the distance metric method, regard each anomaly point as a separate cluster, and recursively merge the clusters with the closest distance using the bottom-up aggregation method of hierarchical clustering until the set level is reached to obtain multiple anomaly point clusters.
7. A TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, characterized in that During the process of generating an attack graph based on the clusters, it includes: Take each anomaly point in the cluster as a node of the attack graph, construct connection edges based on the time sequence, user association, or file dependency between the nodes, and represent the directionality of the edges based on the process of the attack chain; Assign the mean attribute of the clustering feature to each node, and set weights for each edge based on the operation frequency, association strength, or time interval to reflect the association strength between outliers. Identify the critical path in the attack chain based on the association strength, and generate one or more attack paths.
8. A TTPs intelligent learning and analysis method for hidden features and weakly correlated features according to claim 1, characterized in that, The process of randomly sampling the neighbor nodes of the target node in the attack graph includes: Obtain the set of neighbor nodes of the target node, randomly set the sampling number of neighbor nodes. When the sampling number is greater than or equal to the number of nodes in the set, select the entire set of neighbor nodes; otherwise, randomly sample within the set of neighbor nodes.
9. A TTPs intelligent learning and analysis method for concealed features and weakly correlated features according to claim 1, characterized in that The process of integrating the features of the neighbor nodes based on the mean aggregation function includes: Obtain the aggregated neighborhood feature vector based on the feature vector of the target node and the sampled set of neighbor nodes. The mathematical expression for calculating the neighborhood feature vector is: where h N′(v) is the neighborhood feature vector, h u is the feature vector of node u, N′(v) is the set of neighbor nodes, and ∣N′(v)∣ is the number of sampled neighbor nodes; h u is the feature vector of neighbor node u; Combine the self-feature of the target node v with the neighborhood feature to generate an updated feature vector. Among them, is the updated feature vector, σ is the activation function, W is the trainable weight matrix, and CONCAT represents the vector concatenation operation. is the node feature vector of the current layer.
10. A TTPs intelligent learning and analysis method for concealed features and weakly correlated features according to claim 1, characterized in that In the process of comparing the vector representation of the attack graph embedding with the vector representation of the TTPs embedding based on the similarity score and the second threshold, it includes: If the similarity score exceeds the second threshold, it is determined that the current attack graph conforms to the TTPs pattern, marked as a potential APT attack behavior, and output the name and description of the matching TTPs pattern. The second threshold is set between 0.7 and 0.9.
Citation Information
Cited By
APT attack path reconstruction method based on time sequence diagram comparison clustering and medium
CN120474829A