A multi-dimensional TTP feature learning and completion method
Through the multi-dimensional TTP feature learning completion method, the heterogeneous graph convolution network and attention mechanism model are used to solve the problem of multi-dimensional features missing in TTP analysis, and more comprehensive attack behavior modeling and defense strategy optimization are achieved.
Patent Information
- Application Number
- CN202510592719.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing technology lacks the comprehensive analysis and integration capabilities of multi-dimensional features in TTP analysis, which leads to the one-sidedness of the analysis results and the limitations of defense strategies. The problem of missing features in TTP data is common, especially when the lack of labeled data in the face of zero-day attacks makes it difficult to learn features.
The multi-dimensional TTP feature learning completion method is adopted, and heterogeneous graphs are defined by obtaining log data, and multiple dimensions are divided for feature extraction. The heterogeneous graph convolution network HGCN and attention mechanism models are used to complete the missing edges, and a comprehensive TTP attack behavior model is constructed, and the attack chain is described through natural language generation technology.
It improves the analysis and modeling capabilities of multi-dimensional features, improves the accuracy of TTP analysis and the effectiveness of defense systems, supports behavioral analysis of complex attack chains, and enhances the sharing of threat intelligence and the optimization of defense strategies.
Smart Images

Figure CN120123661B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cyberspace security, and in particular to a multi-dimensional TTP feature learning and completion method. Background Art
[0002] TTP (Tactics, Techniques, Procedures) is a framework for describing the behavior patterns of attackers to achieve their goals, including tactics (goals or intentions), techniques (specific methods), and procedures (implementation processes). This framework helps security researchers, enterprises, and defense teams understand the behavior of attackers, thus supporting more accurate protection and response.
[0003] With the increasing diversification and complexity of network attack means, the TTP of attackers is also constantly evolving, and traditional defense methods and strategies are gradually showing their limitations. Therefore, a comprehensive and accurate analysis of attackers' TTP and an in-depth understanding of their behavior patterns have become an important research direction in the field of network security.
[0004] Although the TTP framework provides a powerful analysis tool for network security defense, in the actual application process, the collection and analysis of TTP data face a series of technical problems. These problems have led to an increasingly prominent phenomenon of TTP data missing, directly affecting the understanding of attackers' behavior patterns and possibly resulting in the failure and misjudgment of defense strategies. TTP feature learning is the key link to solve this problem. Feature learning, as a bridge connecting the original data and machine learning models, directly affects the model performance and learning efficiency. Generally speaking, feature learning can be divided into two categories: supervised feature learning and unsupervised feature learning.
[0005] Feature learning automatically extracts and learns meaningful features from the original data through algorithms for tasks such as pattern recognition, classification, and inference. Compared with traditional feature engineering methods, feature learning does not require manual design and selection of features, but autonomously discovers key information from the data, especially showing significant advantages when dealing with complex, large-scale, and high-dimensional data (such as network traffic, log files, or text data).
[0006] Due to the concealment of attack organizations, the technical limitations of monitoring systems, and defects in data processing, TTP data often lacks and is insufficient. Feature learning can mine potential patterns from existing data, infer and complete missing attack behaviors, and provide support for comprehensively depicting attackers' behaviors.
[0007] However, the disadvantages of the existing technologies mainly lie in:
[0008] Existing technologies in TTP analysis are often limited to the study of a single dimension (such as technology or tactics), lacking the ability to comprehensively analyze and integrate multi-dimensional features. Such a one-sided perspective is difficult to fully reveal the behavior patterns and overall portraits of attacking organizations, resulting in the one-sidedness of analysis results and the limitations of defense strategies. In addition, the lack of features in TTP data is a widespread problem. Especially in the face of new types of attacks such as zero-day attacks, the lack of labeled data makes feature learning more difficult. Although existing feature completion methods can handle some missing data, the inference relying on existing data often fails to effectively capture the complex non-linear relationships in attack behaviors. The current expression methods of TTP features are relatively scattered, lacking a unified description standard and a systematic characterization method between different platforms or systems. This inconsistency not only increases the complexity of analysis but also weakens the practical value of threat intelligence in collaboration and sharing, further restricting the overall protection ability of the network security system.
[0009] Therefore, there is an urgent need to develop a solution to address the above problems. Summary of the Invention
[0010] The purpose of the present invention is to provide a multi-dimensional TTP feature learning and completion method, which improves the problems that the current analysis methods of attacking organizations lack the comprehensive modeling ability of multi-dimensional features, there are missing features in TTP data, and the dynamic update and expression of attack behavior features are not intuitive enough.
[0011] A multi-dimensional TTP feature learning and completion method provided by the present invention adopts the following technical solutions:
[0012] A multi-dimensional TTP feature learning and completion method specifically includes the following steps:
[0013] Obtain log data and define a heterogeneous graph, divide it into several dimensions according to attack behaviors, and extract and define features for each dimension.
[0014] Initialize node embeddings, use the heterogeneous graph convolutional network HGCN as the convolutional layer, update the embeddings of nodes according to the information of their neighbor nodes, and use the attention mechanism model to calculate the dynamic relative weights for each pair of neighbor nodes to complete the missing edges.
[0015] A multi-dimensional TTP feature learning and completion method, which obtains log data and defines a heterogeneous graph, divides several dimensions according to different attack behaviors, extracts the features of each dimension, and then constructs a more comprehensive TTP attack behavior model. Using the heterogeneous graph convolutional network HGCN as the convolutional layer, it updates the information of the heterogeneous graph node embedding, introduces an attention mechanism model to calculate the dynamic relative weight for each pair of neighbor nodes, completes the behavior of complementing the attack chain, thereby improving the analysis and modeling ability of multi-dimensional features, effectively supporting the behavior analysis requirements in complex attack chains, inferring the missing attack behavior features from the existing data, complementing the attack chain, and further improving the accuracy of TPP analysis.
[0016] Optionally, the multi-dimensional TTP feature learning and completion method further includes extracting the target attack node and all its associated nodes, and extracting the edges connecting the nodes to construct an attack subgraph according to the nodes and edges; converting the attack subgraph into a natural language text by means of natural language generation.
[0017] In the present invention, the attack subgraph of the heterogeneous graph converts the abstract graph structure information of the attack subgraph into a natural language text by means of natural language generation, which can improve the interpretability of security events, help non-technical personnel evaluate risks through natural language texts without understanding graph structure data, promote decision-making efficiency, and enhance the emergency response coordination ability.
[0018] Optionally, defining a heterogeneous graph includes defining a heterogeneous graph, which includes several nodes and edges for connecting the nodes. G represents the heterogeneous graph, V represents the nodes, including different types of nodes; E represents the edges, which are used to express the relationships between the nodes V, and each relationship corresponds to a different edge type.
[0019] In the present invention, the heterogeneous graph G=(V,E) is defined, where the nodes (V) represent the entities in the graph, including different types of nodes; the edges (E) represent the relationships between the nodes, and each relationship has a different type. The node types in the heterogeneous graph include: Attacker (information such as attacker identifier, ip address, etc.), Technique (identifier of each attack technique and its description information, such as phishing email, remote code execution, etc.), Tool (tools used by the attacker, such as Metasploit, Cobalt Strike, etc.), Target System (the target system of the attacker, such as Windows server, Linux host, etc.), Attack Phases (each step or stage passed by the attacker when launching an attack, such as data stealing, lateral movement, etc.). The edge types in the heterogeneous graph include: attack organization - technique, technique - tool, technique - target system, technique - attack phase, etc.
[0020] Optionally, the attack behavior is divided into several dimensions, and features are extracted and defined for each dimension, including:
[0021] According to the attack behavior, it is divided into tactical dimension, technical dimension, time series dimension and correlation dimension; each dimension is independent of each other, and the tactical features of the tactical dimension, the technical features of the technical dimension, the time series features of the time series dimension and the correlation features of the correlation dimension are extracted separately; the extracted features are defined.
[0022] In the present invention, different aspects of the attack behavior are divided into multiple different dimensions, each dimension is independent of each other, and the features in each dimension are extracted separately. Among them, it is divided into tactical dimensions according to the strategies and methods used by the attacker, and divided into technical dimensions according to the technical level behaviors used by the attacker; it is divided into time series dimensions according to the order and time difference of events; it is divided into correlation dimension according to the correlation between events and logs from multiple sources; each dimension is independent of each other, and the tactical features of the tactical dimension, the technical features of the technical dimension, the time series features of the time series dimension, and the correlation features of the correlation dimension are extracted separately; the extracted features are defined. By dividing into different dimensions, it can help defenders better understand the attacker's attack motivation, strengthen data encryption and access control in a targeted manner, support rapid attribution and threat intelligence sharing, and provide operational guidance for defense.
[0023] Optionally, the extracted features are defined, including defining the extracted tactical features and technical features according to the types of definable nodes; defining the extracted time series features by adding a timestamp feature to each attack behavior node to record the specific time when the behavior occurred; mapping the dependencies between the attacker's behaviors and the relationships between the attacker and the tools and target systems into edges, and defining the extracted association relationship features.
[0024] In the present invention, when extracting tactical features and technical features, the extracted features can be defined according to the defined node types, such as attackers, attack stages, tools, and target systems, and the specific operations can be: "phishing attack" → define the node as an attack tactic (phishing); "Metasploit exploits vulnerabilities for remote execution" → can be directly connected as a tool node (Metasploit) and an attack technology node (remote execution). For the time series dimension, a timestamp feature can be added to each attack behavior node to record the specific time when each attack technology, tool, target system, etc. occurs. The timestamp helps to arrange the attack events in chronological order, thereby providing a basis for subsequent time series analysis and attack prediction.
[0025] For the association relationship, the dependencies between behaviors and the relationships between attackers, tools, and target systems are mapped to edges, that is, the dependencies and interactions of the attack process. For example, the association between the attack phase and the attack technology: the attack phase node (such as "initial access") can be connected to the attack technology node (such as "phishing email") through an edge, indicating that in this phase, the attacker used the "phishing email" technology, and the edge type is obtained: attack phase → technology.
[0026] By extracting and defining features in four dimensions, the effectiveness of the defense system can be significantly improved, detection efficiency can be increased, and defense depth can be enhanced. After the feature library is dynamically updated, the defense system can also automatically adapt to new attack methods. The dimensional division and feature extraction of TTPs for network security defense provides a systematic, quantifiable, and collaborative analysis method. By deconstructing attacker behavior and extracting key features, defenders can achieve more accurate threat identification and more efficient resource allocation, thereby taking the initiative in complex attack and defense confrontations.
[0027] Optionally, node embeddings are initialized, which consists of representing each type of node and edge as a vector, for each node V, whose initial embedding vector is , the embedding vector is initialized by pre-trained word embedding or randomly, and each edge E connects two nodes and .
[0028] Optionally, a heterogeneous graph convolutional network HGCN is used as a convolutional layer, and in the process of updating the embedding of a node according to the information of its neighbor nodes, the updating method includes: neighbor aggregation and cross-type information transmission.
[0029] The aggregation operation of the neighbor aggregation is weighted according to the type of node and the type of connected edge, and the formula is:
[0030] ,
[0031] in, , r is the edge type, R is the edge type set, Is with the node The set of all connected neighbor nodes of type r, is the weight matrix associated with edge type r, the learned feature transformation matrix, is the activation function, Is a node The number of neighbor nodes of type r, Is with the node The set of all connected neighbor nodes of type r, Is a node The number of neighbor nodes of type r, is the node at the layer's embedding vector, is the updated node embedding vector obtained by the node after the aggregation operation of the convolutional layer.
[0032] The cross-type information transfer combines the node type and edge type for cross-type information aggregation. The formula is:
[0033] ,
[0034] The update of each type of node depends on neighbor nodes of different types and edge types. Each type of node uses its own weight matrix for feature transformation. Among them, is the bias term.
[0035] In the present invention, in the TTP heterogeneous graph, the node types and edge types are complex. The heterogeneous graph convolutional network HGCN is used to assign independent parameters or attention weights to different nodes and edge types, which can accurately capture the diversity of attack behaviors. For example, it can distinguish the feature propagation methods of two types of edges, "vulnerability exploitation" and "phishing attack". And when updating the node embedding, it automatically aggregates the heterogeneous information of neighbors (such as the tool types used by attackers and the industries to which the attack targets belong) to generate context-related feature representations.
[0036] In addition, the heterogeneous graph convolutional network HGCN does not require manual definition of the original path. It automatically learns the key path in the attack chain through a hierarchical aggregation mechanism, avoids generating a dense intermediate graph, and directly updates the node embedding through sparse matrix operations, significantly reducing the computational complexity. It is suitable for large-scale TTP heterogeneous graph data and supports incremental learning. When new attack samples or threat intelligence are added, the node embedding can be quickly adjusted to adapt to the evolution of attack techniques.
[0037] Optionally, an attention mechanism model is used to calculate the dynamic relative weights for each pair of neighbor nodes, including after passing through the convolutional layer, introducing an attention mechanism model to calculate the dynamic relative weights for each pair of neighbor nodes, determining the influence degree of each pair of neighbor nodes on the target node representation. The relationship between nodes of edge type r is calculated through the attention weight formula:
[0038] α ij (r) = exp(LeakyReLU( a T [ W r h i || W r h j ])) ∑ kϵ N r (i) exp(LeakyReLU( a T [ W r h i || W r h k ]) ,
[0039] Among them, , k is the neighbor node connected to the node , a is the learned attention weight vector, is the weight matrix of edge type r, represents the concatenation operation, is the node 's embedding vector, is the node 's embedding vector, is the embedding vector of node k.
[0040] Optionally, complete the missing edges, including predicting the missing edges using node embeddings. Given the embedding vectors and of two nodes and , calculate the similarity between them, use cosine similarity to measure the relationship strength between two nodes, and predict the missing edges. The formula is:
[0041] ,
[0042] The predicted missing edges and their similarity scores indicate which node pairs may have potential relationships;
[0043] Define a binary classification loss to represent whether an edge exists: the label {0, 1}, if the edge exists then = 1, otherwise = 0. The formula for the binary cross-entropy loss function is:
[0044] ,
[0045] where, is the set of edges actually existing in the graph, is the set of negative sampled edges; through the inference of the heterogeneous graph neural network, complete the missing nodes and missing edges in the attack chain.
[0046] Optionally, in the way of natural language generation, convert the attack subgraph into natural language text, including mapping the nodes and relationships in the attack subgraph to natural language text according to the defined template by the way of natural language generation, generating a dynamic description of the attack chain, and systematically characterizing the attack behavior.
[0047] The beneficial effects of the present invention are as follows:
[0048] A multi-dimensional TTP feature learning and completion method divides several dimensions according to different attack behaviors, extracts features for each dimension, and then constructs a more comprehensive TTP attack behavior model. The heterogeneous graph convolutional network HGCN is used as the convolutional layer to update the information of heterogeneous graph node embeddings. An attention mechanism model is introduced to calculate the dynamic relative weights for each pair of neighbor nodes to complete the behavior of complementing the attack chain. Moreover, after completion, an attack subgraph of the attack chain can be constructed, and natural language generation technology is combined to generate a complete report on complex attack behaviors, which not only provides an easy-to-understand description of the attack chain for security analysts, but also enables a comprehensive characterization of complex attack behaviors, supports the prediction of forward-looking threat intelligence and the optimization of defense decisions, and improves the problems of the current attack organization analysis method lacking the comprehensive modeling ability of multi-dimensional features, feature missing in TTP data, and the dynamic update and expression of attack behavior features not being intuitive enough. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 FIG. is a schematic flow framework diagram of a multi-dimensional TTP feature learning and completion method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the technical field to which the present invention belongs. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0051] As Figure 1 shown, the embodiments of the present invention provide a multi-dimensional TTP feature learning and completion method, including the following steps:
[0052] S1. Obtain log data and define a heterogeneous graph, divide it into several dimensions according to attack behaviors, and extract and define features for each dimension;
[0053] S2. Initialize node embeddings, use the heterogeneous graph convolutional network HGCN as the convolutional layer, update the embeddings of nodes according to the information of their neighbor nodes, and use the attention mechanism model to calculate the dynamic relative weights for each pair of neighbor nodes to complement the missing edges;
[0054] S3. Extract the target attack nodes and all the nodes associated with them, and extract the edges connecting these nodes. Construct an attack subgraph based on the nodes and edges, and convert the attack subgraph into natural language text by means of natural language generation.
[0055] Further, in some embodiments, when performing step S1, it specifically includes:
[0056] Collect log data from firewalls, intrusion detection systems (IDS), server logs, endpoint detection and response (EDR) tools, etc., remove noise data such as duplicate logs and false alarms, handle missing values, and standardize the log field format. Define a heterogeneous graph G=(V, E), where the heterogeneous graph includes several nodes and edges for connecting the nodes, G represents the heterogeneous graph, the nodes (V) represent the entities in the graph, including different types of nodes; the edges (E) represent the relationships between the nodes, and each relationship corresponds to a different edge type.
[0057] The node types in the heterogeneous graph include: Attacker (information such as attacker identifier, IP address, etc.), Technique (identifier of each attack technique and its description information, such as phishing email, remote code execution, etc.), Tool (tools used by the attacker, such as Metasploit, Cobalt Strike, etc.), Target System (the target system of the attacker, such as Windows server, Linux host, etc.), Attack Phases (each step or stage passed by the attacker when launching an attack, such as data theft, lateral movement, etc.). The edge types in the heterogeneous graph include: Attack Organization - Technique, Technique - Tool, Technique - Target System, Technique - Attack Phase, etc.
[0058] Further, divide the attack behavior into several dimensions, and perform feature extraction and definition for each dimension, specifically including:
[0059] Divide different aspects of the attack behavior into multiple different dimensions, each dimension is independent of each other, and extract the features in each dimension separately. Among them, according to the strategies and methods used by the attacker, it is divided into the tactical dimension; according to the technical - level behaviors used by the attacker, it is divided into the technical dimension; according to the order of occurrence and time difference of events, it is divided into the time - series dimension; according to the correlation relationship between events and logs from multiple sources, it is divided into the correlation - relationship dimension; each dimension is independent of each other, and separately extract the tactical features of the tactical dimension, the technical features of the technical dimension, the time - series features of the time - series dimension, and the correlation - relationship features of the correlation - relationship dimension; define the extracted features.
[0060] The tactical dimension describes the strategies and methods used by attackers, usually focusing on three aspects: the stages, tools, and targets of the attack. By comparing with the MITRE ATT&CK framework, the tactical characteristics used by attackers are identified. Tactical characteristics can be divided into multiple levels to describe the strategies of different attack stages. It is also possible to analyze the attack patterns and historical cases of attack organizations in threat intelligence and extract their commonly used attack tactics.
[0061] The technical dimension describes the technical-level behaviors such as analyzing the API call sequences and network traffic data of malicious samples, and uses static and dynamic analysis techniques to mine the technical characteristics in suspicious files or traffic.
[0062] The time series dimension helps to identify the time patterns of attack behaviors by analyzing the occurrence order and time differences of events. The timestamps of each attack event and the time differences between events can be used as an important dimension for modeling.
[0063] The correlation relationship dimension identifies time series, behavior dependencies, and front-back associations by aggregating and analyzing events and logs from multiple sources. For example, the time dependencies of events and the front-back relationships of behaviors.
[0064] By dividing different dimensions, it can help defenders better understand the attack motives of attackers, strengthen data encryption and access control in a targeted manner, support rapid attribution and threat intelligence sharing, and provide guidance at the operational level for defense.
[0065] Furthermore, the extracted features are defined, including: defining the extracted tactical and technical features according to the types of definable nodes; defining the extracted time series features by adding timestamp features to each attack behavior node to record the specific time when the behavior occurs; defining the extracted correlation relationship features by mapping the dependencies between attacker behaviors and the relationships between attackers, tools, and target systems as edges.
[0066] Specifically, in some embodiments, when extracting tactical features and technical features, the extracted features can be defined according to the defined node types, such as attackers, attack phases, tools, and target systems. The specific operations can be as follows: "Phishing attack" → Define the node as an attack tactic (phishing); "Metasploit uses vulnerabilities for remote execution" → Can be directly connected as a tool node (Metasploit) and an attack technique node (remote execution). For the time series dimension, a timestamp feature can be added to each attack behavior node to record the specific time when each attack technique, tool, target system, etc. occurs. The timestamp helps to arrange attack events in chronological order, thus providing a basis for subsequent time series analysis and attack prediction. For the association relationship, map the dependencies between behaviors and the relationships between attackers, tools, and target systems as edges, that is, the front and back dependencies and interactions in the attack process. For example, the association between the attack phase and the attack technique: The attack phase node (such as "Initial Access") can be connected to the attack technique node (such as "Phishing Email") through an edge, indicating that in this phase, the attacker uses the technique of "Phishing Email", and the edge type is obtained: Attack Phase → Technique.
[0067] By extracting and defining the features in the four dimensions, the effectiveness of the defense system can also be significantly improved, the detection efficiency can be increased, and the defense depth can be enhanced. After the feature library is dynamically updated, the defense system can also automatically adapt to new attack methods. The dimension division and feature extraction of TTP provide a systematic, quantifiable, and collaborative analysis method for network security defense. By deconstructing the attacker's behavior and extracting key features, the defender can achieve more accurate threat recognition and more efficient resource allocation, thus gaining the initiative in complex attack and defense confrontations.
[0068] Furthermore, when performing step S2, it specifically includes:
[0069] Node embedding initialization, including representing each type of node and edge as a vector. For any node V, its initial embedding vector is , and the embedding vector is initialized through pre-trained word embeddings or randomly. Each edge E connects two nodes and .
[0070] Specifically, in some embodiments, the number of layers of the heterogeneous graph is set to two. The first layer is used to aggregate first-order neighbors, and the second layer is used to capture the semantic information of second-order neighbor structures. Node embedding represents each node as a low-dimensional vector, and the embeddings of the two layers of nodes are used to meet the requirements of downstream tasks. Through the operation of initializing node embeddings, the nodes in the network are mapped to the low-dimensional vector space, which is convenient for machine learning models to process. This mapping can maintain the similarity of nodes in the network, so that the node relationships in the embedding space can approximately reflect the structure and properties of the original network.
[0071] At the same time, during the initialization process, the network structure information of the nodes, such as neighbor relationships and node degrees, is retained. According to the requirements of subsequent specific tasks, the embeddings are initialized to adapt to downstream tasks and improve the performance of the model on these tasks. The low-dimensional embeddings of nodes can reduce the computational complexity, accelerate the model training speed, capture the similarity and structure information between nodes, and improve the performance of the model on unseen data. This generalization ability enables the model to better adapt to different graph structures and features.
[0072] Furthermore, in some embodiments, the graph convolutional layer is the core part of the heterogeneous graph neural network. Since the heterogeneous graph contains multiple types of nodes and multiple types of edges, the heterogeneous graph convolutional network HGCN is used as the convolutional layer. In HGCN, the convolution operation takes into account the node type and edge type. During the process of updating the embedding of a node according to the information of its neighbor nodes, the update methods include: neighbor aggregation and cross-type information transfer;
[0073] The aggregation operation of the neighbor aggregation is weighted according to the type of the node and the type of the connected edge. The formula is:
[0074] ,
[0075] where, , r is the edge type, R is the set of edge types, is the set of all neighbor nodes of type r connected to node , is the weight matrix related to the edge type r, a learned feature transformation matrix, is the activation function, is the number of neighbor nodes of type r of node , is the set of all neighbor nodes of type r connected to node , is the number of neighbor nodes of type r of node , is the embedding vector of node at the th layer, is the embedding vector of node The updated node embedding vectors obtained after the aggregation operation in the convolutional layer.
[0076] In some embodiments, through the method of neighbor aggregation, by aggregating the features of neighbor nodes, the central node can perceive the information of its local neighbors, thereby capturing the dependencies between nodes and the structural features of the graph. Through the aggregation operation, the embedding of the central node can be updated to make it contain more context information and improve the expressive power of the node representation. During the aggregation process, the structural information of the heterogeneous graph is retained, such as the degrees of neighbor nodes and the path lengths between nodes.
[0077] The cross-type information transfer combines node types and edge types for cross-type information aggregation, and the formula is:
[0078] ,
[0079] The update of each type of node depends on different types of neighbor nodes and edge types. Each type of node uses its own weight matrix for feature transformation, where, is the bias term.
[0080] In some embodiments, in a heterogeneous graph, there may be multiple relationships between different node types. Through the method of cross-type aggregation, these complex relationships can be modeled to improve the expressive power of the model. In tasks such as node classification and link prediction, cross-type aggregation can utilize information from different types of nodes to improve the prediction performance of the model. Cross-type aggregation can also flexibly handle different types of nodes and edges, adapt to different task requirements, and the model itself has stronger robustness to noise and missing data.
[0081] In some embodiments, in the TTP heterogeneous graph, the node types and edge types are complex. By using the heterogeneous graph convolutional network HGCN to assign independent parameters or attention weights to different node and edge types, the diversity of attack behaviors can be accurately captured. For example, the feature propagation methods of two types of edges, "vulnerability exploitation" and "phishing attack", can be distinguished. And when updating the node embedding, the heterogeneous information of neighbors (such as the tool types used by attackers and the industries to which the attack targets belong) is automatically aggregated to generate context-related feature representations.
[0082] In addition, the heterogeneous graph convolutional network HGCN does not require manual definition of the original path. It automatically learns the key paths in the attack chain through a hierarchical aggregation mechanism, avoids generating a dense intermediate graph, and directly updates the node embedding through sparse matrix operations, significantly reducing the computational complexity. It is suitable for large-scale TTP heterogeneous graph data and supports incremental learning. When new attack samples or threat intelligence are added, the node embedding can be quickly adjusted to adapt to the evolution of attack techniques.
[0083] Furthermore, in some embodiments, an attention mechanism model is used to calculate dynamic relative weights for each pair of neighbor nodes. Specifically, after passing through the convolutional layer, the attention mechanism model is introduced to calculate the dynamic relative weights for each pair of neighbor nodes, determining the influence degree of each pair of neighbor nodes on the representation of the target node. The relationship between nodes of edge type r is calculated through the attention weight Formula calculation:
[0084] α ij (r) = exp(LeakyReLU( a T [ W r h i || W r h j ])) ∑ kϵ N r (i) exp(LeakyReLU( a T [ W r h i || W r h k ]) ,
[0085] wherein, , k is the neighbor node connected to node , a is the learned attention weight vector, is the weight matrix of edge type r, represents the concatenation operation, is the embedding vector of node , is the embedding vector of node , is the embedding vector of node k.
[0086] Specifically, in some embodiments, the feature data of each node is mapped through a learnable linear transformation matrix . For node and its neighbor node , hidden representations and are generated, converting the original feature data into a new space that is more conducive to comparison. By concatenating the transformed features of node and its neighbor node , it becomes [ W r h i || W r h j ] in the form of, and taking the dot product with the learned attention weight vector a, then applying the LeakyReLU activation function to obtain the unnormalized attention scores; then using the Softmax function to normalize the obtained attention scores to ensure that the sum of the attention weights of all neighbors is 1, and calculating K groups of independent attention heads in parallel, each group of heads generating different weight distributions . Finally, the outputs of all heads are concatenated or averaged to enhance the model's ability to capture complex relationships.
[0087] Compared with the traditional GCN graph neural network that aggregates neighbor information using fixed weights, the present invention realizes differential weight distribution through the attention mechanism. The weights are dynamically generated based on node features rather than being fixed and preset, enabling the model to adapt to local patterns of different structures, such as the identification of key users in social networks, and significantly improving the node classification accuracy. Moreover, through the attention mechanism, the model can automatically reduce the weights of irrelevant neighbors, reduce noise interference, filter out low-value interaction behaviors, and use the multi-head mechanism to capture multi-subspace information, improving the prediction accuracy of biochemical properties.
[0088] Further, in some embodiments, to complete the missing edges, it includes predicting the missing edges using node embeddings. Given two nodes and with their embedding vectors and , calculate the similarity between them. Use cosine similarity to measure the relationship strength between the two nodes and predict the missing edges. The formula is:
[0089] ,
[0090] The predicted missing edges and their similarity scores indicate which pairs of nodes may have potential relationships.
[0091] Specifically, in some embodiments, after completing the graph embedding, adopt the self-attention mechanism of Transformer, and simultaneously consider the interaction of node embeddings and time embeddings. For the embedding vectors and of two given nodes and , calculate the similarity between them. Use cosine similarity to measure the relationship strength between the two nodes, calculate the attention scores, and predict the probability of the existence of missing edges.
[0092] Further, in some embodiments, define a binary classification loss to represent whether an edge exists: the label {0, 1}, if the edge exists then =1, otherwise =0. The formula for the binary cross-entropy loss function is:
[0093] ,
[0094] where, is the set of edges actually existing in the graph, It is a negative sampling edge set; through the inference of the heterogeneous graph neural network, the missing nodes and missing edges in the attack chain are complemented. Since there is no accurate answer as to whether this edge exists when predicting the existence probability of the missing edge, a binary classification loss is defined to represent whether the edge exists. By generating positive samples and negative samples, where the negative samples represent non-existent edges; comparing the positive and negative samples to optimize the cross-entropy or margin loss, and it is also possible to backpropagate through the positive and negative samples to continuously optimize the parameters and node embeddings.
[0095] Compared with the static graph model, in the present invention, by the operation method of complementing the missing edges, the TTP can process the dynamic graph structure, adapt to the real-time data stream, and improve the prediction ability for time-series related edges; at the same time, due to its multi-source information fusion characteristics, it can combine node attributes, topological structures, and time features to improve the prediction robustness, avoid the over-smoothing problem of traditional GNNs, and directly model the long-range dependencies between nodes.
[0096] Further, in some embodiments, when performing step S3, it specifically further includes:
[0097] Extract the target attack node and all the nodes associated with it, and extract the edges connecting the nodes to construct an attack subgraph; convert the attack subgraph into a natural language text by means of natural language generation.
[0098] In the present invention, the attack subgraph of the heterogeneous graph converts the abstract graph structure information of the attack subgraph into a natural language text by means of natural language generation, which can improve the interpretability of security events, help non-technical personnel evaluate risks through natural language texts without understanding graph structure data, promote decision-making efficiency, and enhance the emergency response coordination ability.
[0099] In some embodiments, converting the attack subgraph into a natural language text by means of natural language generation includes mapping the nodes and relationships in the attack subgraph to natural language texts according to a defined template by means of natural language generation, generating a dynamic description of the attack chain, and systematically depicting the attack behavior.
[0100] Through the results of the first two stages, the TTP of the attacker is depicted in a more intuitive and systematic manner. The attack subgraph extracts the subgraph related to a specific attack event or attack chain from the entire graph, and deeply analyzes the specific process of the attack behavior, the participating nodes and their mutual relationships. This process can help focus on the specific path of the attack, understand the attacker's tactics, techniques, and goals, and at the same time provide support for attack chain completion and prediction.
[0101] Natural language generation technology can transform attack subgraphs into natural language texts that are easy to understand. By defining templates to map the nodes and relationships of attack subgraphs into natural language. For example, if a certain attacker uses a certain attack technique, a template "The attacker [attacker ID] used the attack technique [technique ID]" can be defined to generate a specific description. Through the nodes and relationships of the attack subgraph, combined with natural language generation technology, a dynamic description of the attack chain can be automatically generated, converting attack behaviors into easy-to-understand texts, and providing an intuitive and systematic portrayal of complex attack behaviors.
[0102] A multi-dimensional TTP feature learning and completion method in the present invention starts from four dimensions of tactics, techniques, time series, and correlation relationships to extract TTP features, constructing a more comprehensive attack behavior model; through heterogeneous graph neural networks for graph embedding, introducing two aggregation methods and an attention mechanism to implement downstream tasks and complete the completion behavior of the attack chain; after completion, an attack subgraph of the attack chain is constructed and combined with natural language generation to generate a complete report on complex attack behaviors, not only providing an easy-to-understand description of the attack chain for security analysts, but also providing comprehensive support for subsequent defense strategies and risk assessments, clearly showing the process of the attacker from initial penetration to final lateral penetration.
[0103] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A multi-dimensional TTP feature learning and completion method, characterized in that, Including: Obtain log data and define a heterogeneous graph, divide it into several dimensions according to attack behaviors, and perform feature extraction and definition for each dimension; Initialize node embeddings, use the heterogeneous graph convolutional network HGCN as the convolutional layer, update the embeddings of nodes according to the information of their neighbor nodes. During the update process, the update methods include: neighbor aggregation and cross-type information transfer; The aggregation operation of the neighbor aggregation is weighted according to the type of the node and the type of the connected edge, and the formula is: , in, , r is the edge type, R is the edge type set, Is with the node The set of all connected neighbor nodes of type r, is the weight matrix associated with edge type r, the learned feature transformation matrix, is the activation function, Is a node The number of neighbor nodes of type r, Is with the node The set of all connected neighbor nodes of type r, Is a node The number of neighbor nodes of type r, Is a node In the The embedding vector of the layer, Is a node The updated node embedding vector obtained after the aggregation operation of the convolutional layer; The cross-type information transfer combines the node type and the edge type for cross-type information aggregation, and the formula is: , The update of each type of node depends on different types of neighbor nodes and edge types. Each type of node uses its own weight matrix for feature transformation, where, is the bias term; the attention mechanism model is used to calculate the dynamic relative weights for each pair of neighbor nodes. After passing through the convolutional layer, the attention mechanism model is introduced to calculate the dynamic relative weights for each pair of neighbor nodes, determining the influence degree of each pair of neighbor nodes on the target node representation. The relationship between nodes of edge type r is calculated through the attention weight formula: , Among them, , k is a neighbor node connected to the node , a is the learned attention weight vector, is the weight matrix of edge type r, represents the concatenation operation, is the embedding vector of node , is the embedding vector of node , is the embedding vector of node k, completing the missing edge.
2. The multi-dimensional TTP feature learning and completion method according to claim 1, wherein It also includes extracting the target attack nodes and all the nodes associated with them, and extracting the edges connecting these nodes. An attack subgraph is constructed according to the nodes and edges. In the way of natural language generation, the attack subgraph is converted into natural language text.
3. A multi-dimensional TTP feature learning and completion method according to claim 2, characterized in that In the way of natural language generation, converting the attack subgraph into natural language text includes, in the way of natural language generation, mapping the nodes and relationships in the attack subgraph to natural language text according to a defined template, generating a dynamic description of the attack chain, and systematically depicting the attack behavior.
4. A multi-dimensional TTP feature learning and completion method according to claim 1, characterized in that Define a heterogeneous graph, including defining a heterogeneous graph, which includes several nodes and edges for connecting the nodes. G represents the heterogeneous graph, V represents the nodes, including different types of nodes; E represents the edges, which are used to express the relationships between the nodes V, and each relationship corresponds to a different edge type.
5. A multi-dimensional TTP feature learning and completion method according to claim 1, characterized in that Divided into several dimensions according to attack behaviors, and perform feature extraction and definition for each dimension, including: Divided into a tactical dimension, a technical dimension, a time series dimension, and an association relationship dimension according to attack behaviors; each dimension is independent of each other, and separately extract the tactical features of the tactical dimension, the technical features of the technical dimension, the time series features of the time series dimension, and the association relationship features of the association relationship dimension; define the extracted features.
6. A multi-dimensional TTP feature learning and completion method according to claim 4, characterized in that, Define the extracted features, including defining the extracted tactical features and technical features according to the type of definable nodes; define the extracted time series features by adding a timestamp feature to each attack behavior node to record the specific time when the behavior occurs; Map the dependencies between attacker behaviors and the relationships between attackers, tools, and target systems to edges, and define the extracted association relationship features.
7. A multi-dimensional TTP feature learning and completion method according to claim 1, characterized in that Node embedding initialization, including representing each type of node and edge as a vector. For node V, its initial embedding vector is . The embedding vector is initialized by pre-trained word embeddings or randomly. Each edge E connects two nodes and .
8. A multi-dimensional TTP feature learning and completion method according to claim 1, characterized in that Complete the missing edges, including predicting the missing edges using node embeddings. Given two nodes and embedding vectors and , calculate the similarity between them, use cosine similarity to measure the strength of the relationship between two nodes, and predict the missing edges. The formula is: , The predicted missing edges and their similarity scores, indicating which pairs of nodes may have potential relationships; Define another binary classification loss to represent whether the edge exists: the label {0, 1}, if the edge exists then = 1, otherwise = 0, and the formula for the binary cross-entropy loss function is: , Among them, is the edge set actually existing in the graph, is the negative sampling edge set; through the inference of the heterogeneous graph neural network, the missing nodes and missing edges in the attack chain are complemented.
Citation Information
Patent Citations
Sample analysis method, apparatus, electronic device, and medium based on missing data
WO2021151305A1