Link prediction reasoning method based on network threat intelligence attack graph

By constructing a heterogeneous attack map in the network threat intelligence attack map and introducing AHP weight sampling, type-aware aggregation and dynamic weighting mechanisms, the shortcomings of traditional technologies in link prediction are solved, and high accuracy prediction of potential attack paths is achieved.

CN120074946APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510412113.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When handling network threat intelligence attack graphs, the prior art lacks the ability to distinguish nodes and relationships in heterogeneous graphs and cannot effectively capture semantic differences. In addition, traditional graph neural networks have shortcomings in multi-relational modeling and information aggregation, resulting in low link prediction accuracy.

Method used

A link prediction inference method based on network threat intelligence attack graph is proposed. By constructing a heterogeneous attack graph, combining AHP weight sampling, type-aware aggregation and dynamic weight mechanisms, effective inference of potential attack paths is achieved. This method combines node content, structure information and neighbor type information, and uses improved graph neural network structure for link prediction.

Benefits of technology

It improves the prediction ability of complex network attack paths, enhances the robustness of the model under diversified structure, improves the quality of node representation, and realizes accurate identification and prediction of potential attack links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074946A_ABST
    Figure CN120074946A_ABST
Patent Text Reader

Abstract

The invention discloses a link prediction reasoning method based on a network threat intelligence attack graph. AHP is used for distributing weights for different types of nodes. And using word2vec to encode the text content of the selected target node and the neighbor nodes thereof, converting the text into vector representation, and for each neighbor type t, aggregating the content embedding of all neighbor nodes of the same type to obtain the final representation of the neighbor nodes of the type. And performing weighted summation on different types of neighbor embedding to obtain aggregation embedding. And calculating the similarity among the target nodes according to the embedded vectors among the target nodes so as to predict unshown edges. Generating a false relation pair for each pair of positive samples through a negative sampling method, and calculating the loss of the positive and negative samples. The input of the method is a network security knowledge graph comprising nodes and edges, each node comprises text and type information, and the edges only comprise type information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and in particular, to a link prediction method using a graph neural network, which is applicable to the task of identifying potential attack paths based on network threat intelligence. Background Art

[0002] Under the background of the increasingly severe current network security situation, network attacks are characterized by strong concealment, complex paths, and diverse technical means. Traditional attack detection means relying on static rules and empirical judgments are difficult to meet the actual combat requirements. As an important means to deal with advanced persistent threat attacks, network threat intelligence has been widely used to construct attack graphs to reveal the potential associations between network entities. However, most of the existing attack graph construction methods stay at the static modeling level and lack the ability to actively reason about the implicit links in the attack paths. At the same time, in the face of heterogeneous network structures, traditional graph neural networks have deficiencies in neighbor sampling, information aggregation, and multi-relationship modeling, resulting in low link prediction accuracy and difficulty in assisting security experts to discover unapparent attack paths or target nodes.

[0003] In summary, the following technical difficulties exist in the prior art when dealing with network threat intelligence attack graphs:

[0004] (1) There are various types of nodes and relationships in the attack graph, which belongs to a typical heterogeneous graph. Traditional graph neural networks lack the ability to distinguish different node types and relationship types and cannot fully capture the semantic differences between heterogeneous information;

[0005] (2) The connections between entities in the attack graph are sparse, and different types of neighbors contribute differently to the prediction task. Conventional aggregation strategies are prone to introducing redundant or noisy information, affecting the model performance;

[0006] (3) Current graph neural networks mostly use fixed aggregation strategies and cannot dynamically adjust the weights of different neighbor types according to the context, lacking a refined representation modeling mechanism;

[0007] (4) Link prediction models often fail to effectively fuse node content information, such as text descriptions, type features, etc., resulting in insufficient prediction ability and difficulty in identifying potential connections hidden in the attack graph.

[0008] Therefore, the present invention proposes a link prediction and reasoning method based on a network threat intelligence attack graph. By constructing a heterogeneous attack graph containing multiple types of nodes and edges and combining an improved graph neural network structure, the method realizes effective reasoning about potential attack paths in the graph. The proposed method can accurately identify unapparent attack links on the basis of fusing node content, structural information, and neighbor type information, thereby enhancing the foresight and pre-judgment ability of the network security system. Summary of the Invention

[0009] The present invention aims to propose a link prediction and reasoning method based on a network threat intelligence attack graph, which can fuse the attack graph structure, node content, and neighbor type information to accurately predict potential attack paths and attack targets. The input of this method is a network security knowledge graph containing nodes and edges, and each node contains two parts of information: text and type, while the edge only contains type information.

[0010] The technical solution provided by the present invention includes the following key steps:

[0011] Step 1. First, use AHP to assign weights to different types of nodes. The weight represents the importance of each type for the link prediction task, and neighbor sampling is performed according to the AHP weight assignment scheme. The purpose of sampling is to ensure that the total number of neighbor sets of each target node is n nei , and neighbors of higher weight types are preferentially sampled. The sampling process is as follows:

[0012] (1) Initialize sampling: Locate the target node in the knowledge graph and set the target number n of neighbor sampling nei .

[0013] (2) Sampling layer by layer: Start sampling from the first-order neighbors of the target node and sample order by order. When sampling each order of neighbors, preferentially select neighbor nodes with high AHP weight types.

[0014] (3) Sampling end condition: Once the total number of neighbors sampled from all levels reaches n nei , stop the sampling process.

[0015] Step 2. Use word2vec to encode the text content of the target nodes and their neighbor nodes selected in Step 1, and convert the text into a vector representation. The text content of node i consists of a sequence of words Text(i) = [w 1 , w 2 ,..., w i ..., w L , and the embedding representation e text (i) of this text content is generated through word2vec. The formula is expressed as: Then, one-hot encoding is used to encode the attribute information of the node. The attribute of node i is Attr(i), and the vector representation e attr (i) of the node attribute is obtained through one-hot encoding. The formula is expressed as: e attr (i) = OneHot(Attr(i)). After the text and attribute information of the node are encoded respectively, the text encoding vector e text (i) and the attribute encoding vector e attr(i) The spliced input is fed into TextCNN for deep encoding to capture the temporal and context dependencies in the node content, generating the final content embedding representation of the node. The final content embedding representation of node i is e(i), and the deep encoding process of TextCNN is as follows: e(i) = TextCNN(concat(e text (i), e attr (i))). The final node content embedding e(i) is the representation after the deep encoding of the node text and attribute information by TextCNN.

[0016] Step 3. For each neighbor type t, aggregate the content embeddings of all neighbor nodes of the same type to obtain the final representation of the neighbor nodes of this type. Let the content embedding of node i ′ be e(i ′ ), the set of neighbor nodes of node i be N(i), and N t (i) be the set of all neighbor nodes of type t. Aggregate the content embeddings of these neighbors through GCN to obtain the neighbor aggregation embedding of target node i corresponding to type t, denoted as h t (i). The formula is expressed as: where |N t (i)| is the number of nodes in the set of neighbor nodes of the same type N t (i) of target node i.

[0017] Step 4. After completing Step 3, perform a weighted sum of the neighbor embeddings of different types to obtain the final aggregated embedding representation of target node i. Design a dynamic weight w t (i) for each type t of neighbor nodes of target node i. This dynamic weight is learned through a fully connected neural network and dynamically adjusted according to the specific context of node i during the training process. The input layer of this fully connected neural network receives the content embedding e(i) of the target node and the neighbor node embedding h t (i) of the current neighbor type t, and concatenates them into a long vector [e(i); h t (i)]; this long vector is passed as input to the first layer of the neural network, and after being processed by weighted sum and bias terms, it undergoes a non-linear transformation through the ReLU activation function; the output of the hidden layer is passed to the second layer, and after being processed by weighted sum and bias terms, the final weight w t (i) is obtained through the Sigmoid activation function. The output weight w t (i) is restricted to the range [0, 1] to ensure that all neighbor type weights are within a reasonable range. The formula can be written as: w t (i) = σ(W 2 ·ReLU(W 1 ·[e(i); h t(i)] + b 1 ) + b 2 ), where W 1 and W 2 are the weight matrices of the first layer and the second layer respectively; b 1 and b 2 are the bias terms. After calculating the weight w t (i), use the weighted summation method to aggregate the features of neighbor nodes, and combine with the content embedding e(i) of the target node itself to obtain the final aggregated embedding h(i): h(i) = e(i) + ∑ t∈T w t (i) · h t (i).

[0018] Step 5. Calculate the similarity between the target nodes according to their embedding vectors, so as to predict the unexposed edges. Use the dot product of node embeddings to calculate the link prediction probability between node pairs. The calculation method of the judgment index is as follows: s ij = σ(h(i) · h(j)), where h(i) and h(j) are the embedding vectors of node i and node j respectively. The inner product of similar vectors in space is larger, and the larger the independent variable of the sigmoid function, the closer the dependent variable is to 1, otherwise it approaches 0. Set the parameter q here. If the actually calculated s is greater than q, it is considered that there is an association relationship between the two nodes.

[0019] Step 6. Generate a pair of false relationship pairs for each pair of positive samples through the negative sampling method, and calculate the losses of positive and negative samples. The loss function includes the losses of positive and negative samples, aiming to minimize the difference between the model prediction and the true label. For a batch of data, the calculation method of the loss function is as follows: where Ω A represents the set of all node pairs in this batch, A is the number of samples; y ij represents the true label of the node pair (i, j), 1 means there is a true edge, that is, a positive sample, and 0 means there is no edge, that is, a negative sample. By minimizing this total loss function, the model will enhance the ability to distinguish positive samples, gradually optimize the node embeddings during the training process, and thus improve the accuracy of link prediction.

[0020] Compared with the prior art, the present invention constructs a heterogeneous attack graph and introduces the AHP weight sampling, type-aware aggregation and dynamic weight mechanism, which improves the prediction ability of complex network attack paths. It has the following advantages compared with the prior art:

[0021] The present invention first introduces the mechanism combining AHP neighbor sampling and type-aware aggregation in the link prediction task, realizing the efficient modeling of heterogeneous neighbor information;

[0022] Through a dynamic weight learning mechanism, the present invention adaptively adjusts the contributions of neighbor types, enhancing the robustness of the model under diverse structures;

[0023] The present invention fully integrates the text content and attribute information of nodes, and combines a deep convolutional network to achieve automatic feature extraction, improving the quality of node representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0025] Figure 1 It is a flowchart of the method of the present invention.

[0026] Figure 2 It is a flowchart of node content encoding.

[0027] Figure 3 It is a flowchart of aggregation of neighbor nodes of the same type. DETAILED DESCRIPTION OF THE INVENTION

[0028] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0029] Embodiment 1

[0030] As Figure 1 shown, Embodiment 1 is a link prediction and reasoning method based on a network threat intelligence attack graph, and this method consists of 8 steps. The specific description is as follows:

[0031] Step 1. This example constructs a cybersecurity knowledge graph based on the publicly available dataset DNRTI in the CTI field. The node types include: HackOrg, OffAct, SamFile, SecTeam, Tool, Time, Purp, Area, Idus, Org, Way, Exp, Features, a total of 13 types. The edge types include: target_industry, target_organization, use_tool, use_way, relate_to, have_file, have_purpose, hacker_active_time, exploit, launch, org_located_at, idus_located_at, tool_have_feature, file_have_feature, sec_active_time, find, used_in, developed_by, associated_with, a total of 19 types. For each directly connected node pair in the knowledge graph, a random node is randomly selected from the non-connected nodes to form a data record containing the content and type of the head node, tail node, and random node. A total of 998 data records are used as the dataset for subsequent model training.

[0032] Step 2. Use AHP to assign weights to the 13 node types. The weights represent the importance of each type for the link prediction task, and neighbor sampling is performed according to the AHP weight assignment scheme. The purpose of sampling is to ensure that the total number of neighbors of each target node is 10, and neighbors with higher weights are preferentially sampled. The sampling process is as follows:

[0033] (1) Initialize sampling: Find the input node in the knowledge graph and set the target number of neighbor sampling to 10.

[0034] (2) Layer-by-layer sampling: Start sampling from the first-order neighbors of the input node and sample step by step. When sampling each order of neighbors, preferentially select neighbor nodes with high AHP weight types.

[0035] (3) Sampling end condition: Once the total number of neighbors sampled from all levels reaches n nei , stop the sampling process.

[0036] Step 3. Use word2vec to encode the text content of the nodes and their neighbor nodes in Step 2, and convert the text into a vector representation. Assume that the text content of node i consists of a sequence of words Text(i) = [w 1 , w 2 ,..., w L , and the embedding representation e of this text content is generated through word2vectext (i). The formula is as follows: Then, the attribute information of the node is encoded using one - hot encoding. Assume the attribute of node i is Attr(i), and the vector representation e of the node attribute is obtained through one - hot encoding attr (i). The formula is as follows: e attr (i) = OneHot(Attr(i)). After encoding the text and attribute information of the node respectively, the text encoding vector e text (i) and the attribute encoding vector e attr (i) are concatenated and input into TextCNN for deep encoding to capture the temporal and context dependencies in the node content, generating the final content embedding representation of the node. The final content embedding representation of node i is e(i), and the deep encoding process of TextCNN is as follows: e(i) = TextCNN(concat(e text (i), e attr (i))). The final node content embedding e(i) is the representation after the deep encoding of the node text and attribute information by TextCNN.

[0037] Step 4. For each neighbor type t, aggregate the content embeddings of all neighbor nodes of the same type to obtain the final representation of the neighbor nodes of this type. Let the content embedding of node i ′ be e(i ′ ), the set of neighbor nodes of node i be N(i), and N t (i) be the set of all neighbor nodes of type t. Aggregate the content embeddings of these neighbors through GCN to obtain the neighbor aggregation embedding of target node i corresponding to type t, denoted as h t (i). The formula is as follows: where, |N t (i)| is the number of nodes in the set N t (i) of neighbor nodes of the same type as node i.

[0038] Step 5. Perform a weighted sum of the neighbor embeddings of different types to obtain the final aggregated embedding representation of node i. A dynamic weight w t (i) is designed for each type t of neighbor nodes of node i. This weight is learned through a fully - connected neural network and dynamically adjusted according to the specific context of node i during the training process. The input layer of this neural network receives the content embedding e(i) of the node and the neighbor node embedding h t (i) of the current neighbor type t, and concatenates them into a long vector [e(i); h t(i); This long vector is passed as input to the first layer of the neural network. After being processed by weighted sums and bias terms, it undergoes a non-linear transformation through the ReLU activation function. The output of the hidden layer is passed to the second layer, and after being processed by weighted sums and bias terms, the final weight w t (i) is obtained. The output weight w t (i) is restricted within the range [0, 1] to ensure that all neighbor type weights are within a reasonable range. The formula can be written as: w t (i) = σ(W 2 ·ReLU(W 1 ·[e(i); h t (i)] + b 1 ) + b 2 ), where W 1 and W 2 are the weight matrices of the first and second layers respectively; b 1 and b 2 are the bias terms. After the weight w t (i) is calculated, the features of neighbor nodes are aggregated using weighted summation, and the final aggregated embedding h(i) is obtained by combining the content embedding e(i) of the node itself:

[0039] Step 6. Calculate the similarity between nodes based on their embedding vectors to predict the unexposed edges. The dot product of node embeddings is used to calculate the link prediction probability between node pairs. The calculation method of the judgment index is as follows: s ij = σ(h(i)·h(j)), where h(i) and h(j) are the embedding vectors of nodes i and j respectively. The inner product of similar vectors in space is larger, and the larger the independent variable of the sigmoid function, the closer the dependent variable approaches 1, and vice versa approaches 0. Here, the parameter q is set. If the actually calculated s is greater than q, it is considered that there is an association relationship between the two nodes.

[0040] Step 7. In each data record mentioned in Step 1, the head node and the tail node form positive samples, and the head node and a random node form negative samples. The dataset in Step 1 is divided into a training set, a validation set, and a test set according to 8:1:1, which are 798, 100, and 100 respectively. The batch training strategy is adopted, Batch Size = 16, the maximum number of training epochs is 50, the learning rate is set to 1e - 4, and Dropou = 0.3. The loss function includes the losses of positive and negative samples and aims to minimize the difference between the model prediction and the true label. For a batch of data, the calculation method of the loss function is as follows: Among them, Ω ADenote the set of all node pairs within this batch, and A is the number of samples; y ij Denote the true label of the node pair (i, j). 1 indicates the true existence of an edge, i.e., a positive sample, and 0 indicates the non-existence of an edge, i.e., a negative sample. By minimizing this total loss function, the model will enhance its ability to distinguish positive samples, gradually optimize the node embeddings during the training process, and thus improve the accuracy of link prediction.

[0041] Step 8. The final model achieves the following performance metrics on the test set: Precision P = 0.7923, Recall R = 0.9100, F1-score = 0.8471, ROC-AUC = 0.8864. Compared with the existing link prediction methods LGNN, ELPH, and Text-enhanced GAT, both F1 and AUC are leading.

[0042] Figure 2 It is the flowchart of node content encoding in Step 3 of Embodiment 1 of the present invention.

[0043] Figure 3 It is the flowchart of aggregating neighbor nodes of the same type in Step 4 of Embodiment 1 of the present invention.

Claims

1. A link prediction reasoning method based on a network threat intelligence attack graph, characterized in that: Step 1. Use AHP to assign weights to different types of nodes; Step 2. Use word2vec to encode the text content of the target node and its neighboring nodes selected in step 1, and convert the text into a vector representation; Step 3. For each neighbor type t, aggregate the content embeddings of all neighbor nodes of the same type to obtain the final representation of the neighbor nodes of that type; Step 4. After completing step 3, perform weighted summation on the neighbor embeddings of different types to obtain the final aggregate embedding representation of the target node i; Step 5. Calculate the similarity between target nodes based on their embedding vectors to predict the undisplayed edges. Step 6. Generate a pair of false relationship pairs for each pair of positive samples through the negative sampling method, and calculate the loss of positive and negative samples.

2. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 1, the node assignment weight represents the importance of each type for the link prediction task, and neighbor sampling is performed according to the AHP weight assignment scheme; the purpose of sampling is to ensure that the total number of neighbor sets of each target node is n nei , and give priority to sampling neighbor types with higher weights; the sampling process is as follows: (1) Initialize sampling: Find the target node in the knowledge graph and set the target number n of neighbor sampling nei ; (2) Layer-by-layer sampling: starting from the first-order neighbors of the target node, sampling layer by layer; When sampling neighbors at each level, neighbor nodes with high AHP weight types are preferred; (3) Sampling end condition: Once the total number of neighbors sampled from all levels reaches n nei , the sampling process stops.

3. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 2, the text content of node i consists of a sequence of words Text(i) = [w1, w2, ..., w i ...,w L ], generate the embedding representation of the text content through word2vec text (i); the formula is: Then the attribute information of the node is encoded using one-hot encoding; the attribute of node i is Attr(i), and the vector representation of the node attribute is obtained by one-hot encoding. attr (i); the formula is: e attr (i) = OneHot(Attr(i)); After encoding the text and attribute information of the node respectively, the text encoding vector e text (i) and the attribute encoding vector e attr (i) The concatenated input is sent to TextCNN for deep encoding to capture the temporal and contextual dependencies in the node content and generate the final content embedding representation of the node; the final content embedding representation of node i is e(i), and the deep encoding process of TextCNN is as follows: e(i) = TextCNN(concat(e text (i),e attr (i))); The final node content embedding e(i) is the representation of the node text and attribute information after deep encoding by TextCNN.

4. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 3, let node i ′ The content is embedded as e(i ′ ), the neighbor node set of node i is N(i), N t (i) is the set of all neighbor nodes of type t. The content embeddings of these neighbors are aggregated through GCN to obtain the aggregated embedding of neighbors of type t corresponding to the target node i, denoted as h t (i); the formula is: Among them, |N t (i)| is the set of neighbor nodes of the same type as the target node i, N t The number of nodes in (i).

5. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 4, a dynamic weight w is designed for each type t of neighbor nodes of target node i t (i), the dynamic weight is learned by a fully connected neural network and dynamically adjusted according to the specific context of node i during the training process; the input layer of the fully connected neural network receives the content embedding e(i) of the target node and the neighbor node embedding h of the current neighbor type t t (i), concatenate them into a long vector [e(i); h t (i)]; this long vector is passed as input to the first layer of the neural network, and after being processed by weights and bias terms, it is activated by the ReLU function to produce a nonlinear transformation; The output of the hidden layer is passed to the second layer, and after being processed by weighted and biased terms, the final weight w is obtained through the Sigmoid activation function. t (i), the output weight w t (i) is limited to the range [0,1], ensuring that all neighbor type weights are within a reasonable range; it can be written as: t (i)=σ(W2·ReLU(W1·[e(i);h t (i)]+b1)+b2), where W1 and W2 are the weight matrices of the first and second layers respectively; b1 and b2 are bias terms; in the weight w t (i) After calculation, the features of neighboring nodes are aggregated using weighted summation, and the final aggregate embedding h(i) is obtained by combining the target node’s own content embedding e(i): h(i) = e(i) + ∑ t∈T w t (i) h t (i).

6. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 5, the dot product of node embeddings is used to calculate the link prediction probability between node pairs; the calculation method of the judgment index is as follows: ij =σ(h(i)·h(j)), h(i) and h(j) are the embedding vectors of node i and node j respectively; the larger the inner product of similar vectors in space, the larger the independent variable of the sigmoid function, the closer the dependent variable is to 1, and vice versa; the parameter q is set here, if the actual calculated s is greater than q, it is considered that the two nodes are associated.

7. The link prediction reasoning method based on network threat intelligence attack graph according to claim 1 is characterized in that: In step 6, the loss function contains the loss of positive samples and negative samples, aiming to minimize the difference between the model prediction and the true label; for a batch of data, the loss function is calculated as follows: Among them, Ω A represents the set of all node pairs in the batch, A is the number of samples; y ij Represents the true label of the node pair (i, j), 1 indicates that the edge actually exists, which is a positive sample, and 0 indicates that the edge does not exist, which is a negative sample. By minimizing this total loss function, the ability to distinguish positive samples is enhanced, and the node embedding is gradually optimized during the training process, thereby improving the accuracy of link prediction.

Citation Information

Cited By

  • Terminal threat event detection method and device, computer equipment and storage medium

    CN120811787A