Threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning
Through the improved method of combining graph neural networks with reinforcement learning, the problem of insufficient inference capability and accuracy in complex threat intelligence knowledge graphs is solved, and more efficient and interpretable knowledge reasoning is achieved, suitable for large-scale and high-complexity data processing.
Patent Information
- Application Number
- CN202510273574.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
When the prior art deals with complex threat intelligence knowledge graphs, the reasoning ability and accuracy need to be improved, and the interpretability is poor, making it difficult to effectively deal with large-scale and high-complexity data.
Using an improved graph neural network (GNN) combined with reinforcement learning, by introducing a variation mechanism and policy network, agents independently learn and optimize inference strategies in the knowledge graph to improve feature extraction and semantic information capture capabilities.
Improve the inference accuracy and efficiency of the knowledge graph, enhance the interpretability and adaptability of the model, and continuously improve the reasoning capabilities in the ever-changing and expanding threat intelligence knowledge graph.
Smart Images

Figure CN120218238A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a knowledge reasoning method for a threat intelligence knowledge graph based on improved GNN and reinforcement learning, and belongs to the technical field of knowledge graph reasoning. Background Art
[0002] In the field of knowledge graph research, knowledge reasoning plays a crucial role, especially in the construction and application of threat intelligence knowledge graphs, where its importance is particularly prominent. Knowledge reasoning aims to utilize the rich information already existing in the knowledge graph to deeply explore those knowledge treasures that have not been revealed, thereby filling the information gaps in the graph and improving its structure. This process can not only enhance the integrity of the knowledge graph but also reveal the complex relationships and laws hidden deep in the data, providing strong support for threat intelligence analysis and promoting new progress in scientific research and practical applications in multiple fields such as network security protection, anti-fraud, and risk management.
[0003] With the rapid development and wide application of network technologies, the amount of data is increasing at an unprecedented speed. Threat intelligence knowledge graphs, with their structured data representation and powerful semantic association capabilities, have become key tools for processing and utilizing this vast amount of data. However, during the construction of knowledge graphs, problems such as incomplete information and the proliferation of noisy data often occur, posing severe challenges to the accuracy and efficiency of knowledge reasoning. Especially when faced with large-scale and highly complex threat intelligence knowledge graphs, traditional knowledge reasoning methods are often unable to cope effectively.
[0004] Traditional knowledge reasoning methods do have many limitations when dealing with such complex graphs. For example, rule-driven reasoning methods can provide accurate and easily interpretable reasoning results, but the formulation of their rules highly depends on the experience and knowledge of domain experts, which limits their expansion ability in new fields. Embedded knowledge reasoning methods, although simple in structure and easy to expand, have poor interpretability and are mainly limited to single-hop reasoning. For complex situations involving multi-hop reasoning paths, their effects are often unsatisfactory. Summary of the Invention
[0005] The purpose of the present invention is to provide a threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning to solve the problems of poor interpretability in knowledge reasoning tasks and the need to improve the reasoning ability and the accuracy of reasoning results in the prior art.
[0006] The technical solution of the present invention is as follows:
[0007] A threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning, comprising the following steps,
[0008] S1. Introduce a variational mechanism to construct an autoencoder based on an improved graph neural network (GNN) to extract features from the original threat intelligence knowledge graph and obtain the node information e at the current moment. t ;
[0009] S2. Construct a policy network. After inputting the node information e at the current moment t into this policy network, output the probability distribution of the actions that the agent will take in the next step in the knowledge graph. The agent interacts with the knowledge graph environment, moves to a new entity node according to the selected action, and updates the current state.
[0010] S3. The soft reward module of reinforcement learning calculates the reward r based on the quality of the historical path that the agent has traversed and feeds the reward r back to the policy network.
[0011] S4. Based on the current state, action, reward, and new state of the agent, update the parameters of the policy network through the policy gradient method.
[0012] S5. Repeat steps S3 and S4 to adjust and optimize the training of the autoencoder based on the improved GNN, the policy network, and the soft reward module until the preset number of training rounds is reached or the inference performance of the agent reaches the expected goal, and complete the construction of the knowledge reasoning model based on the improved GNN and reinforcement learning.
[0013] S6. Obtain the inference result by passing the threat intelligence knowledge graph to be inferred through the constructed knowledge reasoning model based on the improved GNN and reinforcement learning.
[0014] Furthermore, in step S1, the autoencoder based on the improved GNN includes a feature extraction module, an encoder, a variational mechanism, a decoder, and a feature vector generation module.
[0015] Feature extraction module: Extract the original node feature vectors and the original topological structure information from the original threat intelligence knowledge graph.
[0016] Encoder: It includes three layers of graph attention networks. After the first layer of graph attention network performs linear transformation, attention coefficient calculation, and weighted summation on the input original node feature vectors and the original topological structure information, it outputs the node feature vectors that fuse the neighborhood information to the second layer of graph attention network and the third layer of graph attention network respectively. The second layer of graph attention network extracts the information representing the mean of the node distribution, and after transformation, outputs the mean vector μ of the node distribution. The third layer of graph attention network calculates and outputs the standard deviation vector σ of the node distribution.
[0017] Variational mechanism: Based on the principle of Bayesian inference, for the mean vector of the input node distribution and the standard deviation vector of the node distribution, the probability distribution of nodes in the embedding space is constructed in the embedding space, and the probability distributions of the head entity, relation, and tail entity in the embedding space are calculated and output to the decoder;
[0018] Decoder: Sample the embedding vectors of the head entity, relation, and tail entity from the probability distributions of the head entity, relation, and tail entity in the embedding space, integrate these embedding vectors to form a combined vector, and pass the combined vector through the first multi-layer perceptron MLP1, the first Gaussian error linear unit GELU1, the second multi-layer perceptron MLP2, the second Gaussian error linear unit GELU2, and the sigmoid activation function in sequence to output the confidence scores of each triple. Compare the confidence scores of the triples with a set threshold, and retain the triples greater than the set threshold. The retained triples are the semantic information of the output knowledge graph;
[0019] Feature vector generation module: Generate the feature vector of the current state from the semantic information of the knowledge graph combined with the current node of the agent, including the node information e at the current moment t and the structural and semantic features of the surrounding nodes connected to the current node.
[0020] Furthermore, in step S2, the policy network includes a gated recurrent unit GRU, a dimension concatenation unit, a rectified linear unit function layer, i.e., the ReLU function layer, and a softmax function layer.
[0021] Gated recurrent unit GRU: Encode the historical path by the node information e at the current moment t input and the search history encoding information h at the previous moment t-1 , and obtain the search history encoding information h at the current moment t ;
[0022] Dimension concatenation unit: Concatenate the search history encoding information h at the current moment t with the node information e at the current moment t , the given query r q , and output the result to the ReLU function layer after concatenation;
[0023] ReLU function layer: Perform a non-linear transformation according to the weight matrix W1 to obtain the transformed result and output it to the softmax function layer;
[0024] Softmax function layer: Combine the candidate action set and the weight matrix W2 to calculate the probability distribution of the actions to be taken by the agent in the knowledge graph in the next step.
[0025] Furthermore, in step S2, the expression of the policy network is as follows:
[0026]
[0027] where, π(r t |s t ) represents the probability that the agent takes an action in state s t , Softmax represents the normalized exponential function, A t represents the set of candidate actions, W1 and W2 are weight matrices, ReLU represents the rectified linear unit, GRU represents the gated recurrent unit, represents the concatenation operation, e t is the node information at the current moment, r q is the given query, h t-1 is the search history encoding information at the previous moment.
[0028] Further, step S3 is specifically as follows
[0029] S31. Obtain the current action selection a t ;
[0030] S32. When the agent finds the target relationship and entity, give a global reward R g = 1; otherwise, the global reward R g = 0;
[0031] S33. Use the scoring function f and combine it with the global reward R g , and calculate to obtain the reward r;
[0032] S34. Feed back the obtained reward r to the policy network.
[0033] Further, in step S33, using the scoring function f and combining it with the global reward R g , the expression for calculating the reward r is
[0034] r = R g +(1 - R g )f(e s , r q , e t )
[0035] where, e t is the node information at the current moment t, e s is the node information at time s, r q is the given query;
[0036] Further, in step S4, based on the current state, action, reward, and new state of the agent, update the parameters of the policy network through the policy gradient method, specifically
[0037] S41. Collect the experience of the agent when interacting with the knowledge graph in the environment, that is, the state-action-reward-new state (s, a, r, s') quadruple;
[0038] S42. Calculate the cumulative return G at each time step t , that is, the discounted cumulative reward from the current time step to the end:
[0039]
[0040] where γ is the discount factor, k represents the time step, and r t+k represents the reward obtained by the agent at time step k + t;
[0041] S43. Use the policy gradient method to update the parameters of the policy network:
[0042]
[0043] where is the gradient of the performance metric function J(θ) with respect to the parameter θ, and E πθ is the expected cumulative reward, represents the gradient operator, which is used to calculate the gradient of the function with respect to the parameter θ in the gradient ascent method. π θ (a|s) is the probability that the policy network selects action a according to state s, θ is the parameter of the policy network, and G t is the cumulative return;
[0044] S44. Use the gradient ascent method to update the parameter θ of the policy network:
[0045]
[0046] where α is the learning rate, θ = (W1, W2), where W1 and W2 are weight matrices.
[0047] The beneficial effects of the present invention are as follows:
[0048] First, compared with the existing methods, the threat intelligence knowledge graph reasoning method based on the improved GNN and reinforcement learning does not rely on a large amount of labeled data, can better handle high-dimensional and sparse data, and by combining the improved graph neural network GNN with reinforcement learning, can efficiently extract features, can capture rich semantic information, provide more accurate representations and reasoning results, improve the interpretability of the model, and has strong adaptability, and can continuously improve the reasoning ability in the continuously changing and expanding threat intelligence knowledge graph.
[0049] Second, in the present invention, by using the graph neural network GNN to extract the features and semantic information of the original threat intelligence knowledge graph in knowledge reasoning, it can efficiently extract features, capture the complex relationships between nodes and edges, has a powerful expressive ability, can capture rich semantic information, provide more accurate representation and reasoning results, and adaptively adjust model parameters to process noisy and incomplete data; at the same time, GNN has good scalability, is suitable for large-scale graph data, and performs reasoning through the graph structure, improving the interpretability of the model. In addition, the flexibility of GNN enables it to be combined with other models such as reinforcement learning to further enhance the knowledge reasoning ability. These advantages make GNN perform excellently in processing complex structure data, improving the performance and efficiency of reasoning.
[0050] Third, this threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning can enable the agent to autonomously learn and optimize the reasoning strategy by using reinforcement learning for knowledge reasoning. By constructing a reasonable reward mechanism, it can effectively guide the agent to explore paths in the complex threat intelligence knowledge graph, thereby discovering hidden knowledge relationships. This method has strong adaptability, can continuously improve the reasoning ability in the constantly changing and expanding threat intelligence knowledge graph, does not rely on a large amount of labeled data, and can better process high-dimensional and sparse data.
[0051] Fourth, this threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning combines the improved graph neural network GNN with reinforcement learning for the knowledge reasoning task of the threat intelligence knowledge graph, can effectively alleviate the problem of data redundancy in the real world, utilize the characteristics of the graph neural network GNN to discover richer semantic information in the original threat intelligence knowledge graph, lay a foundation for the subsequent knowledge reasoning method based on reinforcement learning, and help effectively handle large-scale and complex threat intelligence knowledge graphs in the subsequent knowledge reasoning task based on reinforcement learning, effectively alleviate the problems of sparse rewards and overfitting of the model, and improve the success rate of path search of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic flow chart of the threat intelligence knowledge graph knowledge reasoning method based on improved GNN + reinforcement learning in an embodiment of the present invention;
[0053] Figure 2 is a schematic illustration of the knowledge reasoning model based on improved GNN and reinforcement learning in the embodiment;
[0054] Figure 3 is a schematic illustration of the autoencoder based on the improved graph neural network GNN in the embodiment;
[0055] Figure 4 is a schematic illustration of the policy network in the embodiment. Specific Embodiment
[0056] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0057] The embodiment provides a threat intelligence knowledge graph reasoning method based on an improved GNN and reinforcement learning, as Figure 1 , including the following steps,
[0058] S1. Introduce a variational mechanism to construct an autoencoder based on an improved graph neural network (GNN), as Figure 3 , and extract the semantic information of the graph spectrum by extracting the original node feature vectors and the original topological structure information.
[0059] In step S1, the autoencoder based on the improved GNN includes a feature extraction module, an encoder, a variational mechanism, a decoder, and a feature vector generation module.
[0060] Feature extraction module: Extract the original node feature vectors and the original topological structure information from the original threat intelligence knowledge graph.
[0061] Encoder: It includes three layers of graph attention networks. After the first layer of graph attention network performs linear transformation, attention coefficient calculation, and weighted summation on the input original node feature vectors and the original topological structure information, it outputs the node feature vectors integrating neighborhood information to the second layer of graph attention network and the third layer of graph attention network respectively. The second layer of graph attention network extracts the information representing the mean of the node distribution, and after transformation, outputs the mean vector μ of the node distribution. The third layer of graph attention network calculates and outputs the standard deviation vector σ of the node distribution.
[0062] Variational mechanism: Based on the Bayesian inference principle, for the input mean vector of the node distribution and the standard deviation vector of the node distribution, construct the probability distribution of the node in the embedding space in the embedding space, calculate the probability distributions of the head entity, relation, and tail entity in the embedding space, and output them to the decoder.
[0063] Decoder: Sample the embedding vectors of the head entity, relation, and tail entity from the probability distributions of the head entity, relation, and tail entity in the embedding space, integrate these embedding vectors to form a combined vector, and pass the combined vector through the first multi-layer perceptron (MLP1), the first Gaussian error linear unit (GELU1), the second multi-layer perceptron (MLP2), the second Gaussian error linear unit (GELU2), and the sigmoid activation function in sequence, and then output the credibility scores of each triple. Compare the credibility scores of the triples with a set threshold, and retain the triples with scores greater than the set threshold. The retained triples are the semantic information of the output graph spectrum.
[0064] In the decoder, data integration is first performed: the embedding vectors of the head, relation, and tail entities of each triple from the threat intelligence knowledge graph are concatenated in order to form a combined vector, enabling subsequent processing to obtain complete triple information and preventing misinterpretation. The combined vector enters a two-layer multi-layer perceptron for processing. After the first-layer multi-layer perceptron MLP adjusts the dimension with a linear transformation, it is enhanced by the first Gaussian error linear unit GELU1 for non-linear expression to initially refine features; then, after the second-layer multi-layer perceptron MLP performs a linear transformation for adaptation and output, the second Gaussian error linear unit GELU deeply explores hidden features, and finally, the credibility score of the triple, which is a value between 0 and 1, is output through the sigmoid activation function. Compared with traditional decoders, this design fits the knowledge graph, enhances feature extraction, provides an intuitive quantification standard, and improves the quality and practicality of reasoning.
[0065] Feature vector generation module: The semantic information of the knowledge graph is combined with the current node of the agent to generate a feature vector of the current state, including the node information e at the current moment t and the structural and semantic features of the surrounding nodes connected to the current node.
[0066] In step S1, by constructing an autoencoder, the task of semantic information mining in the original threat intelligence knowledge graph is completed. The encoder constructed with a three-layer graph attention network GAT processes the complex dependencies between nodes by assigning different attention scores, fully capturing the rich semantic information in the threat intelligence knowledge graph. Compared with traditional encoders, the design of this encoder has significant advantages: on the one hand, the attention mechanism enables the model to adaptively focus on important connections, accurately locating key nodes and connections when dealing with large-scale knowledge graphs, improving the understanding of the knowledge graph; on the other hand, the second and third layers are designed to output the mean and standard deviation for subsequent random sampling of the variational mechanism, providing accurate statistical information for the variational mechanism, making the sampled latent variables fit the original data distribution, enhancing the learning and expression ability of the autoencoder for the semantic information of the knowledge graph, providing a reliable basis for subsequent threat intelligence reasoning, and improving the reasoning accuracy. In the autoencoder, a variational mechanism is introduced to capture the uncertainty in node representations, providing a solid foundation for the subsequent reasoning process. By introducing the variational mechanism, a distribution is output for each node in the embedding space, comprehensively considering uncertainties such as incomplete or noisy node information, making the node representations more robust and accurate.
[0067] S2. Construct a policy network. After inputting the node information e at the current moment t into this policy network, the probability distribution of the actions to be taken by the agent in the next step in the knowledge graph is output. The agent interacts with the knowledge graph environment, moves to a new entity node according to the selected action, and updates the current state.
[0068] In step S2, the policy network includes a gated recurrent unit (GRU), a dimensional concatenation unit, a rectified linear unit function layer (i.e., ReLU function layer), and a softmax layer, as Figure 4 :
[0069] Gated recurrent unit (GRU): Encodes the historical path using the current node information e t at the current time and the historical search encoding information h t-1 at the previous time to obtain the historical search encoding information h t at the current time;
[0070] In the gated recurrent unit (GRU), encoding the historical path specifically involves
[0071] S21. The gated recurrent unit (GRU) includes a reset gate R t for capturing short-term dependencies in the path and an update gate Z t for capturing long-term dependencies in the path. The calculation formulas for the reset gate R t and the update gate Z t are as follows:
[0072] R t = sigmoid(e t W xr + h t-1 W hr + b r ) (1)
[0073] Z t = sigmoid(e t W xz + h t-1 W hz + b z ) (2)
[0074] where h t-1 is the historical search encoding information at the previous time, W xr , W hr , W xz , W hz are weight parameters, and b r , b z are bias parameters;
[0075] S22. Calculate the obtained reset gate R t and update gate Z t with the historical search encoding information h t-1 at the previous time, and process through the update gate to obtain the historical search encoding information h t at the current time. The calculation formula is as follows:
[0076]
[0077] Among them, W xh , W hh is the weight parameter, b h It is a biased parameter.
[0078] Dimension splicing unit: Search the historical coding information h at the current moment t and the node information e at the current moment t , given a query r q After concatenation, the output is sent to the ReLU function layer. The dimension concatenation unit uses functions such as TensorFlow and PyTorch to complete the dimension concatenation operation.
[0079] ReLU function layer: Perform nonlinear transformation according to the weight matrix W1 to obtain the transformed result and output it to the normalized exponential function layer softmax;
[0080] Normalized exponential function layer softmax: Combine the candidate action set and the weight matrix W2 to calculate the probability distribution of the next action taken by the agent in the knowledge graph; the candidate action set is the set of all possible actions that the agent can perform in a certain state, and the calculation uses dot multiplication or matrix multiplication, depending on the representation of the action set.
[0081] In step S2, the expression of the policy network is as follows:
[0082]
[0083] Among them, π(r t |s t ) means in state s t The probability of taking an action under , Softmax represents the normalized exponential function, A t represents the candidate action set, that is, the set of all possible actions that the agent can perform in a certain state. W1 and W2 are weight matrices. ReLU represents the rectified linear unit, and GRU represents the gated recurrent unit. represents the splicing operation, e t is the node information at the current moment, r q For a given query, h t-1 Search historical encoding information for the last moment.
[0084] In step S2, a reinforcement learning policy network is constructed to complete the selection of candidate actions for the agent in the current entity, realize the dynamic movement of the agent in the knowledge graph, and provide the possibility for exploring potential threat associations.
[0085] S3. The soft reward module of reinforcement learning calculates the reward r based on the quality of the historical path of the agent's movement and feeds the reward r back to the policy network.
[0086] S31. Obtain the current action selection a t ;
[0087] S32. When the agent finds the target relationship and entity, give the global reward R g = 1; otherwise, the global reward R g = 0;
[0088] S33. Adopt the scoring function f and combine it with the global reward R g , and calculate the reward r:
[0089] r = R g +(1 - R g ) f(e s , r q , e t ) (6)
[0090] where e t is the node information at the current time t, e s is the node information at time s, and r q is the given query;
[0091] S34. Feed the obtained reward r back to the policy network.
[0092] In step S3, by constructing a soft reward mechanism of reinforcement learning, set soft rewards are given according to the quality of the historical path of the agent's movement, and positive feedback (reward) or negative feedback (punishment) is given to each action, prompting the agent to make the correct choice at each step; the agent is encouraged to explore high-quality and low-risk paths while avoiding getting stuck in inefficient or wrong reasoning paths, completing the key link of the agent's path optimization. Through a series of rewards fed back by the reward function, the agent can gradually learn to select the optimal path in the knowledge graph, thereby achieving more accurate knowledge reasoning. The agent is an entity that explores and operates in the threat intelligence knowledge graph environment. It interacts with the knowledge graph, selects actions according to the action probability distribution obtained from the policy network in step S2, and moves in the graph. The goal of the agent is to find an optimal path or make an optimal decision according to the guidance of the policy network to achieve the task.
[0093] S4. Based on the current state, action, reward, and new state of the agent, update the parameters of the policy network through the policy gradient method.
[0094] S41. Collect the experience of the agent when interacting with the knowledge graph in the environment, that is, the state-action-reward-new state (s, a, r, s') quadruple;
[0095] S42. Calculate the cumulative return G for each time step t , that is, the discounted cumulative reward from the current time step to the end point:
[0096]
[0097] where γ is the discount factor, k represents the time step, and r t+k represents the reward obtained by the agent at time step k + t;
[0098] S43. Update the parameters of the policy network using the policy gradient method:
[0099]
[0100] where ▽J(θ) is the gradient of the performance metric function J(θ) with respect to the parameter θ, and E πθ is the expected cumulative reward, and ▽ θ represents the gradient operator, which is used to calculate the gradient of the function with respect to the parameter θ in the gradient ascent method. π θ (a|s) is the probability that the policy network selects action a according to state s, θ is the parameter of the policy network, and G t is the cumulative return;
[0101] S44. Update the parameter θ of the policy network using the gradient ascent method to maximize the cumulative return:
[0102] θ ← θ + α▽J(θ) (9)
[0103] where α is the learning rate, which is a hyperparameter that determines the step size of each update. θ = (W1, W2), where W1 and W2 are weight matrices.
[0104] In step S4, the parameters of the policy network are updated using the policy gradient method to optimize the path selection strategy of the agent, making the path selection strategy of the agent more in line with expectations and efficiently and accurately deriving the key information in the threat intelligence.
[0105] S5. Repeat steps S3 and S4 to adjust and optimize the training of the autoencoder, policy network, and soft reward module based on the improved graph neural network GNN until the preset number of training rounds is reached or the inference performance of the agent reaches the expected goal, completing the construction of the knowledge inference model based on the improved GNN and reinforcement learning, as Figure 2 .
[0106] In step S5, the autoencoder, policy network, and soft reward module based on the improved graph neural network (GNN) are adjusted and optimized for training to optimize the path selection strategy of the agent, improve the inference performance and efficiency, and finally complete the construction of the knowledge inference model based on the improved GNN and reinforcement learning. It can capture more graph semantic information, optimize the policy network and reward mechanism, integrate and test each module to ensure efficient collaborative work, and guarantee the excellent performance of the constructed model.
[0107] S6. Obtain the inference result by passing the threat intelligence knowledge graph to be inferred through the constructed knowledge inference model based on the improved GNN and reinforcement learning.
[0108] This threat intelligence knowledge graph inference method based on the improved GNN and reinforcement learning deeply embeds the knowledge graph through the improved GNN, fully excavates the semantic information in the graph, and at the same time uses reinforcement learning to optimize the path selection strategy to gradually explore and infer potential threat intelligence knowledge. Introducing the knowledge inference model based on the graph neural network (GNN) and reinforcement learning into the threat intelligence knowledge graph not only significantly improves the inference accuracy and efficiency but also endows the model with strong scalability and adaptability, enabling it to easily handle large-scale and high-complexity threat intelligence knowledge graph inference tasks.
[0109] This threat intelligence knowledge graph inference method based on the improved GNN and reinforcement learning aims at the problems that traditional knowledge graph construction methods are difficult to effectively capture deep semantic relationships and cope with the rapid update of security knowledge when dealing with large-scale network security data. By improving the autoencoder structure and introducing a variational mechanism, it outputs a distribution rather than a single point for each node in the embedding space, thereby capturing the uncertainty in the node representation and effectively capturing deep semantic relationships. The autoencoder integrates GNN layers and incorporates edge features into the calculation, capable of fully capturing the rich semantic information in real-world network security data. To optimize the path inference of the agent in the knowledge graph, this invention constructs a policy network that can select the next action according to the current state. At the same time, by designing a reward mechanism, a set soft reward is given according to the quality of the historical path that the agent traverses, and its behavior is optimized through positive feedback (reward) or negative feedback (punishment) to achieve the goal. This reward mechanism particularly considers the characteristics of the network security and threat intelligence scenarios to ensure that the agent can accurately and efficiently explore and utilize relevant knowledge. In addition, the policy network continuously updates the module and parameters, interacts and feeds back with the agent, and finally enables the agent to gradually learn the effective path inference from the starting entity (such as a known threat source) to the target entity (such as a potential security vulnerability). This not only improves the inference ability of the threat intelligence knowledge graph but also realizes more effective path exploration, providing stronger support for the network security and threat intelligence scenarios.
[0110] This threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning applies the knowledge reasoning model based on improved GNN and reinforcement learning to the threat intelligence knowledge graph, which is an important way to enhance knowledge reasoning ability, improve the structure of the knowledge graph, and reveal potential threat relationships, and is of great significance for promoting the development and application of threat intelligence analysis technology.
[0111] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning, characterized by: The following steps are included: S1. Introduce the variational mechanism, build an autoencoder based on the improved graph neural network GNN, extract features from the original threat intelligence knowledge graph to obtain the node information e at the current moment t ; S2, build a strategy network, and convert the node information e at the current moment t After inputting the policy network, the probability distribution of the next action taken by the agent in the knowledge graph is output. The agent interacts with the knowledge graph environment, moves to the new entity node according to the selected action, and updates the current state; S3, the soft reward module of reinforcement learning calculates the reward r according to the quality of the historical path of the agent, and feeds the reward r back to the policy network; S4. Update the parameters of the policy network using the policy gradient method based on the agent’s current state, action, reward, and new state. S5. Repeat steps S3 and S4 to adjust and optimize the training of the autoencoder, policy network and soft reward module based on the improved graph neural network GNN until the preset number of training rounds is reached or the reasoning performance of the agent reaches the expected goal, completing the construction of the knowledge reasoning model based on the improved GNN and reinforcement learning; S6. Obtain the reasoning results of the threat intelligence knowledge graph to be inferred through the knowledge reasoning model based on improved GNN and reinforcement learning.
2. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning as claimed in claim 1, characterized in that: In step S1, the autoencoder based on the improved graph neural network GNN includes a feature extraction module, an encoder, a variational mechanism, a decoder and a feature vector generation module. Feature extraction module: extracts the original threat intelligence knowledge graph to obtain the original node feature vector and original topological structure information; Encoder: It includes three layers of graph attention networks. The first layer of graph attention network performs linear transformation, attention coefficient calculation and weighted summation on the original input node feature vector and original topological structure information, and then outputs the node feature vector integrating neighborhood information to the second layer of graph attention network and the third layer of graph attention network respectively. The second layer of graph attention network extracts the information representing the mean of node distribution and outputs the mean vector μ of node distribution after transformation. The third layer of graph attention network calculates and outputs the standard deviation vector σ of node distribution. Variational mechanism: Based on the Bayesian inference principle, the mean vector of the input node distribution and the standard deviation vector of the node distribution are used to construct the probability distribution of the node in the embedding space. The probability distribution of the head entity, relationship, and tail entity in the embedding space is calculated and output to the decoder. Decoder: Sample the embedding vectors of the head entity, relationship and tail entity from the probability distribution of the embedding space, integrate the data of these embedding vectors to form a combination vector, pass the combination vector through the first layer of multi-layer perceptron MLP1, the first Gaussian error linear unit GELU1, the second layer of multi-layer perceptron MLP2, the second Gaussian error linear unit GELU2 and the S-type activation function, and output the credibility score of each triplet, compare the credibility score of the triplet with the set threshold, retain the triplet greater than the set threshold, and the retained triplet is the semantic information of the output graph; Feature vector generation module: The feature vector of the current state is generated by combining the semantic information of the graph with the current node of the intelligent agent, including the node information at the current moment. t As well as the structural and semantic features of the surrounding nodes connected to the current node.
3. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning as claimed in claim 1, characterized in that: In step S2, the policy network includes a gated recurrent unit GRU, a dimension concatenation unit, a rectified linear unit function layer, i.e., a ReLU function layer, and a normalized exponential function layer, Softmax. Gated recurrent unit GRU: input by the current node information e t Search the historical encoding information h at the previous moment t-1 , encode the historical path and obtain the current search history encoding information h t ; Dimension splicing unit: Search the historical coding information h at the current moment t and the node information e at the current moment t , given a query r q After splicing, output to the ReLU function layer; ReLU function layer: Perform nonlinear transformation according to the weight matrix W1 to obtain the transformed result and output it to the normalized exponential function layer softmax; Normalized exponential function layer softmax: Combine the candidate action set and the weight matrix W2 to calculate the probability distribution of the next action taken by the agent in the knowledge graph.
4. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning according to any one of claims 1 to 3, characterized in that: In step S2, the policy network is expressed as follows: Among them, π(r t |s t ) indicates that the agent is in state s t The probability of taking an action under , Softmax represents the normalized exponential function, A t represents a set of candidate actions, W1 and W2 are weight matrices, ReLU represents a rectified linear unit, GRU represents a gated recurrent unit, represents the splicing operation, e t is the node information at the current moment, r q For a given query, h t-1 Search historical encoding information for the last moment.
5. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning according to any one of claims 1 to 3, characterized in that: Step S3, specifically, S31. Get the current action selection a t ; S32. When the agent finds the target relationship and entity, it gives a global reward R g =1; Otherwise, the global reward R g =0; S33, using the scoring function f and combining it with the global reward R g , calculate the reward r; S34: Feedback the obtained reward r to the strategy network.
6. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning as claimed in claim 5, characterized in that: In step S33, the scoring function f is used in combination with the global reward R g , the reward r is calculated as: r=R g +(1-R g )f(e s ,r q ,e t ) Among them, e t is the node information at the current time t, e s is the node information at time s, r q for a given query.
7. The threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning according to any one of claims 1 to 3, characterized in that: In step S4, based on the current state, action, reward and new state of the agent, the parameters of the policy network are updated by the policy gradient method, specifically, S41, collect the experience of the agent when interacting with the knowledge graph in the environment, that is, the state-action-reward-new state (s, a, r, s') quadruple; S42. Calculate the cumulative return G for each time step t , which is the discounted cumulative reward from the current time step to the end point: Where γ is the discount factor, k represents the time step, and r t+k represents the reward obtained by the agent at time step k+t; S43. Update the parameters of the policy network using the policy gradient method: in, is the gradient of the performance index function J(θ) with respect to the parameter θ, E πθ To expect cumulative rewards, Represents the gradient operator, which is used to calculate the gradient of the function with respect to the parameter θ in the gradient ascent method, π θ (a|s) is the probability that the policy network selects action a based on state s, θ is the parameter of the policy network, G t is the cumulative return; S44. Update the parameters θ of the policy network using the gradient ascent method: Among them, α is the learning rate, θ = (W1, W2), where W1 and W2 are weight matrices.
Citation Information
Cited By
Intelligent monthly settlement task scheduling method and system based on reinforcement learning
CN120803681A
A method and system for intelligent scheduling of monthly tasks based on reinforcement learning
CN120803681B
Event allocation method based on collaborative reasoning of multi-modal data and knowledge graph
CN120950699A
Scientific and technological intelligence agent training method and system based on reinforcement learning
CN121328607A
AI agent security threat disposal-self-healing collaborative integrated system fusing knowledge graph and reinforcement learning
CN122174241A