Time sequence knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning
By introducing hierarchical knowledge embedding and reinforcement learning in knowledge graph reasoning, the problems of poor interpretability and lack of knowledge in knowledge graph reasoning are solved, and more accurate and interpretable knowledge reasoning effects are achieved.
Patent Information
- Application Number
- CN202510206584.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
AI Technical Summary
The existing knowledge graph reasoning methods are limited in interpreting the reasoning process, and it is difficult to present the reasoning basis intuitively, and there is a problem of lack of knowledge, which affects the effectiveness of downstream tasks.
The timing knowledge graph inference method based on hierarchical knowledge embedding and reinforcement learning is adopted to obtain knowledge feature representations through embedded learning at the sub-graph level and global graph level, and the action scoring function and path-based soft rewards are designed in the reinforcement learning inference model to improve the inference effect.
A more accurate representation of potential features of knowledge graphs is achieved, which improves the accuracy and interpretability of knowledge reasoning, alleviates the problem of reward sparseness, and improves the effectiveness of downstream tasks.
Smart Images

Figure CN120124745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temporal knowledge graph reasoning, and more specifically to a temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning. Background Art
[0002] A knowledge graph is a multi-relational graph that reflects the internal connections of real-world events. It is represented by triples containing entities and relationships and stores a large amount of information about the real world. Knowledge graphs have achieved remarkable success in many downstream applications, such as question-answering systems and recommendation systems. However, whether constructed manually or extracted automatically, knowledge graphs inevitably suffer from incompleteness. This phenomenon of missing knowledge directly and significantly affects the performance of many downstream tasks that rely on knowledge reasoning.
[0003] To address this challenge, the academic community has explored diverse KG reasoning strategies aimed at compensating for its incompleteness. These include, but are not limited to, those based on translational vector models, which simulate the association between entities and relationships through translational transformations in vector spaces; advanced methods based on neural networks, which utilize the powerful learning ability of neural networks to capture complex semantic information; and methods based on tensor decomposition, which attempt to reveal hidden association patterns in the data by decomposing high-order tensors. However, these strategies are still limited in explaining the reasoning process and are difficult to intuitively present the reasoning basis. Summary of the Invention
[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning to solve the technical problems raised in the background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning, characterized by comprising the following steps:
[0006] Step S1, establish a hierarchical knowledge embedding model, and the hierarchical knowledge embedding model obtains knowledge feature representations through embedding learning at two levels: the sub-graph level and the global graph level;
[0007] Step S2, establish a reinforcement learning reasoning model, and the reinforcement learning reasoning model introduces a weighted action scoring mechanism to design a policy network;
[0008] Step S3, train and optimize the reinforcement learning reasoning model;
[0009] Step S4, combine the hierarchical knowledge embedding model and the reinforcement learning reasoning model to conduct temporal knowledge graph reasoning experiments;
[0010] Step S5: Analyze the experimental results, conduct ablation experiments, perform interpretability analysis, and complete the temporal knowledge graph reasoning based on hierarchical knowledge embedding and reinforcement learning.
[0011] In a preferred embodiment, in step S1, the sub-graph level is semantic dependency knowledge embedding. The model uses a relational graph neural network as a semantic aggregator to obtain the dynamic semantic embedding of each entity in the sub-graph The formula for obtaining the dynamic semantic embedding is In the formula represents the neighbor nodes of e in the knowledge graph i at time t, is the set of neighbor nodes of e, σ(·) is the activation function, and are trainable weight parameter matrices, and the initial entity feature embedding h and e are set to the static features X and e After passing through w layers of convolution, the feature representation of entity e at time t considering the semantic dependency with its neighbors can be obtained i For the relationship r between entities at time t its feature vector is represented by i and is calculated by aggregating the features of entities with relationship r at time t The calculation formula is: i In the formula represents the set of entities related to relationship r at time t Meanpooling(·) is the average pooling operation, which acts on the set of entities related to relationship r at time t i i
[0012] In a preferred embodiment, in step S1, the global graph level is time-dependent knowledge embedding. First, the knowledge graphs at different times are connected through common entities to construct a global graph. Then, the influence of the time interval on the relationship between adjacent nodes and in the global graph is considered. The edge between adjacent nodes and is regarded as the time-related relationship of the same entity and is represented by r τ The feature embedding of each relationship r τ is represented by the feature embedding formula. The feature embedding formula is In the formula, |t i - t j | represents the absolute value of the time interval, represents the relationship r τ Static feature embedding, where φ(·) is the time encoding function.
[0013] In a preferred embodiment, in step S1, the hierarchical knowledge embedding model uses an attention mechanism to weight the relevance between entities in the global graph and neighbor nodes with a weight coefficient calculated as follows. The formula for the weight coefficient is where each initial input entity feature embedding representation is the entity feature output at the subgraph level is the neighbor set of t in P, α ∈ R 3d and are learnable weight parameters, σ(·) is the activation function, ·T represents the transpose, and || is the concatenation operation.
[0014] In a preferred embodiment, after the hierarchical knowledge embedding model uses the attention mechanism, by adaptively aggregating features from all neighbors, the feature embedding of the entity in the global graph is obtained. The formula for obtaining the feature embedding is where σ(·) is the activation function, and are weight parameter matrices for aggregation and self-loop. After β-layer operations at the global graph level, the feature embedding representation of the entity can be obtained The hierarchical knowledge embedding model uses to represent the feature embedding representation of entity e output at the global graph level.
[0015] In a preferred embodiment, in step S2, the agent and the environment continuously interact through the quadruple of the Markov decision process to establish a reinforcement learning inference model. The quadruple of the Markov decision process is the state space, action space, state transition, and reward. The reinforcement learning inference model uses an action scoring function to score each candidate action and calculate the probability of state transition. The reinforcement learning inference model first uses two multi-layer perceptrons to encode the state information containing the query problem and outputs the entity feature embedding representation of the predicted action and the feature embedding representation of the outgoing edge. The formula for the entity feature embedding representation is: The formula for the feature embedding representation is
[0016] In a preferred embodiment, in step S2, the policy network calculates the entity similarity and relationship similarity between the predicted action and the candidate action. By performing a weighted sum of these two similarities, a scoring function φ(a n , s l ) is used to obtain the final candidate action score. The calculation formula of the scoring function φ(a n , s l ) is where W 6 , W e , W r and W β are learnable matrices, || is the concatenation operation, σ(·) is the activation function, and <·> represents the similarity calculation operation. are the feature embedding representations of the candidate action entity e n and the relationship r n calculated by the hierarchical knowledge embedding model, respectively; are the feature embedding representations of e l , r l at time t l calculated by the hierarchical knowledge embedding model, respectively; are the feature embedding representations of the query problem entity e q and the relationship r q calculated by the hierarchical knowledge embedding model, respectively. After scoring all the candidate actions in the candidate action set, the reinforcement learning inference model obtains the policy network through the softmax function.
[0017] In a preferred embodiment, in step S4, four publicly available TKG datasets, ICEWS14, ICEWS18, WIKI, and YAGO, are used in the experiment to evaluate the performance of the temporal knowledge graph reasoning, and widely used evaluation metrics are adopted for the performance evaluation of the temporal knowledge graph. The evaluation metrics include the mean reciprocal rank and Hits@k. For each quadruple q = (e s , r, e o , t) in the test set, two query problems are evaluated: q o = (e s , r,?, t) and q s = (?, r, e o , t). The calculation formula of the mean reciprocal rank is where f test is the set of quadruples in the test set, |f test | is the number of quadruples in the test set, rank(e o |q o ) represents the ranking of the predicted tail entity, and rank(e s |qs ) It represents the ranking of the predicted head entity. The larger the MRR value, the better the model performance. Hits@k refers to the percentage of the number of times the target entity appears among the top k candidate entities after ranking, where k takes values of 1, 3, and 10.
[0018] In a preferred embodiment, in step S4, parameter settings need to be performed before the experiment. When setting parameters, the dimension d of the entity and relationship feature embedding vectors is set to 100. During the inference experiment process, the reinforcement learning inference model selects the latest N outgoing edges as candidate actions at each step. Among them, N for the ICEWS14 and ICEWS18 datasets takes a value of 50, 60 for the WIKI dataset, and 30 for the YAGO dataset. The length L of the inference path is set to 3, the discount factor γ of the REINFORCE algorithm is set to 0.95, the batch size during training is set to 512, the Adam optimizer is used to optimize the parameters, the learning rate is set to 0.001, and in the hierarchical knowledge embedding model module, the number of sub-graph level network layers w and the number of global graph level network layers β are set to {0, 1, 2, 3, 4} and experiments are carried out.
[0019] The technical effects and advantages of the present invention:
[0020] 1. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning of the present invention introduces a hierarchical knowledge embedding model. Through two-level knowledge embedding, it fully captures semantic dependencies and temporal evolution information to obtain a more accurate latent feature representation of the knowledge graph. An action scoring function is designed in the reinforcement learning inference model, and at the same time, a path-based soft reward is introduced to achieve reward shaping, alleviating the reward sparsity problem to further improve the inference effect;
[0021] 2. The present invention conducts entity prediction tasks and conducts comparative analysis on the results. The results show the effectiveness of the method. Through ablation experiments, it can be seen that fully capturing the semantic dependencies and temporal evolution information of entities and relationships to obtain a more accurate knowledge feature embedding representation is particularly important for improving the knowledge inference effect. It will be considered to further optimize the candidate action space and scoring function in the inference path in reinforcement learning, thereby improving the accuracy of action selection and achieving a better knowledge inference effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic flowchart of the temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning of the present invention.
[0023] Figure 2 It is a schematic diagram of the temporal knowledge graph reasoning model based on hierarchical knowledge embedding and reinforcement learning of the present invention.
[0024] Figure 3Schematic diagram of the number of sub - graph level network layers w of the present invention.
[0025] Figure 4 Schematic diagram of the number of global graph - level network layers β of the present invention. Detailed implementation manners
[0026] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. In addition, the forms of each structure described in the following embodiments are merely examples. The method for temporal knowledge graph reasoning based on hierarchical knowledge embedding and reinforcement learning involved in the present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0027] Refer to Figure 1 And Figure 2 , the present invention provides a method for temporal knowledge graph reasoning based on hierarchical knowledge embedding and reinforcement learning, including the following steps:
[0028] Step S1, establish a hierarchical knowledge embedding model, and the hierarchical knowledge embedding model obtains knowledge feature representations through embedding learning at two levels: sub - graph level and global graph level;
[0029] Step S2, establish a reinforcement learning reasoning model, and the reinforcement learning reasoning model introduces a weighted action scoring mechanism to design a policy network;
[0030] Step S3, train and optimize the reinforcement learning reasoning model;
[0031] Step S4, combine the hierarchical knowledge embedding model and the reinforcement learning reasoning model to conduct temporal knowledge graph reasoning experiments;
[0032] Step S5, analyze the experimental results, conduct interpretability analysis of reasoning, and complete temporal knowledge graph reasoning based on hierarchical knowledge embedding and reinforcement learning.
[0033] In the embodiments of the present application, through the knowledge embedding representation learning at two levels, the present application fully captures semantic dependencies and temporal evolution information to obtain more accurate potential feature representations of the knowledge graph. The knowledge embedding representation learning at the sub - graph level focuses on capturing semantic dependencies between concurrent facts, uses relational graph neural networks to fully capture the semantic relevance between entities and their neighbors, and uses high - order neighbor information in concurrent facts to enhance the semantic representation of entities at each timestamp. The knowledge embedding representation learning at the global graph level is committed to mining temporal evolution laws and modeling the temporal dependence relationships of entities, and uses the attention mechanism to integrate diverse historical information into the knowledge representation to achieve more accurate embedding representation learning of the dynamic features of entities over time.
[0034] Further, in the step S1, the subgraph level is semantic dependency knowledge embedding, and the model uses a relational graph neural network as a semantic aggregator to obtain the dynamic semantic embedding of each entity in the subgraph The formula for obtaining the dynamic semantic embedding is In the formula represents the neighbor nodes of e in the knowledge graph i at time t and is the set of neighbor nodes of e, σ(·) is the activation function and are trainable weight parameter matrices, and the initial entity feature embedding h e and are set to the static feature X e and After w layers of convolution, the feature representation of entity e at time t i considering the semantic dependency with its neighbors can be obtained For the relationship r between entities at time t i , its feature vector is represented by and is calculated by aggregating the features of entities with relationship r at time t i . The calculation formula is: In the formula represents the set of entities related to relationship r at time t i , and Meanpooling(·) is the average pooling operation, which acts on the set of entities related to relationship r at time t i
[0035] In the embodiments of the present application, for concurrent facts, entities usually have strong semantic correlations with their neighbors. Therefore, the hierarchical knowledge embedding model first considers capturing the semantic dependency relationships between these concurrent facts to obtain the feature embedding representation of each entity e in the knowledge graph i at time t and the feature embedding representation of relationship r The knowledge embedding representation learning at the subgraph level focuses on capturing the semantic dependency relationships between concurrent facts, uses a relational graph neural network to fully capture the semantic correlations between entities and their neighbors, and utilizes high-order neighbor information in concurrent facts, aiming to enhance the semantic representation of entities at each timestamp
[0036] Further, in the step S1, the global graph level is time-dependent knowledge embedding. First, the connection between knowledge graphs at different times is realized through common entities to construct a global graph, and then the time interval is considered for adjacent nodes in the global graph and The influence of the relationship between adjacent nodes and The edge between them is regarded as a time-related relationship of the same entity, denoted by r τ For each relationship r τ The feature embedding is represented by the feature embedding formula, and the feature embedding formula is In the formula, |t i -t j | represents the absolute value of the time interval, represents the static feature embedding of the relationship r τ , φ(·) is a time encoding function, and in the step S1, the hierarchical knowledge embedding model uses an attention mechanism to calculate the weight coefficient for the correlation between the entity and its neighbor nodes The calculation formula of the weight coefficient is In the formula, each initial input entity feature embedding representation is the entity feature output at the subgraph level is the neighbor set in P t , α ∈ R 3d and are learnable weight parameters, σ(·) is an activation function, ·T represents transpose, and || is a concatenation operation.
[0037] In the embodiment of the present application, after completing the semantic dependency modeling of knowledge embedding at the subgraph level, the hierarchical knowledge embedding model can obtain the semantic dependency feature embedding representation i of each entity node e in the knowledge graph at time t and the feature embedding representation r of the relationship To further capture the time-dependent relationship between entities at different time points, the hierarchical knowledge embedding model performs message propagation and aggregation operations based on the semantic-level output, thereby modeling the time-dependent relationship between entities, connecting knowledge graphs at different times through common entities, and constructing a global graph. For any and If they have a common entity e s , it is assumed that there is an edge between and . In this way, knowledge graphs at different times can be connected through common entities. Therefore, the knowledge graph sequence can be transformed into a multi-relational graph P t , that is, the global graph, where each can be regarded as its subgraph, while and are regarded as two different nodes therein;
[0038] Next, consider the influence of the time interval on the relationship between adjacent nodes in the global graph P t in and The edges between are regarded as time-related relationships of the same entity, denoted by r and Similar to each semantic relationship r, this paper also converts r τ into a d-dimensional embedding vector. At this level, the feature embedding representation of each relationship r τ can be calculated by the following formula τ where |t
[0039]
[0040] - t i - t j | represents the absolute value of the time interval, represents the static feature embedding of relationship r τ and φ(·) is the time encoding function. The definition formula of the time encoding function φ(·) is where w, p ∈ R d are learnable parameter vectors;
[0041] To more accurately capture the feature information of entities evolving over time, the hierarchical knowledge embedding model uses the attention mechanism
[32] to calculate the weights of the correlation between entities and neighbor nodes in the global graph. The weight coefficient is specifically calculated as The feature embedding representation of each initial input entity is the entity feature output at the subgraph level is in P t the neighbor set, α ∈ R 3d and are learnable weight parameters, σ(·) is the activation function, ·T represents the transpose, and || is the concatenation operation.
[0042] Furthermore, after the hierarchical knowledge embedding model adopts the attention mechanism, by adaptively aggregating the features from all neighbors, the feature embedding of the entity in the global graph is obtained. The calculation formula for obtaining the feature embedding is where σ(·) is the activation function, and are the weight parameter matrices for aggregation and self-loop. After β-layer operations at the global graph level, the feature embedding representation of the entity can be obtained The hierarchical knowledge embedding model uses z e,tiTo represent the feature embedding representation of entity e for the global graph level output.
[0043] In the embodiments of the present application, compared with the feature embedding of entities in the global graph, the embedding representation of the semantic relationship r∈R is relatively stable over a long period of time. Therefore, the relationship feature embedding representation output at the sub-graph level As t i The final feature embedding representation of relationship r at time t That is
[0044] Furthermore, in step S2, the intelligent agent and the environment are continuously interacted through the quadruple of the Markov decision process to establish a reinforcement learning inference model. The quadruple of the Markov decision process is the state space, action space, state transition, and reward. The reinforcement learning inference model uses an action scoring function to score each candidate action and calculate the probability of state transition. The reinforcement learning inference model first encodes the state information containing the query problem using two multi-layer perceptrons and outputs the entity feature embedding representation of the predicted action And the feature embedding representation of the outgoing edge Entity feature embedding representation The calculation formula is: Feature embedding representation The calculation formula is
[0045] In the embodiments of the present application, in the state space of the quadruple of the Markov decision process, let S represent the state space. A state is represented by a quintuple s l =(e l ,t l ,e q ,t q ,r q )∈S, where (e l ,t l ) is the node visited by the intelligent agent at the l-th step, and (e q ,t q ,r q ) are the elements in the query problem, where e q represents the entity in the query problem, r q represents the relationship in the query problem, and (e q ,t q ,r q ) can be regarded as global information, while (e l ,t l ) is local information. The intelligent agent starts from the source node of the query. Therefore, the initial state is s 0 =(e q ,t q ,e q ,tq , r q ) ∈ S;
[0046] Let A denote the action space. A l represents the set of optional actions at the l-th step. A l is composed of the outgoing edges of node e l . Specifically, A l should be {(r′, e′, t′) | (e l , r′, e′, t′) ∈ P t}, but an entity usually has many related historical facts, resulting in a large number of optional actions. Therefore, the final set of candidate actions A l is sampled from the above set of outgoing edges;
[0047] State transition: Transfer to a new node through the edge selected by the agent. The transition function δ: S × A → S is defined as δ(s l , A l ) = s l+1 = (e l+1 , t l+1 , e q , t q , r q ), where A l is the sampled outgoing edge of e l ;
[0048] In the quadruple of the Markov decision process, when designing the reward, the reinforcement learning inference model designs a path reward function R p that fully considers the relationship between the query problem and the inference path. The expression formula is In the formula, D(·) represents the cosine similarity function. To ensure that the similarity calculation result is non-negative, the absolute value of the calculation result is taken. is the feature embedding representation of e L at time t L calculated by the HKEM model; are the feature embedding representations of the query problem entity e q and the relationship r q calculated by the HKEM model respectively. Finally, the reinforcement learning inference model obtains a new reward function R, and the expression formula is R = R h + (1 - R h )R p . In the formula, R h is the hit reward. Through the above reward shaping, the agent can obtain corresponding rewards according to the similarity relationship between the inference path and the query problem during the inference process, improving the problem of sparse rewards.
[0049] Further, in step S2, the policy network calculates the entity similarity and relationship similarity between the predicted action and the candidate actions, and by performing a weighted sum of these two similarities, the scoring function φ(a n , s l ) is used to obtain the final candidate action score. The calculation formula of the scoring function φ(a n , s l ) is where W 6 , W e , W r and W β are learnable matrices, || is the concatenation operation, σ(·) is the activation function, and <·> represents the similarity calculation operation. are the feature embedding representations of the candidate action entity e n and the relationship r n calculated by the hierarchical knowledge embedding model, respectively; are the feature embedding representations of e l , r l at time t l calculated by the hierarchical knowledge embedding model, respectively; are the feature embedding representations of the query problem entity e q and the relationship r q calculated by the hierarchical knowledge embedding model, respectively. After scoring all the candidate actions in the candidate action set, the reinforcement learning inference model obtains the policy network through the softmax function.
[0050] In the embodiment of the present application, the policy network π θ (a l | s l ) = P(a l | s l ; θ) is used to simulate the behavior of the agent in the continuous space, where a l ∈A l , θ is the model parameter. The reinforcement learning inference model designs an action scoring function to score each candidate action and calculate the probability of state transition, and let a n = (e n , t n , r n ) ∈ A l represent a candidate action at the l-th step. Future events are often full of uncertainties, and some queries lack a clear causal logic chain. Therefore, the relevance between the candidate action and the query problem is particularly crucial. In this context, the reinforcement learning inference model first uses two multi-layer perceptrons to encode the state information containing the query problem and outputs the entity feature embedding representation of the predicted action and the feature embedding representation Based on this, the agent further calculates the entity similarity and relationship similarity between the predicted action and the candidate action. Finally, by weighted summing these two similarities, the agent obtains the final candidate action score, and the scoring function φ(a n ,s l ) is And the calculation of the weighting factor β n is obtained according to the formula where W 6 , W e , W r and W β are learnable matrices, ‖ is the concatenation operation, σ(·) is the activation function, and <·> represents the similarity calculation operation. are the feature embedding representations of the candidate action entity e n and the relationship r n calculated by the HKEM model respectively; are the feature embedding representations of e l , r l at time t l calculated by the HKEM model respectively; are the feature embedding representations of the query problem entity e q and the relationship r q calculated by the HKEM model respectively. After scoring all candidate actions in A l , the reinforcement learning inference model can obtain π θ (a l |s l ) through the softmax function.
[0051] Referring to Figure 3 and Figure 4 , in step S4, four publicly available TKG datasets are used in the experiment: ICEWS14, ICEWS18, WIKI, and YAGO to evaluate the performance of temporal knowledge graph reasoning, and widely used evaluation metrics are adopted for temporal knowledge graph performance evaluation. The evaluation metrics include mean reciprocal rank and Hits@k, and for each quadruple q = (e s ,r,e o ,t) in the test set, two query problems are evaluated: q o =(e s ,r,?,t) and q s =(?,r,e o ,t). The formula for mean reciprocal rank is where f test is the set of quadruples in the test set, |f test| is the number of quadruples in the test set, rank(e o |q o ) represents the ranking of the predicted tail entity, rank(e s |q s ) represents the ranking of the predicted head entity. The larger the MRR value, the better the model performance. Hits@k refers to the percentage of the number of times the target entity appears among the top k candidate entities after ranking. k takes values of 1, 3, and 10. In step S4, parameter settings need to be performed before the experiment. When setting parameters, the dimension d of the entity and relationship feature embedding vectors is set to 100. During the inference experiment process, the reinforcement learning inference model selects the latest N outgoing edges as candidate actions at each step. Among them, N for the ICEWS14 and ICEWS18 data sets takes a value of 50, 60 for the WIKI data set, and 30 for the YAGO data set. The length L of the inference path is set to 3, the discount factor γ of the REINFORCE algorithm is set to 0.95, the batch size is set to 512 during training, the Adam optimizer is used to optimize the parameters, and the learning rate is set to 0.001. In the hierarchical knowledge embedding model module, the number of subgraph-level network layers w and the number of global graph-level network layers β are set to {0, 1, 2, 3, 4} and experiments are conducted. For the hierarchical knowledge embedding model, the number of subgraph-level network layers w and the number of global graph-level network layers β are respectively set to {0, 1, 2, 3, 4} for experiments.
[0052] In the embodiments of the present application, for the hierarchical knowledge embedding model, when the number of subgraph-level network layers w and the number of global graph-level network layers β are respectively set to {0, 1, 2, 3, 4} for experiments, the optimal w and β parameter combinations are determined through experiments. On the ICEWS14, ICEWS18, WIKI, and YAGO data sets, they are (2, 3), (2, 2), (2, 2), and (2, 3) respectively. In this article, these are used as the default settings for the parameters w and β. The mean reciprocal rank MRR is used as the evaluation index. Ablation experiments are conducted on the hierarchical knowledge embedding model and the reinforcement learning inference model on the ICEWS14, ICEWS18, WIKI, and YAGO data sets. The number of subgraph-level network layers w and the number of global graph-level network layers β of the hierarchical knowledge embedding model are respectively set to {0, 1, 2, 3, 4} for experiments, and the optimal w and β values are determined through experiments. Figure 3 shows the change in the MRR of THKERL when β is fixed at 3 and w takes values of {0, 1, 2, 3, 4}; while Figure 4 shows the influence of adjusting the value of β on the model performance when w is fixed at 2. Using reinforcement learning technology, the temporal knowledge graph reasoning is defined as a Markov decision process, and knowledge reasoning is performed based on the trained policy network. At the same time, the interpretability of the reasoning is enhanced by visualizing the reasoning path.
[0053] This paper proposes a temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning. By introducing a hierarchical knowledge embedding model, semantic dependencies and temporal evolution information are fully captured through two levels of knowledge embedding to obtain a more accurate representation of the latent features of the knowledge graph. In the reinforcement learning reasoning model, an action scoring function is designed, and a path-based soft reward is introduced to achieve reward shaping, which alleviates the problem of sparse rewards to a certain extent and further improves the reasoning effect. Through entity prediction tasks and comparative analysis of the results, the effectiveness of the method is demonstrated. Through ablation experiments, it can be seen that fully capturing the semantic dependencies and temporal evolution information of entities and relationships to obtain a more accurate knowledge feature embedding representation is particularly important for improving the knowledge reasoning effect. In future research, further optimization of the candidate action space and scoring function in the reasoning path in reinforcement learning will be considered to improve the accuracy of action selection and achieve better knowledge reasoning effects.
[0054] Example 2
[0055] When training and optimizing in step S3, the reinforcement learning reasoning model fixes the search path length to L and generates a trajectory of length L from the policy network π θ :{a 1 ,a 2 ,…,a L}. The policy network is trained by maximizing the expected reward over all training samples f train , and the objective function formula is Subsequently, the reinforcement learning reasoning model uses the policy gradient algorithm REINFORCE to optimize the policy. For all quadruples in f train , θ is updated using stochastic gradient, and the specific calculation formula is
[0056] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The units and algorithm steps described in the embodiments can be implemented by electronic hardware or a combination of computer software and electronic hardware.
[0057] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0058] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
[0059] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning, characterized by: The following steps are involved: Step S1, establishing a hierarchical knowledge embedding model, which obtains knowledge feature representation through embedding learning at two levels: subgraph level and global graph level; Step S2: Establish a reinforcement learning reasoning model, and introduce a weighted action scoring mechanism into the reinforcement learning reasoning model to design a strategy network; Step S3: training and optimizing the reinforcement learning inference model; Step S4: Combining the hierarchical knowledge embedding model with the reinforcement learning reasoning model to conduct a temporal knowledge graph reasoning experiment; Step S5: Analyze the experimental results and conduct ablation experiments, perform reasoning interpretability analysis, and complete temporal knowledge graph reasoning based on hierarchical knowledge embedding and reinforcement learning.
2. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized by: In step S1, the subgraph level is semantic dependency knowledge embedding, and the model uses the relational graph neural network as a semantic aggregator to obtain the subgraph The dynamic semantic embedding of each entity in is obtained by In the formula Indicates t i Moment Knowledge Graph The neighbor nodes of e in is the set of neighbor nodes of e, σ(·) is the activation function, W1 l and is a trainable weight parameter matrix, and the initial entity feature embedding h e and Set as static feature X e and After w layers of convolution, we can obtain entity e at time t i The feature representation that always takes into account the semantic dependencies between its neighbors For t i The relationship between entities at a certain moment, r, Its eigenvector is It means that by aggregating t i The calculation is based on the characteristics of the entity that has relationship r at the moment. The calculation formula is: In the formula Indicates t i The entity set related to relation r at time t. Meanpooling(·) is the average pooling operation, which acts on t i The set of entities related to relation r at a given moment 3. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized in that: In step S1, the global graph level is time-dependent knowledge embedding. First, the knowledge graphs at different times are connected through common entities to build a global graph. Then, the time interval is considered to connect the adjacent nodes in the global graph. and The influence of the relationship between adjacent nodes and The edges between them are regarded as time-related relationships of the same entity, and r τ It means that every relationship r τ The feature embedding of is expressed by the feature embedding formula, which is: Where |t i -t j | represents the absolute value of the time interval, Represents the relationship τ is the static feature embedding of , and φ(·) is the temporal encoding function.
4. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized in that: In step S1, the hierarchical knowledge embedding model uses an attention mechanism to With neighbor nodes Correlation weight coefficient Calculation, weight coefficient The calculation formula is Where each initial input entity feature embedding representation It is the entity feature output at the sub-graph level yes In P t The set of neighbors in R 3d and is a learnable weight parameter, σ(·) is an activation function, ·T represents transposition, and || is a concatenation operation.
5. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized in that: After the hierarchical knowledge embedding model adopts the attention mechanism, it obtains the feature embedding of entities in the global graph by adaptively aggregating features from all neighbors. The calculation formula for obtaining the feature embedding is: Where σ(·) is the activation function, and It is the weight parameter matrix for aggregation and self-loop. After the global graph-level β layer operation, the feature embedding representation of the entity can be obtained. Hierarchical knowledge embedding model usage To represent the feature embedding representation of entity e output at the global graph level.
6. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized by: In step S2, the agent is trained to continuously interact with the environment through the four-tuple of the Malv decision process, and a reinforcement learning reasoning model is established. The four-tuple of the Malv decision process is state space, action space, state transition and reward. The reinforcement learning reasoning model uses an action scoring function to score each candidate action and calculate the probability of state transition. The reinforcement learning reasoning model first uses two multi-layer perceptrons to encode the state information containing the query question and outputs the entity feature embedding representation of the predicted action. and the feature embedding representation of outward edges Entity Feature Embedding Representation The calculation formula is: Feature Embedding Representation The calculation formula is 7. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized by: In step S2, the policy network calculates the entity similarity and relationship similarity between the predicted action and the candidate action, and uses the scoring function φ(a n ,s l ) to obtain the final candidate action score, the scoring function φ(a n, s l ) is calculated as Where W6, W e , W r and W β is a learnable matrix, || is a concatenation operation, σ(·) is an activation function, <·> represents a similarity calculation operation, are the candidate action entities e calculated by the hierarchical knowledge embedding model respectively. n and the relationship n The feature embedding representation of are respectively calculated by the hierarchical knowledge embedding model at t l Moment l 、r l The feature embedding representation of are the query question entities e calculated by the hierarchical knowledge embedding model respectively. q and relationship q The feature embedding representation of , after scoring all candidate actions in the candidate action set, the reinforcement learning reasoning model obtains the policy network through the softmax function.
8. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized by: In step S4, the experiment uses four public TKG datasets: ICEWS14, ICEWS18, WIKI and YAGO to evaluate the temporal knowledge graph reasoning performance, and adopts widely used evaluation indicators to evaluate the temporal knowledge graph performance. The evaluation indicators include average reciprocal ranking and Hits@k, and the average reciprocal ranking and Hits@k are for each quadruple q=(e s ,r,e o ,t), evaluate two query questions: q o =(e s ,r,?,t) and q s =(?,r,e o ,t), the calculation formula of the average reciprocal ranking is: Where f test is the set of four-tuples in the test set, |f test | is the number of quadruplets in the test set, rank(e o |q o ) indicates the ranking of the predicted tail entity, rank(e s |q s ) represents the ranking of the predicted head entity. The larger the MRR value, the better the model performance. Hits@k refers to the percentage of the number of times the target entity appears in the first k candidate entities after ranking. The values of k are 1, 3, and 10.
9. The temporal knowledge graph reasoning method based on hierarchical knowledge embedding and reinforcement learning according to claim 1 is characterized by: In step S4, parameter setting is required before conducting the experiment. When setting the parameters, the dimension d of the entity and relationship feature embedding vector is set to 100. During the reasoning experiment, the reinforcement learning reasoning model selects the latest N outgoing edges as candidate actions at each step, where the value of N for the ICEWS14 and ICEWS18 data sets is 50, the value for the WIKI data set is 60, and the value for the YAGO data set is 30. The length L of the reasoning path is set to 3, the discount factor γ of the REINFORCE algorithm is set to 0.95, the batch size is set to 512 during training, the Adam optimizer is used to optimize the parameters, the learning rate is set to 0.001, and in the hierarchical knowledge embedding model module, the number of network layers w at the subgraph level and the number of network layers β at the global graph level are set to {0, 1, 2, 3, 4} and the experiment is conducted.