Knowledge graph completion method based on reinforcement learning
By obtaining the adjacency subgraph features and semantic correlations of candidate actions, and optimizing the search path decision of reinforcement learning, the problem of lack of foresight and semantic supervision in the existing methods is solved, and efficient and credible knowledge graph completion is achieved.
Patent Information
- Application Number
- CN202510626965.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-12
AI Technical Summary
The existing knowledge graph completion method based on reinforcement learning lacks predictability for future candidate nodes, which leads to relatively short-sighted in the inference model, prone to local optimal traps, and insufficient semantic correlation, affecting the quality of path generation.
By obtaining the adjacency subgraph features and semantic correlation of candidate actions, calculating the transfer probability of candidate actions, combining Markov decision-making process and reinforcement learning, optimizing search path decisions, introducing semantic supervision mechanisms, and improving the structural perception and semantic correlation of paths.
It improves the inference efficiency of knowledge graph completion and the self-consistent path, enhances the semantic correlation of generated paths, and improves the credibility of knowledge graph completion.
Smart Images

Figure CN120471158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of reinforcement learning and knowledge graph technology, and in particular to a knowledge graph completion method based on reinforcement learning. Background Art
[0002] Knowledge graphs aim to provide a good data organization form for the storage, management, query and retrieval of massive amounts of information data on the Internet. Although the number and scale of current knowledge graphs are constantly growing, such as DBpedia, Freebase, NELL, YAGO, etc., they are often incomplete due to their own knowledge capacity. Therefore, it is of great significance to use appropriate automated knowledge graph completion algorithms to mine potential ternary relationships in knowledge graphs. Among them, link prediction for entities and relationships is one of the core technologies of knowledge graph completion, which aims to predict the relationship between entities and relationships through a given query. , that is, entity Relationship Yes, combine the existing knowledge in the knowledge graph to predict possible tail entities , in order to realize the potential ternary relationship in the knowledge graph of excavation.
[0003] In recent years, knowledge graph completion methods have made great progress. In particular, knowledge reasoning methods based on multi-hop paths of reinforcement learning have achieved good reasoning performance due to their advantages of both performance and interpretability. Typical methods of this type include: Reference [1] defines this link prediction task as a finite-horizon deterministic partial observation Markov decision process, and trains an intelligent agent to continuously search on the knowledge graph by comparing the query and the strategy until it reaches the possible tail entity. Reference [2] separates the walking agent into a relationship agent and an entity agent, which are responsible for relationship selection and entity selection respectively. Through agent separation and space pruning strategies, the search space is significantly reduced and the reasoning efficiency is improved. Reference [3] introduces the Monte Carlo tree search strategy in reinforcement learning, uses a deterministic state transition model to generate high-quality trajectories, and incorporates non-policy optimal paths into the learning through Q learning, so that the policy network can be indirectly improved through Q learning on non-policy trajectories, which to some extent alleviates the problem of reward sparsity. Reference [4] directly incorporates the consideration of path rationality into the reward function, penalizes false paths of poor quality, and introduces the idea of curriculum learning to dynamically balance the accuracy and rationality of the generated path, thereby improving the semantic relevance and self-consistency of the reasoning path to a certain extent.
[0004] However, existing knowledge graph completion methods based on reinforcement learning mainly focus on decision optimization on search path sequences. Although some literature [5] uses the attention mechanism to achieve the integration and perception of nodes around the path, it is still a re-modeling of historical paths and historical decisions, lacking the foresight of future candidate nodes, resulting in a relatively short-sighted decision perception ability of the reasoning model, which makes it easy to fall into unnecessary local optimal traps or wrong path backtracking problems during the reasoning process, damaging the model's reasoning efficiency and the quality of the generated path. In addition, the existing methods lack semantic supervision measures for the generated path, do not directly incorporate semantic relevance into the decision process, and do not have direct supervision of the path semantic quality, resulting in poor semantic relevance between the generated path and the query pair.
[0005] References:
[0006] [1]Das R, Dhuliawala S, Zaheer M, et al. Go for a Walk and Arrive at the Answer: Reasoning Over Paths in Knowledge Bases using ReinforcementLearning[C] / / International Conference on Learning Representations. 2018. DOI:10.48550 / arXiv.1711.05851.
[0007] [2]Lei D, Jiang G, Gu X, et al. Learning Collaborative Agents withRule Guidance for Knowledge Graph Reasoning[C] / / Proceedings of the 2020Conference on Empirical Methods in Natural Language Processing. 2020: 8541-8547. DOI: 10.18653 / v1 / 2020.emnlp-main.688.
[0008] [3]Shen Y, Chen J, Huang PS, et al. M-walk: Learning to walk overgraphs using monte carlo tree search[J]. Advances in Neural InformationProcessing Systems, 2018, 31. DOI: 10.48550 / arXiv.1802.04394. [4]Jiang C, ZhuT, Zhou H, et al. Path spuriousness-aware reinforcement learning for multi-hop knowledge graph reasoning[C] / / Proceedings of the 17th Conference of theEuropean Chapter of the Association for Computational Linguistics. 2023:3181-3192. DOI: 10.18653 / v1 / 2023.eacl-main.232.
[0009] [5]Wang H, Li S, Pan R, et al. Incorporating graph attentionmechanism into knowledge graph reasoning based on deep reinforcement learning[C] / / Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). 2019: 2623-2631. DOI: 10.18653 / v1 / D19-1264. Summary of the Invention
[0010] In order to solve the above problems, the present invention proposes a knowledge graph completion method based on reinforcement learning, which includes:
[0011] S1. Given an input query pair , the time step of the initial Markov decision process is , search path , and based on the knowledge graph Get Status , candidate action space ;
[0012] S2. Obtain the adjacency subgraph of the candidate action and the adjacency structure information features of the candidate action;
[0013] S3, calculating the semantic relevance between the candidate action and the query pair;
[0014] S4, the policy network calculates the transition probability of all candidate actions at the current time step based on the current state information and search path information, as well as the structural characteristics and semantic relevance of the candidate actions;
[0015] S5, the transfer function is based on the transfer probability from the candidate action space Select the dynamic path to join the search path, and then the time step Add 1, update the state, and update the candidate action space;
[0016] S6. Loop through S2-S5 until the maximum time step is reached, and output the entity at the end of the final search path as the prediction result, and add it to the knowledge graph.
[0017] Furthermore, reinforcement learning works using Markov decision processes.
[0018] Furthermore, S1 includes:
[0019] S11. Get the current time step The state below ;
[0020] S12. Get the candidate action space at the current time step and state ;
[0021] S13. Define transfer function , including transfer probability distribution and transfer strategy;
[0022] S14. Set the reward function for reinforcement learning training.
[0023] Furthermore, in S12, the candidate action space in the current state Include all nodes in the knowledge graph related to the current entity A collection of directly connected relationships and entities, as well as pause actions.
[0024] Furthermore, in S14, the reward function plays a feedback role in the training process and provides corresponding positive and negative feedback for the final state of the Markov decision process.
[0025] Furthermore, S2 includes:
[0026] S21. Extract the entity nodes of each candidate action in the knowledge graph The adjacent subgraph centered at ;
[0027] S22, use the aggregation function to encode the adjacent subgraph and extract features;
[0028] S23. Obtain the most significant feature in the adjacent subgraph.
[0029] Furthermore, in S22, the aggregation operation is instantiated as a graph neural network.
[0030] Furthermore, S3 includes:
[0031] S31. Pre-train the embedding representation model and reuse the relational embedding matrix ;
[0032] S32. Calculate the semantic relevance between the relationship in the query pair and the relationship in the candidate action according to the relationship embedding matrix.
[0033] Furthermore, S4 includes:
[0034] S41, the current time step The state below Convert them into embedding vectors respectively and concatenate them into vectors ;
[0035] S42, the search path under the current time step Use LSTM to encode the hidden state of the last time step of LSTM The output is a feature vector representation of the entire search path;
[0036] S43, receiving the features of the adjacent subgraph encoding of each candidate action in S2 ;
[0037] S44, receiving the semantic relevance score weight of each candidate action and query calculated in S3 ;
[0038] S45, the current state vector obtained Character vector with search path As a condition, the structural information, semantic information and representation vector of the candidate action space are matched and calculated, and the transition probability of all candidate actions at the current time step is calculated based on the correlation score between the two. The calculation method is as follows:
[0039]
[0040] in, is the probability normalization function, represents the Hadamard dot product, Represents the concatenation operation of the vector, ReLU is the activation function, 、 is a learnable parameter, is the semantic relevance weight between the candidate action space and the query pair, is the feature representation of the current state, is the feature representation of the search path, The transition probability distribution of the candidate action space is calculated based on information such as the current state, and the probability is then given to the transition function.
[0041] Furthermore, the parameters in S2-S4 need to be trained and learned through backpropagation. The training objectives are:
[0042]
[0043] The above formula represents the maximization Maximize the rewards obtained through reinforcement learning for the true expectation , and the search path Optimization.
[0044] The beneficial effects of the present invention are as follows:
[0045] (1) This paper proposes a novel forward-looking mechanism for the reinforcement learning knowledge reasoning method. By modeling and capturing the subgraphs around candidate nodes, it greatly improves the structural perception and foresight capabilities of the model, avoids the backtracking problem in the reasoning path, and improves the reasoning efficiency and self-consistency.
[0046] (2) This paper proposes a semantic matching paradigm, which incorporates the semantic prior between the query relation and the candidate relation into the decision-making considerations of the strategy module, emphasizes the direct modeling of the semantic quality of the path, and improves the semantic relevance and rationality of the generated path.
[0047] (3) The present invention achieves excellent performance on two standard datasets, WN18RR and FB15K-237, and improves the credibility of knowledge graph completion results. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0049] Figure 1 2. A flowchart of a method for completing a knowledge graph based on reinforcement learning according to an embodiment of the present invention;
[0050] Figure 2 A schematic diagram of a process for obtaining adjacency structure information features of candidate actions according to an embodiment of the present invention;
[0051] Figure 3 A schematic diagram of a process for calculating the semantic relevance between a candidate action and a query pair according to an embodiment of the present invention;
[0052] Figure 4 FIG. 4 is a flow chart of calculating transition probability according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be reviewed and fully described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] The present invention provides a knowledge graph completion method based on reinforcement learning, such as Figure 1 As shown, the method includes the following steps:
[0055] S1. Given an input query pair , the time step of the initial Markov decision process is , search path , and based on the knowledge graph Get Status , candidate action space ;
[0056] S2. Obtain the adjacency subgraph of the candidate action and the adjacency structure information features of the candidate action;
[0057] S3, calculating the semantic relevance between the candidate action and the query pair;
[0058] S4, the policy network calculates the transition probability of all candidate actions at the current time step based on the current state information and search path information, as well as the structural characteristics and semantic relevance of the candidate actions;
[0059] S5, the transfer function is based on the transfer probability from the candidate action space Select an action Add to search path , and then the time step Add 1 and update the status to , update the candidate action space to ;
[0060] S6. Loop through S2-S5 until the maximum time step is reached, and output the entity at the end of the final search path as the prediction result.
[0061] In step S1, reinforcement learning works using a Markov decision process (MDP), which consists of the following components: state ,action , transfer and rewards Given an input query pair , reinforcement learning to query the entity nodes Start the policy search as the starting point and initialize the time step to , search path , the specific steps include:
[0062] S11. Get the current time step The state below ;
[0063] S12. Get the candidate action space at the current time step and state ;
[0064] S13. Define transfer function , including transfer probability distribution and transfer strategy;
[0065] S14. Set the reward function for reinforcement learning training.
[0066] In step S11, initially, the time step , the initial state contains the entity of the input query pair , input query pair relationship And the entity node where the initial time step is located , together constitute .
[0067] In step S12, the candidate action space in the current state Include all nodes in the knowledge graph related to the current entity A collection of directly connected relationships and entities, namely: , and additionally includes a special Pause Action , together constitute the optional transfer action space of the current time step .
[0068] In step S13, the transfer function It means that the transition probability distribution corresponding to each candidate action is obtained according to the current state and action space, and the transition strategy adopted is specified. The transition strategy can be to randomly select the candidate action space according to the transition probability distribution. An action in Select Add to search path and update the path to .
[0069] In step S14, the reward function only plays a feedback role in the training process and provides corresponding positive and negative feedback for the final state of the Markov decision process. Specifically, if the search path can reach the correct tail entity corresponding to the query pair , the reward function gives a reward value of The correct answer is the feedback, otherwise according to the entity that finally arrives With query The relevance of The reward value between , the reward function is defined as follows:
[0070]
[0071] In the above formula, For the scoring function part of models such as TransE, ComplEx, ConvE, etc., it is used to give reinforcement learning below weak feedback signal.
[0072] In step S2, the adjacent subgraphs around the candidate action are encoded and extracted to provide structural information features for the reinforcement learning policy network (the policy network is the network that comes with the Markov decision process and is an existing technology) to enhance the structural perception ability of the Markov decision process, provide a larger perception field around the candidate action, and avoid falling into the local optimal trap or wrong path backtracking during the search process. Figure 2 As shown, specifically including:
[0073] S21. Extract the entity nodes of each candidate action in the knowledge graph The adjacent subgraph centered at ;
[0074] S22, use the aggregation function to encode the adjacent subgraph and extract features;
[0075] S23. Obtain the most significant feature in the adjacent subgraph.
[0076] In step S21, a candidate action is selected from the knowledge graph. Entities in The adjacent subgraph composed of one-hop adjacent edges and nodes as the center:
[0077]
[0078] in, Represents the set of triples in the knowledge graph, Represents the entity node part in the candidate action, Representation and entity nodes Adjacent edges and nodes.
[0079] In step S22, the subgraph is aggregated using the aggregation function Encode and extract the most significant features:
[0080]
[0081] In the above formula, Represents an aggregation operation and can be instantiated as graph neural networks such as graph attention network (GAT), graph convolutional network (GCN), and deep Bellman-Ford network (NBF).
[0082] If GAT is used as the aggregation function, the aggregation calculation can be expressed as:
[0083]
[0084]
[0085]
[0086] in, Represents a nonlinear activation function, which can be instantiated as LeakyReLU. Represents the vector concatenation operation, 、 、 are all learnable parameters, is the entity or relationship feature vector of the aggregated subgraph.
[0087] If GCN is used as the aggregation function, the aggregation calculation can be expressed as:
[0088]
[0089] In the above formula, represents a nonlinear activation function, are learnable parameters, For loop-related operations on vectors, is the entity or relationship feature vector of the aggregated subgraph.
[0090] If NBF is used as the aggregation function, the aggregation calculation can be expressed as:
[0091]
[0092] In the above formula, Represents the aggregation operation of the graph network, which can be instantiated as operations such as summation, averaging, pooling, and maximum value. Represents the message function of the graph network, which can be instantiated as the scoring function part of TransE, ComplEx, ConvE, etc. For the relationship between The associated learnable vector, is the initial representation of the central entity, is the entity or relationship feature vector of the aggregated subgraph.
[0093] In step S23, a single-layer feedforward neural network (MLP) is used to extract the most significant features of the adjacent subgraph and the candidate action in the center to efficiently express the node and structure information of the adjacent subgraph. The feature representation is as follows:
[0094]
[0095] The parameters of the network and MLP in the above aggregation function can be adjusted in reverse during the training process.
[0096] In step S3, the semantic prior knowledge between knowledge graph relations is introduced to calculate the semantic relevance between query pairs and candidate actions, providing semantic information factors for the reinforcement learning policy network and realizing direct supervision of the semantic quality of the search path, such as Figure 3 As shown, specifically including:
[0097] S31. Pre-training models such as TransE, ComplEx, or ConvE for embedding representation and reusing the relational embedding matrix ;
[0098] S32. Calculate the semantic relevance between the relationship in the query pair and the relationship in the candidate action according to the relationship embedding matrix.
[0099] In step S31, all triples in the knowledge graph are used as training samples to pre-train the knowledge graph model for embedding representation and extract the relation embedding matrix. It is retained as a feature matrix rich in semantic prior knowledge for reuse. The parameters in the knowledge graph embedding representation model can be reversely adjusted during the training process.
[0100] In step S32, from the relation embedding matrix Extract the feature vector rows corresponding to the relationship between the query pair and the relationship between the candidate actions, and then calculate the cosine similarity distance between the two vectors as the semantic relevance score between the two:
[0101]
[0102] In the above formula, is a nonlinear activation function, Indicates the relationship between Corresponding rows and relations The corresponding column value, The weight of the relevance score between the query pair's relationship and the candidate action's relationship.
[0103] In step S4, Figure 4 As shown, the policy network decision module of reinforcement learning comprehensively considers the current time step The state, search path, candidate action space, structural information factors and semantic information factors under the state are used to calculate the corresponding transition probability for all actions in the candidate action space. The specific method includes:
[0104] S41, the current time step The state below Convert them into embedding vectors respectively and concatenate them into vectors ;
[0105] S42, the search path path under the current time step Use LSTM to encode the hidden state of the last time step of LSTM The output is a feature vector representation of the entire search path;
[0106] S43, receiving the features of the adjacent subgraph encoding of each candidate action in S2 ;
[0107] S44, receiving the semantic relevance score weight of each candidate action and query calculated in S3 ;
[0108] S45, the current state vector obtained Character vector with search path As a condition, the structural information, semantic information and representation vector of the candidate action space are matched and calculated, and the transition probability of all candidate actions at the current time step is calculated based on the correlation score between the two. The calculation method is as follows:
[0109]
[0110] in, is the probability normalization function, represents the Hadamard dot product, Represents the concatenation operation of the vector, ReLU is the activation function, 、 are all learnable parameters, is the semantic relevance weight between the candidate action space and the query pair, is the feature representation of the current state, is the feature representation of the search path, It is the transition probability distribution of the candidate action space calculated based on the current state and other information, and then the probability is given to the transition function.
[0111] In step S5, the transfer function is based on the transfer probability distribution and the random sampling transfer strategy from the candidate action space Select an action Add to the search path, that is: , and then update the time step to , update status to , update the candidate action space to , thus completing a search. For example: given the query pair "(Diamond Head Mountain, located)", at the initial time step, the corresponding state , search path ,and is the executable action space in this state.
[0112] In step S6, loop steps S2-S5 until the maximum time step is reached, and then obtain the entity at the end of the final search path , and with the query pair Form a new triple Output. For example: given the query pair "(Diamond Head Mountain, located)", the reasoning process starts from the entity node "Diamond Head Mountain" of the query pair, and then and the candidate action space Calculate the transition probability of all candidate actions and then select the action according to the probability { Add to the search path, that is: The above process is then repeated, and the final search path is "Diamond Head, located in, Oahu, located in Honolulu, located in Hawaii", where "Hawaii" is the output target entity, thus generating a new triple "(Diamond Head, located in, Hawaii)", and there is a search path as the inference evidence, which is then added to the knowledge graph.
[0113] During reinforcement learning training, the parameters of the policy network and aggregation function in steps S2-S4 need to be trained and learned through backpropagation. The training objectives are as follows:
[0114]
[0115] The above formula represents the maximization Maximize the rewards obtained through reinforcement learning for the true expectation , and the search path Finally, the output target entity With query Composed of The newly discovered triples are inserted into the original knowledge graph, and the search path is used as evidence of the authenticity of the triples, thus completing the process of interpretability of the knowledge graph.
[0116] Experimental verification
[0117] In one embodiment, the datasets used are the WN18RR dataset and the FB15K-237 dataset. The WN18RR dataset is a more challenging dataset obtained by removing multiple test leak source relationships and triples from the original WN18 dataset. The FB15K-237 dataset is a more realistic dataset obtained by removing additional interdependent relationships from the FB15K dataset.
[0118] The present invention effectively enhances the performance of reinforcement learning knowledge reasoning on the WN18RR dataset and the FB15K-237 dataset. Specifically, the present invention uses reinforcement learning for experiments.
[0119] This paper experimentally validates the proposed method on two standard datasets, using MINERVA, M-walk, Multihop-KG, RuleGuider, RARL, PAAR, HiAM, RL-MHR, CURL, PSRL, and RKLE as baseline models for comparison on four metrics: MRR, Hits@1, Hits@3, and Hits@10. The experimental results are shown in Tables 1 and 2, respectively. The best result is in bold, and the second-best result is underlined.
[0120] Table 1 Results on the WN18RR dataset
[0121]
[0122] Table 2 Results on the FB15k-237 dataset
[0123]
[0124] Those skilled in the art will understand that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A knowledge graph completion method based on reinforcement learning, characterized in that: The method comprises: S1. Given an input query pair , the time step of the initial Markov decision process is , search path , and based on the knowledge graph Get Status , candidate action space ; S2. Obtain the adjacency subgraph of the candidate action and the adjacency structure information features of the candidate action; S3, calculating the semantic relevance between the candidate action and the query pair; S4, the policy network calculates the transition probability of all candidate actions at the current time step based on the current state information and search path information, as well as the structural characteristics and semantic relevance of the candidate actions; S5, the transfer function is based on the transfer probability from the candidate action space Select the dynamic path to join the search path, and then the time step Add 1, update the state, and update the candidate action space; S6. Loop through S2-S5 until the maximum time step is reached, and output the entity at the end of the final search path as the prediction result, and add it to the knowledge graph.
2. The knowledge graph completion method according to claim 1, characterized in that: Reinforcement learning works using Markov decision processes.
3. The knowledge graph completion method according to claim 2, characterized in that: S1 includes: S11. Get the current time step The state below ; S12. Get the candidate action space at the current time step and state ; S13. Define transfer function , including transfer probability distribution and transfer strategy; S14. Set the reward function for reinforcement learning training.
4. The knowledge graph completion method according to claim 3, characterized in that: In S12, the candidate action space in the current state Include all nodes in the knowledge graph related to the current entity A collection of directly connected relationships and entities, as well as pause actions.
5. The knowledge graph completion method according to claim 3, characterized in that: In S14, the reward function plays a feedback role in the training process and provides corresponding positive and negative feedback for the final state of the Markov decision process.
6. The knowledge graph completion method according to claim 2, characterized in that S2 include: S21. Extract the entity nodes of each candidate action in the knowledge graph The adjacent subgraph centered at ; S22, use the aggregation function to encode the adjacent subgraph and extract features; S23. Obtain the most significant feature in the adjacent subgraph.
7. The knowledge graph completion method according to claim 6, characterized in that: In S22, the aggregation operation is instantiated as a graph neural network.
8. The knowledge graph completion method according to claim 2, characterized in that S3 include: S31. Pre-train the embedding representation model and reuse the relational embedding matrix ; S32. Calculate the semantic relevance between the relationship in the query pair and the relationship in the candidate action according to the relationship embedding matrix.
9. The knowledge graph completion method according to claim 2, characterized in that S4 include: S41, the current time step The state below Convert them into embedding vectors respectively and concatenate them into vectors ; S42, the search path under the current time step Use LSTM to encode the hidden state of the last time step of LSTM The output is a feature vector representation of the entire search path; S43, receiving the features of the adjacent subgraph encoding of each candidate action in S2 ; S44, receiving the semantic relevance score weight of each candidate action and query calculated in S3 ; S45, the current state vector obtained Character vector with search path As a condition, the structural information, semantic information and representation vector of the candidate action space are matched and calculated, and the transition probability of all candidate actions at the current time step is calculated based on the correlation score between the two. The calculation method is as follows: ,in, is the probability normalization function, represents the Hadamard dot product, Represents the concatenation operation of the vector, ReLU is the activation function, 、 is a learnable parameter, is the semantic relevance weight between the candidate action space and the query pair, is the feature representation of the current state, is the feature representation of the search path, The transition probability distribution of the candidate action space is calculated based on information such as the current state, and the probability is then given to the transition function.
10. The knowledge graph completion method according to claim 2, characterized in that: The parameters in S2-S4 need to be trained and learned through back propagation. The training objectives are: , the above formula represents the maximization Maximize the rewards obtained through reinforcement learning for the true expectation , and the search path Optimization.