Knowledge graph reasoning method based on prior knowledge enhancement

By introducing the prior knowledge enhancement of large language models into the knowledge graph reasoning method based on reinforcement learning, the problem of inaccurate rewards caused by insufficient prior knowledge of the knowledge graph is solved, the performance and generalization ability of the model are improved, and the training efficiency is improved.

CN120671820APending Publication Date: 2025-09-19BEIJING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510735988.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the knowledge graph reasoning method based on reinforcement learning, the lack of prior knowledge of the knowledge graph leads to inaccurate logical rationality rewards, which affects the generalization and performance of the model.

Method used

This approach uses a priori knowledge enhancement method based on a large language model to improve the accuracy of rewards by incorporating the internal knowledge of the LLM into the calculation of logical rationality rewards. This method identifies path importance and only uses the LLM to calculate rewards for important paths, thereby reducing the number of LLM calls and improving efficiency.

Benefits of technology

It effectively alleviates the problem of inaccurate rewards caused by insufficient prior knowledge of the knowledge graph, improves the performance and generalization ability of the model, reduces the number of calls to the large language model, and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671820A_ABST
    Figure CN120671820A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph inference method based on prior knowledge enhancement, and mainly solves the problem of inaccurate reward caused by insufficient knowledge graph prior knowledge in the existing knowledge graph inference method based on reinforcement learning, thereby optimizing a model more effectively. The method comprises the following steps: reasoning path searching based on reinforcement learning; calculating an answer correctness reward; logic rationality reward calculation; a reward enhancement strategy based on path importance; performing logic rationality reward enhancement based on large language model context learning; and optimizing the model. According to the method, on the basis of an existing knowledge graph reasoning method based on reinforcement learning, huge internal knowledge of a large language model is fused into reward calculation in an efficient and information loss resistant mode, and the problem that rewards are inaccurate due to insufficient prior knowledge of a knowledge graph is solved. The performance of the knowledge graph reasoning method based on priori knowledge enhancement is remarkably improved compared with that of an existing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a knowledge graph reasoning method based on prior knowledge enhancement, which belongs to knowledge graph reasoning. Background Art

[0002] In the era of artificial intelligence and big data, knowledge graphs (KGs), as a structured knowledge representation, can efficiently organize and manage massive amounts of data, providing powerful knowledge support for applications such as information retrieval, intelligent question answering, and data traceability. In a KG, entities are represented as nodes, and relationships between entities are represented as edges. A triple consisting of two connected entities and their relationship is a fact. Together, these facts form a semantically structured knowledge network. Based on explicit facts in the KG, knowledge graph reasoning (KGR) can mine latent facts from the KG, empowering AI systems with powerful reasoning capabilities and further enhancing their intelligence. For example, in a data traceability scenario, based on the KG triples <device B, download, important file A> and <device B, abnormal data transmission, device C>, it can be assumed that <device C, leak, important file A> has a high probability of occurring. Among the many KGR methods, those based on reinforcement learning (RL) are highly favored due to their excellent interpretability. Based on the RL framework, this approach models KGR as a path search process on a knowledge graph, generating reasoning paths that can provide reasonable explanations for prediction results. This approach not only achieves good reasoning accuracy but also provides a clear logical basis for the reasoning results, further promoting the development of KGR technology.

[0003] In the RL-based KGR method, the model searches for a path starting from the source entity based on a given query, and the end point of the path is the predicted target entity. During the training process, after the model searches for the reasoning path, it learns an effective path search strategy under the guidance of rewards that can reflect the effectiveness of the reasoning path. Due to the lack of path labels, methods such as MINERVA and MultiHop only reward the model based on the correctness of the target entity to which the reasoning path leads. However, this will result in the reasoning path that accidentally leads to the correct target entity also bringing higher rewards, thereby damaging the generalization of the model. For example, in Figure 1In the query <Tom, nationality,?>, the model may search for path 3. The meaning of this path, "two people with the same occupation also have the same nationality", is illogical, but it will bring a very high reward because it accidentally leads to the correct target entity "United States". Guided by this reward, the model learns the strategy of "inferring nationality based on occupation", which is difficult to be used to infer the correct target entity of other queries. To alleviate this problem, PSRL, RuleGuider, etc. use heuristic methods based on KG prior knowledge to calculate the logical rationality reward of the reasoning path. For example, in order to calculate Figure 1 To calculate the logical rationality reward for path 3, we first extract all paths from the KG that contain the meaning "two people with the same occupation also have the same nationality," namely paths 3 and 4. The correct proportion of target entities led by these paths is then used as the logical rationality reward for path 3, which is 0.5. However, insufficient prior knowledge in the KG may affect the accuracy of this reward. For example, path 2 states that "two people with the same birthplace have the same nationality," which is obviously more reasonable than path 3, but the calculated logical rationality reward is also 0.5. The fact that there are too few paths in the KG that have the same meaning as paths 1, 2, 3, and 4 leads to statistical deviations, which in turn affects the accuracy of the reward.

[0004] Large Language Model (LLM) is an emerging natural language processing technology that, after being trained on massive amounts of text data, can understand and generate natural language. It is therefore widely used in a variety of language processing tasks, including question-answering, translation, and writing. One of the core features of LLM is its vast internal knowledge base, which covers multiple fields such as science, history, and culture. This knowledge comes from a wealth of training data, including publicly available text resources on the internet, such as encyclopedias, news articles, academic papers, books, code libraries, and social media content. The knowledge contained in this data is encoded as model parameters during the training process and is flexibly called upon when the model is inferring and generating, enabling LLM to demonstrate powerful language understanding and generation capabilities. Summary of the Invention

[0005] This paper addresses the issue of inaccurate logical plausibility rewards in RL-based KGR methods, often caused by insufficient prior knowledge in the KG. Specifically, it proposes a KGR method based on enhanced prior knowledge. This method incorporates the extensive internal knowledge of the LLM into the calculation of the logical plausibility reward in an efficient and information-loss-resistant manner, alleviating the inaccurate reward problem caused by insufficient prior knowledge. This reward can more effectively guide model optimization and further improve its performance.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is a KGR method based on LLM to enhance prior knowledge (PKA-KGR). Given a KG dataset, represented as G = { <e head ,r,e tail >|e head ∈E,r∈R,e tail ∈E}, where E and R represent entity set and relationship set respectively, e head and e tail They represent the head and tail entities of the triple, and r is the relationship connecting the two entities. PKA-KGR can s ∈E and query relation r q ∈R, search for a T-hop reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, thereby predicting the target entity e T , it can be used with e s 、r q Forming potential triplets <e s ,r q ,e T >. Among them, <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >∈G. The training process of PKA-KGR on the KG dataset G is as follows Figure 2 The specific steps are as follows:

[0007] Step (1) Reasoning path search based on reinforcement learning; for each training sample in the KG dataset G <e s ,r q ,e o >,e s 、r q and e o Represent the source entity, query relation and correct target entity respectively. The PKA-KGR model adopts the reinforcement learning framework and introduces an intelligent agent as the path search subject. The intelligent agent takes the source entity e s As the starting point, new actions (composed of relations and entities) are selected at each time step, and finally the reasoning path is expanded after T time steps. <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >. Among them, r t (1≤t≤T) and e t(1≤t≤T) represent the relationship and entity selected at time step t, <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >∈G. Specifically, at time step t, the agent is located at entity e t-1 , its action space Actions t By e t-1 Adjacent relationships and entities, i.e. Actions t ={a t |a t =(r′,e′), <e t-1 ,r′,e′>∈G}. The agent selects action a according to the following strategy network t : a t ~π θ (a t )=σ(A t ×W2ReLU(W1[V(e t-1 );h t ;U(r q )])) Among them, π θ (a t ) represents the agent’s policy network, which is used to select action a t ; σ and ReLU represent the Softmax function and the linear rectification function respectively; W1 and W2 represent the fully connected neural network; V and U represent the entity embedding layer and the relationship embedding layer respectively, V(e t-1 ) and U(r q ) represent e t-1 and r q Embedding vector of A t Actions t Each action a t The embedding vector x t The stacked matrix, and x t Composed of a t The embedding vectors U(r′) and V(e′) of the relation r′ and entity e′ are concatenated, that is, x t =[U(r′);V(e′)];h t It is the historical information embedding, which is obtained by encoding the historical actions through a long short-term memory network, i.e. h t =LSTM(h t-1 ,x t-1 ), where x t-1 is the embedding vector of the action selected at time step t-1.

[0008] Step (2) Calculation of reward for correct answer: After searching for the reasoning path, the effectiveness of the reasoning path needs to be evaluated, and the evaluation result will be used as a reward for optimizing the model. The reward based on the correct answer can evaluate the effectiveness of the path search strategy from the perspective of whether the reasoning result is correct or not. Specifically, if the predicted target entity e T Equal to the correct target entity e o , then the correct answer reward R a is equal to 1. Otherwise, we will be represented by the source entity e s , query relation r q and the predicted target entity e T The triplet <e s ,r q ,e T > Input the pre-trained KG embedding model ConvE. ConvE will output the authenticity score of the triple, which reflects the probability that the triple really exists, so it is used as the answer correctness reward R a .

[0009] Step (3) Calculation of logical rationality reward: In addition to the reward based on the correctness of the answer, it is also necessary to evaluate the effectiveness of the reasoning path from the perspective of logical rationality and give the model a logical rationality reward. Usually, a heuristic method based on KG prior knowledge is used to calculate the logical rationality reward. Specifically, for the searched reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, where <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >Belongs to the KG dataset G, and uses the following formula to calculate the logical rationality reward R of the path l : in, <e s ′,r1,e1′>∧…∧ <e T-1 ′,r T ,e T ′> represents any KG dataset G that satisfies the condition And contains relations r1,r2,…,r T The path, <e s ′,r1,e1′>,…, <e T-1 ′,r T ,e T ′>∈G,r q It is a query relationship; <e s″,r1,e1″>∧…∧ <e T-1 ″,r T ,e T "> means any KG dataset G that meets the conditions And contains relations r1,r2,…,r T The path, <e s ″,r1,e1″>,…, <e T-1 ″,r T ,e T ″>∈G,r q is the query relation. Therefore, R l In fact, the path calculated by using the empirical samples from the KG prior knowledge can infer the empirical probability of the correct target entity, which reflects the logical rationality of the reasoning path. In addition, it is difficult to perform path statistics on the entire KG, so we directly count S from the paths searched in the previous training iteration. all (r q ,r1,r2,…,r T ) and S correct (r q ,r1,r2,…,r T ).

[0010] Step (4) Reward enhancement strategy based on path importance; reward R for logical rationality l is the empirical probability, and the accuracy of the empirical probability depends heavily on the empirical samples from the KG prior knowledge. So when the prior knowledge in KG is insufficient, R l is inaccurate. PKA-KGR uses the vast internal knowledge of LLM to enhance the logical rationality reward, but frequent reward calculations will lead to a large number of LLM calls, which will seriously increase the training time overhead. In order to reduce the number of LLM calls while ensuring model performance, an effective strategy is to prioritize the use of LLM to enhance the rewards of paths that have a significant impact on model training, while the remaining paths still use rewards based on KG prior knowledge that do not require LLM calls. Specifically, for the searched reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, where <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T > belongs to the KG dataset G. If the corresponding |S all (r q ,r1,r2,…,r T )| is less than the threshold α and R lIf the value of α and β is greater than or equal to the threshold value β, the inference path is considered an important path. α and β are hyperparameters, and their optimal values ​​are determined by grid search. all (r q ,r1,r2,…,r T )| is the number of experience samples used to calculate the experience probability. The smaller the number of experience samples, the less accurate the logical rationality reward based on the experience probability. l The larger the value, the stronger the reward feedback to the model. Therefore, the path determined based on the above conditions will bring inaccurate and strong feedback to the model, which will have a significant impact on model training.

[0011] Step (5) Enhancement of logical rationality reward based on LLM context learning; LLM’s vast internal knowledge can effectively enhance the calculation of logical rationality reward. For example Figure 1 The logical rationality rewards of paths 2 and 3 in the figure are both equal to 0.5, which is obviously inaccurate. However, based on the common sense that "occupation and nationality are not directly related in most cases" and "place of birth and nationality are the same in most cases", LLM can determine that path 3 in the figure should have a lower logical rationality reward. Therefore, for the important paths determined in the previous step, LLM is used to enhance their logical rationality rewards. The steps of this process for LLM are as follows: Figure 3 As shown. First, because the unclear semantics of structured data such as triples will affect the semantic understanding of LLM, resulting in the loss of semantic information of triples. Therefore, in order to alleviate this problem, it is necessary to use context learning to convert the triples in the logical rationality reward calculation into natural language text with clearer semantics. Specifically, for triples <e s ,r1,e1>, sample 20 triples with relation r1 from KG dataset G, and ask LLM to summarize the meaning of triples with relation r1. This can guide LLM to use its context learning ability to identify the pattern of triples with relation r1, which contains the complete semantics of triples. Then, ask LLM to summarize the triples according to the summarized meaning. <e s ,r1,e1> converted into natural language text Text(e s ,r1,e1). When the triple pattern is known, LLM can fully express the semantics of the triple in the form of natural language text, thereby enhancing its ability to capture the semantic information of the triple. Through the above process, the reasoning path can be <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >The triples in the text are converted into natural language text (e s ,r1,e1),…,Text(eT-1 ,r T ,e T ), and will be queried by <e s ,r q ,? > and the predicted target entity e T The triplet <e s ,r q ,e T 〉Convert to natural language text Text(e s ,r q ,e T ). The former is regarded as the "explanation" and the latter as the "conclusion". LLM is required to give a score between 0 and 10 according to the rationality of the "explanation". Finally, the score is extracted from the LLM's answer through regular expression and divided by 10 to obtain the enhanced logical rationality reward R l ′.

[0012] Step (6) Model optimization; get the answer correctness reward R a and the enhanced logical rationality reward R l ′, the final reward R is obtained by calculating the average of the two: R=0.5·R a +0.5·R l ' Finally, the model is optimized by maximizing the following objective function: Among them, E represents expectation; <e s ,r q ,e o > represents any training sample in the KG dataset G; π θ (a t ) represents the policy network in step (1); a1, a2, ... a T is the process of reasoning through π θ (a t ) selects the action of T time steps; J(θ) represents the objective function, that is, according to any training sample in G <e s ,r q ,e o >, through the policy network π θ (a t ) is the expected value of the final reward R obtained by the path inferred; θ represents all the learnable parameters of the model, including the policy network π θ (a t ) entity embedding layer V, relation embedding layer U, W1, W2 and LSTM in. This optimization is achieved by the REINFORCE algorithm, which iteratively updates θ according to the following formula: in, and Denote J(θ) and logπ respectively θ (a t ) is the gradient of θ; θ′ represents the updated θ. Beneficial effects This method, based on reinforcement learning-based knowledge graph reasoning, incorporates the vast internal knowledge of large language models into reward calculations in an efficient and information-loss-resistant manner, alleviating the problem of inaccurate rewards caused by insufficient prior knowledge of the knowledge graph. The performance of this knowledge graph reasoning method, enhanced with prior knowledge, significantly improves that of existing reinforcement learning-based knowledge graph reasoning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Schematic diagram of knowledge graph reasoning

[0014] Figure 2 Flowchart of this method

[0015] Figure 3 Cue words used for logical plausibility reward enhancement based on contextual learning of large language models DETAILED DESCRIPTION

[0016] The purpose of this invention is to propose a KG reasoning method based on prior knowledge enhancement. On the basis of KG reasoning based on reinforcement learning, LLM internal knowledge is integrated into reward calculation in an efficient and information loss-resistant way, thereby providing more accurate reward feedback for the model and performing more effective model optimization.

[0017] To achieve the above objectives, the technical solution adopted by the present invention is a KG reasoning method based on LLM to enhance prior knowledge. Given a KG dataset, it can be expressed as G = { <e head ,r,e tail >|e head ∈E,r∈R,e tail ∈E}, where E and R represent entity set and relationship set respectively, e head and e tail They represent the head and tail entities of the triple respectively, and r is the relationship connecting the two entities. The model can be based on the source entity e s ∈E and query relation r q ∈R, search for a T-hop reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, thus predicting that it can be compared with es and r q Constructing potential triples <e s ,r q ,e T >Target entity e T The training process of the model on the KG dataset G is as follows Figure 2 The specific steps are as follows:

[0018] Step (1) Reasoning path search based on reinforcement learning; for a training sample in the KG dataset G <e s ,r q ,e o >,e s 、r q and e o Represent the source entity, query relation and correct target entity respectively. The model adopts the reinforcement learning framework and introduces an intelligent agent as the path search subject. The intelligent agent is based on e s As the starting point, a new action (composed of relations and entities) is selected through the policy network at each time step, and finally the reasoning path is expanded after T time steps. <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >. Among them, r t (1≤t≤T) and e t (1≤t≤T) represent the relationship and entity selected at time step t, <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >∈G. Specifically, the agent is located at entity e at time step t t , its action space Actions t By e t-1 Adjacent relationships and entities, i.e. Actions t ={a t |a t =(r′,e′), <e t-1 ,r′,e′>∈G}. The agent selects action a according to the following strategy network t : a t ~π θ (a t )=σ(A t ×W2ReLU(W1[e t ;h t ; r q ])) Among them, πθ (a t ) represents the policy network; σ and ReLU represent the Softmax function and the linear rectification function respectively; W1 and W2 represent the fully connected neural network; V and U represent the entity embedding layer and the relationship embedding layer respectively, V(e t-1 ) and U(r q ) represent e t-1 and r q Embedding vector of A t Actions t Each action in a t The embedding vector x t The stacked matrix, and x t Composed of a t The embedding vectors U(r′) and V(e′) of the relation r′ and entity e′ are concatenated, that is, x t =[U(r′);V(e′)];h t It is the historical information embedding, which is obtained by encoding the historical actions through a long short-term memory network, i.e. h t =LSTM(h t-1 ,x t-1 ), where x t-1 is the embedding vector of the action selected at time step t-1.

[0019] Step (2) Calculation of reward for answer correctness: After searching for the reasoning path, the validity of the reasoning path is evaluated from the perspective of whether the reasoning result is correct or not, and the reward is fed back to the model, thereby guiding the model to learn an effective path search strategy. Specifically, if the predicted target entity e T is the correct target entity e o , then the correct answer reward R a is equal to 1. Otherwise, the source entity e is embedded in the pre-trained KG model ConvE. s , query relation r q and e T The triplet <e s ,r q ,e T > Score. This score reflects the probability that the triple actually exists, so it is used as the correctness reward R for the answer a .

[0020] Step (3) Calculation of logical rationality reward: In addition to the reward based on the correctness of the answer, it is also necessary to evaluate the effectiveness of the reasoning path from the perspective of logical rationality and give the model logical rationality reward. Specifically, for the searched reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T,e T >, where <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >Belongs to the KG dataset G, and uses the following formula to calculate the logical rationality reward R of the path l : in, <e s ′,r1,e1′>∧…∧ <e T-1 ′,r T ,e T ′> represents any KG dataset G that satisfies the condition And contains relations r1,r2,…,r T The path, <e s ′,r1,e1′>,…, <e T-1 ′,r T ,e T ′>∈G,r q It is a query relationship; <e s ″,r1,e1″>∧…∧ <e T-1 ″,r T ,e T "> means any KG dataset G that meets the conditions And contains relations r1,r2,…,r T The path, <e s ″,r1,e1″>,…, <e T-1 ″,r T ,e T ″>∈G,r q is the query relation. Therefore, R l In fact, the path calculated by using the empirical samples from the KG prior knowledge can infer the empirical probability of the correct target entity, which reflects the logical rationality of the reasoning path. In addition, it is difficult to perform path statistics on the entire KG, so we directly count S from the paths searched in the previous training iteration. all (r q ,r1,r2,…,r T ) and S correct (r q ,r1,r2,…,r T ).

[0021] Step (4) Reward enhancement strategy based on path importance: Before using the vast internal knowledge of LLM to enhance the calculation of logical rationality rewards, it is necessary to first determine whether the reasoning path will have a significant impact on the model. Because using LLM to calculate the rewards of all paths will result in a large number of LLM calls, which will seriously increase the training time overhead. Therefore, by prioritizing the calculation of logical rationality rewards for important paths, the number of LLM calls can be reduced while ensuring model performance. Specifically, for <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, if the corresponding |S all (r q ,r1,r2,…,r T )| is less than the threshold α and R l If the value of α and β is greater than or equal to the threshold value β, the inference path is considered an important path. α and β are hyperparameters, and their optimal values ​​are determined by grid search. all (r q ,r1,r2,…,r T )| is the number of experience samples used to calculate the experience probability. The smaller the number of experience samples, the less accurate the logical rationality reward based on the experience probability. l The larger the value, the stronger the reward feedback to the model. Therefore, the path determined based on the above conditions will bring inaccurate and strong feedback to the model, which will have a significant impact on model training.

[0022] Step (5) Enhancement of logical rationality reward based on LLM context learning: For the important paths identified, LLM is used to enhance their logical rationality reward. The steps of this process for LLM are as follows: Figure 3 As shown in Figure 1. First, context learning is used to convert triples in the calculation of logical rationality rewards into natural language text with clearer semantics. This is because the semantics of structured data such as triples are usually unclear, which will affect the semantic understanding of LLM and lead to semantic loss. Specifically, for triples <e s ,r1,e1>, sample 20 triples with relation r1 from KG dataset G, and ask LLM to summarize the meaning of triples with relation r1. This can guide LLM to use its context learning ability to identify the pattern of triples with relation r1, which contains the complete semantics of triples. Then, ask LLM to summarize the triples according to the summarized meaning. <e s ,r1,e1> converted into natural language text Text(e s,r1,e1). When the triple pattern is known, LLM can fully express the semantics of the triple in the form of natural language text, thereby enhancing its ability to capture the semantic information of the triple. Through the above process, the reasoning path can be <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >The triples in the text are converted into natural language text (e s ,r1,e1),…,Text(e T-1 ,r T ,e T ), will be queried by <e s ,r q ,? > and the predicted target entity e T The triplet <e s ,r q ,e T >Converted into natural language textText(e s ,r q ,e T ). The former is regarded as the "explanation" and the latter as the "conclusion". LLM is required to give a score between 0 and 10 according to the rationality of the "explanation". Finally, the score is extracted from the LLM's answer through regular expression and divided by 10 to obtain the enhanced logical rationality reward R l ′.

[0023] Step (6) Model optimization; get the answer correctness reward R a and the enhanced logical rationality reward R l ′, the final reward R is obtained by calculating the average of the two: R=0.5·R a +0.5·R l ' Finally, the model is optimized by maximizing the following objective function: Among them, E represents expectation; <e s ,r q ,e o > represents any training sample in the KG dataset G; π θ (a t ) represents the policy network in step (1); a1, a2, ... a T is the process of reasoning through π θ (a t ) selects the action of T time steps; J(θ) represents the objective function, that is, according to any training sample in G <e s ,rq ,e o >, through the policy network π θ (a t ) is the expected value of the final reward R obtained by the path inferred; θ represents all the learnable parameters of the model, including the policy network π θ (a t ) entity embedding layer V, relation embedding layer U, W1, W2 and LSTM in. This optimization is achieved by the REINFORCE algorithm, which iteratively updates θ according to the following formula: in, and Denote J(θ) and logπ respectively θ (a t ) is the gradient of θ; θ′ represents the updated θ.

[0024] Step (7) Method parameter setting: We divide the KG dataset CoDEx-S into training set, validation set and test set in a ratio of 8:1:1, and train the knowledge graph reasoning model based on prior knowledge enhancement (PKA-KGR) on this dataset. We select the optimal values ​​of hyperparameters α and β through grid search. The value range of α is {0.3, 0.5, 0.7}, and the value range of β is {10, 20, 30, 40}. The process of grid search to select the optimal value is as follows: First, for each value combination of α and β, the model is trained for 100 iterations. Then, the mean reciprocal rank (MRR) is used to evaluate the model performance corresponding to each value combination on the validation set. Finally, the value combination with the best model performance is selected as the optimal value, and the final settings of α and β are 0.5 and 10 respectively. The remaining model hyperparameters were set as follows: training batch size of 256; number of iterations of 100; entity and relation embedding dimensions of 200 dimensions; 3 layers of LSTM with 200 hidden units; maximum timestep of 3; learning rate of 0.001; and Llama-3-8B for the large language model, with a temperature of 0 to avoid stochasticity. We evaluated the performance of PKA-KGR on the test set using the Hits@1, Hits@10, and MRR metrics. Table 1 shows the results of a performance comparison experiment between PKA-KGR and existing reinforcement learning-based knowledge graph reasoning methods. As can be seen from the table, PKA-KGR significantly outperforms existing methods in terms of Hits@1, Hits@10, and MRR. Table 2 shows the number of large language model calls during PKA-KGR training. PKA-KGR without I indicates that path importance identification is not performed and rewards are calculated using the large language model for all paths. The results in the table show that prioritizing reward calculations for important paths using the large language model significantly reduces the number of large language model calls. This enables PKA-KGR to more efficiently incorporate large language models into reward calculations. In summary, these experimental results demonstrate that PKA-KGR can incorporate the internal knowledge of large language models into reward calculations in an efficient and information-loss-resistant manner, thereby facilitating model optimization and improving model performance. Table 1: Experimental results Table 2: Number of large language model calls method CoDEx-S PKA-KGR w / o I 73387 PKA-KGR 2101

Claims

1. Knowledge graph reasoning method based on prior knowledge enhancement, characterized by The following steps are involved: Step (1) Reasoning path search based on reinforcement learning: Given a KG dataset, represented as G = { <e head ,r,e tail >|e head ∈E,r∈R,e tail ∈E}, where E and R represent entity set and relationship set respectively, e head and e tail Represent the head and tail entities of the triple respectively, and r is the relationship connecting the two entities; for each training sample in the KG dataset G <e s ,r q ,e o >,e s 、r q and e o Represent the source entity, query relation and correct target entity respectively; the PKA-KGR model adopts the reinforcement learning framework and introduces an intelligent agent as the path search subject; the intelligent agent takes the source entity e s As the starting point, a new action (composed of relations and entities) is selected at each time step, and finally a T-hop reasoning path is expanded after T time steps. <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >; Among them, r t (1≤t≤T) and e t (1≤t≤T) represent the relationship and entity selected at time step t, <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >∈G; Specifically, at time step t, the agent is located at entity e t-1 , its action space Actions t By and e t-1 Adjacent relationships and entities, i.e. Actions t ={a t |a t =(r′,e′), <e t-1 ,r′,e′>∈G}; the agent selects action a according to the following strategy network t : a t ~π θ (a t )=σ(A t ×W2ReLU(W1[V(e t-1 );h t ;U(r q )])) Among them, π θ (a t ) represents the agent’s policy network, which is used to select action a t ; σ and ReLU represent the Softmax function and the linear rectification function respectively; W1 and W2 represent the fully connected neural network; V and U represent the entity embedding layer and the relationship embedding layer respectively, V(e t-1 ) and U(r q ) represent e t-1 and r q Embedding vector of A t Actions t Each action a t The embedding vector x t The stacked matrix, and x t Composed of a t The embedding vectors U(r′) and V(e′) of the relation r′ and entity e′ are concatenated, that is, x t =[U(r′);V(e′)];h t It is the historical information embedding, which is obtained by encoding the historical actions through a long short-term memory network, i.e. h t =LSTM(h t-1 ,x t-1 ), where x t-1 is the embedding vector of the action selected at time step t-1; Step (2) Calculation of reward for correct answer: After searching for the reasoning path, the effectiveness of the reasoning path needs to be evaluated, and the evaluation result will be used as a reward for optimizing the model; the reward based on the correct answer can evaluate the effectiveness of the path search strategy from the perspective of whether the reasoning result is correct or not; specifically, if the predicted target entity e T Equal to the correct target entity e o , then the correct answer reward R a is equal to 1; otherwise, we will be represented by the source entity e s , query relation r q and the predicted target entity e T The triplet <e s ,r q ,e T > Input the pre-trained KG embedding model ConvE; ConvE will output the authenticity score of the triple, which reflects the probability that the triple really exists, so it is used as the answer correctness reward R a ; Step (3) Calculation of logical rationality reward: In addition to the reward based on the correctness of the answer, it is also necessary to evaluate the effectiveness of the reasoning path from the perspective of logical rationality and give the model a logical rationality reward. Usually, a heuristic method based on KG prior knowledge is used to calculate the logical rationality reward. Specifically, for the searched reasoning path e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, where <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T >Belongs to the KG dataset G, and uses the following formula to calculate the logical rationality reward R of the path l : in, <e s ′,r1,e1′>∧…∧ <e T-1 ′,r T ,e T ′> represents any KG dataset G that satisfies the condition And contains relations r1,r2,…,r T The path, <e s ′,r1,e1′>,…, <e T-1 ′,r T ,e T ′>∈G,r q It is a query relationship; <e s ″,r1,e1″>∧…∧ <e T-1 ″,r T ,e T "> means any KG dataset G that meets the conditions And contains relations r1,r2,…,r T The path, <e s ″,r1,e1″>,…, <e T-1 ″,r T ,e T ″>∈G,r q is the query relation; therefore, R l In fact, the path calculated by using the empirical samples from the KG prior knowledge can infer the empirical probability of the correct target entity, which reflects the logical rationality of the reasoning path. In addition, it is difficult to perform path statistics on the entire KG, so we directly count S from the paths searched in the previous training iteration. all (r q ,r1,r2,…,r T ) and S correct (r q ,r1,r2,…,r T ); Step (4) Reward enhancement strategy based on path importance; reward R for logical rationality l is the empirical probability, and the accuracy of the empirical probability depends heavily on the empirical samples from the KG prior knowledge; so when the prior knowledge in KG is insufficient, R l is inaccurate; PKA-KGR uses the vast internal knowledge of LLM to enhance the logical rationality reward, but frequent reward calculations will lead to a large number of LLM calls, which will seriously increase the training time overhead; In order to reduce the number of LLM calls while ensuring model performance, an effective strategy is to give priority to using LLM to enhance the rewards of paths that have a significant impact on model training, while the remaining paths still use rewards based on KG prior knowledge that do not require calling LLM; Specifically, for the searched reasoning path <e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >, where <e s ,r1,e1>,<e1,r2,e2> ,…, <e T-1 ,r T ,e T > belongs to KG dataset G; if the corresponding |S all (r q ,r1,r2,…,r T )| is less than the threshold α and R l If it is greater than or equal to the threshold β, then the inference path is an important path; α and β are hyperparameters, and their optimal values ​​are determined by grid search; |S all (r q ,r1,r2,…,r T )| is the number of experience samples used to calculate the experience probability. The smaller the number of experience samples, the less accurate the logical rationality reward based on the experience probability. l The larger the value, the stronger the reward feedback to the model. Therefore, the path determined based on the above conditions will bring inaccurate and strong feedback to the model, which will have a significant impact on model training. Step (5) Logical rationality reward enhancement based on LLM context learning; LLM's huge internal knowledge can effectively enhance the calculation of logical rationality reward; for example, the logical rationality rewards of path 2 and path 3 in Figure 1 are both equal to 0.5, which is obviously inaccurate; and LLM can judge that path 3 in the figure should have a lower logical rationality reward based on the two common senses that "occupation and nationality are not directly related in most cases" and "place of birth and nationality are the same in most cases"; therefore, for the important paths judged in the previous step, LLM is used to enhance its logical rationality reward. The prompt steps of this process for LLM are shown in Figure 3; first, because the unclear semantics of structured data such as triples will affect the semantic understanding of LLM, resulting in the loss of triple semantic information; so in order to alleviate this problem, it is necessary to use context learning to convert the triples in the logical rationality reward calculation into natural language text with clearer semantics; specifically, for triples <e s ,r1,e1>, sample 20 triples with relation r1 from the KG dataset G, and ask LLM to summarize the meaning of the triples with relation r1; this can guide LLM to use its context learning ability to identify the pattern of triples with relation r1, which contains the complete semantics of the triples; then, ask LLM to summarize the triples according to the summarized meaning. <e s ,r1,e1> converted into natural language text Text(e s ,r1,e1); When the triple pattern is known, LLM can fully express the semantics of the triple in the form of natural language text, thereby enhancing its ability to capture the semantic information of the triple; through the above process, the reasoning path e s ,r1,e1>∧<e1,r2,e2> ∧…∧ <e T-1 ,r T ,e T >The triples in the text are converted into natural language text (e s ,r1,e1),…,Text(e T-1 ,r T ,e T ), and will be queried by <e s ,r q ,? > and the predicted target entity e T The triplet <e s ,r q ,e T >Converted into natural language textText(e s ,r q ,e T ); take the former as "explanation" and the latter as "conclusion", and ask LLM to give a score between 0 and 10 based on the rationality of the "explanation"; finally, extract the score from the LLM's answer through regular expression and divide it by 10 to obtain the enhanced logical rationality reward R l '; Step (6) Model optimization; get the answer correctness reward R a and the enhanced logical rationality reward R l ′, the final reward R is obtained by calculating the average of the two: R=0.5·R a +0.5·R l ′ Finally, the model is optimized by maximizing the following objective function: Among them, E represents expectation; <e s ,r q ,e o > represents any training sample in the KG dataset G; π θ (a t ) represents the policy network in step (1); a1, a2, ... a T is the process of reasoning through π θ (a t ) selects the action of T time steps; J(θ) represents the objective function, that is, according to any training sample in G <e s ,r q ,e o >, through the policy network π θ (a t ) is the expected value of the final reward R obtained by the path inferred; θ represents all the learnable parameters of the model, including the policy network π θ (a t ) in the entity embedding layer V, relation embedding layer U, W1, W2 and LSTM; The optimization is achieved by the REINFORCE algorithm, which iteratively updates θ according to the following formula: in, and Denote J(θ) and logπ respectively θ (a t ) is the gradient of θ; θ′ represents the updated θ.

2. The knowledge graph reasoning method based on prior knowledge enhancement according to claim 1 is characterized by: Steps (4-6) first identify the paths that bring strong and inaccurate reward feedback to the model through a reward enhancement strategy based on path importance; Then, based on the context learning strategy, the large language model's ability to capture the semantic information of triples in the path is enhanced; then, the large language model is driven to use its internal knowledge to enhance the logical rationality rewards of these paths; finally, the logical rationality rewards and answer correctness rewards are combined in a weighted manner to guide model optimization and improve model performance.

Citation Information

Cited By

  • Government affair problem processing method and device based on large model driving

    CN121052388A

  • Government affair problem processing method and device based on large model driving

    CN121052388B