Loyal question and answer system for knowledge graph

By adopting a tree-based reasoning method in the knowledge graph question and answer system, the Monte Carlo tree search algorithm is used to integrate the language model and knowledge graph, and the problems of low efficiency and insufficient accuracy of multi-hop reasoning are solved, and a more efficient and interpretable reasoning process is achieved.

CN120181235APending Publication Date: 2025-06-20YUNNAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510302408.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing knowledge graph question and answer system is inefficient, has high error rate, and lacks interpretability in multi-hop inference tasks. It is difficult to ensure the accuracy and credibility of the inference results in scenarios where the inference path is complex and the data is sparse.

Method used

A tree-based inference method (RwT) is adopted, and a large-scale language model (LLM) and knowledge graph (KG) are integrated through the Monte Carlo Tree Search (MCTS) algorithm to build a framework for discrete decision-making problems, and iteratively optimize the inference path to improve accuracy and interpretability.

Benefits of technology

It significantly improves the accuracy and stability of the knowledge graph question and answer system in multi-hop depth inference tasks, reduces computational complexity, and improves the interpretability and transparency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181235A_ABST
    Figure CN120181235A_ABST
Patent Text Reader

Abstract

The invention discloses a mapping knowledge domain-oriented loyalty question-answering system, which comprises the following modules: a query analysis module for receiving and analyzing a natural language query input by a user; the knowledge graph acquisition module is used for extracting related information from the knowledge graph and establishing a mapping relation between user query and the knowledge graph; the inference engine module adopts a tree-based inference method, collaboratively integrates a large language model and a knowledge graph through Monte Carlo tree search, constructs a discrete decision framework, and iteratively optimizes an inference path to improve inference performance and interpretability; the answer generation module is used for converting a result of the inference engine into a natural language answer; and the credibility evaluation module is used for evaluating the credibility of the reasoning path according to the evidence in the knowledge graph and the context information in the reasoning process so as to ensure the reliability of the final answer. The method has the advantages that the accuracy and efficiency of the multi-hop reasoning task are effectively improved, and the reasoning ability and interpretability of the LLM in knowledge graph questions and answers are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a faithful question answering system for knowledge graphs. Background Art

[0002] With the rapid development of artificial intelligence technology, the knowledge graph (KG), as an important way of representing structured data, is widely used in fields such as natural language processing (NLP), question answering systems, and recommendation systems. The knowledge graph provides rich external knowledge for machines by representing entities and the relationships between them, which helps to improve the machine's understanding and reasoning abilities. In question answering systems, especially in knowledge graph question answering (KGQA), how to effectively utilize the structured information in the knowledge graph to improve multi-hop reasoning ability and accuracy has become an important research issue.

[0003] Currently, question answering systems based on knowledge graphs mainly face the following problems:

[0004] 1. Complexity of multi-hop reasoning

[0005] Knowledge graph question answering systems often involve multi-hop reasoning, that is, starting from an entity, through multi-step relational reasoning, finally reaching the answer to the question. In multi-hop reasoning tasks, the system needs to find the correct reasoning path in a huge graph, usually accompanied by a large amount of calculation and complex reasoning processes. How to perform multi-hop reasoning efficiently, especially in deep reasoning scenarios, has become a challenge.

[0006] 2. Sparsity and incompleteness of knowledge graphs

[0007] Although knowledge graphs contain a large number of entities and relationships, due to the problems of information sparsity and incompleteness in the process of graph construction, the knowledge in some fields may not be fully covered. Therefore, in these sparse regions, how to use the existing knowledge for reasonable reasoning and make up for the deficiencies of the graph is an urgent problem to be solved.

[0008] 3. Inference ability limitations of existing models

[0009] Existing knowledge graph question answering models mostly rely on embedded learning and semantic parsing methods. Although they can achieve good results in some tasks, these methods usually have difficulty dealing with complex multi-hop reasoning tasks. Especially in complex reasoning paths, traditional methods are prone to misinference and lack sufficient interpretability.

[0010] 4. The combination problem of large language models (LLMs) and knowledge graphs

[0011] Large language models (such as GPT-3, BERT, etc.) perform excellently in many NLP tasks due to their powerful language understanding capabilities. However, LLMs usually lack the ability to directly access knowledge graphs. Therefore, how to effectively combine knowledge graphs with LLMs so that LLMs can fully utilize the external knowledge in the graphs for reasoning is an important topic in current research.

[0012] To overcome the above problems, researchers have proposed various innovative methods to improve the performance of knowledge graph question-answering systems. For example, some methods attempt to convert structured knowledge graph information into text form so that LLMs can process and utilize this information. However, existing knowledge graph-based reasoning methods still have problems such as low efficiency in searching for reasoning paths, high reasoning error rates, and non-explainability of the reasoning process when dealing with multi-hop reasoning problems.

[0013] Therefore, how to design an efficient and interpretable reasoning method that can perform accurate multi-hop reasoning in large-scale knowledge graphs, especially in scenarios with complex reasoning paths and sparse data, and still ensure the accuracy and reliability of the reasoning results, remains a technical problem to be solved urgently. Summary of the Invention

[0014] In view of the defects of the prior art, the present invention provides a faithful question-answering system for knowledge graphs. The reasoning method based on tree (RwT) combines the structural information of the knowledge graph with large language models through an innovative reasoning framework, effectively improving the accuracy and stability of reasoning, especially showing excellent performance in multi-hop deep reasoning tasks.

[0015] To achieve the above invention objectives, the technical solutions adopted by the present invention are as follows:

[0016] A faithful question-answering system for knowledge graphs, comprising the following modules:

[0017] Query parsing module: responsible for receiving the natural language query input by the user and parsing it into a query representation form suitable for knowledge graph retrieval.

[0018] Knowledge graph acquisition module: extracts relevant information from the knowledge graph and establishes a mapping between the user query and the knowledge graph.

[0019] Inference Engine Module: The inference engine is the core of the system. The inference engine module adopts the tree-based inference RwT method. The RwT method collaboratively integrates the large language model (LLM) and the knowledge graph (KG) through Monte Carlo tree search (MCTS), constructs a discrete decision problem framework, and improves the inference performance and interpretability by iteratively optimizing the inference path. Among them, the RwT method includes the following steps:

[0020] Select relevant entities from the knowledge graph as seed nodes according to the entities of the given problem;

[0021] Select appropriate relationships for exploration through the prior probability predicted by the LLM and the Q value simulated by MCTS;

[0022] Through the iterative MCTS process, continuously update the inference path and optimize the decision-making according to the value of the evaluation model.

[0023] Answer Generation Module: Responsible for converting the results generated by the inference engine module into natural language answers and returning them to the user.

[0024] Confidence Evaluation Module: Evaluate the credibility of each inference path, and judge the reliability of the final answer based on the evidence in the knowledge graph and the context information in the inference process.

[0025] Furthermore, the inference engine module includes an expansion model. The expansion model uses a pre-trained language model (PLM) to iteratively train the problem and optimize the inference path according to the training data.

[0026] Furthermore, the RwT method includes the following steps:

[0027] S1. Problem Parsing and Entity Recognition: Extract key entities and relationships from the user input problem, identify the entities in the problem through natural language processing technology, and match them with the entities in the knowledge graph.

[0028] S2. Construct Query Tree: Construct a query tree according to the structured information in the knowledge graph, where the root node is the entity or concept in the user problem, and other nodes represent related entities, relationships or attributes, and the edges represent the connections between them.

[0029] S3. Initialize the Inference Path of Tree Nodes: Initialize the inference path for each node in the query tree. The inference path starts from the root node and gradually expands to the child nodes. Each path represents an inference process from the current entity to the target answer.

[0030] S4. Monte Carlo Tree Search (MCTS): Execute the Monte Carlo tree search algorithm on each node of the query tree, simulate the inference path multiple times through the "selection - expansion - simulation - backtracking" process, estimate the value of each path, and find the optimal inference path.

[0031] S5. Path Expansion and Optimization: Each inference path obtained through MCTS simulation is expanded and optimized according to its result; if the path derives the correct answer, the path is further expanded, otherwise, it backtracks and adjusts the inference direction.

[0032] S6. Optimize Inference through Policy Model: Use the trained policy model to score each inference path, measure the reliability of the path, and select the optimal inference path based on the evaluation result of the policy model.

[0033] S7. Inference Result Evaluation and Backtracking: During the search process, evaluate the result of each path according to the reliability of the path and the coherence of the inference process, and backtrack when a suitable inference path is found to ensure that each hop of the path meets the inference conditions.

[0034] S8. Generate the Final Answer: Generate the answer most relevant to the user's question according to the final inference path, and convert the answer into a natural language format through the language generation model and return it to the user.

[0035] Furthermore, MCTS uses the PUCT algorithm to traverse the tree, and selects the optimal action based on the action a in the current state s, the prior probability P(s, a), and the posterior value, and calculates the selection criterion of the action through the following formula: t in the action a t the prior probability P(s t , a t ), and the posterior value Select the optimal action, and calculate the selection criterion of the action through the following formula:

[0036] where P(s t , a) is the prior score provided by the policy model, is the estimated reward of the state-action pair, N parent (a) and N(s t , a) are the visit counts of the parent node and the current state-action pair respectively, and c puct is a constant for balancing exploration and exploitation.

[0037] Furthermore, the RwT method uses an evaluation function that combines future rewards, state relevance, and actual results to evaluate newly expanded nodes. The evaluation function is defined as:

[0038] where V roll-out (s t ) is generated by the expansion model, and predicts the future reward through the pre-trained LLM; V critical (s t ) is calculated by the key model to evaluate whether the current node is a reasonable answer to the question; r(·) is the actual result reward, representing the result of a simulation starting from the current state.

[0039] Furthermore, after the RwT method completes the evaluation, it propagates the values of the leaf nodes back into the tree, updating the value estimates and visit counts of all ancestor nodes on the path. The update rules are as follows:

[0040] N(s,a)←N(s,a)+1

[0041] where is an indicator function that is 1 when the state-action pair (s,a) lies on the path leading to state s t and 0 otherwise. This update mechanism ensures that the value estimate of each node reflects the cumulative experience of all its descendant nodes.

[0042] Compared with the prior art, the advantages of the present invention are as follows:

[0043] 1. Through the iterative optimization of the MCTS (Monte Carlo Tree Search) algorithm, the present invention can effectively find the most promising inference paths in the knowledge graph, thereby improving the accuracy of inference, especially showing superiority in dealing with multi-hop inference tasks.

[0044] 2. By using a fixed LLM as the policy model and focusing on training the evaluation model, the present invention enables the model to more efficiently utilize the knowledge of the LLM and the information of the external knowledge graph, thereby significantly enhancing the inference ability in the knowledge graph question answering (KGQA) task.

[0045] 3. In the faithful question answering system, RwT balances the relationship between exploration and exploitation through the PUCT algorithm, making the inference process more efficient, reducing the need for large-scale datasets, and dynamically updating the value of nodes during the inference process, reducing the computational complexity.

[0046] 4. RwT in the faithful question answering system performs particularly well in dealing with multi-hop deep inference tasks with long logical chains, especially in question answering tasks involving complex semantic understanding, and can provide accurate and reasonable answers.

[0047] 5. RwT in the faithful question answering system is designed as a modular inference framework, which can be conveniently combined with different types of LLMs (such as ChatGPT, LLaMA, etc.) to improve the inference effect of the LLM, and there is no need to completely retrain the LLM, having good flexibility and scalability.

[0048] 6. By fine-tuning the PLM (pre-trained language model) and inserting it into the LLM as the evaluation model, the present invention can significantly improve the inference efficiency, especially in scenarios where the external knowledge is not fully covered, and can effectively reduce the dependence on the background knowledge of the LLM.

[0049] 7. The inference path of RwT in a faithful question-answering system can be visualized through a tree structure, which helps users understand the decision logic of each step in the inference process and improves the interpretability and transparency of the model. Description of the Drawings

[0050] Figure 1 is the framework diagram of the RwT method in an embodiment of the present invention;

[0051] Figure 2 are the diagrams of four processing steps of the RwT method in an embodiment of the present invention;

[0052] Figure 3 is the performance comparison diagram of RwT in an embodiment of the present invention and a comparison method under the same fine-tuning samples;

[0053] Figure 4 is the case analysis diagram of RwT in an embodiment of the present invention. Detailed Embodiment

[0054] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following provides further detailed descriptions of the present invention according to the drawings and by listing embodiments.

[0055] I. Faithful Question-Answering System for Knowledge Graph

[0056] The present invention provides a faithful question-answering system for knowledge graph, including the following modules:

[0057] Query parsing module: This module is responsible for receiving the natural language query input by the user and parsing it into a query representation form suitable for knowledge graph retrieval. It uses natural language processing technologies to perform operations such as word segmentation, part-of-speech tagging, and syntactic analysis on the query, and extracts key information (such as entities, relationships, etc.) in the query.

[0058] Knowledge graph acquisition module: This module extracts relevant information from the knowledge graph and establishes a mapping between the user query and the knowledge graph. It filters out possible nodes, edges, and relevant entity-relationship pairs from the knowledge graph according to the parsed query, providing basic data for subsequent inference.

[0059] Inference Engine Module: The inference engine is the core of the system. The inference engine module adopts a tree-based reasoning method (RwT). The RwT is iteratively optimized through the Monte Carlo Tree Search (MCTS) framework and dynamically updates the external knowledge graph information used in the reasoning process to improve the accuracy and interpretability of the reasoning. It mainly uses the Monte Carlo Tree Search (MCTS) algorithm for reasoning. It conducts decision-making searches through a tree structure and dynamically selects the optimal reasoning path. This module combines the structural information of the knowledge graph and the reasoning ability of the large language model (LLM), selects the optimal answer from multiple reasoning paths, and ensures the fidelity of the reasoning process to avoid generating "hallucinations" or inaccurate answers.

[0060] Answer Generation Module: This module is responsible for converting the results generated by the inference engine module into natural language answers and returning them to the user. It generates concise and accurate natural language answers through language generation technology based on the reasoning results.

[0061] Confidence Evaluation Module: This module evaluates the confidence of each reasoning path, judges the reliability of the final answer based on the evidence in the knowledge graph and the context information in the reasoning process. It can also provide the source information of the answer to the user to enhance the interpretability of the system.

[0062] User Feedback Module: The user feedback module allows users to evaluate the answer results and optimizes the reasoning process and model performance based on user feedback. By collecting user feedback, the system can continuously improve its own reasoning ability and accuracy.

[0063] Knowledge Graph Update and Maintenance Module: Responsible for managing and updating the content in the knowledge graph. It regularly checks and updates the information of entities, relationships, etc. in the knowledge graph to ensure that the system can adapt to new knowledge and changes, and improve the long-term stability and reliability of the system.

[0064] II. Tree-based Reasoning Method (RwT)

[0065] 2.1 Overview

[0066] As Figure 1 shown, in RwT, the MCTS algorithm starts exploring the knowledge graph from the initial state s0, which corresponds to the entity mentioned in the given question q. Each node in MCTS represents a state s t , defined by the current entity in the knowledge graph. From each state, the possible actions a t are the relationships connecting to other entities. Taking an action will transfer to the next entity through the selected relationship, leading to a new state s t+1 . The goal of this exploration is to find a knowledge graph path leading to the correct answer entity for the given question q.

[0067] Each state-action (entity-relation) pair is evaluated using a value model that estimates future rewards and the relevance of the current entity to the problem. As Figure 2 shown, MCTS in RwT finds the most promising path in the knowledge spectrum by iteratively selecting, expanding, evaluating, and backpropagating node values. The detailed process is described in Algorithm 1 in the appendix. Our method uses a fixed LLM as the policy model and only focuses on training the evaluation model. The evaluation model includes an expansion and a key model, which are iteratively trained to improve their accuracy in guiding the search process.

[0068] 2.2 MCTS Planning

[0069] For each action a of state s, the method stores a set of statistics {P(s,a), Q(s,a), N(s,a)}, where P is the prior score (i.e., the likelihood score from the policy model), Q is the action value, and N is the number of visits. Q(s,a) can be regarded as the posterior score, considering the future impact of a. Specifically, consider a complete inference path consisting of T planning steps. At a given time t, the state st represents the current entity in the knowledge graph, encapsulating the results of all previous inference steps. The possible subsequent inference steps are represented as actions at, corresponding to selecting the relations connecting to the current entity. The planning algorithm ends after a fixed simulation budget. Each simulation includes the following phases:

[0070] (1) Selection

[0071] In the selection phase, the algorithm starts from the root node (the given problem entity) and traverses the tree to select a child node according to the PUCT (Policy + UCT) algorithm. RwT adopts PUCT instead of the standard UCB algorithm because it is more suitable for integrating with neural network priors, which is crucial for methods using an LLM as the policy model. PUCT incorporates the prior probability into the process, enabling the method to utilize the knowledge of the LLM. The selection criterion for action a in state s t is given by the following formula:

[0072]

[0073] where represents the estimated state-action value, representing the expected future reward for taking action a in state s t . represents the MCTS tree constructed in the k-th round. In addition, c puct is a constant that helps balance exploration and exploitation. The probability π policy (a∣s t ) represents the likelihood of selecting action a in state s determined by the policy model t . The LLM is a policy model that provides the likelihood of selecting action a in state s tThe likelihood score of taking action a (selection relationship) in the (current entity). Additionally, N parent (a) and N(s t , a) denote the number of visits to the parent node and the state-action pair respectively.

[0074] The selection phase continues until a leaf node is reached, ensuring the exploration of the most promising nodes based on prior knowledge and cumulative statistics.

[0075] (2) Expansion

[0076] After reaching a leaf node, it is expanded by generating all possible actions from that node. Each action corresponds to a transition to a new state. For a knowledge graph, the possible actions are the relationships connected to the current entity, leading to new entities. Thus, all un-explored entities connected to the current entity are added as leaf nodes of the current node to the search tree.

[0077] (3) Evaluation

[0078] The newly expanded nodes are evaluated using an evaluation function that integrates future rewards, state relevance, and actual results. The value function is defined as:

[0079] where V roll-out (s t ) is generated by the expansion model, which predicts the future reward (Q-value) of the state based on a pre-trained LLM by simulating future steps. V critical (s t ) is calculated by the key model, which uses the BERT model to evaluate whether the entity of the current node is a semantically reasonable answer to the question. The term represents the actual result, where and represent the action and state sampled by the policy model in the i-th simulation. r(·|s t ) is the reward of the result of a single simulation starting from state s t . The parameter μ is set to the indicator function I(s t ). If the expanded node is a terminal node, the reward is used; otherwise, the model relies on the estimated value. During inference, μ is set to 0 to ensure consistency and always relies on the value model for node evaluation, including terminal nodes.

[0080] (4) Backpropagation

[0081] After evaluation, the values of the leaf nodes are propagated back up the tree, updating the value estimates and visit counts of all ancestor nodes on the path. The backpropagation process ensures that the information obtained from evaluating the leaf nodes influences higher-level decisions in the tree. The update rules are as follows:

[0082] N(s, a) ← N(s, a) + 1

[0083] where is an indicator function that is 1 when the state - action pair (s, a) lies on the path leading to state s t and 0 otherwise. This update mechanism ensures that the value estimate of each node reflects the cumulative experience of all its descendant nodes.

[0084] III. Experimental Section

[0085] To verify the effectiveness of the tree - based reasoning method (RwT) in the present invention, the present invention tested and compared its reasoning performance in the following two experiments:

[0086] 3.1 Experimental Setup

[0087] (1) Datasets

[0088] In this embodiment, the reasoning performance of RwT was evaluated on two benchmark KGQA (Knowledge Graph Question Answering) datasets:

[0089] WebQuestionSP (WebQSP)

[0090] Complex WebQuestions (CWQ)

[0091] These two datasets contain multi - hop KG reasoning questions with a maximum of 4 hops and are both based on the Freebase knowledge graph.

[0092] (2) Evaluation Protocol

[0093] To evaluate the performance of the reasoning method, in this embodiment, referring to previous studies, the task was constructed as a ranking problem. Specifically, for each question, we ranked the candidate entities according to the answer scores, and then used the Hits@1 metric to measure the accuracy of the top - ranked answer. Considering that a question may have multiple correct answers, this embodiment also introduced the F1 metric for a more comprehensive evaluation.

[0094] (3) Baselines

[0095] In this embodiment, RwT was compared with the following four types of baseline methods:

[0096] Traditional embedding and semantic parsing methods: such as KV - Mem, EmbedKGQA, NSM, TransferNet, KGT5, SPARQL, QGG, ArcaneQA, RnG - KBQA.

[0097] Retrieval methods: such as GraftNet, PullNet, SR + NSM, SR + NSM + E2E.

[0098] Directly using LLMs for inference: such as Flan-T5-xl, Alpaca-7B, LLaMA2-Chat-7B, ChatGPT, GPT4o, ChatGPT+CoT.

[0099] LLMs+KGs methods: such as KD-CoT, UniKGQA, DECAF (DPR+FiD-3B), StructureGPT, ReasoningLM, RoG.

[0100] 2.2 Implementation details

[0101] (1) Expansion model

[0102] For the expansion model in RwT, LLaMA3-Base-8B is adopted as the basic PLM (Pre-trained Language Model). This model is iteratively trained on the training sets of WebQSP and CWQ, with a learning rate of 4e-5 and the AdamW optimizer used for training.

[0103] (2) Key model

[0104] A fine-tuned Sentence-BERT model is used to measure the semantic similarity between the entities found in the current search and the answers predicted by the policy model.

[0105] (3) Generating training data through MCTS

[0106] As shown in Table 1, the value model training program of this method is trained using a 4-round iterative method. In each round, the experiment generates 8 Monte Carlo Search Trees (MCTS) for each question-answer pair. Up to 5 correct and 5 incorrect reasoning paths are extracted from these trees, maintaining an approximate 1:1 ratio of positive and negative examples.

[0107] Table 1. Performance comparison of different baselines on two KGQA datasets

[0108]

[0109]

[0110] 3.3 Main results

[0111] Such as Figure 3As shown, the experimental results show that RwT outperforms all baseline methods in all datasets and metrics. Compared with the current state-of-the-art models (such as ReasoningLM and RoG), RwT shows an average performance improvement of 9.81% and 5.77% on two datasets respectively. In addition, RwT does not require full retraining of the LLM, but only fine-tunes a smaller PLM and then inserts it into any LLM.

[0112] Specifically, on the CWQ dataset, RwT performs excellently, especially when dealing with multi-hop deep knowledge reasoning tasks with long logical chains, showing obvious advantages. Traditional embedding and semantic parsing methods perform limitedly on these complex problems, while RwT can effectively handle them.

[0113] 3.4 Performance Evaluation as a Module Plugin

[0114] To verify the performance of RwT when integrated as a module into different LLMs, this experiment integrated RwT as a plugin with different LLMs such as ChatGPT and Qwen to evaluate its performance improvement. As shown in Table 2, the experimental results show that after integrating RwT as a plugin with other LLMs, the average Hits@1 of ChatGPT and Qwen increased by 55.88% and 66.33% respectively.

[0115] Table 2. Effects of Integrating the RwT Reasoning Module with Different LLMs

[0116]

[0117]

[0118] Even in a smaller model (such as LLaMA-2, with 7B parameters), RwT significantly outperforms larger LLMs. In contrast, the combination method of RwT improves by 31.57% compared to ChatGPT+CoT. This shows that using the RwT method to fine-tune a smaller LLM into an evaluation model can effectively enhance the LLM reasoning performance, especially in scenarios involving external knowledge not covered by the LLM training data.

[0119] 3.5 Tuning Efficiency

[0120] As Figure 4 shown, compared with other models that require full-scale retraining, RwT requires approximately 80% less tuning data to achieve stable reasoning performance. RwT only needs to fine-tune the PLMs for the evaluation of reasoning tasks, while other methods require more parameter adjustment and retraining.

[0121] This result indicates that the RwT method provides a more efficient training approach and effectively addresses the challenges of LLMs in multi-hop reasoning tasks.

[0122] 3.6 Ablation Study

[0123] To further verify the effectiveness of each component in RwT, an ablation study was conducted in this embodiment. In the experiment, when only using LLMs as the sole component, the performance of the model was poor, with Hits@1 being 53.1%. After adding MCTS, the performance increased to 59.2%, indicating that mapping KGQA to a Markov decision process is effective. When combining all three components, the accuracy of the model increased to 69.6%.

[0124] IV. Conclusion

[0125] The present invention proposes a tree-based reasoning method (RwT), which enhances the reasoning ability of LLMs by integrating external knowledge graph data. Experimental results show that RwT significantly improves the reasoning performance of LLMs without the need for additional data annotation, especially performing excellently in multi-hop reasoning tasks. As a plug-and-play reasoning module, RwT can effectively reduce the dependence of LLMs on background knowledge, reduce domain-specific knowledge bias, and improve the accuracy and interpretability of reasoning results.

[0126] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the implementation methods of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A faithful question answering system for knowledge graph, characterized in that: Includes the following modules: Query parsing module: responsible for receiving natural language queries input by users and parsing them into query representations suitable for knowledge graph retrieval; Knowledge graph acquisition module: extracts relevant information from the knowledge graph and establishes a mapping between user queries and the knowledge graph; Reasoning engine module: The reasoning engine is the core of the system. The reasoning engine module adopts the tree-based reasoning RwT method. The RwT method integrates the large language model LLM with the knowledge graph KG through Monte Carlo tree search MCTS to build a discrete decision problem framework, and improves the reasoning performance and interpretability by iteratively optimizing the reasoning path. The RwT method includes the following steps: According to the entities of the given question, relevant entities are selected from the knowledge graph as seed nodes; Select appropriate relationships for exploration using the prior probability predicted by LLM and the Q value simulated by MCTS; Through the iterative MCTS process, the reasoning path is continuously updated and the decision is optimized based on the value of the evaluation model; Answer generation module: responsible for converting the results generated by the inference engine module into natural language answers and returning them to the user; Credibility assessment module: Evaluate the credibility of each reasoning path and judge the reliability of the final answer based on the evidence in the knowledge graph and the contextual information in the reasoning process.

2. The faithful question-answering system for knowledge graph according to claim 1 is characterized in that: The reasoning engine module includes an unfolding model, which uses a pre-trained language model PLM to iteratively train the problem and optimize the reasoning path according to the training data.

3. The faithful question-answering system for knowledge graph according to claim 1 is characterized in that: The RwT method comprises the following steps: S1. Question parsing and entity recognition: Extract key entities and relationships from the questions input by users, identify entities in the questions through natural language processing technology, and match them with entities in the knowledge graph; S2. Build a query tree: Based on the structured information in the knowledge graph, build a query tree, where the root node is the entity or concept in the user's question, other nodes represent related entities, relationships or attributes, and edges represent the connections between them; S3. Initialize the reasoning path of the tree node: Initialize the reasoning path for each node in the query tree. The reasoning path starts from the root node and gradually expands to the child nodes. Each path represents a reasoning process from the current entity to the target answer. S4, Monte Carlo Tree Search MCTS: Execute the Monte Carlo Tree Search algorithm on each node of the query tree, simulate the reasoning path multiple times through the "select-expand-simulate-backtrack" process, estimate the value of each path, and find the optimal reasoning path; S5, path expansion and optimization: Each time the reasoning path obtained through MCTS simulation is expanded and optimized according to its results; if the path derives the correct answer, the path is further expanded, otherwise it is backtracked and the reasoning direction is adjusted; S6. Optimize reasoning through policy model: Use the trained policy model to score each reasoning path, measure the reliability of the path, and select the optimal reasoning path based on the evaluation results of the policy model; S7, Reasoning result evaluation and backtracking: During the search process, the results of each path are evaluated according to the reliability of the path and the coherence of the reasoning process, and backtracking is performed when a suitable reasoning path is found to ensure that each hop of the path meets the reasoning conditions; S8. Generate final answer: Generate the answer most relevant to the user's question based on the final reasoning path, and convert the answer into natural language format through the language generation model and return it to the user.

4. The faithful question-answering system for knowledge graph according to claim 3 is characterized in that: MCTS uses the PUCT algorithm to traverse the tree, based on the current state s t Action a t The prior probability P(s t ,a t ) and the posterior value Select the best action and calculate the action selection criteria using the following formula: Among them, P(s t ,a) is the prior score provided by the policy model, is the estimated reward for the state-action pair, N parent (a) and N(s t ,a) are the number of visits to the parent node and the current state-action pair, c puct A constant that balances exploration and exploitation.

5. The faithful question-answering system for knowledge graph according to claim 3 is characterized in that: The RwT method evaluates the newly expanded nodes using an evaluation function that combines future rewards, state relevance, and actual results. The evaluation function is defined as: Among them, V roll-out (s t ) is generated by the unfolded model, predicting future rewards through the pre-trained LLM; V critical (s t ) is calculated by the key model to evaluate whether the current node is a reasonable answer to the question; r(·) is the actual result reward, which represents the result of a simulation starting from the current state.

6. The faithful question-answering system for knowledge graph according to claim 3 is characterized in that: After the evaluation is completed, the RwT method propagates the value of the leaf node back to the tree and updates the value estimation and visit count of all ancestor nodes on the path. The update rules are as follows: N(s,a)←N(s,a)+1 in is an indicator function that indicates that when the state-action pair (s, a) is on the path to state s t This update mechanism ensures that the value estimate of each node reflects the accumulated experience of all its descendant nodes.

Citation Information

Cited By

  • Multi-hop RAG question and answer method and system based on path exploration and hyperbolic refining

    CN120611030A

  • A multi-hop RAG question-answering method and system based on path exploration and hyperbolic refinement

    CN120611030B

  • Multi-round reasoning question answering method and system based on query graph driving

    CN120687579A