Iterative retrieval reasoning enhanced knowledge graph RAG method based on introspection mechanism

By introducing iterative retrieval inference and Self-RAG introspection technology into the knowledge graph RAG method, the problems of incomplete information and inaccurate answers in complex inference tasks are solved, and a more efficient and accurate knowledge acquisition and reasoning process is achieved.

CN120216732AActive Publication Date: 2025-06-27HARBIN INST OF TECH

Patent Information

Application Number
CN202510276963.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The prior art has problems such as incomplete information, broken reasoning chains or inaccurate answers in complex reasoning and multi-domain Q&A tasks, especially in areas such as technical document Q&A, medical diagnosis and legal reasoning.

Method used

The knowledge graph RAG method enhanced by iterative search inference based on the introspection mechanism is adopted, combined with chain reasoning and self-feedback mechanism, through multiple rounds of search and gradual expansion of the search scope, the human reasoning process is simulated, and the quality of generated content is evaluated through Self-RAG introspection technology.

Benefits of technology

It significantly improves the completeness of information and the accuracy of answers in complex tasks, ensures that the generated results are more coherent and comprehensive, and is suitable for scenarios such as intelligent question and answer, technical document analysis and intelligent decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216732A_ABST
    Figure CN120216732A_ABST
Patent Text Reader

Abstract

The invention discloses an iterative retrieval reasoning enhanced knowledge graph RAG method based on an introspection mechanism, and relates to the technical field of artificial intelligence and natural language processing. The knowledge retrieval process is enhanced by combining iterative retrieval reasoning with chain reasoning, and it is ensured that most relevant knowledge paragraphs are obtained in multiple rounds of retrieval and reasoning. And through alternate implementation of recursive reasoning and retrieval, the retrieval range is dynamically expanded, and the correlation and integrity of information are optimized. A Self-RAG introspection technology is introduced, and the retrieved paragraphs and the generated content are reflected through a trans-province mark, so that a language model can be subjected to controllable operation in an inference stage, and behaviors of the language model are customized according to different task requirements. The efficiency and accuracy of a large-scale knowledge graph retrieval and generation model in a complex reasoning task are improved by combining an iterative retrieval reasoning technology and a Self-RAG technology in an RAG enhancement method and utilizing chain reasoning and a self-inverse provincial mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and natural language processing, and specifically to a knowledge graph RAG method based on an introspective mechanism for iterative retrieval inference enhancement. Background Art

[0002] In the fields of artificial intelligence and natural language processing (NLP), in recent years, with the rapid development of large-scale pre-trained models, the models have demonstrated unprecedented performance in tasks such as text generation, question-answering systems, and knowledge reasoning. However, although these models have powerful language understanding and generation capabilities, their knowledge reserves mainly come from training data, resulting in the models often showing incomplete information, broken reasoning chains, or inaccurate answers when facing complex reasoning, multi-domain question-answering, or tasks that require external knowledge supplementation. This limitation has become an important factor affecting the model's performance and generalization ability in practical applications. Especially in scenarios such as technical document question-answering, medical diagnosis, and legal reasoning, where high requirements are placed on the rigor and integrity of knowledge, the answers generated by existing models may lack technical details or key steps.

[0003] To address this issue, Retrieval-Augmented Generation (RAG) technology has emerged. RAG technology combines knowledge retrieval with a generation model, dynamically invoking an external knowledge base during the generation process and using the retrieved relevant information as supplementary context to enhance the model's knowledge coverage and generation quality. This approach effectively avoids the limitations of the model generating answers solely based on internal parameters and greatly improves the accuracy of answering complex questions. However, most traditional RAG systems perform single-round retrieval based on a static knowledge base and directly input the retrieved results into the generation model for answer generation. This method performs well in simple question-answering scenarios but often proves inadequate in complex tasks that require multi-step reasoning or multi-level knowledge association. The static retrieval mechanism is prone to one-sided information acquisition and cannot fully capture the multi-dimensional features of the question, thereby affecting the integrity and accuracy of the final generation result.

[0004] As a structured knowledge representation method, knowledge graphs have been widely used in multiple fields in recent years, especially in natural language understanding and reasoning tasks. Different from traditional text data, knowledge graphs represent information through explicit entities and relationships, enabling machines to reason about and query complex world knowledge. Introducing a knowledge graph into a RAG system can significantly improve the precision and diversity of retrieval. Especially in tasks that require explicit entity and relationship reasoning, the knowledge graph can provide more accurate and rigorous answers.

[0005] To address the above problems, iterative retrieval inference technology has been proposed to further enhance the retrieval ability and information integrity of the RAG system in complex inference tasks. The core lies in introducing a chain reasoning mechanism, which simulates the step-by-step reasoning process of humans when solving complex problems by means of multi-round retrieval and gradually expanding the retrieval scope. The specific process includes first conducting a preliminary retrieval based on the user's question to obtain initially relevant document paragraphs; subsequently, the model generates an intermediate reasoning (CoT) based on the retrieved information, uses this intermediate reasoning as a new retrieval query to further expand the retrieval scope, and continuously obtains more relevant paragraphs. This process is iterated until the retrieval results meet the knowledge requirements of the question or reach the set termination conditions. In this way, iterative retrieval inference ensures that the model can gradually improve the knowledge acquisition path in complex problems, avoid information loss or retrieval bias caused by the limitations of single-round retrieval, and make the answers finally generated by the model more coherent and comprehensive.

[0006] Although iterative retrieval inference improves the information coverage and relevance during the retrieval process, the answers generated by the generation model based on the multi-round retrieval results may still have problems such as uncertainty or incomplete reasoning chains. This is because the generation process depends on the content of the existing knowledge base, and the knowledge base itself may have knowledge gaps or unaddressed areas. Therefore, to further improve the accuracy and integrity of the generated results, the Self-RAG introspection technology is introduced.

[0007] Based on traditional RAG and iterative retrieval inference, Self-RAG adds a self-feedback mechanism. This technology can reflect on the retrieved paragraphs and the content generated by itself according to requirements through special markers (introspection markers). Generating introspection markers enables the language model to perform controllable operations during the inference stage, thereby customizing its behavior according to different task requirements. The Self-RAG introspection technology enables the generation model to have better task adaptability, stronger factuality and accuracy when dealing with complex tasks, ensuring that the final output results are closer to the actual requirements.

[0008] In summary, the inventors have proposed a knowledge graph RAG method integrating iterative retrieval inference technology and Self-RAG introspection technology, aiming to comprehensively improve the inference ability and generation quality of the model in complex tasks by constructing a knowledge graph, combining chain multi-round retrieval and self-feedback mechanism. In practical applications, this method can not only effectively address the deficiencies of traditional RAG systems in multi-round inference tasks, but also play an important role in scenarios such as intelligent question-and-answer systems, technical document parsing, and intelligent decision-making support, providing an efficient and reliable solution for large-scale knowledge question-and-answer and complex inference tasks. Summary of the Invention

[0009] To solve the problems of possible knowledge gaps, uncertainties, and incomplete reasoning steps in the existing retrieval-augmented generation technology during the generation process, the present invention provides a knowledge graph RAG method for iterative retrieval inference enhancement based on an introspection mechanism. By combining the iterative retrieval inference technology and Self-RAG technology in the RAG enhancement method, and using chain reasoning and self-reflection mechanisms, the efficiency and accuracy of large-scale knowledge graph retrieval and generation models in complex reasoning tasks are improved.

[0010] To achieve the above object, the present invention adopts the following technical solutions: A knowledge graph RAG method for iterative retrieval inference enhancement based on an introspection mechanism, comprising the following steps:

[0011] Step 1: Construction of the knowledge graph

[0012] Collect and clean multi-source heterogeneous data and preprocess the data through natural language processing. Then, identify entities from the text through named entity recognition, and extract the relationships between entities to construct triples (h, r, e), where h is the head entity, r is the relationship, and e is the tail entity;

[0013] Generate a semantic vector v for each entity through the Sentence-BERT model i , expressed as: v i = Sentence-BERT(h), store v i into the entity. At the same time, construct a knowledge graph based on the triples, use the Leiden algorithm to identify the community structure formed by the entities, and use the BERT language model to generate community summaries;

[0014] Step 2: Retrieval based on iterative retrieval inference technology

[0015] Extract the entities in the initial question Q0 input by the user through named entity recognition, retrieve the extracted entities through a graph database, encode the initial question Q0 using the Sentence-BERT model to obtain a query vector q, calculate the similarity between the query vector q and the semantic vector v of each entity i , and based on the similarity score, initially retrieve K entities Retriever R K Return at most K entities, expressed as

[0016] Enter a multi-round iterative retrieval and reasoning process until the termination condition is met. In each round of iteration, first use the chain reasoner T to combine the current question Q i and the information carried by the retrieved entities P i K to generate a new CoT sentence C iThe reasoning summary for the current question and retrieved paragraphs, denoted as C i = T(P i K , Q i ), and use the CoT sentence C i as the new question query Q i+1 , denoted as Q i+1 = NextQuery(C i ), and use the retriever R K to retrieve the top K entities most relevant to Q i+1 from the knowledge graph, denoted as Then the iterative retrieval reasoning algorithm is expressed as:

[0017]

[0018] where n represents the maximum number of iterations;

[0019] Step 3: Enhancement generation based on the Self-RAG introspection algorithm

[0020] For the generation model, combine the retrieved K paragraphs with the question, generate K continuations at each timestamp t, score the K continuations, and select the output with the highest score as the final result;

[0021] When evaluating the quality f(y t , d, Critique) of each continuation, consider the generation probability p(y t |x, d, y < t) of this paragraph, and introduce the critique score S(Critique), where y t is the currently generated candidate paragraph, d is the relevant knowledge retrieved from the knowledge graph, x is the initial question Q0, y is the content already generated before timestamp t, and the calculation formula is as follows:

[0022] f(y t , d, Critique) = p(y t |x, d, y < t) + S(Critique)

[0023] In the generation model, by introducing introspective Tokens as marker Tokens, the generation model will give the generation probability of each Token in the logprobs, and calculate the critique score S(Critique) based on these probabilities;

[0024] Step 4: Construct an iterative retrieval reasoning knowledge graph RAG framework with an introspection mechanism

[0025] Combine the iterative retrieval reasoning technology with the Self-RAG introspection algorithm to construct a new knowledge graph RAG framework. Obtain knowledge fragments through iterative retrieval reasoning, and optimize the generation process through the Self-RAG introspection algorithm. Select the optimal continuation as the output to generate the final answer.

[0026] Furthermore, when constructing the knowledge graph in step one, attach statements or attribute information about entities to the triples to enrich the structured knowledge, and store this structured knowledge in the Neo4j graph database. In the knowledge graph, nodes represent entities, edges represent relationships, and the triples and their attached information are stored as attributes of nodes and edges.

[0027] Furthermore, the termination conditions for the multi-round iterative retrieval and reasoning process in step two are divided into two types:

[0028] Condition one: Set a maximum number of iterations. If the maximum number of iterations is not reached, continue with the next round of iteration. If the maximum number of iterations is reached, stop the iteration.

[0029] Condition two: After each round of iteration, generate the current CoT sentence C i and the CoT sentence C i-1 from the previous round. Convert these two CoT sentences into semantic vectors respectively through the pre-trained model Sentence-BERT, and calculate the distance between these two semantic vectors through the Euclidean distance formula to judge the similarity. If the Euclidean distance is less than the preset threshold, stop the iteration.

[0030] Furthermore, in step three, the calculation formula for the critique score S(Critique) is as follows:

[0031]

[0032] In the formula, G represents the critique tag group, and W G represents the weight hyperparameter of each critique tag group, represents the score of each critique tag group within the time stamp t.

[0033] Furthermore, the calculation formula is as follows:

[0034]

[0035] In the formula, represents the probability that the most expected reflection tag appears within the time stamp t, P t (r i ) represents the probability of generating all possible tags within the time stamp t, and N G is the number of reflection tags in the critique tag group.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] 1. Enhance retrieval quality and accuracy: In traditional retrieval methods, the retrieval process may lack sufficient relevance and is prone to retrieving incomplete information. The present invention introduces an iterative retrieval and reasoning technique, which alternately performs recursive reasoning and retrieval to obtain more relevant and complete information. Through multiple rounds of retrieval and reasoning, it effectively avoids retrieving irrelevant or one-sided information, improving the relevance and integrity of the retrieval results. In addition, by generating CoT sentences and using them as the basis for subsequent queries, it can expand the breadth of information, helping to understand the problem more comprehensively from multiple perspectives, thereby improving the accuracy of the answers;

[0038] 2. Enhancement of the reasoning and generation process: Traditional generation models are prone to generating inaccurate or incomplete answers during the generation process. The present invention adopts the Self-RAG introspection algorithm, which evaluates the quality of each generated paragraph through a critical scoring mechanism, and evaluates by combining the generation probability and critical markers (such as supportiveness, usability, etc.), thereby improving the accuracy and factuality of the generated content. In addition, by introducing introspective Tokens, the model can perform reflective evaluation to ensure that the generated content not only conforms to logic but also meets diverse requirements, thus optimizing the generation process;

[0039] 3. Progressive information integration: In multi-step reasoning tasks, the model faces problems such as insufficient knowledge or inaccurate reasoning. The present invention combines the iterative retrieval and reasoning technique with Self-RAG, enabling the generation model to continuously improve information based on each reasoning and retrieval, gradually approaching the correct answer. The alternating of chain reasoning and generation effectively avoids problems of knowledge gaps and incomplete information, ensuring the accuracy of the reasoning process and the reliability of the answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is the architecture diagram of the method of the present invention;

[0041] Figure 2 is the schematic diagram of the principle of the method of the present invention;

[0042] Figure 3 is the result diagram of Example 1;

[0043] Figure 4 is the result diagram of Example 2. DETAILED DESCRIPTION OF THE INVENTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] As Figures 1 to 2 shown, the knowledge graph RAG method with enhanced iterative retrieval reasoning based on the introspection mechanism includes the following steps:

[0046] Step 1: Construction of the knowledge graph

[0047] Under the condition of ensuring data quality, collect and clean multi-source heterogeneous data (such as technical documents, papers, and forum Q&A), and preprocess the data using natural language processing (NLP) technology. Then, use named entity recognition (NER) technology to identify entities from the text and extract the relationships between entities to construct triples (h, r, e), where h is the head entity, r is the relationship, and e is the tail entity. Generate the semantic vector v of each entity through the Sentence-BERT model i , expressed as: v i = Sentence-BERT(h), where v i is a high-dimensional vector, and store v i into the entity. At the same time, attach statements or attribute information about the entity to the triple to enrich the structured knowledge, and store this structured knowledge into the Neo4j graph database to construct the knowledge graph. In the knowledge graph, nodes represent entities, edges represent relationships, and the triples and their attached information are stored as attributes of nodes and edges.

[0048] To further mine the potential information in the knowledge graph, use the Leiden algorithm to analyze the graph and find communities composed of closely related entity groups. The objective function of the Leiden algorithm is expressed as:

[0049]

[0050] In the formula, A ij represents the edge weight between nodes i and j, k i and, k j respectively represent the degrees of nodes i and j, m represents the total edge weight of the network, and δ(c i , c j ) represents 1 when nodes i and j belong to the same community, otherwise 0.

[0051] Through the Leiden algorithm, the community structure in the graph can be efficiently identified. Then, the BERT language model is used to extract and summarize the information of each community to generate a community summary.

[0052] Step 2: Conduct retrieval based on the iterative retrieval and reasoning technique

[0053] The iterative retrieval and reasoning technique realizes multi-round iterative retrieval by combining retrieval and chain reasoning, gradually expanding relevant knowledge, and finally generating a complete and accurate answer. The detailed process of retrieval enhancement based on the iterative retrieval and reasoning algorithm is as follows:

[0054] First, the user inputs the initial question Q0. Then, the named entity recognition (NER) technique is used to extract the entities in the question, and the extracted entities are retrieved through the graph database. Since the community has been constructed, not only a single entity directly related to this query needs to be searched, but also it is run on each community related to the query. The Sentence-BERT model is used to encode the initial question Q0 of the query to obtain the query vector q, and calculate the semantic vector v of the query vector q and each entity i The similarity is expressed as follows:

[0055]

[0056] According to the similarity score, the K most relevant entities are retrieved: N represents all entities in the knowledge graph, and the retriever R K returns at most K entities, denoted as This process can be expressed as where are the entities obtained from the preliminary retrieval.

[0057] Next, the system enters the multi-round iterative retrieval and reasoning process. In each round of iteration, first, the chain reasoner T combines the current question Q i and the information carried by the retrieved entities P i K to generate a new CoT sentence C i , which represents the reasoning summary of the current question and the retrieved passage and is used to guide the next round of retrieval. This process can be expressed as C i = T(P i K , Q i ). Then, the CoT sentence C i is used as the new question query Q i+1 , that is, Q i+1 = NextQuery(C i ).

[0058] Use the retriever R KRetrieve the K entities most relevant to Q from the knowledge graph and denote them as i+1 This process can be expressed as

[0059] The iterative retrieval inference algorithm is expressed as:

[0060]

[0061] where n represents the maximum number of iterations.

[0062] After each round of iteration, the system checks whether the termination condition is met. The termination condition is divided into two conditions:

[0063] Condition 1: Set a maximum number of iterations (usually set to 10 times). If the maximum number of iterations is not reached, continue with the next round of iteration. If the maximum number of iterations is reached, stop the iteration;

[0064] Condition 2: After each round of iteration, generate the current CoT sentence C i and the CoT sentence C i-1 from the previous round. Convert these two CoT sentences into semantic vectors through the pre-trained model Sentence-BERT, calculate the distance between these two semantic vectors using the Euclidean distance formula, and judge the similarity, which is expressed as follows:

[0065]

[0066] where C i-1,l and C i,l represent the quantities of the semantic vectors of C i-1 and C i on the l-th dimension respectively. o represents the dimension of the semantic vector, which is 768 in this article.

[0067] If the Euclidean distance is less than the preset threshold, it means that the answers of the two rounds of iteration are very close, indicating that the system has converged, and the iteration is stopped.

[0068] Step 3: Enhancement generation based on the Self-RAG introspection algorithm

[0069] For the enhanced part of the generation model, use the self-introspection technology of Self-RAG. The generation model takes the knowledge fragments retrieved by the iterative retrieval inference technology as the context input. For the generation model, combine the retrieved K paragraphs with the question, generate K continuations at each timestamp t, score the K continuations, and select the output of the continuation with the highest score as the final result. When evaluating the quality f(y t ,d,Critique), in addition to considering the generation probability p(y t ​|x, d, y < t), a critique score S(Critique) is also introduced, where y t is the currently generated candidate paragraph, d is the relevant knowledge retrieved from the knowledge graph, x is the initial question Q0, y is the content generated before the timestamp t, and the formula for calculating the segment score is as follows:

[0070] f(y t , d, Critique) = p(y t |x, d, y < t) + S(Critique)

[0071] In the generation model, by introducing new "introspective Tokens" as marker Tokens, the generation model will give the generation probability of each Token in the logprobs, and calculate the critique score S(Critique) based on these probabilities. In the calculation of the critique score S(Critique), G is the critique marker group, which contains {ISREL, ISSUP, ISUSE}, and each critique marker has several reflection markers. For example, ISREL includes {[Relevant], [Irrelevant]}, which respectively indicate whether the retrieved thing is relevant to the question; ISSUP includes {[Fully supported], [Partially supported], [No support]}, which respectively indicate the degree of support of the output content by the retrieved knowledge; ISUSE includes [Utility:x], which indicates the effectiveness of the generated content for answering the question, divided into 1-5 points; the score of each critique marker group within the timestamp t is denoted as S t G , and the weight hyperparameter of each critique marker group is denoted as W G , then the formula for calculating the critique score S(Critique) is as follows:

[0072]

[0073] In , is the probability of the most expected reflection marker appearing within the timestamp t, P t (r i ) is the probability of generating all possible markers within the timestamp t, N G is the number of reflection markers in the critique marker group, then the formula for

[0074]

[0075] Through the self-reflection technology and multi-faceted criticism mechanism of Self-RAG, it is possible to improve the factuality and quality of the language model without sacrificing its diversity.

[0076] Step 4: Construct an iterative retrieval and reasoning knowledge graph RAG framework for the self-introspection mechanism

[0077] Combine the iterative retrieval and reasoning technology with the Self-RAG self-introspection algorithm to construct a new knowledge graph RAG framework. The core idea of this framework is to obtain high-quality knowledge fragments through iterative retrieval and reasoning, and optimize the generation process through the Self-RAG self-introspection algorithm, so as to generate a high-quality final answer.

[0078] In the retrieval part, the system uses the iterative retrieval and reasoning technology for iterative retrieval and reasoning. After the user inputs the initial question Q0, the system passes through the retriever R K Obtain relevant paragraphs from the knowledge source, and combine the chain reasoning engine T to generate new queries, gradually expanding the relevant knowledge fragments. This process continuously optimizes the retrieval results through multiple rounds of iteration to ensure that the generation model can obtain high-quality context inputs. The core of retrieval enhancement lies in dynamically expanding the knowledge scope through iterative retrieval and reasoning, so as to provide more comprehensive and accurate information support for the generation model.

[0079] In the generation part, the system introduces the Self-RAG self-introspection algorithm to enhance the generation model through the self-reflection mechanism. The generation model uses the retrieved paragraph set and the user's question as context inputs, generates multiple continuation candidates at each time stamp t, and dynamically evaluates them. By introducing "self-introspection Tokens" and the critique score S(Critique), the generation model can evaluate the relevance, supportiveness, and practicality of the content during the generation process, so as to select the optimal continuation as the output. The core of generation enhancement lies in dynamically optimizing the quality of the generated content through the self-reflection mechanism to ensure that the generated results are both factual and diverse.

[0080] Finally, the system combines the high-quality paragraph set obtained by retrieval enhancement and the generation result optimized by generation enhancement to generate the final answer.

[0081] Example 1

[0082] Based on the method of the present invention, the experimental environment of this example is a high-performance laboratory server equipped with 8 NVIDIA GeForce RTX 4090 GPUs. This server has powerful parallel computing capabilities and can efficiently support the training and reasoning tasks of large-scale deep learning models. In this experiment, we focused on comparing the performance differences between the framework combining self-introspection technology and iterative retrieval and reasoning and the self-introspection algorithm + traditional knowledge graph RAG technology.

[0083] CombineFigure 3 As shown, on different datasets, we found that the iterative retrieval and reasoning knowledge graph RAG framework with self-introspection mechanism showed significant advantages in multiple key performance indicators. Specifically, in terms of retrieval recall rate, the iterative retrieval and reasoning knowledge graph RAG framework with self-introspection mechanism improved by an average of 15.3% compared to the self-introspection algorithm + traditional knowledge graph RAG technology. This result indicates that with the help of the iterative retrieval and reasoning enhancement mechanism, it can capture knowledge fragments related to the question more comprehensively and accurately, dynamically expand the retrieval scope, and gradually mine deeper relevant knowledge, thus significantly improving the retrieval effect.

[0084] Example 2

[0085] Based on the method of the present invention, the experimental environment of this embodiment is a high-performance laboratory server equipped with 8 NVIDIA GeForce RTX 4090 GPUs. This server has powerful parallel computing capabilities and can efficiently support the training and reasoning tasks of large-scale deep learning models. In this experiment, we focused on comparing the question-answering accuracy of the framework combining self-introspection technology and iterative retrieval and reasoning with the frameworks of different generative models (including Llama2-7B and Alpaca-7B) + iterative retrieval and reasoning technology on multiple datasets to comprehensively evaluate the performance of different generative models under the retrieval enhancement framework.

[0086] Combined with Figure 4 As shown, on different datasets, we found that the iterative retrieval and reasoning knowledge graph RAG framework with self-introspection mechanism showed significant advantages in multiple key performance indicators. Specifically, in terms of question-answering accuracy, the iterative retrieval and reasoning knowledge graph RAG framework with self-introspection mechanism achieved higher accuracy rates on all datasets compared to iterative retrieval and reasoning + Llama2-7B and iterative retrieval and reasoning + Alpaca-7B. These results indicate that the self-introspection mechanism of Self-RAG further optimizes the factual accuracy and semantic coherence of the generated results by dynamically evaluating the relevance, supportiveness, and practicality of the generated content. In contrast, although Llama2-7B and Alpaca-7B have strong generation capabilities, due to the lack of a self-introspection mechanism, the quality and accuracy of the generated results are relatively low.

[0087] As can be seen from the above, the method of the present invention enhances the knowledge retrieval process by means of iterative retrieval and reasoning techniques, combined with the way of chain reasoning, to ensure the acquisition of the most relevant knowledge paragraphs in multiple rounds of retrieval and reasoning. Through the alternating process of recursive reasoning and retrieval, the system can dynamically expand the retrieval scope, gradually explore deeper relevant knowledge, thereby optimizing the relevance and integrity of information, avoiding the retrieval of irrelevant knowledge, and significantly improving the accuracy of the final reasoning result. To further improve the quality of the generated answers, the method of the present invention introduces the Self-RAG introspection technique, which reflects on the retrieved paragraphs and the generated content through special markers (called introspection markers). These introspection markers enable the language model to perform controllable operations during the reasoning stage, so as to customize its behavior according to different task requirements. Through this self-feedback mechanism, the system can identify and fill the knowledge gaps in the generated answers, initiate new retrievals, and finally generate more accurate and complete answers. This closed-loop optimization mechanism not only realizes self-correction, but also continuously improves the accuracy and reliability of the generated answers, and is widely applicable to fields such as intelligent question answering, decision support, and automatic knowledge completion.

[0088] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent conditions of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention.

[0089] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. The knowledge graph RAG method based on iterative retrieval reasoning enhancement based on introspection mechanism is characterized by: The following steps are involved: Step 1: Construction of knowledge graph Collect and clean multi-source heterogeneous data and pre-process the data through natural language processing. Then identify entities from the text through named entity recognition, extract the relationship between entities and construct triples (h, r, e), where h is the head entity, r is the relationship, and e is the tail entity. Generate the semantic vector v of each entity through the Sentence-BERT model i , expressed as: v i = Sentence-BERT(h), where v i Stored in entities, at the same time, build a knowledge graph based on triples, use the Leiden algorithm to identify the entity community structure, and use the BERT language model to generate community summaries; Step 2: Search based on iterative search reasoning technology For the initial question Q0 input by the user, named entity recognition is used to extract the entities in the question, the extracted entities are retrieved through the graph database, the initial question Q0 is encoded using the Sentence-BERT model to obtain the query vector q, and the query vector q and the semantic vector v of each entity are calculated. i Based on the similarity score, K entities are initially retrieved Retriever K Returns at most K entities, represented as Enter multiple rounds of iterative retrieval and reasoning process until the termination condition is met. In each round of iteration, the chain reasoner T is first used to combine the current question Q i and the retrieved entity P i K Carrying information, generate a new CoT sentence C i Used to summarize the reasoning of the current question and the retrieval paragraph, represented by C i =T(P i K ,Q i ), the CoT sentence C i As a new question query Q i+1 , denoted as Q i+1 =NextQuery(C i ), using the retriever R K Retrieve and Q from knowledge graph i+1 The most relevant K entities are represented as The iterative retrieval reasoning algorithm is expressed as: In the formula, n represents the maximum number of iterations; Step 3: Enhanced generation based on the Self-RAG introspection algorithm For the generative model, K continuations are generated at each timestamp t based on the K retrieved paragraphs combined with the question, and the K continuations are scored, and the output of the scored continuation is selected as the final result; In evaluating the quality of each continuation f(y t ,d,Critique), consider the generation probability p(y t |x,d,y<t), and introduce the critical score S(Critique), where y t is the candidate paragraph currently generated, d is the relevant knowledge retrieved from the knowledge graph, x is the initial question Q0, and y is the content generated before timestamp t. The calculation formula is as follows: f(y t ,d,Critical)=p(y t |x,d,y (t)+S(Critical) In the generative model, by introducing introspective tokens as marked tokens, the generative model will give the generation probability of each token in logprobs, and calculate the critical score S (Critique) based on these probabilities; Step 4: Constructing the RAG framework of iterative retrieval and reasoning knowledge graph with introspection mechanism Combining iterative retrieval reasoning technology with the Self-RAG introspection algorithm, a new knowledge graph RAG framework is constructed. Knowledge fragments are obtained through iterative retrieval reasoning, and the generation process is optimized through the Self-RAG introspection algorithm. The optimal continuation is selected as the output to generate the final answer.

2. The knowledge graph RAG method based on iterative retrieval and reasoning enhancement based on introspection mechanism according to claim 1 is characterized by: When constructing the knowledge graph in step 1, statements or attribute information about entities are added to triples to enrich structured knowledge, and the structured knowledge is stored in the Neo4j graph database. In the knowledge graph, nodes represent entities, edges represent relationships, and triples and their additional information are stored as attributes of nodes and edges.

3. The knowledge graph RAG method based on iterative retrieval and reasoning enhancement based on introspection mechanism according to claim 1 is characterized by: There are two types of termination conditions for the multi-round iterative retrieval and reasoning process in step 2: Condition 1: Set a maximum number of iterations. If the maximum number of iterations is not reached, continue to the next round of iterations. If the maximum number of iterations is reached, stop the iteration. Condition 2: After each round of iteration, generate the current CoT sentence C i and the CoT sentence C from the previous round i-1 ,The two CoT sentences are converted into semantic vectors through the pre-trained model Sentence-BERT, and the distance between the two semantic vectors is calculated by the Euclidean distance formula to judge the similarity. If the Euclidean distance is less than the preset threshold, the iteration is stopped.

4. The knowledge graph RAG method based on iterative retrieval and reasoning enhancement based on introspection mechanism according to claim 1 is characterized by: In step 3, the calculation formula of the critical score S (Critique) is as follows: Where G represents the critical marker group, W G represents the weight hyperparameter for each critical label group, represents the score of each critical marking group at timestamp t.

5. The knowledge graph RAG method based on iterative retrieval and reasoning enhancement based on introspection mechanism according to claim 4 is characterized by: Said The calculation formula is as follows: In the formula, represents the most expected reflection mark at timestamp t The probability of occurrence, P t (r i ) represents the probability of generating all possible tags within timestamp t, N G The number of reflective marks in the critical marking group.

Citation Information

Patent Citations

  • Knowledge graph construction method and management system based on large language model

    CN117556054A

  • Knowledge graph automatic construction method based on self-check retrieval enhancement generation and instruction expansion

    CN117591677A

  • Large model reasoning method and device based on knowledge graph retrieval enhancement

    CN118939783A

  • Intelligent traditional Chinese medicine auxiliary diagnosis system based on knowledge graph retrieval enhanced generation

    CN119170258A

  • Method and apparatus for automatically generating inference questions and answers

    WO2021184311A1

Cited By

  • Multi-agent collaborative query optimization method and system based on knowledge sharing information pool

    CN120930782A

  • Knowledge base engine construction method and system based on large language model

    CN121119086A

  • Knowledge base engine construction method and system based on large language model

    CN121119086B

  • Digital research and development demand automatic disassembling method based on knowledge reasoning

    CN121390756A

  • Dynamic retrieval enhancement generation method based on reinforcement learning

    CN121638478A