Intelligent agent-based big language model retrieval enhancement generation system and method

Through the multi-layer structure of the agent and multiple iterative decomposition technology, the timeliness and model illusion problems of large language models in the processing of complex problems are solved, accurate and comprehensive answers are generated, and the autonomy and security of the system are improved.

CN120470088APending Publication Date: 2025-08-12ECCOM NETWORK SYST CO LTD +1

Patent Information

Application Number
CN202510546403.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Large language model (LLM) has problems with timeliness and model illusions when dealing with complex problems. Traditional retrieval enhancement generation (RAG) technology cannot generate accurate answers and cannot meet user needs.

Method used

Adopt a multi-layer structure based on the agent, including the planning layer, execution layer, answer detection module and dynamic decision module, and decompose complex tasks through multiple iterations, generate atomic queries, call external knowledge base data, and perform multi-dimensional detection and adaptive adjustment to achieve collaborative optimization.

Benefits of technology

It improves the application performance and reliability of large language models in complex scenarios, generates accurate and comprehensive results, and improves user satisfaction and system autonomy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470088A_ABST
    Figure CN120470088A_ABST
Patent Text Reader

Abstract

The invention provides an agent-based large language model retrieval enhancement generation system and method, and the system comprises a planning layer which is used for receiving user query, carrying out the multi-round iterative decomposition of a complex task through a task planning agent, and generating an atomic query or a direct response; the execution layer is used for executing the atomic query generated by the planning layer in parallel, calling a search module to obtain external knowledge base data, and caching an intermediate result through a memory module; the answer detection module is used for performing multi-dimensional detection on the generated result, including preference, accuracy, integrity and logicality; the dynamic decision-making module is used for adaptively adjusting a subsequent retrieval strategy and a task planning process according to a detection result and user feedback; and a cross-layer interaction mechanism enables the planning layer and the execution layer to realize collaborative optimization through context sharing and iterative feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models and retrieval enhancement generation, and in particular to an agent-based large language model retrieval enhancement generation system and method. Background Art

[0002] With the explosive growth of AI computing power, large language models have become a research focus. Large language models (LLMs) with high parameter counts have been continuously released and widely used in various industries, with the most popular application being intelligent question-answering assistants. Due to the limitations of artificial neural networks and deep learning underlying technologies, the knowledge reserves of LLMs are fixed in the form of model parameters after training. Due to the limited coverage of the training corpus, LLMs have two major limitations: (1) timeliness. LLMs cannot update internal prior knowledge at any time. For new factual knowledge, they can only make inferences and summaries based on existing knowledge (model capabilities); (2) model hallucinations. The model makes inferences based on statistical laws in the training data, but sometimes produces inaccurate or absurd results due to data bias, knowledge boundaries, or insufficient understanding of context.

[0003] Traditional RAG technology introduces an external knowledge base as the knowledge source for LLM. Specifically, when a user asks a question, the system first searches the knowledge base. It then filters, sorts, and assembles the retrieved data fragments (chunks) and passes them to the LLM as context along with the user input. This solves the problem of the model being unable to generate the latest content due to timeliness. At the same time, context can effectively reduce model hallucinations. Traditional RAG can already generate good answers to simple questions. Such questions usually obtain complete contextual information through a single query and can be effectively answered without multiple rounds of induction. However, when faced with complex questions, traditional RAG may not be able to generate accurate answers.

[0004] Patent application document CN118708678A discloses a retrieval enhancement generation system and method with intelligent optimization and security enhancement, including: a retrieval generation module, a performance evaluation module, a security protection module and an interactive module; the interactive module is used to integrate the retrieval results, the text generated by the large language model, the optimization suggestions and the text reviewed by the security protection module and output them to the user, and the user can adjust the generation parameters or process the generated text according to the optimization suggestions. Finally, the user adjusts or resubmits the task based on the results and suggestions fed back by the system, and the system performs iterative optimization based on user feedback; by real-time evaluation of text quality and retrieval efficiency during the retrieval enhancement generation process, while ensuring that the generated content meets security and user needs, it provides users with high-quality, efficient, safe and reliable text generation services. However, this patent cannot completely solve the current technical problems, nor can it meet the needs of the present invention. Summary of the Invention

[0005] In view of the defects in the prior art, the purpose of the present invention is to provide an agent-based large language model retrieval enhancement generation system and method.

[0006] The agent-based large language model retrieval enhancement generation system provided by the present invention includes:

[0007] The planning layer receives user queries and uses the task planning agent to perform multiple rounds of iterative decomposition of complex tasks to generate atomic queries or direct responses.

[0008] The execution layer is used to execute the atomic queries generated by the planning layer in parallel, call the search module to obtain external knowledge base data, and cache the intermediate results through the memory module;

[0009] The answer detection module is used to perform multi-dimensional detection on the generated results, including preference, accuracy, completeness and logic;

[0010] Dynamic decision-making module, which adaptively adjusts subsequent retrieval strategies and task planning processes based on detection results and user feedback;

[0011] The cross-layer interaction mechanism enables the planning layer and the execution layer to achieve collaborative optimization through context sharing and iterative feedback.

[0012] Preferably, the task planning agent includes:

[0013] The multi-round decomposition unit, based on the reasoning capabilities of the large language model, breaks down complex queries into multiple atomic queries and dynamically adjusts the decomposition strategy based on historical search results;

[0014] A termination determination unit is used to determine whether the iteration termination conditions are met, including an information integrity threshold, a maximum number of iterations, or a user-initiated termination instruction;

[0015] The tool call interface supports three methods: predefined prompt words, function call or tool call to guide LLM to generate planning results.

[0016] Preferably, the memory module includes:

[0017] Structured cache unit, which stores user queries, atomic queries, search results, and intermediate answers in a standardized format;

[0018] The context assembly unit automatically associates historical cache data with the current input each time the LLM is called to build a dynamic context;

[0019] Feedback learning unit optimizes caching strategy and retrieval priority based on user satisfaction ratings of answers.

[0020] Preferably, the answer detection module implements multi-dimensional detection in the following manner:

[0021] Preference detection: associate historical user feedback data, search for positive / negative feedback records for similar questions, and generate a preference confidence score;

[0022] Accuracy testing: comparing the key facts of search results with the generated answers, and quantifying the deviation value through semantic matching algorithms;

[0023] Completeness check: analyze whether the answer covers all sub-results of the atomic query and verify the closure of the logical chain;

[0024] Logicality testing: evaluating the causal coherence and contradictions of the answers through a pre-trained logical reasoning model.

[0025] Preferably, the search module includes:

[0026] The offline processing unit divides and vectorizes the documents, and builds a semantic index that is stored in the vector database;

[0027] The real-time query unit vectorizes user queries or atomic queries and retrieves the top-K relevant data blocks from the vector database through similarity matching;

[0028] Cross-modal extension interface supports structured and unstructured data retrieval that integrates text, images, and tables.

[0029] Preferably, the dynamic decision module includes:

[0030] Feedback optimization unit, which automatically adjusts the weight distribution and query granularity of subsequent retrieval based on the answer detection results;

[0031] User profile unit, which records user preferences and behavior patterns for personalized task planning and answer generation;

[0032] The exception handling unit triggers re-planning or manual intervention when contradictory or low-confidence results are detected.

[0033] The agent-based large language model retrieval enhancement generation method provided by the present invention includes:

[0034] Step S1: Receive user input query, perform multiple rounds of iterative decomposition through the task planning agent, and generate atomic query or direct response;

[0035] Step S2: execute atomic queries in parallel, call the search module to obtain external knowledge base data, and cache the results in the memory module;

[0036] Step S3: If the iteration termination condition is not met, re-plan the subtasks based on the cached results and return to step S2; if the condition is met, integrate all sub-results to generate the final answer;

[0037] Step S4: Perform multi-dimensional testing on the final answer. If it fails the test, re-planning is triggered. If it passes, the answer is output and the user feedback is updated to the memory module.

[0038] Preferably, the multiple rounds of iterative decomposition in step S1 include:

[0039] Parse the structured JSON instructions output by LLM into atomic query lists or direct answers;

[0040] When parsing into atomic queries, each query is assigned a unique identifier and a priority weight;

[0041] When parsed into direct answers, historical search data is correlated to generate context-enhanced responses.

[0042] Preferably, step S2 includes:

[0043] Documents are segmented and vectorized, and semantic indexes are constructed and stored in a vector database. User queries or atomic queries are vectorized, and the top-K relevant data blocks are retrieved from the vector database through similarity matching.

[0044] User queries, atomic queries, search results, and intermediate answers are stored in a standardized format. Each time the LLM is called, historical cached data is automatically associated with the current input to build a dynamic context. Cache strategies and retrieval priorities are optimized based on the user's satisfaction rating of the answer.

[0045] Preferably, the multi-dimensional detection in step S4 includes:

[0046] Set the confidence threshold for each detection dimension. If the score of any dimension is lower than the threshold, it will be marked as failed.

[0047] Generate correction suggestions for unsuccessful answers, including supplementing search keywords, adjusting query order, or expanding the scope of the knowledge base;

[0048] The detection results are correlated with user feedback for analysis and iterative optimization of the training answer detection model.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] (1) This paper uses the reasoning and tool execution capabilities of the Large Language Model (LLM) to construct an intelligent agent that automatically processes complex queries, thereby forming the ability to plan and reflect on tasks. It iteratively plans independent subtasks for complex problems and guides the LLM to generate accurate and comprehensive results by integrating the query results of each subtask.

[0051] (2) This invention achieves automatic processing and efficient decomposition of complex queries by constructing an intelligent agent with task planning and reflection capabilities, overcoming the limitations of traditional RAG technology in processing multi-step queries, inductive analysis, and comprehensive problems. Through the agent's dynamic decision-making, task decomposition, and adaptive learning and feedback optimization mechanism, complex problems are iteratively planned into independent subtasks, and the query results of each subtask are integrated, thereby guiding the large language model to generate accurate, comprehensive, and user-friendly results.

[0052] (3) This invention is dedicated to improving the application performance and reliability of large language models in complex scenarios, promoting their efficient implementation, and providing innovative technical solutions for building knowledge service systems with complex reasoning capabilities, real-time interactivity, and high credibility, while promoting the development of human-computer collaborative intelligence in a more autonomous and secure direction.

[0053] (4) The present invention designs an AgenticRAG system with a planning-execution two-layer structure, and gives a detailed description of the function and structure of each module in the system. Finally, the overall algorithm flow is described. As the third paradigm after traditional RAG and GraphRAG, AgenticRAG can effectively solve the pain points of the former two that are unable to face complex problems and multi-step reasoning scenarios, and improve the performance of LLM in various intelligent question-answering scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0055] Figure 1 This is the overall framework diagram of the system;

[0056] Figure 2 Plan a flow chart for the task;

[0057] Figure 3 This is the memory module structure diagram;

[0058] Figure 4 Flowchart for answer detection;

[0059] Figure 5 It is the structure diagram of the search module;

[0060] Figure 6 The flowchart of the overall algorithm is shown in Figure 2. DETAILED DESCRIPTION

[0061] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0062] Example

[0063] The present invention provides an agent-based large language model retrieval enhancement generation method, comprising:

[0064] First, architecture design

[0065] like Figure 1 The overall framework of AgenticRAG can be divided into the planning layer and the execution layer. The planning layer agent is responsible for understanding, analyzing and decomposing the task, and then generates the corresponding atomic query or outputs the complete response result through LLM. The execution layer agent detects and judges the response result, outputs the final result or continues iterative optimization.

[0066] Planning layer: Decomposes complex tasks through multiple rounds of task planning. Specifically, in each round, user queries are planned, and the tasks to be executed in the current round are planned based on the complexity of the task and the search results of the previous round. For complex questions or questions with missing information, an iterative search plan is executed to plan the atomic query nodes that can be executed in the current round; for simple questions or questions with complete information, answers are directly generated.

[0067] The execution layer searches multiple query nodes generated by the planning layer in parallel, consolidates the search results, and returns to the planning layer to continue decomposing the task. The memory module is also called to maintain and cache the input and output of the planning process, as well as the search data.

[0068] Second, module function description

[0069] (1) Task Planning Agent

[0070] like Figure 2, providing iterative task planning capabilities for the current user query. When insufficient information is available, it continues planning subtasks, known as atomic queries. When the information meets or reaches the maximum iteration limit, it terminates task planning and outputs the answer. Based on the LLM's support for tool calls, it guides the LLM in executing planned tasks through predefined prompts, functions, and tools. The context includes the cached search data for each round (empty in the initial round) and the original user query. By parsing the LLM output, it decides whether to continue planning the task or directly output the answer.

[0071] (2) Memory module

[0072] like Figure 3 The memory module maintains historical context and processes and caches the Q&A data and intermediate data for each round in a standardized format. Each time the LLM is called, the Q&A data for each round is retrieved from the cache queue and assembled with the current user input to form a message, obtaining structured data through the LLM.

[0073] (3) Answer Detection Agent

[0074] like Figure 4 , providing the function of detecting the results generated by LLM, and the detection dimensions include four aspects: (1) preference, (2) accuracy, (3) completeness, and (4) logic. Each detection dimension predefines a corresponding prompt template. When detecting the answer, it is necessary to associate the user's historical feedback data and the search context information. For the user's historical feedback data, the search template uses the original user query as input, searches the data marked in the user's history (user marked positive feedback or negative feedback), obtains similar questions and answers, and then uses the searched historical question and answer results, the original user query and the current answer to build a context as the input of the large language model to generate the preference detection result; for the cached search data, it is used to build a context with the original user query and the current answer as the input of the large language model to generate the accuracy, completeness and logic detection results.

[0075] (4) Search module

[0076] like Figure 5 The search module provides offline document processing, storage, and real-time query capabilities. During the offline phase, documents are parsed and segmented through preprocessing. The embedding model is then used to generate semantic vectors for the offline documents. Finally, an index is constructed for these document fragments and stored in a vector database. During the real-time query phase, user queries are first vectorized, and language vectors are generated using the embedding model. Then, vector similarity is used to search for and return the K most similar data blocks in the vector database.

[0077] Third, the overall algorithm process

[0078] like Figure 6 , algorithm flow:

[0079] Step 1, Task Planning: Obtain the user input query, obtain the current search results from the memory module (initially empty), assemble them into a context, construct a task planning task (in the form of prompt or tool call, etc.), and call the LLM to generate the response result;

[0080] Step 2: Analyze the planning results: Analyze the LLM generated results. If the LLM output is an atomic query, proceed to step 3; if the LLM output is an answer, proceed to step 5.

[0081] Step 3: Execute atomic query: obtain atomic query input, call the search module to obtain search results, and call the memory module to cache the search results;

[0082] Step 4, re-planning: Determine the termination condition of the planning iteration. If it is not met, re-execute step 1. When the termination condition is met, obtain historical data from the memory module, assemble it with the user query into context, and directly call the LLM to generate the answer.

[0083] Step 5, Answer Verification: The user's original query and answer are retrieved, the memory module is invoked to obtain search results for each round, and the search module is invoked to obtain historical feedback data for similar questions and answers. The test iteration termination condition is determined. If it is not met, the answer is verified; otherwise, it is directly output. During answer verification, the above data is assembled into a context, and the LLM is invoked to perform a test on four dimensions: preference, accuracy, completeness, and logic. If the test passes, the final answer is directly output. If the test fails, the memory module is invoked to cache the test result, and then the process proceeds to Step 4.

[0084] Algorithm Description:

[0085] (1) Task planning: Based on the characteristics of LLM, there are three ways to guide LLM to execute task planning: (1) prompt; (2) function call; (3) tools call. The prompt template format is as follows:

[0086] ```md

[0087] ##Task

[0088] {{task info}}

[0089] ##Context

[0090] {{search result}}

[0091] {{check result}}

[0092] ##Output Format

[0093] {{output format}}

[0094] Input

[0095] ##{{query}}

[0096] ```

[0097] {{}} tags embed corresponding text information:

[0098] task info: task description;

[0099] search result: search results;

[0100] check result: answer detection result;

[0101] output format: subtask output JSON format;

[0102] query: user input query;

[0103] (2) Answer detection: preset a prompt word template for each detection dimension, guide LLM to output a confidence score s∈(0, 1), set a threshold t, and when s>=t, it is considered to have passed the specified detection dimension.

[0104] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.

[0105] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. An agent-based large language model retrieval enhancement generation system, characterized by: include: The planning layer receives user queries and uses the task planning agent to perform multiple rounds of iterative decomposition of complex tasks to generate atomic queries or direct responses. The execution layer is used to execute the atomic queries generated by the planning layer in parallel, call the search module to obtain external knowledge base data, and cache the intermediate results through the memory module; The answer detection module is used to perform multi-dimensional detection on the generated results, including preference, accuracy, completeness and logic; Dynamic decision-making module, which adaptively adjusts subsequent retrieval strategies and task planning processes based on detection results and user feedback; The cross-layer interaction mechanism enables the planning layer and the execution layer to achieve collaborative optimization through context sharing and iterative feedback.

2. The agent-based large language model retrieval enhancement generation system according to claim 1, characterized in that: The task planning agent includes: The multi-round decomposition unit, based on the reasoning capabilities of the large language model, breaks down complex queries into multiple atomic queries and dynamically adjusts the decomposition strategy based on historical search results; A termination determination unit is used to determine whether the iteration termination conditions are met, including an information integrity threshold, a maximum number of iterations, or a user-initiated termination instruction; The tool call interface supports three methods: predefined prompt words, function call or tool call to guide LLM to generate planning results.

3. The agent-based large language model retrieval enhancement generation system according to claim 1, characterized in that: The memory module includes: Structured cache unit, which stores user queries, atomic queries, search results, and intermediate answers in a standardized format; The context assembly unit automatically associates historical cache data with the current input each time the LLM is called to build a dynamic context; Feedback learning unit optimizes caching strategy and retrieval priority based on user satisfaction ratings of answers.

4. The agent-based large language model retrieval enhancement generation system according to claim 1, characterized in that: The answer detection module implements multi-dimensional detection in the following ways: Preference detection: associate historical user feedback data, search for positive / negative feedback records for similar questions, and generate a preference confidence score; Accuracy testing: comparing the key facts of search results with the generated answers, and quantifying the deviation value through semantic matching algorithms; Completeness check: analyze whether the answer covers all sub-results of the atomic query and verify the closure of the logical chain; Logicality testing: evaluating the causal coherence and contradictions of the answers through a pre-trained logical reasoning model.

5. The agent-based large language model retrieval enhancement generation system according to claim 1, characterized in that: The search module includes: The offline processing unit divides and vectorizes the documents, and builds a semantic index that is stored in the vector database; The real-time query unit vectorizes user queries or atomic queries and retrieves the top-K relevant data blocks from the vector database through similarity matching; Cross-modal extension interface supports structured and unstructured data retrieval that integrates text, images, and tables.

6. The agent-based large language model retrieval enhancement generation system according to claim 1, characterized in that: The dynamic decision module includes: Feedback optimization unit, which automatically adjusts the weight distribution and query granularity of subsequent retrieval based on the answer detection results; User profile unit, which records user preferences and behavior patterns for personalized task planning and answer generation; The exception handling unit triggers re-planning or manual intervention when contradictory or low-confidence results are detected.

7. A large language model retrieval enhancement generation method based on an intelligent agent, characterized in that: include: Step S1: Receive user input query, perform multiple rounds of iterative decomposition through the task planning agent, and generate atomic query or direct response; Step S2: execute atomic queries in parallel, call the search module to obtain external knowledge base data, and cache the results in the memory module; Step S3: If the iteration termination condition is not met, re-plan the subtasks based on the cached results and return to step S2; if the condition is met, integrate all sub-results to generate the final answer; Step S4: Perform multi-dimensional testing on the final answer. If it fails the test, re-planning is triggered. If it passes, the answer is output and the user feedback is updated to the memory module.

8. The agent-based large language model retrieval enhancement generation method according to claim 7, characterized in that: The multiple rounds of iterative decomposition in step S1 include: Parse the structured JSON instructions output by LLM into atomic query lists or direct answers; When parsing into atomic queries, each query is assigned a unique identifier and a priority weight; When parsed into direct answers, historical search data is correlated to generate context-enhanced responses.

9. The agent-based large language model retrieval enhancement generation method according to claim 7, characterized in that: The step S2 comprises: Documents are segmented and vectorized, and semantic indexes are constructed and stored in a vector database. User queries or atomic queries are vectorized, and the top-K relevant data blocks are retrieved from the vector database through similarity matching. User queries, atomic queries, search results, and intermediate answers are stored in a standardized format. Each time the LLM is called, historical cached data is automatically associated with the current input to build a dynamic context. Cache strategies and retrieval priorities are optimized based on the user's satisfaction rating of the answer.

10. The agent-based large language model retrieval enhancement generation method according to claim 7, characterized in that: The multi-dimensional detection in step S4 includes: Set the confidence threshold for each detection dimension. If the score of any dimension is lower than the threshold, it will be marked as failed. Generate correction suggestions for unsuccessful answers, including supplementing search keywords, adjusting query order, or expanding the scope of the knowledge base; The detection results are correlated with user feedback for analysis and iterative optimization of the training answer detection model.

Citation Information

Patent Citations

  • Intelligent optimization and security enhancement retrieval enhancement generation system and method

    CN118708678A

Cited By

  • Multi-intelligent-agent search enhancement generation method and system

    CN120804377A

  • A multi-agent retrieval augmented generation method and system

    CN120804377B

  • Database interaction method and device based on multi-agent cooperation, equipment and medium

    CN120821738A

  • Security event association analysis and question and answer method and system and medium

    CN120910224A

  • A security event correlation analysis and question answering method, system and medium

    CN120910224B