Retrieval enhancement generation optimization method and device and electronic equipment

By using a question-answering system with four lightweight intelligent agents working together, and by using different large language models to process and review questions, this system solves the problems of insufficient utilization of multimodal data and inaccurate answer results in traditional retrieval enhancement generation technology, and achieves efficient and flexible question-answering service.

CN121501946APending Publication Date: 2026-02-10CHINA MOBILE (XIONGAN) ICT CO LTD +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511628663.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional search enhancement generation techniques cannot fully utilize multimodal data, struggle to capture specific relationships between documents, lack depth and logic in their results, and lack a result verification mechanism, making them prone to generating incorrect answers and lacking flexibility.

Method used

The system utilizes four lightweight agents working together: the first, second, third, and fourth large language models, which are responsible for reasoning, information filtering, reasoning planning, and review, respectively. This optimizes the question-and-answer process, and the fourth agent reviews the target answer to ensure accuracy and flexibility.

Benefits of technology

It significantly improves the ability to handle complex problems and the efficiency and accuracy of information retrieval, and the answer results are both highly accurate and highly flexible, comprehensively enhancing the reliability of the question-and-answer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501946A_ABST
    Figure CN121501946A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation optimization method and device and electronic equipment. The method comprises the steps of obtaining a to-be-answered question input by a user; calling a first agent, a second agent or a third agent to answer the to-be-answered question, and generating a target answer result; using a fourth agent to check the target answering result, and if the check result is passed, outputting the target answering result as a final answering result; according to the technical scheme, through cooperative operation of the four lightweight intelligent agents and optimization of the whole question and answer process, the complex question processing capacity and the information retrieval efficiency and accuracy are remarkably improved, the method depends on a multi-strategy self-adaptive processing mechanism and a structure correction mechanism, the answer result has high accuracy and high flexibility, and the method is suitable for large-scale popularization and application. And the reliability of the question-answer service is comprehensively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, and electronic device for search enhancement and generation optimization. Background Technology

[0002] In the context of the deep application of artificial intelligence, users have higher requirements for the response speed, accuracy of answer results and timeliness of knowledge of question answering systems. Traditional large-scale language models are limited by the time window of training data and suffer from the problem of forgetting professional knowledge. Retrieval Augmentation Generation (RAG) technology provides contextual support for LLM by retrieving relevant knowledge fragments from external document libraries, which makes up for this deficiency and has become an important technical solution for real-time knowledge question answering scenarios.

[0003] The standard RAG, serving as the foundational architecture for retrieval-enhanced generation technology, primarily comprises two core phases: the retrieval phase and the generation phase. In the retrieval phase, user questions and external document fragments are transformed into feature vectors. Similarity is calculated using these vectors to filter relevant document fragments. In the generation phase, relevant document fragments are input into the LLM (Learning Resource Management) system along with the user question, and an answer is generated by combining contextual knowledge with general LLM knowledge.

[0004] Standard RAG relies on vector retrieval of document blocks, which cannot fully utilize multimodal data such as images and tables, easily leading to information loss. Relying solely on textual semantic similarity makes it difficult to capture specific relationships between documents, resulting in solutions lacking depth and logical coherence. Its multi-hop reasoning and complex logic processing capabilities are limited, and the lack of a result verification mechanism makes it prone to generating incorrect solutions due to noise. It cannot effectively identify non-knowledge base questions, relies on fixed responses, and lacks flexibility. Summary of the Invention

[0005] This invention provides a retrieval enhancement generation optimization method, device, and electronic device. Through the collaborative operation of four lightweight intelligent agents, the entire question-and-answer process is optimized, which not only significantly improves the ability to handle complex questions and the efficiency and accuracy of information retrieval, but also relies on a multi-strategy adaptive processing mechanism and a structure correction mechanism to make the answer results have both high accuracy and high flexibility, thus comprehensively enhancing the reliability of the question-and-answer service.

[0006] According to one aspect of the present invention, a retrieval enhancement generation optimization method is provided, the method comprising:

[0007] Get the user's input of the unanswered question;

[0008] The system invokes a first intelligent agent, a second intelligent agent, or a third intelligent agent to answer the question and generate a target answer result; wherein, the first intelligent agent is generated based on a first large-scale language model; the second intelligent agent is generated based on a second large-scale language model; and the third intelligent agent is generated based on a third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other.

[0009] The fourth intelligent agent reviews the target solution result. If the review result is satisfactory, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on a fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first, second, and third large-scale language models.

[0010] According to another aspect of the present invention, a retrieval enhancement generation optimization apparatus is provided, the apparatus comprising:

[0011] The unanswered question acquisition module is used to acquire unanswered questions input by the user;

[0012] The target solution result generation module is used to call a first intelligent agent, a second intelligent agent, or a third intelligent agent to answer the question to be answered and generate a target solution result; wherein, the first intelligent agent is generated based on a first large-scale language model; the second intelligent agent is generated based on a second large-scale language model; the third intelligent agent is generated based on a third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other.

[0013] The final solution result determination module is used to review the target solution result using a fourth intelligent agent. If the review result is passed, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on a fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first, second, and third large-scale language models.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the retrieval enhancement generation optimization method according to any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the retrieval enhancement generation optimization method according to any embodiment of the present invention.

[0017] The technical solution of this invention obtains the user-input question, calls a first, second, or third intelligent agent to answer the question, generates a target answer, and then uses a fourth intelligent agent to review the target answer. If the review is successful, the target answer is output as the final answer. This technical solution optimizes the entire question-and-answer process through the collaborative operation of four lightweight intelligent agents, significantly improving the ability to handle complex questions and the efficiency and accuracy of information retrieval. Furthermore, relying on a multi-strategy adaptive processing mechanism and a structural correction mechanism, the answer results possess both high accuracy and high flexibility, comprehensively enhancing the reliability of the question-and-answer service.

[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of the retrieval enhancement generation optimization method provided in Embodiment 1 of the present invention;

[0021] Figure 2 This is a diagram of a multi-agent collaborative architecture provided in Embodiment 1 of this application;

[0022] Figure 3 This is a schematic diagram of the retrieval enhancement generation optimization process provided in Embodiment 2 of the present invention;

[0023] Figure 4 This is a schematic diagram of the retrieval enhancement generation optimization device provided in Embodiment 3 of the present invention;

[0024] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the retrieval enhancement generation optimization method of the present invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] Example 1

[0028] Figure 1 This is a flowchart of a retrieval enhancement generation optimization method according to Embodiment 1 of the present invention. This embodiment is applicable to retrieval enhancement generation optimization. The method can be executed by a retrieval enhancement generation optimization device, which can be implemented in hardware and / or software and can be configured in a device. For example, the device can be a backend server or other device with communication and computing capabilities. Figure 1 As shown, the method includes:

[0029] S110. Obtain the unanswered question input by the user.

[0030] In this scheme, an unanswered question refers to a question that has been asked but not yet answered. The unanswered question is retrieved in response to user input.

[0031] S120. Invoke the first intelligent agent, the second intelligent agent, or the third intelligent agent to answer the question to be answered and generate the target answer result; wherein, the first intelligent agent is generated based on the first large-scale language model; the second intelligent agent is generated based on the second large-scale language model; the third intelligent agent is generated based on the third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other.

[0032] In this scheme, the first agent is a reasoning agent, the second agent is an information filtering agent, and the third agent is a reasoning computation agent. Each agent is responsible for different tasks, and they work together to achieve efficient question-and-answer processing.

[0033] Specifically, the reasoning agent is responsible for assessing the complexity of the problem to be solved and whether retrieval is necessary, determining the problem-solving strategy. The information filtering agent is responsible for processing and filtering the retrieved information, extracting content suitable for the LLM (Large Language Model). The reasoning planning agent is responsible for evaluating the reasoning progress in the multi-step reasoning strategy, deciding whether further retrieval is needed or whether to directly generate the solution.

[0034] In this embodiment, efficient collaboration among intelligent agents can be achieved through the following mechanisms: Direct response mechanism: For general knowledge-based questions, the first large-scale language model directly generates the answer. Single retrieval mechanism: For simple questions, the answer is obtained through a single retrieval-filtering loop, i.e., the answer is generated through the second large-scale language model. Multi-step reasoning mechanism: For complex questions, information is gradually gathered through multiple retrieval-filtering loops, and the answer is finally generated, i.e., the answer is generated through the third large-scale language model.

[0035] Furthermore, Figure 2 This is a diagram of the multi-agent cooperative architecture provided in Embodiment 1 of this application, such as... Figure 2 As shown, this system is built on a multi-agent collaborative architecture, with its core consisting of a reasoning agent, an information filtering agent, and a reasoning planning agent. Each agent has a clear division of labor and works closely together. After receiving a user's input question, the reasoning agent first evaluates the question and determines the processing strategy. If retrieval support is needed, the information filtering agent extracts key information. For complex questions requiring multiple steps, the reasoning planning agent coordinates the reasoning process. The entire architecture aims to efficiently respond to and accurately solve diverse user questions. Through the orderly interaction of the agents, it optimizes and upgrades traditional Retrieval-Augmented Generation (RAG) technology.

[0036] In this embodiment, by invoking the first, second, or third intelligent agent, the first, second, or third large-scale language model they are equipped with can be used to predict and analyze the problem to be solved, and finally generate the target solution.

[0037] Among them, the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other, and are all trained based on user-input questions.

[0038] S130. The fourth intelligent agent reviews the target solution result. If the review result is passed, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on the fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first large-scale language model, the second large-scale language model and the third large-scale language model.

[0039] In this scheme, the fourth agent is the review agent, which is responsible for reviewing the final generated solution to ensure its matching degree and accuracy with the user's question.

[0040] In this embodiment, as Figure 2 As shown, the reviewing agent reviews the final generated target solution result. If the review result is passed, the target solution result is output as the final solution result to ensure the quality of the final solution result.

[0041] Optionally, a fourth intelligent agent is used to review the target solution result. If the review result is satisfactory, the target solution result is output as the final solution result, including:

[0042] The matching degree between the target solution result and the question to be solved is obtained using a fourth intelligent agent, as well as the accuracy of the target solution result;

[0043] If the matching degree reaches a preset first threshold and the accuracy reaches a preset second threshold, then the target solution result will be output as the final solution result.

[0044] The first and second thresholds are set based on the review requirements of the target solution results.

[0045] In this scheme, after the reviewing agent receives the target solution, it will employ multiple evaluation methods such as semantic matching and logical verification. Semantic matching accurately determines the degree of matching between the target solution and the question to be answered; logical verification verifies the accuracy of the logical structure and reasoning process of the target solution.

[0046] In this embodiment, if the matching degree corresponding to the target answer result reaches a preset first threshold and the accuracy reaches a preset second threshold, the review agent will output the target answer result as the final answer result and feed it back to the user.

[0047] Furthermore, if the review confirms that the matching degree corresponding to the target answer result does not reach the preset first threshold, and / or the accuracy does not reach the preset second threshold, it indicates that the target answer result has problems such as inaccurate content or irrelevance to the question to be answered. At this time, the target answer result is fed back to the reasoning agent, driving the system to restart the reasoning process until a target answer result that meets the quality standards is generated, thereby effectively improving the reliability of the question answering system's output results and user satisfaction.

[0048] By introducing a multi-step reasoning strategy, the complexity of a problem can be dynamically assessed, and the multi-step reasoning process can be automatically triggered when necessary, thereby efficiently handling various complex problems.

[0049] The technical solution of this invention obtains the user-input question, calls a first, second, or third intelligent agent to answer the question, generates a target answer, and then uses a fourth intelligent agent to review the target answer. If the review is successful, the target answer is output as the final answer. By implementing this technical solution, the collaborative operation of four lightweight intelligent agents optimizes the entire question-and-answer process, significantly improving the ability to handle complex questions and the efficiency and accuracy of information retrieval. Furthermore, relying on a multi-strategy adaptive processing mechanism and a structural correction mechanism, the answer results possess both high accuracy and high flexibility, comprehensively enhancing the reliability of the question-and-answer service.

[0050] Example 2

[0051] Figure 3 This is a schematic diagram of the retrieval enhancement generation optimization process provided in Embodiment 2 of the present invention. The relationship between this embodiment and the above embodiments is a detailed description of the target solution result generation process. Figure 3 As shown, the method includes:

[0052] S310. Obtain the unanswered question input by the user.

[0053] S320. Call the first intelligent agent to answer the question to be answered and generate the first answer result.

[0054] In this scheme, the reasoning agent is the first step in the question-answering system to process user questions. It is used to assess the complexity of the question to be answered and to determine whether the retrieval process needs to be initiated.

[0055] Specifically, the reasoning agent conducts a comprehensive analysis of the unanswered question input by the user, taking into account core elements such as the question's specificity, complexity, and clarity, and then determines whether the question falls within the existing knowledge coverage of the first large-scale language model.

[0056] In this scheme, if the question to be answered falls within the scope of general knowledge or the existing knowledge coverage of the first large-scale language model, the question to be answered is directly passed to the first large-scale language model, triggering the direct answer strategy, and the first large-scale language model uses its internal knowledge to quickly generate the first answer result.

[0057] Optionally, a first intelligent agent is invoked to answer the question to be answered, generating a first solution result, including:

[0058] If the first solution result meets the preset solution conditions, then the first solution result is determined as the target solution result.

[0059] The solution conditions are the constraints that govern the solution results.

[0060] In this scheme, if the first solution result meets the preset solution conditions, it indicates that the problem to be solved is within the existing knowledge coverage of the first large-scale language model, and the first solution result is determined as the target solution result.

[0061] By having the first intelligent agent process the questions to be answered, the solution generation cycle is significantly shortened, and the complex process of multi-step screening is simplified. This not only improves the overall efficiency and reliability of the solution, but also ensures that the output results meet both rapid response and preset quality requirements.

[0062] S330. When the first solution result does not meet the preset solution conditions, the second intelligent agent is invoked to solve the problem to be solved, and a second solution result is generated.

[0063] In this scheme, if the first solution does not meet the preset solution conditions, it means that the question to be answered may involve specific information such as the latest events or sub-fields outside the knowledge scope of the first large language model. At this time, a single search strategy will be activated, and the second agent will be called to answer the question to be answered and generate a second solution.

[0064] Specifically, the second large-scale language model in the second intelligent agent is used to solve the problem to be solved, and a second solution result is generated.

[0065] Optionally, if the first solution result does not meet the preset solution conditions, a second intelligent agent is invoked to solve the problem to be solved, generating a second solution result, including:

[0066] When the first solution result does not meet the preset solution conditions, obtain the first document data corresponding to the question to be answered;

[0067] The second intelligent agent is mobilized to answer the question to be answered based on the first document information, and a second answer result is generated.

[0068] In this embodiment, if the first solution result does not meet the preset solution conditions, it means that the question to be answered may involve specific information such as the latest events or sub-fields outside the knowledge scope of the first large-scale language model. At this time, a single search strategy will be activated, and the first intelligent agent will generate a precise query based on the core requirements of the question to obtain the corresponding first document information.

[0069] Among them, the information filtering agent is used to process and filter the first document data, and extract the key content adapted to the processing of the second large language model. Its workflow needs to be flexibly executed according to the different strategies adopted by the system.

[0070] Furthermore, under the single-retrieval strategy, the information filtering agent will combine the question to be answered with the first document, and conduct an in-depth analysis of the first document from multiple dimensions such as content relevance and information timeliness. By judging the degree of matching between the first document and the question to be answered, the first document with the strongest relevance will be accurately selected and provided to the second large-scale language model, ensuring that the information received by the second large-scale language model is targeted and effective.

[0071] By analyzing the retrieved first-level documents, the most relevant parts to the question and the current reasoning goal are extracted, reducing interference from noise. This precise information extraction ensures that the content passed to the second, larger language model is of high quality, thereby improving the accuracy of the generated answer.

[0072] Specifically, the second large-scale language model of the second intelligent agent will develop an answer to the question based on the first document data, and then generate a second answer result.

[0073] In this scheme, when the system adopts a multi-step reasoning strategy, the information filtering agent, in addition to inputting the question and retrieving documents, will also take into account the current reasoning objective for comprehensive consideration. The information filtering agent will prioritize the selection of documents that both fit the current reasoning objective and are highly relevant to the question, providing accurate support for information processing at each stage of multi-step reasoning and ensuring the coherence and accuracy of the entire reasoning process.

[0074] When the first solution does not meet the preset conditions, the system automatically matches the corresponding first document and calls the second agent to provide a second solution. This not only makes up for the limitations of a single agent's solution, but also provides accurate evidence for the solution with the help of targeted documents. This effectively improves the accuracy, comprehensiveness and adaptability of the solutions to the questions to be answered, and enhances the reliability and flexible response capability of the intelligent solution system.

[0075] Optionally, obtain first document data corresponding to the question to be answered, including:

[0076] The first document data corresponding to the question to be answered is obtained by the retrieval device in the first intelligent agent.

[0077] In this scheme, the first intelligent agent interacts with the retrieval device to retrieve the first document information corresponding to the question to be answered.

[0078] Specifically, after receiving the user's question, the first agent converts it into a vector representation. This vector is then passed to the retrieval system, which uses it to search external knowledge bases or document sets. The retrieval system uses similarity calculations, such as cosine similarity, to match the most relevant documents.

[0079] Among them, the search engine has efficient search capabilities, enabling it to quickly and accurately find information related to the question from a large amount of document data.

[0080] Through the collaborative operation of the first intelligent agent and the retrieval device, the first document information highly relevant to the question to be answered can be quickly obtained from massive document data, providing reliable data support for the accurate answering of subsequent questions and effectively improving the efficiency and accuracy of question answering.

[0081] Optionally, a second intelligent agent is invoked to answer the question based on the first document information, generating a second answer result, including:

[0082] If the second solution result meets the preset solution conditions, then the second solution result is determined as the target solution result.

[0083] The solution conditions are the constraints that govern the solution results.

[0084] In this scheme, if the second solution result meets the preset solution conditions, it indicates that the problem to be solved is within the existing knowledge coverage of the second large-scale language model, and the second solution result is determined as the target solution result.

[0085] By using a second intelligent agent to process the questions to be answered, the solution generation cycle is significantly shortened, and the complex process of multi-step screening is simplified. This improves the overall efficiency and reliability of the solution, and also ensures that the output results meet both rapid response and preset quality requirements.

[0086] S340. When the second solution result does not meet the preset solution conditions, a third agent is invoked to solve the problem to be solved, a third solution result is generated, and the third solution result is determined as the target solution result.

[0087] In this scheme, the inference planning agent is responsible for the overall control of the inference process. It continuously tracks the inference progress and makes dynamic decisions based on the actual situation: whether to continue supplementing the retrieved information or directly output the target solution.

[0088] Furthermore, during the reasoning process, the third AI will first check the documents accumulated from previous retrieval-filtering cycles to determine whether the existing information meets the requirements for answering the question. At the same time, it will strictly align with the high-level reasoning roadmap generated by the third large-scale language model, verifying step by step whether the current reasoning stage has achieved the preset goal.

[0089] In this scheme, if the first document cannot cover the information requirements of each sub-goal in the reasoning roadmap, the third agent will initiate a new retrieval command. It will generate targeted sub-queries around the missing information, initiating a new round of retrieval-filtering process to further search the second document. Only when the third agent confirms that the second document covers the core information of all sub-goals will it pass this data to the third large-scale language model. The third large-scale language model will then conduct in-depth analysis based on the integrated data, ultimately outputting a complete and accurate solution to the target problem, ensuring that complex issues are properly resolved.

[0090] For complex problems such as logical reasoning and contextual reasoning, multi-step reasoning strategies can be activated to guide the system to start a more refined and comprehensive reasoning process. Through multi-dimensional analysis and step-by-step deduction, it can accurately address these types of problems that require comprehensive consideration.

[0091] Optionally, when the second solution result does not meet the preset solution conditions, a third intelligent agent is invoked to solve the problem to be solved, generating a third solution result, and the third solution result is determined as the target solution result, including:

[0092] When the second solution result does not meet the preset solution conditions, a second document corresponding to the question to be answered is obtained; wherein, the second document is different from the first document.

[0093] A third intelligent agent is mobilized to answer the question to be answered based on the second document and the first document, generating a third answer result, and the third answer result is determined as the target answer result.

[0094] In this embodiment, if the second solution result does not meet the preset solution conditions, it means that the question to be answered may involve specific information such as the latest events or sub-fields outside the knowledge scope of the second large language model. At this time, a multi-step reasoning strategy will be launched, and the first intelligent agent will generate a precise query based on the core requirements of the question to obtain the corresponding second document information.

[0095] Specifically, after receiving the user's question, the first agent converts it into a vector representation. This vector is then passed to the retrieval system, which uses it to search external knowledge bases or document sets. The retrieval system uses similarity calculations, such as cosine similarity, to match the most relevant second-level documents.

[0096] Furthermore, the third large-scale language model in the third intelligent agent generates a third solution result by solving the problem to be solved using the second document data and the first document data, and determines the third solution result as the target solution result.

[0097] The third-party intelligent agent combines two types of differentiated data to provide a comprehensive answer and determine the target answer. This not only expands the information source dimensions for answering the problem, but also improves the comprehensiveness, accuracy, and reliability of the answer by cross-validating and supplementing multiple data. It effectively avoids the problem of answer deviation or information limitation that may occur when supported by a single data, and ensures that the problem to be answered finally obtains a high-quality answer that meets the preset requirements.

[0098] S350. The fourth intelligent agent reviews the target solution result. If the review result is passed, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on the fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first large-scale language model, the second large-scale language model and the third large-scale language model.

[0099] The technical solution of this invention involves acquiring a user-input question, calling a first intelligent agent to answer the question, and generating a first answer result. If the first answer result does not meet preset answer conditions, a second intelligent agent is called to answer the question, generating a second answer result. If the second answer result does not meet preset answer conditions, a third intelligent agent is called to answer the question, generating a third answer result, which is then determined as the target answer result. A fourth intelligent agent then reviews the target answer result; if the review is successful, the target answer result is output as the final answer result. By implementing this technical solution, the collaborative operation of four lightweight intelligent agents optimizes the entire question-and-answer process, significantly improving the ability to handle complex questions and the efficiency and accuracy of information retrieval. Furthermore, relying on a multi-strategy adaptive processing mechanism and a structural correction mechanism, the answer result possesses both high accuracy and high flexibility, comprehensively enhancing the reliability of the question-and-answer service. The system can adaptively select a direct answer strategy, a single-retrieval strategy, or a multi-step reasoning strategy based on the complexity of the question. This mechanism can flexibly handle different types of user questions, improving the system's adaptability and flexibility.

[0100] In this solution, the reasoning agent, information filtering agent, reasoning planning agent, and review agent form a closed-loop collaborative architecture, fully covering the entire process of problem evaluation, information processing, reasoning planning, and answer verification. The reasoning agent is responsible for dynamically deciding on problem-solving strategies, accurately matching direct answers, single-step retrieval, or multi-step reasoning modes; the information filtering agent efficiently filters core related information based on the reasoning objective; the reasoning planning agent controls the retrieval process of multi-step reasoning; and the review agent verifies the semantic matching degree and logical accuracy of the target answer result. The agents support bidirectional interaction and iterative cycles: when the review agent detects an answer deviation, it can directly feed back to the reasoning agent to restart the reasoning process; the reasoning planning agent can dynamically trigger new retrieval-filtering cycles based on the reasoning progress, ultimately achieving closed-loop optimization and continuous iteration of evaluation, retrieval, reasoning, and verification.

[0101] In this embodiment, a dynamic response strategy based on question complexity is employed: for general knowledge questions, the first large-scale language model directly outputs the answer, efficiently responding to basic information needs. Simple questions utilize a single-retrieval-filtering loop mechanism to quickly obtain core information and filter redundant content, accurately matching simple query requests. For complex questions, a multi-step reasoning strategy is activated, achieving full scenario coverage through multiple retrieval-filtering loops and reasoning roadmap planning, applicable to everything from common sense questions to multi-hop logical reasoning. The reasoning agent verifies the reasoning objectives stage by stage based on the reasoning roadmap generated by the large-scale language model. Sub-queries are automatically triggered to supplement key information; after all sub-objectives are completed, the information is integrated to generate the final answer. For example, when processing "how to optimize urban traffic congestion," a step-by-step retrieval can be performed according to "current situation analysis - policy cases - technical solutions," ensuring the systematic and complete nature of the answer.

[0102] Furthermore, information filtering agents: collaborative processing of structured and unstructured data.

[0103] Single-search scenario: Precisely filter highly relevant content from a document collection to quickly locate core information. Multi-step reasoning scenario: Targeted filtering of key paragraphs in documents based on the current reasoning objective (e.g., "the implementation effect of a certain policy"). Ensure that the information acquired by the large language model combines structured facts with contextual semantics, effectively improving the completeness and readability of the answer.

[0104] Inference Agents and LLM: Deeply Coupled Hybrid Inference Mechanisms

[0105] Direct Answer Strategy: LLM autonomously generates answers by calling its internal knowledge base, efficiently responding to simple queries. Retrieval Enhancement Strategy: LLM deeply participates in reasoning roadmap planning, with the reasoning agent dynamically adjusting its retrieval direction based on its reasoning logic. This achieves a hybrid reasoning model combining "symbolic reasoning (knowledge graph) + semantic reasoning (LLM)". For example, when dealing with the question "What is the difference between quantum computing and classical computing?", it retrieves structured principles through the knowledge graph while leveraging LLM to integrate cross-domain semantic interpretations, balancing the accuracy and understandability of the answer.

[0106] The reviewing agent will verify the quality of the answers from multiple dimensions:

[0107] Semantic matching: Determine whether the answer accurately addresses the core of the question. For example, when a user asks "the ethical risks of AI", the answer should revolve around key points such as "privacy leakage and algorithmic bias".

[0108] Logical verification: Check for contradictions in the reasoning process. For example, when answering "the advantages of a certain technology", avoid conflicting descriptions.

[0109] Information consistency: Compare the content of the searched documents with the answer content to ensure that the factual information is accurate and effectively avoid the problems of LLM.

[0110] If the review agent finds a problem with the answer, it will immediately feed back the problem and the corresponding answer to the reasoning agent, triggering a full restart: the reasoning agent re-evaluates the problem response strategy, the information filtering agent re-screens relevant documents, and the reasoning planning agent adjusts the reasoning steps until the correct answer is generated, forming a closed-loop quality control mechanism for verification and correction.

[0111] This solution supports a hybrid retrieval mode combining knowledge graphs and unstructured documents. Leveraging an information filtering agent, it deeply analyzes textual information (including chart descriptions and image annotations) within multimodal data, integrating the advantages of both vector and graph retrieval. It is widely adaptable to diverse scenarios such as customer service Q&A, enterprise knowledge management, and professional field research. For example, in the legal field, it can simultaneously retrieve legal texts and legal knowledge graphs to generate accurate answers that combine factual basis and logical reasoning.

[0112] The reasoning agent possesses the ability to identify non-knowledge base questions. If it encounters a query such as "What's the weather like today?" which is not included in the knowledge base, it will automatically trigger the LLM direct answer strategy. Relying on the model's built-in common sense knowledge, it generates natural and fluent responses without relying on external retrieval. The processing method is more flexible. Compared with the traditional approach that requires rigid processing through a fixed prompt, this proposal achieves policy-level intelligent adaptation.

[0113] Example 3

[0114] Figure 4 This is a schematic diagram of the retrieval enhancement generation optimization device provided in Embodiment 3 of the present invention. Figure 4 As shown, the apparatus includes: a retrieval enhancement generation optimization method, characterized in that it includes:

[0115] The unanswered question acquisition module 410 is used to acquire unanswered questions input by the user;

[0116] The target solution result generation module 420 is used to call a first intelligent agent, a second intelligent agent, or a third intelligent agent to answer the question to be answered and generate a target solution result; wherein, the first intelligent agent is generated based on a first large-scale language model; the second intelligent agent is generated based on a second large-scale language model; the third intelligent agent is generated based on a third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other.

[0117] The final solution result determination module 430 is used to review the target solution result using a fourth intelligent agent. If the review result is passed, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on a fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first large-scale language model, the second large-scale language model, and the third large-scale language model.

[0118] Optionally, the target solution result generation module 420 includes:

[0119] The first solution result generation submodule is used to call the first intelligent agent to answer the question to be answered and generate the first solution result;

[0120] The second solution result generation submodule is used to call the second intelligent agent to solve the problem to be solved and generate a second solution result when the first solution result does not meet the preset solution conditions.

[0121] The target solution result determination submodule is used to call a third agent to solve the problem to be solved when the second solution result does not meet the preset solution conditions, generate a third solution result, and determine the third solution result as the target solution result.

[0122] Optionally, the first solution result generation submodule is used for:

[0123] If the first solution result meets the preset solution conditions, then the first solution result is determined as the target solution result.

[0124] Optionally, the second solution result generation submodule is specifically used for:

[0125] When the first solution result does not meet the preset solution conditions, obtain the first document data corresponding to the question to be answered;

[0126] The second intelligent agent is mobilized to answer the question to be answered based on the first document information, and a second answer result is generated.

[0127] Optionally, the second solution result generation submodule is also used for:

[0128] The first document data corresponding to the question to be answered is obtained by the retrieval device in the first intelligent agent.

[0129] Optionally, the second solution result generation submodule is also used for:

[0130] If the second solution result meets the preset solution conditions, then the second solution result is determined as the target solution result.

[0131] Optionally, the final solution determination module 430 is specifically used for:

[0132] When the second solution result does not meet the preset solution conditions, a second document corresponding to the question to be answered is obtained; wherein, the second document is different from the first document.

[0133] A third intelligent agent is mobilized to answer the question to be answered based on the second document and the first document, generating a third answer result, and the third answer result is determined as the target answer result.

[0134] Optionally, the final solution determination module 430 is specifically used for:

[0135] The matching degree between the target solution result and the question to be solved is obtained using a fourth intelligent agent, as well as the accuracy of the target solution result;

[0136] If the matching degree reaches a preset first threshold and the accuracy reaches a preset second threshold, then the target solution result will be output as the final solution result.

[0137] The retrieval enhancement generation optimization device provided in the embodiments of the present invention can execute the retrieval enhancement generation optimization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0138] Example 4

[0139] Figure 5A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0140] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0141] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0142] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as retrieving augmented generative optimization methods.

[0143] In some embodiments, the retrieval enhancement generation optimization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the retrieval enhancement generation optimization method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the retrieval enhancement generation optimization method by any other suitable means (e.g., by means of firmware).

[0144] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0148] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0149] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0150] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0151] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A retrieval enhancement generation optimization method, characterized in that, include: Get the user's input of the unanswered question; The system invokes a first intelligent agent, a second intelligent agent, or a third intelligent agent to answer the question and generate a target answer result; wherein, the first intelligent agent is generated based on a first large-scale language model; the second intelligent agent is generated based on a second large-scale language model; and the third intelligent agent is generated based on a third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other. The fourth intelligent agent reviews the target solution result. If the review result is satisfactory, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on a fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first, second, and third large-scale language models.

2. The method according to claim 1, characterized in that, Invoke a first intelligent agent, a second intelligent agent, or a third intelligent agent to solve the problem to be solved, and generate a target solution result, including: The first intelligent agent is invoked to answer the question to be answered, and a first answer result is generated; If the first solution result does not meet the preset solution conditions, the second intelligent agent is invoked to solve the problem to be solved and generate a second solution result. If the second solution does not meet the preset solution conditions, a third agent is invoked to solve the problem to be solved, a third solution is generated, and the third solution is determined as the target solution.

3. The method according to claim 2, characterized in that, Invoke the first intelligent agent to solve the problem to be solved, and generate a first solution result, including: If the first solution result meets the preset solution conditions, then the first solution result is determined as the target solution result.

4. The method according to claim 2, characterized in that, When the first solution does not meet the preset solution conditions, a second agent is invoked to solve the problem and generate a second solution, including: When the first solution result does not meet the preset solution conditions, obtain the first document data corresponding to the question to be answered; The second intelligent agent is mobilized to answer the question to be answered based on the first document information, and a second answer result is generated.

5. The method according to claim 4, characterized in that, Obtain the first document data corresponding to the question to be answered, including: The first document data corresponding to the question to be answered is obtained by the retrieval device in the first intelligent agent.

6. The method according to claim 4, characterized in that, The second intelligent agent is invoked to answer the question to be answered based on the first document information, generating a second answer result, including: If the second solution result meets the preset solution conditions, then the second solution result is determined as the target solution result.

7. The method according to claim 4, characterized in that, When the second solution does not meet the preset solution conditions, a third agent is invoked to solve the problem, generating a third solution, which is then identified as the target solution, including: When the second solution result does not meet the preset solution conditions, a second document corresponding to the question to be answered is obtained; wherein, the second document is different from the first document. A third intelligent agent is mobilized to answer the question to be answered based on the second document and the first document, generating a third answer result, and the third answer result is determined as the target answer result.

8. The method according to claim 1, characterized in that, The fourth agent reviews the target solution result. If the review result is satisfactory, the target solution result is output as the final solution result, including: The matching degree between the target solution result and the question to be solved is obtained using a fourth intelligent agent, as well as the accuracy of the target solution result; If the matching degree reaches a preset first threshold and the accuracy reaches a preset second threshold, then the target solution result will be output as the final solution result.

9. A retrieval enhancement and optimization device, characterized in that, include: The unanswered question acquisition module is used to acquire unanswered questions input by the user; The target solution result generation module is used to call a first intelligent agent, a second intelligent agent, or a third intelligent agent to answer the question to be answered and generate a target solution result; wherein, the first intelligent agent is generated based on a first large-scale language model; the second intelligent agent is generated based on a second large-scale language model; the third intelligent agent is generated based on a third large-scale language model; the first large-scale language model, the second large-scale language model, and the third large-scale language model are independent training models that are different from each other. The final solution result determination module is used to review the target solution result using a fourth intelligent agent. If the review result is passed, the target solution result is output as the final solution result. The fourth intelligent agent is generated based on a fourth large-scale language model. The fourth large-scale language model is an independently trained model that is independent of the first, second, and third large-scale language models.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the retrieval enhancement generation optimization method according to any one of claims 1-8.