Quick-response method and system for enhancing generation of questions and answers through intelligent agent text retrieval

By building a large model pool that combines localization and cloud and using tool models to deal with user problems, the problems of privacy protection, cost optimization and search accuracy in the existing technology are solved, and a smart text retrieval enhancement generation question-and-answer system with fast response, efficient retrieval and high matching degree is realized.

CN120162403APending Publication Date: 2025-06-17GUANGZHOU BAIJIEYUN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192034.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing technology has shortcomings in privacy protection, cost optimization and retrieval accuracy, especially the lack of privacy in cloud inference, high local deployment costs, lack of uniformity in the agent architecture and inability to effectively solve the proxy problem.

Method used

By building a large model pool that combines localization and the cloud, reduce local deployment costs and ensure data privacy. Use tool models to deal with user problems, and use parallel computing to call local knowledge base search tools and network search tools to speed up information retrieval and integration. At the same time, the search results are optimized through the prompt word template and data enhancement module, and the referential problem is solved in combination with historical records.

Benefits of technology

It realizes a fast-responsive intelligent text retrieval enhancement generation question-and-answer system, reduces local deployment costs, ensures data privacy, improves retrieval efficiency and matching results, and meets the needs of high-concurrency tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162403A_ABST
    Figure CN120162403A_ABST
Patent Text Reader

Abstract

The invention discloses a quick-response agent text retrieval enhanced question and answer generation method and system, and the method comprises the following steps: inputting a user question, receiving the question input by the user, processing a tool large model, preprocessing the user question through the locally deployed tool large model, generating parameters to call one or more tools, and calling the tools. According to the parameters generated by the tool large model, corresponding tools including a local knowledge base retrieval tool, a networking search tool and the like are called, data are enhanced, information returned by the called tools is recombined and enhanced, and enhanced information is formed, so that the situation that the paint needs to be manually taken out for comparison is avoided; the problems that real-time color monitoring and adjustment in the paint color mixing process cannot be achieved through the method, in a traditional color mixing method, if the color deviation is large, the steps of color mixing and comparison possibly need to be repeated for multiple times, and therefore resources are wasted, and the production efficiency is reduced are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text retrieval, and in particular to a method and system for intelligent agent text retrieval enhanced generation and question answering with fast response. Background Art

[0002] With the rapid development of large language models and intelligent agent technologies, the Retrieval-Augmented Generation (RAG) technology solves the problem of lagging update of the model knowledge base through external knowledge retrieval and information enhancement. The intelligent agent technology further expands the functions of the language model, adding memory, tool invocation, and planning capabilities, and can be applied to fields such as Internet of Things control, navigation, and task planning.

[0003] The existing technologies still face various problems. First, the privacy of cloud inference is insufficient, making it difficult for users with sensitive data to accept. Second, the cost of local deployment is high, making it difficult for ordinary enterprises to afford. In addition, the intelligent agent architecture lacks unity, making it difficult to support high-concurrency tasks and multi-tool invocations. More importantly, the current retrieval enhancement technology cannot effectively solve the anaphora problem, resulting in a mismatch between the generated results and the user's questions, affecting the actual application effect. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems in the related technologies to a certain extent.

[0005] For this purpose, the object of the present invention is to propose a method and system for intelligent agent text retrieval enhanced generation and question answering with fast response. By constructing a large model pool that combines localization and the cloud, the local deployment cost is reduced while ensuring data privacy. By optimizing the tool invocation decision and parallel computing method, the retrieval and generation efficiency is improved to meet the high-concurrency requirements. Using the tool large model to handle the anaphora problem and data enhancement, the matching degree between the retrieval result and the user's question is improved.

[0006] To achieve the above object, the present invention proposes a method and system for intelligent agent text retrieval enhanced generation and question answering with fast response, including the following steps:

[0007] S1. User question input: Receive the question input by the user.

[0008] S2. Tool large model processing: Preprocess the user's question through the locally deployed tool large model to generate parameters for invoking one or more tools.

[0009] S4. Tool invocation: Invoke the corresponding tools according to the parameters generated by the tool large model, including local knowledge base retrieval tools, Internet search tools, etc.

[0010] S4. Data enhancement: Recombine and enhance the information returned by the invoked tools to form enhanced information.

[0011] S5. Agent Text Retrieval Enhanced Generation Q&A Method and System with Quick Response to Prompt Words Generation: Based on the preset agent text retrieval enhanced generation q&a method and system with quick response to prompt words and enhanced information, generate the input for the generation large model to use;

[0012] S6. Generate Q&A: Process the agent text retrieval enhanced generation q&a method and system with quick response to prompt words through the generation large model, and output the final q&a result;

[0013] S7. Output Result: Return the generated q&a result to the user;

[0014] Among them, the tool large model adopts a large model with a low parameter scale and supports the parallel operation mode of tool invocation to improve the response speed and reduce the cost; the generation large model can be deployed locally or in the cloud according to privacy requirements.

[0015] The quick response agent text retrieval enhanced generation q&a method and system of the present invention constructs a large model pool that combines localization and the cloud, selects local or cloud models to run according to privacy requirements, reduces the local deployment cost and ensures data privacy; uses the tool large model to process user questions, and utilizes the parallel operation mode to call the local knowledge base retrieval tool and the network search tool, which speeds up the information retrieval and integration speed and meets the requirements of high-concurrency tasks; at the same time, optimizes the retrieval results through the prompt word template and the data enhancement module, and solves the anaphora problem by combining historical records, significantly improving the matching degree between the retrieval results and user questions, and overcoming the deficiencies of the prior art in privacy protection, cost optimization, and retrieval accuracy.

[0016] In addition, the quick response agent text retrieval enhanced generation q&a method and system proposed above according to the present invention may also have the following additional technical features:

[0017] Specifically, the parallel operation mode of tool invocation includes simultaneously invoking multiple tools and organizing, rearranging, and enhancing the information returned by them.

[0018] Specifically, the tool large model deployed locally generates question parameters by combining historical records to solve the anaphora problem in user input.

[0019] Specifically, the intelligent agent text retrieval enhanced generation Q&A method and system with prompt word fast response includes the basic intelligent agent text retrieval enhanced generation Q&A method and system with fast response and the customized intelligent agent text retrieval enhanced generation Q&A method and system with fast response. The basic intelligent agent text retrieval enhanced generation Q&A method and system with fast response includes the intelligent agent text retrieval enhanced generation Q&A method and system with system prompt word fast response, historical record placeholder, tool information placeholder, and the intelligent agent text retrieval enhanced generation Q&A method and system with user prompt word fast response.

[0020] Specifically, the enhanced information is generated by calculating the matching degree between the user question and the information returned by the tool through a vector database, and then trimming and merging similar information.

[0021] Specifically, the generation large model includes a locally deployed large model and a cloud-deployed large model, which are uniformly managed through a large model pool, and the calling method is selected according to privacy requirements.

[0022] An intelligent agent text retrieval enhanced generation Q&A system with fast response, the system includes:

[0023] a. User question input module, used to receive user questions;

[0024] b. Tool large model module, used to generate parameters for calling tools;

[0025] c. Tool call module, used to call local knowledge base retrieval tools, network search tools, etc.;

[0026] d. Data enhancement module, used to organize and enhance the information returned by the tool;

[0027] e. Intelligent agent text retrieval enhanced generation Q&A method and system module with prompt word fast response, used to generate the input for the generation large model;

[0028] f. Generation large model module, used to generate the final Q&A result;

[0029] g. Result output module, used to return the Q&A result to the user.

[0030] Specifically, the tool large model module adopts a low-parameter scale model, and the tool call efficiency is accelerated through parallel computing. The generation large model module supports local deployment and cloud deployment to meet high privacy requirements.

[0031] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings

[0032] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:

[0033] Figure 1 is a schematic structural diagram of the dialogue workflow of the present invention;

[0034] Figure 2 is a schematic structural diagram of the AI engine architecture of the present invention;

[0035] Figure 3 is a schematic structural diagram of the tool workflow of the present invention. Detailed Embodiments

[0036] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention. On the contrary, the embodiments of the present invention include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0037] The following describes a method and system for fast-response intelligent agent text retrieval enhanced generation of questions and answers according to embodiments of the present invention in conjunction with the accompanying drawings.

[0038] As Figures 1 - 3 shown, the method and system for fast-response intelligent agent text retrieval enhanced generation of questions and answers according to embodiments of the present invention may include the following steps:

[0039] S1. User question input: Receive the question input by the user;

[0040] S2. Tool large model processing: Preprocess the user question through a locally deployed tool large model to generate parameters for invoking one or more tools;

[0041] S4. Tool invocation: Invoke the corresponding tools according to the parameters generated by the tool large model, including local knowledge base retrieval tools, network search tools, etc.;

[0042] S4. Data enhancement: Reorganize and enhance the information returned by the invoked tools to form enhanced information;

[0043] S5. Generation of the method and system for fast-response intelligent agent text retrieval enhanced generation of questions and answers: Generate the input for the generation large model based on a preset method and system for fast-response intelligent agent text retrieval enhanced generation of questions and answers and the enhanced information;

[0044] S6. Generate Q&A: Use the intelligent agent text retrieval enhanced generation Q&A method and system that processes the prompt words through a generative large model to quickly respond, and output the final Q&A result;

[0045] S7. Output the result: Return the generated Q&A result to the user;

[0046] Among them, the tool large model uses a large model with a low parameter scale and supports the parallel operation mode of tool invocation to improve the response speed and reduce costs; the generative large model can be deployed locally or in the cloud according to privacy requirements.

[0047] Further, the parallel operation mode of tool invocation includes simultaneously invoking multiple tools and organizing, rearranging, and enhancing the information returned by them.

[0048] Further, the locally deployed tool large model generates question parameters by combining historical records to solve the anaphora problem in the user input.

[0049] Further, the intelligent agent text retrieval enhanced generation Q&A method and system with fast response to prompt words includes the basic intelligent agent text retrieval enhanced generation Q&A method and system with fast response and the customized intelligent agent text retrieval enhanced generation Q&A method and system with fast response. The basic intelligent agent text retrieval enhanced generation Q&A method and system with fast response includes the intelligent agent text retrieval enhanced generation Q&A method and system with fast response to system prompt words, historical record placeholders, tool information placeholders, and the intelligent agent text retrieval enhanced generation Q&A method and system with fast response to user prompt words.

[0050] Further, the enhanced information is generated by calculating the matching degree between the user question and the information returned by the tool through a vector database, and then trimming and merging similar information.

[0051] Further, the generative large model includes a locally deployed large model and a cloud-deployed large model, which are uniformly managed through a large model pool, and the invocation method is selected according to privacy requirements.

[0052] An intelligent agent text retrieval enhanced generation Q&A system with fast response, the system includes:

[0053] a. User question input module, used to receive user questions;

[0054] b. Tool large model module, used to generate parameters for invoking tools;

[0055] c. Tool invocation module, used to invoke local knowledge base retrieval tools, online search tools, etc.;

[0056] d. Data enhancement module, used to organize and enhance the information returned by the tool;

[0057] e. An intelligent agent text retrieval enhanced generation Q&A method and system module with prompt word quick response, used to generate inputs for the generation large model;

[0058] f. A generation large model module, used to generate the final Q&A result;

[0059] g. A result output module, used to return the Q&A result to the user.

[0060] Furthermore, the tool large model module adopts a low-parameter scale model and accelerates the tool call efficiency through parallel computing. The generation large model module supports local deployment and cloud deployment to meet high privacy requirements.

[0061] Specifically, the present invention is a quick response intelligent agent text retrieval enhanced generation Q&A method and system. The method proposed by the present invention is a generative artificial intelligence engine system that can reduce the computing power cost while taking into account quick response. At the same time, combined with the retrieval enhanced generation technology, it reduces the hallucination rate of the large language model and improves the knowledge breadth of the large language model.

[0062] In common retrieval enhanced generation applications, including the official tutorials of langchain, the retrieved content is usually embedded in the context of the prompt word for the large model to understand. However, the common method is to directly input the user's question into the vector database for retrieval. There is a drawback to this retrieval method, that is, anaphora. When the user asks a question, they will use it to refer to something in the historical record. At this time, there will be a problem that the matching result does not match the question. The method to solve this anaphora problem is called anaphora resolution, and anaphora resolution is usually also completed by the large model. By inserting an intermediate link and using the form of a prompt word to let the large model regenerate the question and then retrieve it, this solution is a commonly used method at present. However, there is a better solution now, which is to use the intelligent agent to take retrieval as a working node, so that this process can be made easier to manage.

[0063] Generally, an intelligent agent is divided into four parts: a large model, memory, planning, and tools. Among them, the large model is responsible for text Q&A and decision-making, the memory provides context information for the large model, the planning formulates the working process of the intelligent agent, and the tools provide the ability for the large model to interact with the outside world.

[0064] In the usage scenario of the present invention, it is usually necessary to meet high concurrency, high reliability, and high scalability, and at the same time meet the requirement of low cost. For this, the intelligent agent generation working process formulated by the present invention is as Figure 1As shown in the figure. User questions will first be uniformly processed by a locally deployed tool large model, which determines the user and sorts out the user questions, calls one or more tools, and by default, actively calls the function of retrieving the local knowledge base. The tool returns information, which is added to the formulated prompt template, and finally submitted to the generation large model to generate the result.

[0065] Among them, in the intelligent agent application, the large language model with large parameters occupies most of the computing power. Therefore, to meet the low-cost requirement, the present invention will adopt a solution of local deployment plus cloud deployment of the large model to save costs.

[0066] In the present invention, based on the retrieval-augmented generation technology, the data storage and retrieval analysis are locally deployed. The information retrieved for user questions is usually fragmented, and even if sent to the cloud through chat messages, not much information will be leaked, thus ensuring data privacy.

[0067] Based on the intelligent agent permission grading, the conversations of the locally deployed large model and the cloud-deployed large model are diverted. Highly private generation tasks are handed over to the locally deployed large model, and the rest are handed over to the cloud large model. Multiple local large models and cloud large models work simultaneously. Even if a certain large model stops working, it does not affect the generation of conversations, thus meeting the high reliability and high concurrency of the generated conversations.

[0068] Knowledge base retrieval, online search, etc. are provided to the large model as tools for calling. The tools called are decided by the tool large model. The tool large model can summarize and generate the tool input parameters based on the historical records to solve the anaphora problem. If there are new requirements from users, tools can also be added or reduced as needed to achieve high scalability.

[0069] Finally, the AI engine architecture of the present invention is as Figure 2 shown in the figure. Here, the conventionally defined intelligent agents are reconstructed. In the intelligent agent definition, the present invention divides the intelligent agents into two categories. One category is the intelligent agents responsible for internal decision-making of the engine, and the other category is the intelligent agents used by users. Both of these two categories of intelligent agents will selectively adjust the tool definition, prompt module, large model, and work plan according to needs. The role of the internal decision-making intelligent agent is to make quick decisions and responses to user questions, so it belongs to a part of the conversation workflow.

[0070] In the agent architecture, there are two types of tool definitions: tool classes and tool information. Tool classes refer to specific class objects. By passing parameters to different tool classes, two types of tasks can be completed: one is to execute tasks, such as sending emails; the other is to retrieve tasks, such as retrieving information from a local knowledge base and from a web search engine. Among them, the retrieval task is based on retrieval enhancement and is part of the retrieval-enhanced generation technology. In this module, the tool will store all the retrieved information in a vector database, calculate the matching degree between the information chunks and the user's question in various ways, and finally return the reorganized and rearranged text information as tool information. The process of reorganizing the retrieved information is called "enhancement".

[0071] The prompt module is an enhancement of the original memory link. In the prompt module, there are two types: basic templates and custom templates. In the basic templates, system prompt templates, historical record placeholders, tool information placeholders, and user prompt templates are defined. To enhance the memory ability of the large model, this part also needs to trim the historical records and merge similar information. The custom templates are the agent settings defined by the user and are based on the basic templates.

[0072] In the large model pool, local large models and cloud large models are defined, and the traffic is split according to the model name defined by the agent. This module initially defines a chat large model, providing a chat function, a tool input interface, and a prompt template input interface for the agent. The large model providing the chat function can be various large models, including common text large models, tool large models, or multimodal large models. Classified by the deployment method, it can be further divided into locally deployed large models and cloud-deployed large models. The locally deployed large models will create a connection pool, which can connect to multiple local deployment nodes simultaneously. When called, it will be evenly distributed to each node to ensure the fluency and reliability of generation and meet the high-concurrency requirements. In addition, cloud large models can also select large model services provided by different vendors according to needs, as long as they are defined according to the interface specifications.

[0073] Finally, the planning module refers to the work process planning of the agent. The dialogue workflow is as Figure 1 shown, and the tool workflow is as Figure 3 shown.

[0074] The working process of the internal decision-making agent is the tool workflow. Currently, it is mainly used for tool decision-making and parameter generation, replacing the position of the tool large model in Figure 1 . The main role of the tool agent is to accelerate reasoning and provide a tool call function for some large models that do not support using tools.

[0075] Since the larger the parameter scale of the large model, the higher the computing power required, reducing the parameter scale can effectively accelerate the inference speed, but it will also reduce the intelligence of the large model. Therefore, in the internal decision-making agent, in order to balance the inference speed and intelligence, a large model with a parameter scale of 7B is used in the present invention.

[0076] As Figure 3 shown. The user input, combined with the chat history, the defined tool list, and other necessary parameters, is input into the tool large model in combination with the set prompt template, and the tool parameters are generated and returned by the tool large model. The main program further processes the returned parameters, calls the corresponding tool class according to the name and parameters of the called tool, and supplements other privacy parameters that cannot be input by the large model. Finally, the return information of the class is obtained and encapsulated into tool information to complete the tool call.

[0077] In this part, since the number of tool calls is completely determined by the large model, the number of tools will increase or decrease according to the user's question during operation. When there are multiple tools, the parallel operation method is adopted instead of the serial operation method of the conventional retrieval-augmented generation agent, which greatly increases the response speed.

[0078] In summary, the fast-response intelligent agent text retrieval-augmented generation question-answering method and system according to the embodiments of the present invention build a large model pool that combines localization and the cloud, select local or cloud models to run according to privacy requirements, reduce the local deployment cost and ensure data privacy; use a tool large model to process user questions, and use the parallel operation method to call local knowledge base retrieval tools and online search tools, which speeds up the information retrieval and integration speed and meets the requirements of high-concurrency tasks; at the same time, the retrieval results are optimized through the prompt template and data augmentation module, and the anaphora problem is solved in combination with the historical record, which significantly improves the matching degree between the retrieval results and the user's questions, and overcomes the deficiencies of the prior art in terms of privacy protection, cost optimization, and retrieval accuracy.

[0079] In the description of this specification, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0080] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0081] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A fast-response agent text retrieval enhanced generation question answering method, characterized in that: The following steps are involved: S1. User question input: receiving questions input by users; S2. Tool big model processing: Pre-process the user's problem through the locally deployed tool big model and generate parameters to call one or more tools; S4, tool calling: according to the parameters generated by the tool model, call the corresponding tool, including local knowledge base retrieval tool, network search tool, etc.; S4, data enhancement: reorganize and enhance the information returned by the calling tool to form enhanced information; S5. Generating an intelligent agent text retrieval enhancement generation question-answering method and system for prompt word rapid response: Based on the preset intelligent agent text retrieval enhancement generation question-answering method and system for prompt word rapid response and enhancement information, generating input for generating a large model; S6. Generate question and answer: Generate question and answer method and system by generating a large model to process prompt words and quickly respond to intelligent agent text retrieval, and output the final question and answer result; S7, output result: return the generated question and answer result to the user; Among them, the tool big model adopts a large model with low parameter scale and supports parallel computing of tool calls to improve response speed and reduce costs; the generated big model can be deployed locally or in the cloud according to privacy requirements.

2. The fast-response agent text retrieval enhanced generation question-answering method according to claim 1, characterized in that: The parallel operation mode of tool calling includes calling multiple tools at the same time, and sorting, rearranging and enhancing the information returned by them.

3. The fast-response agent text retrieval enhanced generation question-answering method according to claim 1 or 2, characterized in that: The locally deployed tool model generates question parameters by combining historical records to solve the proxy problem in user input.

4. The fast-response agent text retrieval enhanced generation question-answering method according to claim 1, characterized in that: The intelligent agent text retrieval enhanced generation question and answer method and system with prompt word quick response includes a basic intelligent agent text retrieval enhanced generation question and answer method and system with quick response and a customized intelligent agent text retrieval enhanced generation question and answer method and system with quick response. The basic intelligent agent text retrieval enhanced generation question and answer method and system with quick response includes an intelligent agent text retrieval enhanced generation question and answer method and system with quick response to system prompt words, a history record placeholder, a tool information placeholder and an intelligent agent text retrieval enhanced generation question and answer method and system with quick response to user prompt words.

5. The fast-response agent text retrieval enhanced generation question-answering method according to claim 1, characterized in that: The enhanced information is generated by calculating the matching degree between the user question and the information returned by the tool through a vector database, and by cutting and merging the same type of information.

6. The rapid response agent text retrieval enhanced generation question answering method according to claim 1 is characterized in that: The generation of the big model includes locally deploying the big model and deploying the big model in the cloud, which are uniformly managed through a big model pool, and the calling method is selected according to privacy requirements.

7. A fast-response intelligent agent text retrieval enhanced generation question answering system, characterized in that: The system includes: a. User question input module, used to receive user questions; b. Tool model module, used to generate parameters for calling tools; c. Tool calling module, used to call local knowledge base retrieval tools, online search tools, etc.; d. Data enhancement module, used to organize and enhance the information returned by the tool; e. A method and system module for enhancing the generation of question-answering by intelligent agent text retrieval with fast response to prompt words, which is used to generate input for generating large models; f. Generate a large model module to generate the final question-answering results; g. Result output module, used to return the question and answer results to the user.

8. The fast-response intelligent agent text retrieval enhanced generation question-answering system according to claim 7, characterized in that: The tool large model module adopts a low-parameter scale model and speeds up the tool calling efficiency through parallel computing. The large model generation module supports local deployment and cloud deployment to meet high privacy requirements.